HDFS distributed log anomaly online detection method, system and electronic equipment based on double model cooperation

CN122839043APending Publication Date: 2026-09-29BEIJING ZHANGSHU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610648934.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,现有日志异常检测方法大多侧重于对历史日志数据进行离线分析,仅能在异常行为发生之后进行识别,缺乏对未来潜在异常的预测能力,难以满足分布式系统对实时性和预防性的要求

Benefits of technology

与现有技术相比,本发明的有益效果主要体现在以下几个方面:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839043A_ABST
    Figure CN122839043A_ABST
Patent Text Reader

Abstract

This application discloses an online anomaly detection method, system, and electronic device for HDFS distributed logs based on dual-model collaboration, belonging to the interdisciplinary field of artificial intelligence and communication technology. The method includes the following steps: structural processing of unstructured HDFS distributed logs and conversion into log event templates; sequence partitioning of the log event templates according to the identifiers of the distributed system tasks to which the log events belong, using session grouping technology to obtain log event sequences; and word frequency statistics of events in the log event sequences to construct a dictionary. This application achieves collaborative log sequence prediction and anomaly detection through the construction of structured logs and event sequences, combined with word vector optimization and sliding window processing; it introduces Informer and weighted loss functions to improve prediction and interpretation capabilities, thereby enabling early warning and blocking of potential anomalies and improving system reliability and operational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the interdisciplinary field of artificial intelligence and communication technology, specifically to an online detection method, system, and electronic device for HDFS distributed log anomalies based on dual-model collaboration. Background Technology

[0002] With the rapid development of cloud computing, big data, and distributed computing technologies, an increasing number of application systems are being deployed using distributed architectures based on large server clusters. In these systems, each node continuously generates a large amount of log data during operation. These logs, in unstructured text format, record system operating status, event information, error alerts, and user behavior, becoming crucial data for system monitoring, fault diagnosis, and performance analysis. To address the analysis needs of distributed system logs, existing technologies typically transform raw logs into a data format suitable for machine learning processing through log structuring, event template extraction, and sequence modeling. This is further combined with deep learning models (such as recurrent neural networks, convolutional neural networks, and Transformer-like models) to achieve log anomaly detection. Furthermore, some research has introduced time series analysis methods to model log events, thereby improving the accuracy of anomaly identification.

[0003] The existing technology has the following shortcomings: However, most existing log anomaly detection methods focus on offline analysis of historical log data, enabling identification only after abnormal behavior occurs. They lack the ability to predict potential future anomalies, failing to meet the real-time and preventative requirements of distributed systems. In online detection scenarios, due to the high dimensionality, strong temporal dependence, and dynamic changes of log data, accurately predicting future log event sequences based on a limited set of observed log sequences and effectively judging their anomalies remains a significant challenge. Furthermore, existing methods typically process log sequence prediction and anomaly detection tasks independently, lacking a unified coordination mechanism. This results in prediction results failing to fully serve the anomaly detection process, limiting the improvement of overall detection performance and system early warning capabilities.

[0004] Therefore, there is an urgent need for an online method that can integrate log sequence prediction and anomaly detection to enable early warning and intervention for abnormal behavior in distributed systems.

[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this application is to provide an online detection method, system, and electronic device for HDFS distributed log anomalies based on dual-model collaboration, in order to solve the problems mentioned in the background art.

[0007] To achieve the above objectives, this application provides the following technical solution: an online detection method for HDFS distributed log anomalies based on dual-model collaboration, comprising the following steps: The unstructured text of the HDFS distributed log is structured and converted into log event templates. The log event templates are then divided into sequences using session grouping technology according to the identifiers of the distributed system tasks to which the log events belong, resulting in a log event sequence. Word frequency statistics are performed on events in the log event sequence to construct a dictionary. The word vector dimension is determined by the unitary invariance of word embedding and matrix perturbation theory. The log event sequence is then converted into the corresponding log vector sequence based on the word embedding algorithm. A sliding window and sliding step size are set according to the length distribution of the log vector sequence. Log vector sequences with a length greater than the sliding window are split and assigned the same category label. Log vector sequences with a length less than the sliding window are padded with zeros to obtain the log time series dataset. A log anomaly detection model is constructed and trained based on a log time-series dataset to obtain a log anomaly detection model for log time-series anomaly identification. A log sequence multivariate prediction model was constructed and trained using a log time series dataset to obtain a log sequence multivariate prediction model for predicting future log event sequences. For a new log event sequence, the existing log vector sequence is input into the log sequence multivariate prediction model to obtain the predicted event sequence. The existing log vector sequence and the predicted event sequence are concatenated to form new log time series data, which is then input into the log anomaly detection model for anomaly identification, thereby completing the online detection of HDFS distributed log anomalies.

[0008] Preferably, the unstructured text of the HDFS distributed log is structured and time-series log data is constructed, including: The HDFS distributed logs of unstructured text are structured and converted into a unified log event template using regular expressions. Log event templates are grouped according to task identifiers to form a log event sequence; Word frequency statistics are performed on events in the log event sequence to construct a dictionary, and the word vector dimension is determined based on the unitary invariance of word embedding and matrix perturbation theory, thus converting log events into word vectors; Based on the length distribution of log event sequences, a sliding window and sliding step size are set. Log event sequences with a length greater than the sliding window are divided into multiple subsequences and assigned the same category label. Log event sequences with a length less than the sliding window are padded with zeros to obtain a log time series dataset.

[0009] Preferably, the log anomaly detection model is a deep learning model built on a temporal convolutional network and a self-attention mechanism.

[0010] Preferably, the log sequence multivariate prediction model is a deep learning model built on Informer, and a weighted loss function is used for optimization during model training.

[0011] Preferably, online anomaly detection is performed using a trained log sequence multivariate prediction model and a log anomaly detection model, including: The unstructured logs belonging to the same task identifier are structured and converted into log event sequences using regular expressions; Convert the events in the log event sequence into word vectors to obtain the log time series; When the length of the log time series reaches a preset threshold, it is input into the log series multivariate prediction model to obtain the subsequent event series; The sequence of events that have occurred is concatenated with the sequence of subsequent events to form a new log time series, which is then input into the log anomaly detection model to obtain the anomaly detection results. Based on the anomaly detection results, determine whether the event sequence of the corresponding task identifier is abnormal, and block the subsequent execution of the corresponding task if it is determined to be abnormal.

[0012] The HDFS distributed log anomaly online detection system based on dual-model collaboration includes a log data acquisition module, a log data processing module, a log time-series data generation module, a log anomaly detection model construction and training module, a log sequence multivariate prediction model construction and training module, and an online log anomaly detection module. The log data acquisition module performs structured processing on unstructured HDFS distributed logs and converts them into log event templates. Based on the identifier of the distributed system task to which the log event belongs, the module uses session grouping technology to divide the log event templates into sequences to obtain log event sequences. The log data processing module performs word frequency statistics on events in the log event sequence to construct a dictionary, determines the word vector dimension through the unitary invariance of word embedding and matrix perturbation theory, and converts the log event sequence into the corresponding log vector sequence based on the word embedding algorithm. The log time-series data generation module sets a sliding window and sliding step size according to the length distribution of the log vector sequence. Log vector sequences with a length greater than the sliding window are split and assigned the same category label. Log vector sequences with a length less than the sliding window are padded with zeros to obtain the log time-series dataset. The log anomaly detection model construction and training module builds and trains a log anomaly detection model based on the log time-series dataset to obtain a log anomaly detection model for log time-series anomaly identification. The log sequence multivariate prediction model construction and training module uses the log time series dataset to build and train a log sequence multivariate prediction model to obtain a log sequence multivariate prediction model for predicting future log event sequences. The online log anomaly detection module, for a new log event sequence, inputs the already occurred log vector sequence into the log sequence multivariate prediction model to obtain the predicted event sequence, concatenates the already occurred log vector sequence with the predicted event sequence to form new log time series data, and inputs it into the log anomaly detection model for anomaly identification, thereby completing the online detection of HDFS distributed log anomalies.

[0013] Preferably, the log data processing module is used for: The unstructured text of the HDFS distributed log is structured, and the structured log is converted into log event templates using regular expressions. Log event templates are grouped according to task identifiers to form a log event sequence.

[0014] Preferably, the log time-series data generation module is used for: Word frequency statistics are performed on log event sequences to construct a dictionary, and the word vector dimension is determined based on the unitary invariance of word embedding and matrix perturbation theory. Log events are converted into word vectors using the skip-gram algorithm in the word embedding model word2vec. Log event sequences are divided according to the sliding window and sliding step size, and the shorter sequences are padded with zeros to obtain the log time series dataset.

[0015] Preferably, the log anomaly detection model building and training module is used to build a deep learning model based on temporal convolutional networks and self-attention mechanisms.

[0016] Preferably, the log sequence multivariate prediction model construction and training module is used to build an Informer-based deep learning model and to train and optimize the model using a weighted loss function.

[0017] Preferably, the online log anomaly detection module is used for: The new unstructured logs are structured and converted into a sequence of log events; Convert the log event sequence into a log time series; Predicting subsequent event sequences using a log sequence multivariate prediction model; The sequence of events that have occurred and the sequence of events that are predicted are concatenated and then input into the log anomaly detection model for anomaly detection. The system determines whether there are any anomalies based on the detection results, and blocks the subsequent execution of the corresponding tasks if an anomaly is found.

[0018] The technical effects and advantages provided by this application in the above technical solution are as follows: Compared with the prior art, the beneficial effects of the present invention are mainly reflected in the following aspects: The online log anomaly detection method proposed in this invention can predict future event sequences in real time based on a small number of already occurred event sequences, thereby identifying potential anomalies or signs of abnormal behavior in the system with high accuracy. By predicting future behavioral trends in advance, it can assist system maintenance personnel in timely identifying potential security risks and taking corresponding protective measures in advance, effectively preventing system failures and network attacks. This provides a technical solution with significant reference value for the field of log detection and analysis.

[0019] This invention introduces the working mode of the Informer model in the field of time series multivariate prediction into log sequence prediction tasks, constructing a log sequence multivariate prediction model. Simultaneously, considering the characteristics of log event prediction, a weighted loss function scheme is designed, jointly weighting the mean squared error between the predicted and true sequences, the cross-entropy loss between the predicted and true event sequences, and the classification error of the log anomaly detection model on the predicted sequence. The weighting coefficients are derived from a trainable Dirichlet distribution. This method improves the model's classification performance and enhances the interpretability of the prediction results while maintaining the accuracy of log event sequence prediction, providing an effective approach to solving the problem of difficulty in predicting potential future threats.

[0020] This invention proposes a method for structured processing and log event transformation of unstructured text-based HDFS distributed logs. By using regular expressions to uniformly map structured logs into standardized log event templates, and grouping them according to task identifiers, a log event sequence dataset with clear structure, explicit semantics, and temporal correlation characteristics is constructed, providing a high-quality data foundation for subsequent modeling and analysis.

[0021] To address the problem of word vector dimension selection, this invention proposes a dimension determination method based on the unitary invariance of word embeddings and matrix perturbation theory. This method adaptively selects the optimal word vector dimension by minimizing the pairwise inner product loss. While effectively reducing data dimensionality, it preserves key semantic information, thereby improving the computational efficiency and feature representation capabilities of subsequent models and enhancing overall system performance.

[0022] The sliding window and sliding step size setting method proposed in this invention fully considers the actual length distribution characteristics of HDFS log sequences, enabling most log event sequences to be directly used for model training and detection without truncation or padding, effectively improving the consistency of data preprocessing and the completeness of model input. Even with uneven log data distribution, the constructed log time-series dataset still supports the model in achieving high prediction accuracy and classification accuracy, verifying the effectiveness and applicability of this data processing method.

[0023] This invention proposes an online log anomaly detection method, system, device, and storage medium based on a dual-model collaborative approach using HDFS distributed logs in unstructured text format. By introducing a collaborative mechanism between a log sequence multivariate prediction model and a log anomaly detection model, it achieves early prediction and effective blocking of potential abnormal behaviors in executing tasks. Considering the characteristics of distributed system operation, log event sequences under the same task identifier exhibit significant temporal dependencies and behavioral continuity during execution. This invention collects new logs in real time and converts them into event sequences. It then uses a prediction model to predict the event sequences corresponding to multiple future execution actions online. The actual sequences and predicted sequences are concatenated and input into the anomaly detection model to construct an anomaly identification mechanism oriented towards future states. When an abnormal trend appears in the predicted sequence, the subsequent execution of the corresponding task can be blocked in a timely manner, thereby preventing the further spread of abnormal behavior. This online detection method, which shifts from in-process monitoring to pre-event warning, not only effectively reduces resource consumption caused by abnormal tasks in distributed systems but also significantly improves the service reliability and operational automation level of the HDFS platform. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0025] Figure 1 This is a flowchart illustrating an online detection method for anomalies in HDFS distributed logs based on a dual-model collaborative approach, according to an embodiment of the present invention. Figure 2This is a schematic diagram of the structure of an online detection system for distributed log anomalies based on dual-model collaborative HDFS according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.

[0027] This application provides, as follows: Figure 1 The online anomaly detection method for HDFS distributed logs based on dual-model collaboration, as shown, includes the following steps: The unstructured text of the HDFS distributed log is structured and converted into log event templates. The log event templates are then divided into sequences using session grouping technology according to the identifiers of the distributed system tasks to which the log events belong, resulting in a log event sequence. Word frequency statistics are performed on events in the log event sequence to construct a dictionary. The word vector dimension is determined by the unitary invariance of word embedding and matrix perturbation theory. The log event sequence is then converted into the corresponding log vector sequence based on the word embedding algorithm. A sliding window and sliding step size are set according to the length distribution of the log vector sequence. Log vector sequences with a length greater than the sliding window are split and assigned the same category label. Log vector sequences with a length less than the sliding window are padded with zeros to obtain the log time series dataset. A log anomaly detection model is constructed and trained based on a log time-series dataset to obtain a log anomaly detection model for log time-series anomaly identification. A log sequence multivariate prediction model was constructed and trained using a log time series dataset to obtain a log sequence multivariate prediction model for predicting future log event sequences. For a new log event sequence, the existing log vector sequence is input into the log sequence multivariate prediction model to obtain the predicted event sequence. The existing log vector sequence and the predicted event sequence are concatenated to form new log time series data, which is then input into the log anomaly detection model for anomaly identification, thereby completing the online detection of HDFS distributed log anomalies. Specific implementation method one: The HDFS distributed logs for each unstructured text in the log dataset are acquired and structured. The structured logs are converted into event templates using regular expressions. The generated log event templates are then grouped according to different identifiers to form a log event sequence dataset.

[0029] In this invention, the log dataset refers to an unstructured text HDFS distributed log dataset, which includes important information such as the behavior, state changes, and events of a distributed system or software at a specific point in time, recorded in text format. This type of log data is typically generated collaboratively by multiple nodes in a distributed system and is characterized by dispersed sources, inconsistent formats, and complex semantics. The log content not only covers normal operational information generated during system operation but also includes abnormal information, warning messages, and user behavior records, making it crucial for system operation status analysis and anomaly detection.

[0030] From a functional perspective, log types are generally divided into transaction logs and operation logs. Transaction logs are mainly used to describe the execution logic of user requests or system transactions. They usually contain strong causal relationships and temporal dependencies, such as the entire process of a task from submission, scheduling, execution to completion. Operation logs are mainly used to record independent events during system operation, such as heartbeat detection, resource scheduling information, and message dispatch behavior. These events are usually independent of each other but reflect the overall system operating status. These two types of logs are intertwined in distributed systems, together forming a complete system operation trajectory.

[0031] Furthermore, log data is typically collected and recorded continuously in chronological order. Under the same task identifier, different log events often exhibit significant time-series correlations and behavioral continuity. For example, in the HDFS system, the processing of a data block usually involves multiple stages such as receiving, copying, verification, and storage, and the corresponding log events have a strict execution order relationship. Therefore, serialization modeling of logs is an important prerequisite for subsequent anomaly detection.

[0032] In the specific implementation process, the unstructured text of the HDFS distributed log is first processed into a structured format using text parsing and feature extraction techniques. Specifically, this includes extracting key fields from the raw log, such as generation date, timestamp, task identifier (e.g., blk_id), process source, event level (INFO, WARN, ERROR, etc.), and action execution description, and converting them into a structured data format (e.g., key-value pairs or tables). This process can be achieved through regular expression matching, string splitting, and field positioning, thereby transforming the originally difficult-to-process free text into a clearly structured and semantically explicit log.

[0033] After obtaining the structured logs, the structured logs are further mapped to a unified format log event template using event template extraction technology. In this process, key constant information in the logs (such as operation type or action description) is retained, while variable fields (such as IP address, port number, data block number, etc.) are uniformly replaced with wildcard symbols "" to eliminate surface differences between different log instances, thereby extracting representative log patterns. For example, the log "Receiving block blk_123 src:10.0.0.1 dest:10.0.0.2" can be converted into the template "<>Receiving block<>src:<>dest:<*>".

[0034] Based on the above rules, by summarizing and generalizing the log data, a set of predefined regular expression templates can be constructed. For example, 30 representative regular expression patterns can be summarized to cover the vast majority of log types. These regular expressions are used to match and process structured logs, mapping each log entry to a corresponding event template set E=(E1, E2, ..., E30). Each event template Ei corresponds to a specific type of system behavior. For example, when a log entry matches the regular expression "<>Receiving block<>src:<>dest:<>", it is marked as event Ei, thus achieving the transformation from raw logs to discrete events.

[0035] This event template extraction process can not only significantly reduce the complexity of log data, but also achieve an abstract expression of log behavior patterns while retaining key semantic information, providing standardized input for subsequent sequence modeling and deep learning processing.

[0036] After mapping logs to event templates, the log events are further grouped according to task identifiers. Specifically, using the task identifier in the distributed system (e.g., blk_id in HDFS) as the dividing criterion, log events belonging to the same task execution process are sorted in chronological order and combined to form a log event sequence dataset. Since logs corresponding to the same task identifier typically describe a complete execution flow, this grouping method can effectively preserve the temporal structure information within the task.

[0037] For example, for the task identifier blk_1, its corresponding log event sequence can be represented as [E1, E4, E8, ...], where each event Ei represents a specific operation or state change that occurs during the execution of the task. Different task identifiers correspond to different log event sequences, thus forming a dataset containing multiple sequence samples. This log event sequence dataset not only reflects the behavioral patterns within a single task but also reflects the differences between different tasks, providing a rich data foundation for subsequent anomaly detection models.

[0038] Furthermore, in practical applications, additional processing can be applied to log event sequences, such as removing noisy events, filtering low-frequency events, and standardizing time granularity, to improve data quality. In addition, to enhance the model's generalization ability, data augmentation can be performed on the log sequences, such as by using a sliding window to generate subsequence samples, thereby expanding the training data scale.

[0039] By performing structuring processing on unstructured HDFS distributed logs, extracting event templates, and constructing sequences based on task identifiers, raw, complex log data can be transformed into a log event sequence dataset with clear structure, unified semantics, and temporal characteristics. This dataset not only provides a standardized input format for subsequent word vector representation and deep learning modeling but also lays a solid data foundation for achieving high-precision log anomaly detection and prediction. Specific Implementation Method Two: The word frequencies of events in the statistical log event sequence dataset are used to construct a dictionary. The optimal dimension of the word vectors is determined using the theories of unitary invariance of word embeddings and matrix perturbation. This is then used to construct the log time-series dataset. Specifically: First, a log event dictionary is constructed by performing a global statistical analysis on all events in the log event sequence dataset to calculate the frequency of each log event across all sequences. This dictionary uses discrete events as the basic unit, mapping each unique log event template to an index identifier for subsequent vectorization. Word frequency statistics not only determine the size of the event set but also provide basic input for subsequent word embedding models. Furthermore, in practical implementations, word frequency information can be used to filter or reduce the weight of low-frequency noise events, thereby improving the stability and generalization ability of the model training.

[0041] After constructing the dictionary, log events are represented using word embeddings. During word embedding, the choice of word vector dimension significantly impacts the final representation. Too low a dimension leads to insufficient semantic information, making it difficult to depict the complex relationships between different log events; too high a dimension introduces redundant information, increases model training time and computational cost, and may cause overfitting. Therefore, determining the optimal word vector dimension while ensuring semantic expressiveness is a crucial issue in log vectorization.

[0042] To address the aforementioned issues, this invention introduces the theory of unitary invariance of word embeddings and the theory of matrix perturbation. It evaluates the word embedding matrix using pairwise inner product loss (PIPloss), thereby achieving adaptive selection of the word vector dimension. Specifically, PIPloss measures the distance between the current word embedding matrix and the ideal optimal word embedding matrix, and its calculation formula is as follows: Here, bias represents the bias term, used to describe the signal loss caused by truncating high-dimensional information; var1 represents the estimation error of the matrix spectrum size, reflecting the impact of noise on eigenvalue estimation; and var2 represents the estimation error of the matrix direction, reflecting the degree of noise perturbation on the eigenvector direction. By summing the above three parts, the overall error level of word embedding representation under different dimension selections can be comprehensively evaluated.

[0043] Furthermore, the parameter bias originates from the information loss caused by the truncation of the signal in the (k+1)th dimension and beyond in the original high-dimensional space when the word embedding dimension is chosen as k. As the dimension k gradually increases, the amount of information retained increases, thus the bias value gradually decreases. On the other hand, var1 and var2 originate from the influence of noise on the spectral decomposition of the signal matrix. As the dimension k increases, the model's sensitivity to noise increases, leading to changes in the estimation error. Therefore, there is a trade-off between bias and var terms under different dimension choices. By iterating through or estimating the PIPloss values ​​under different dimensions, the dimension k that minimizes the PIPloss can be found, thus determining the optimal word vector dimension.

[0044] In the HDFS distributed log data scenario addressed in this invention, experimental analysis and theoretical calculations revealed that the PIPloss reaches its minimum when the word vector dimension is 4. At this point, the word embedding representation achieves the optimal balance between information preservation and noise suppression. Therefore, a dimension of 4 is selected as the vector representation dimension for log events in subsequent processing.

[0045] After determining the word vector dimensions, a word embedding model is further used to vectorize the log events. Specifically, the skip-gram training strategy from the word2vec model is adopted, treating each log event as a "word" and the sequence of log events as a "sentence." By maximizing the co-occurrence probability between the central event and its context events, a low-dimensional dense vector representation for each event is learned. This vector effectively characterizes the semantic similarity and correlation between log events, thus providing high-quality input features for subsequent time-series modeling.

[0046] In actual training, the model can be optimized by setting parameters such as window size, negative sampling rate, and learning rate. For example, window size controls the context scope, thus affecting the ability to model the correlation between events; negative sampling strategies reduce computational complexity and improve training efficiency; and the learning rate controls the model's convergence speed. Furthermore, hierarchical softmax or improved sampling strategies can be used to further enhance training performance.

[0047] After vectorizing log events, it is necessary to convert log event sequences of different lengths into time-series data in a uniform format to meet the fixed input length requirement of deep learning models. Since the lengths of log event sequences vary significantly across different tasks, directly inputting them into the model would lead to inconsistent computational complexity and affect model training stability; therefore, the sequences need to be standardized.

[0048] Specifically, the length distribution of all sequences in the log event sequence dataset is first statistically analyzed to obtain the longest sequence length Emax and the shortest sequence length Emin. The range [Emin, Emax] is then divided into multiple sub-intervals, and the number of sequences within each interval is counted to analyze the distribution characteristics of log sequence lengths. In the HDFS log data targeted by this invention, statistical results show that the length of the vast majority of log event sequences is less than 30. Therefore, 30 is chosen as the sliding window length to cover the majority of sequence samples while avoiding information loss due to excessive padding or truncation.

[0049] After determining the sliding window size, the sliding step size is further set. In this invention, a sliding step size of 4 is selected, meaning the window slides forward 4 events each time, thus ensuring sample diversity while avoiding the generation of too many redundant subsequences. By combining the sliding window and the sliding step size, a long log event sequence can be split into multiple fixed-length subsequences, and these subsequences can be used as independent samples in model training, while inheriting the category labels of the original sequence, thereby improving data utilization.

[0050] For log event sequences shorter than the sliding window size, a zero-padding strategy is used, which involves adding zero vectors to the beginning of the sequence to make its length equal to the sliding window size. This strategy can convert short sequences into fixed-length inputs without disrupting the original time order, while avoiding future information interference caused by padding at the end of the sequence.

[0051] Through the sliding window splitting and sequence padding processes described above, the original irregular-length log event sequence dataset can be transformed into a log time-series dataset with a uniform structure. Each sample in this dataset is a fixed-length vector sequence, which preserves the temporal dependencies between the original log events while also meeting the input requirements of subsequent deep learning models.

[0052] Furthermore, in practical applications, log time-series datasets can be normalized or standardized to improve the stability of model training. Additionally, auxiliary information such as time interval features and event frequency features can be introduced and fused with word vector representations to further enhance the model's ability to express complex time-series patterns.

[0053] By constructing a log event dictionary, introducing a word vector dimension selection method based on unitary invariance and matrix perturbation theory, using the skip-gram model for log event vectorization, and employing a sliding window and sequence filling strategy to construct a unified log time-series dataset, this invention achieves a complete transformation process from raw log data to high-quality time-series feature representation, providing an efficient, stable, and semantically expressive data foundation for subsequent log anomaly detection models and log sequence prediction models. Specific implementation method three: A log anomaly detection model was constructed and trained on a log time-series dataset.

[0055] Construct a log anomaly detection model, including: Based on the characteristics of log time-series datasets, a deep learning model based on temporal convolutional networks and self-attention mechanisms is constructed as a log anomaly detection model.

[0056] The log anomaly detection model employs a deep learning network structure combining a Temporal Convolutional Network (TCN) and a Self-Attention mechanism. The log anomaly detection model structure built in this implementation consists of an input layer, a TCN layer, a Self-Attention layer, and a linear layer, as detailed below: Input layer: This layer is responsible for transforming the dimension of the input log sequence x, using a 1D convolutional layer to transform the dimension 4 of sequence x into 256 dimensions.

[0057] The TCN layer is used to capture the temporal correlation of log sequences. The TCN layer consists of 6 temporal blocks. Each temporal block contains 2 dilated convolutional layers, a one-dimensional pruning layer (Chomp1d), and residual connections. The Chomp1d layer ensures that the input and output dimensions are consistent, which facilitates the addition operation of the residual connections. The residual connections help alleviate the gradient vanishing problem.

[0058] Let the dilation coefficient list be d=[1, 2, 4, 8, 16, 32], then the dilation coefficient of the i-th TemporalBlock is d[i]. The filter size of each dilated convolutional layer is 64, and the padding is defined as follows: Where kernel_size is the size of the convolution kernel.

[0059] Self-Attention layer: Used to learn the relationships between different parts of a sequence. The calculation formula is as follows: in, , , These are three equal-dimensional matrices, representing the query vector, key vector, and value vector, respectively, where dk is the dimension of the key vector. The division operation helps prevent the inner product from becoming too large, causing the softmax function to enter the saturation region and thus impairing the model's learning ability.

[0060] Output layer: This layer projects the vector output by the Self-Attention layer onto the final prediction value. It consists of three linear layers (256-128-64), and finally outputs the judgment result by the sigmoid function.

[0061] The model was tuned and trained using a log time-series dataset.

[0062] A multivariate prediction model for log sequences was constructed and trained on a log time-series dataset.

[0063] Constructing a multivariate prediction model for log sequences includes: Based on the characteristics of log time-series datasets, an Informer-based deep learning model is constructed as a multivariate prediction model for log sequences, and a weighted loss function is used to optimize the training process.

[0064] The Informer model consists of an encoder and a decoder, both of which are composed of multiple layers. Each layer contains the following two sub-layers: Probability-based self-attention (PSA) is used to capture global dependencies in an input sequence. Based on the sparsity assumption, the PSA only performs full attention computation on queries that significantly deviate from a uniform distribution, thereby filtering out important queries to reduce computational cost. The calculation formula is as follows: Where Q is the query matrix. It is the query matrix obtained through sparse sampling, where K is the key matrix, V is the value matrix, and dk is the dimension of the key vector.

[0065] Feedforward Network: Used to perform non-linear transformations on the representation at each location.

[0066] This implementation pertains to time series forecasting, with the primary goal of understanding the meaning of the input sequence and predicting the sequence at multiple future time steps. The log series multivariate forecasting model constructed in this implementation consists of an input layer, an encoder layer, and a decoder, as detailed below: Input layer: Each log time series s={s1,…,s30} is divided into two subsequences: sen={s1,…,s20} and sde={s21,…,s30}. The subsequence sen serves as the input to the encoder, while the subsequence sde is further divided into two parts: the prior sequence stoken={s21,…,s25} and the prediction sequence sτ={s26,…,s30}. The sequence sτ is replaced with a zero vector 0τ={026,…,030}, and then stoken and 0τ are concatenated to obtain {stoken, 0τ}, which serves as the input to the decoder.

[0067] Encoder: The main function of the encoder is to convert the input sequence into a semantic representation. The encoder consists of three identical layers, each of which mainly includes position encoding, probabilistic self-attention machine, residual connection and feedforward neural network, etc. The encoding layer outputs the semantic and temporal feature matrix sfeed_de of sen.

[0068] Decoder: The main function of the decoder is to generate the target sequence based on the encoder's output. The decoder consists of three identical layers, each mainly containing a mask self-attention mechanism, an encoder-decoder self-attention mechanism, residual connections, and a feedforward neural network. The decoder accepts a real time sequence `stoken` along with `0τ` ({stoken, 0τ}) as input before the predicted sequence. Its semantic and temporal features are extracted by `ProbSparseAttention`, then attention is calculated with `sfeed_de`, and finally, the predicted sequence is output through a fully connected layer.

[0069] In this invention, the weighted loss function refers to the weighted combination of the mean squared error (LMSE) between the predicted sequence and the true sequence, the cross-entropy loss (LCE) between the predicted event sequence and the true event sequence, and the classification error (LBCE) of the log anomaly detection model on the predicted sequence during the training process. The calculation formula is as follows: in, These are trainable weights that follow a Dirichlet distribution.

[0070] The model was tuned and trained using a log time-series dataset. Specific implementation method four: An online log anomaly detection system is implemented using a pre-trained log sequence multivariate prediction model and a log anomaly detection model.

[0072] Online log anomaly detection is achieved using a trained log anomaly detection model and a log sequence multivariate prediction model, including the following process: Multiple unstructured log messages belonging to the same task identifier are structured to obtain multiple structured logs belonging to the same task identifier. These structured logs are then converted into corresponding log event sequences using regular expressions. Specifically, during system operation, nodes in the distributed environment continuously generate new log data. These logs exist in unstructured text form, containing information such as timestamps, task identifiers, and event descriptions. Using the same structured processing method as before, these newly generated logs undergo field parsing and standardization. Predefined regular expression templates are then used to map the structured logs to corresponding log events, thus forming a log event sequence consistent with historical data. Maintaining consistency in data processing ensures the stability of the model input distribution and improves the reliability of online detection results.

[0073] The events in a new sequence of occurred log events are converted into word vectors, resulting in a log time series corresponding to the new sequence of occurred log events. In this process, a pre-built log event dictionary and word embedding model are invoked to map each discrete log event into a corresponding low-dimensional vector representation, thus transforming the original event sequence into a continuous vector sequence. This log time series not only preserves the semantic information of the events but also reflects the similarity and correlation between events through distance relationships in the vector space. Furthermore, since the word vector dimensions have been optimized using unitary invariance and matrix perturbation theory, this representation maintains high expressive power while possessing high computational efficiency, making it suitable for real-time processing scenarios.

[0074] Furthermore, when the number of events in a new log time series equals or exceeds a threshold, it is input into a trained log sequence multivariate prediction model to predict in real time the event sequences corresponding to multiple subsequent actions belonging to the same task identifier. In the specific implementation, this threshold is used to ensure that the input sequence has sufficient contextual information, thereby improving prediction accuracy. The log sequence multivariate prediction model learns the temporal dependencies between events based on historical log time series. After receiving the currently occurring event sequence, it can predict log events that may occur in multiple future time steps. The prediction results include not only event category information but also the corresponding probability distribution or confidence level, thus providing richer reference for subsequent anomaly detection. This prediction process is continuously executed during system operation, thereby achieving dynamic prediction of future system behavior.

[0075] Subsequently, the sequence of events that have occurred and the predicted sequence of events are concatenated to form a new log time series with a sliding window size. This new time series is then input into the trained log anomaly detection model to obtain the output recognition result. In this step, the predicted future event sequence is no longer used as an independent result but is processed uniformly with the currently occurring event sequence. The concatenation operation constructs a complete time series input that includes "past state + future prediction." This input format allows the log anomaly detection model to not only make judgments based on historical behavior but also to comprehensively consider possible future behavioral patterns, thereby achieving anomaly recognition oriented towards future states. The sliding window mechanism ensures that the input sequence length is consistent, enabling the model to run stably and maintain a data distribution consistent with the training phase.

[0076] Within the anomaly detection model, feature extraction and pattern recognition are performed on the input log time series, outputting corresponding anomaly judgment results. These results can be binary labels (normal or abnormal) or continuous probability values, representing the likelihood of an anomaly occurring in the current sequence. In practical applications, a judgment threshold can be set according to specific needs, thereby transforming the model output into a clear anomaly judgment result.

[0077] Finally, based on the output identification results, it is determined whether a new event sequence belonging to the same task identifier is abnormal. If abnormal, the subsequent execution of that task identifier is blocked in advance. During this process, when the detection results indicate an abnormal trend in the current sequence, control measures can be immediately taken for the corresponding task, such as terminating task execution, triggering an alarm mechanism, recording anomaly logs, or activating security protection strategies, thereby preventing the further spread of abnormal behavior. Because anomaly judgment is based on both actual and predicted events, the system can intervene before abnormal behavior fully occurs, realizing a shift from traditional post-event detection to pre-event warning.

[0078] Furthermore, in practical engineering applications, this online anomaly detection process can be integrated with the scheduling, monitoring, and security modules of a distributed system. For example, when a task is identified as potentially abnormal, the detection result can be fed back to the resource scheduling system, thereby preventing resources from being allocated to abnormal tasks. Simultaneously, the anomaly information can be uploaded to the monitoring platform for visualization and operational analysis. In addition, the anomaly detection results can be used as feedback signals for continuous model optimization and online learning, thereby continuously improving the overall system performance.

[0079] Through the above process, this invention realizes an online log anomaly detection method based on dual-model collaboration. Its core lies in the collaborative design of a log sequence prediction model and an anomaly detection model, allowing the prediction results to directly participate in the anomaly judgment process, thereby constructing an anomaly detection mechanism oriented towards future behavior. Compared with traditional methods, this method can not only identify anomalies that have already occurred but also detect potential anomaly trends in advance, thus significantly improving the security and stability of distributed systems.

[0080] By processing newly generated log data in real time, constructing event sequences, representing word vectors, predicting future events, and judging anomaly detection results, this invention constructs a complete online log anomaly detection process, realizing dynamic perception and early warning of the operating status of distributed systems, and has high practical application value and promotion prospects.

[0081] In this invention, an HDFS distributed log dataset of unstructured text to be analyzed and identified is obtained, the log dataset is preprocessed, and the generated log event templates are sequenced using session grouping technology based on the identifier of the distributed system task to which the log event belongs, to form a log event sequence.

[0082] A dictionary is constructed by statistically analyzing the word frequencies of events in all log event sequences. The optimal word vector dimension is determined using the unitary invariance of word embeddings and a dimension selection technique based on matrix perturbation theory. Log events are then vectorized using a word embedding algorithm. Based on the lengths of the longest and shortest log event sequences, a sliding window and sliding step size are used to divide the log event sequences. Longer sequences are split into multiple subsequences and assigned the same category label, while shorter log sequences are padded with leading zeros, thus obtaining a log time-series dataset. Based on the characteristics of the log time-series dataset, a deep learning model based on temporal convolutional networks and self-attention mechanisms is constructed as a log anomaly detection model. The model is then tuned and trained using the log time-series dataset. Furthermore, based on the characteristics of the log time-series dataset, an Informer-based log sequence multivariate prediction model is constructed. This model is also tuned and trained using the log time-series dataset, and a weighted loss function is designed to optimize the training process. Using the trained log anomaly detection model and the log sequence multivariate prediction model, an online log anomaly detection system is implemented. This system predicts and identifies anomalies in new log sequences, and if an anomaly is detected, the subsequent execution of that task identifier is prematurely blocked.

[0083] This invention enables the detection system to identify abnormal events in advance by predicting event sequences that may occur within a future period, and to promptly block subsequent actions of abnormal behavior, thus achieving the purpose of online log detection and early prevention of potential threats. At the same time, by performing event template conversion based on regular expressions and grouping processing by task identifier on unstructured HDFS distributed logs, a log event sequence dataset with clear structure and semantics is constructed.

[0084] This paper proposes an optimal word vector dimension selection technique based on the unitary invariance of word embeddings and matrix perturbation theory, which effectively reduces data dimensionality while preserving key semantic information. Log event sequence datasets are divided according to sliding windows and sliding step sizes to obtain log time-series datasets. A deep learning model based on temporal convolutional networks and self-attention mechanisms is constructed, demonstrating high accuracy in detecting log time-series data. An Informer-based deep learning model is also constructed, capable of predicting future event sequences in log time-series datasets with high precision.

[0085] An online log anomaly detection system was constructed, and real-time prediction and online anomaly detection of log time series were achieved using a pre-trained log sequence multivariate prediction model and a log anomaly detection model.

[0086] This application provides, as follows: Figure 2The HDFS distributed log anomaly online detection system shown is based on dual-model collaboration and includes a log data acquisition module, a log data processing module, a log time-series data generation module, a log anomaly detection model construction and training module, a log sequence multivariate prediction model construction and training module, and an online log anomaly detection module. The log data acquisition module performs structured processing on unstructured HDFS distributed logs and converts them into log event templates. Based on the identifier of the distributed system task to which the log event belongs, the module uses session grouping technology to divide the log event templates into sequences to obtain log event sequences. The log data processing module performs word frequency statistics on events in the log event sequence to construct a dictionary, determines the word vector dimension through the unitary invariance of word embedding and matrix perturbation theory, and converts the log event sequence into the corresponding log vector sequence based on the word embedding algorithm. The log time-series data generation module sets a sliding window and sliding step size according to the length distribution of the log vector sequence. Log vector sequences with a length greater than the sliding window are split and assigned the same category label. Log vector sequences with a length less than the sliding window are padded with zeros to obtain the log time-series dataset. The log anomaly detection model construction and training module builds and trains a log anomaly detection model based on the log time-series dataset to obtain a log anomaly detection model for log time-series anomaly identification. The log sequence multivariate prediction model construction and training module uses the log time series dataset to build and train a log sequence multivariate prediction model to obtain a log sequence multivariate prediction model for predicting future log event sequences. The online log anomaly detection module, for a new log event sequence, inputs the already occurred log vector sequence into the log sequence multivariate prediction model to obtain the predicted event sequence, concatenates the already occurred log vector sequence with the predicted event sequence to form new log time series data, and inputs it into the log anomaly detection model for anomaly identification, thereby completing the online detection of HDFS distributed log anomalies.

[0087] The HDFS distributed log anomaly online detection method based on dual-model collaboration provided in this embodiment of the invention is implemented through the aforementioned HDFS distributed log anomaly online detection system based on dual-model collaboration. For details of the specific methods and processes of the HDFS distributed log anomaly online detection system based on dual-model collaboration, please refer to the embodiments of the HDFS distributed log anomaly online detection method based on dual-model collaboration described above, which will not be repeated here.

[0088] like Figure 3As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned online anomaly detection methods for HDFS distributed logs based on dual-model collaboration. Specifically: The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The memories 310 store at least one computer program 330, which is loaded and executed by the processors 320 to enable the electronic device 300 to implement any of the online anomaly detection methods for HDFS distributed logs based on dual-model collaboration provided in the above embodiments. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. It may also include other components for implementing device functions, which will not be elaborated here. Specifically, the electronic device may be a computer, etc.

[0089] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-mentioned online anomaly detection methods for HDFS distributed logs based on dual-model collaboration.

[0090] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0091] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the above-described online anomaly detection methods for HDFS distributed logs based on a dual-model collaborative approach.

[0092] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0093] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0094] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0095] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or modules may be electrical, mechanical, or other forms.

[0097] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0098] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0099] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An online anomaly detection method for HDFS distributed logs based on dual-model collaboration, characterized in that, Includes the following steps: The unstructured text of the HDFS distributed log is structured and converted into a log event template. The log event template is then divided into sequences using session grouping technology according to the identifier of the distributed system task to which the log event belongs, thereby obtaining the log event sequence. The events in the log event sequence are counted by word frequency to construct a dictionary. The word vector dimension is determined by the unitary invariance of word embedding and matrix perturbation theory. The log event sequence is then converted into a corresponding log vector sequence based on the word embedding algorithm. A sliding window and sliding step size are set according to the length distribution of the log vector sequence. Log vector sequences with a length greater than the sliding window are split and assigned the same category label. Log vector sequences with a length less than the sliding window are padded with zeros to obtain a log time series dataset. A log anomaly detection model is constructed and trained based on the log time series dataset to obtain a log anomaly detection model for log time series anomaly discrimination. A log sequence multivariate prediction model is constructed and trained using the log time series dataset to obtain a log sequence multivariate prediction model for predicting future log event sequences. For a new log event sequence, the already occurred log vector sequence is input into the log sequence multivariate prediction model to obtain a predicted event sequence. The already occurred log vector sequence and the predicted event sequence are concatenated to form new log time-series data, which is then input into the log anomaly detection model for anomaly identification, thus completing the online detection of HDFS distributed log anomalies.

2. The online anomaly detection method for HDFS distributed logs based on dual-model collaboration according to claim 1, characterized in that, The unstructured text of the HDFS distributed log is structured and time-series log data is constructed, including: The HDFS distributed logs of unstructured text are structured and converted into a unified log event template using regular expressions. Log event templates are grouped according to task identifiers to form a log event sequence; The events in the log event sequence are counted by word frequency to construct a dictionary, and the word vector dimension is determined based on the unitary invariance of word embedding and matrix perturbation theory, so as to convert the log events into word vectors. Based on the length distribution of log event sequences, a sliding window and sliding step size are set. Log event sequences with a length greater than the sliding window are divided into multiple subsequences and assigned the same category label. Log event sequences with a length less than the sliding window are padded with zeros to obtain a log time series dataset.

3. The online anomaly detection method for HDFS distributed logs based on dual-model collaboration according to claim 2, characterized in that, The log anomaly detection model is a deep learning model built on a temporal convolutional network and a self-attention mechanism.

4. The online anomaly detection method for HDFS distributed logs based on dual-model collaboration according to claim 3, characterized in that, The log sequence multivariate prediction model is a deep learning model built on Informer, and a weighted loss function is used for optimization during model training.

5. The online detection method for HDFS distributed log anomalies based on dual-model collaboration according to claim 4, characterized in that, Online anomaly detection is performed using a pre-trained log sequence multivariate prediction model and a log anomaly detection model, including: The unstructured logs belonging to the same task identifier are structured and converted into log event sequences using regular expressions; Convert the events in the log event sequence into word vectors to obtain the log time series; When the length of the log time series reaches a preset threshold, it is input into the log series multivariate prediction model to obtain the subsequent event series; The sequence of events that have occurred is concatenated with the sequence of subsequent events to form a new log time series, which is then input into the log anomaly detection model to obtain the anomaly detection result. Based on the anomaly detection results, determine whether the event sequence of the corresponding task identifier is abnormal, and block the subsequent execution of the corresponding task if it is determined to be abnormal.

6. An online anomaly detection system for HDFS distributed logs based on dual-model collaboration, characterized in that, It includes a log data acquisition module, a log data processing module, a log time-series data generation module, a log anomaly detection model building and training module, a log sequence multivariate prediction model building and training module, and an online log anomaly detection module; The log data acquisition module performs structured processing on the unstructured text HDFS distributed logs and converts them into log event templates. According to the identifier of the distributed system task to which the log event belongs, the module uses session grouping technology to divide the log event templates into sequences to obtain a log event sequence. The log data processing module performs word frequency statistics on events in the log event sequence to construct a dictionary, determines the word vector dimension through the unitary invariance of word embedding and matrix perturbation theory, and converts the log event sequence into a corresponding log vector sequence based on the word embedding algorithm. The log time-series data generation module sets a sliding window and sliding step size for the length distribution of the log vector sequence, splits the log vector sequence with a length greater than the sliding window and assigns it the same category label, and performs forward zero-padding on the log vector sequence with a length less than the sliding window to obtain the log time-series dataset. The log anomaly detection model construction and training module constructs and trains a log anomaly detection model based on the log time-series dataset to obtain a log anomaly detection model for log time-series anomaly discrimination. The log sequence multivariate prediction model construction and training module uses the log time series dataset to construct and train a log sequence multivariate prediction model to obtain a log sequence multivariate prediction model for predicting future log event sequences. The online log anomaly detection module, for a new log event sequence, inputs the already occurred log vector sequence into the log sequence multivariate prediction model to obtain a predicted event sequence, concatenates the already occurred log vector sequence with the predicted event sequence to form new log time-series data, and inputs it into the log anomaly detection model for anomaly identification, thus completing the online detection of HDFS distributed log anomalies.

7. The HDFS distributed log anomaly online detection system based on dual-model collaboration according to claim 6, characterized in that, The log data processing module is used for: The unstructured text of the HDFS distributed log is structured, and the structured log is converted into log event templates using regular expressions. Log event templates are grouped according to task identifiers to form a log event sequence.

8. The HDFS distributed log anomaly online detection system based on dual-model collaboration according to claim 6, characterized in that, The log time-series data generation module is used for: Word frequency statistics are performed on log event sequences to construct a dictionary, and the word vector dimension is determined based on the unitary invariance of word embedding and matrix perturbation theory. Log events are converted into word vectors using the skip-gram algorithm in the word embedding model word2vec. Log event sequences are divided according to the sliding window and sliding step size, and short sequences are padded with zeros to obtain a log time series dataset. The log anomaly detection model building and training module is used to build a deep learning model based on temporal convolutional networks and self-attention mechanisms; The log sequence multivariate prediction model building and training module is used to build an Informer-based deep learning model and to train and optimize the model using a weighted loss function.

9. The HDFS distributed log anomaly online detection system based on dual-model collaboration according to claim 8, characterized in that, The online log anomaly detection module is used for: The new unstructured logs are structured and converted into a sequence of log events; Convert the log event sequence into a log time series; Predicting subsequent event sequences using a log sequence multivariate prediction model; The sequence of events that have occurred and the sequence of events that are predicted are concatenated and then input into the log anomaly detection model for anomaly detection. The system determines whether there are any anomalies based on the detection results, and blocks the subsequent execution of the corresponding tasks if an anomaly is found.

10. An electronic device, characterized in that, The device includes a processor and a memory, wherein the memory stores a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any one of claims 1 to 5.