Abnormal Detection Model Training, Abnormal Detection Method and System for Cloud-Edge-Terminal System

Through the cloud-edge collaborative pre-training plus transfer learning method, the abnormal detection model is trained for massive terminal applications, and the problems of low training efficiency and difficulty in obtaining labeled data in the existing technology are solved, and efficient abnormal detection model training and adaptation are achieved.

CN115766518BActive Publication Date: 2025-06-03709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211474687.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-23
Publication Date
2025-06-03
Estimated Expiration
2042-11-23

AI Technical Summary

Technical Problem

The existing technology cannot efficiently train anomaly detection model for massive terminal applications, and requires a large amount of labeling data. The cost of obtaining labeling is high, and traditional cloud computing platforms cannot adapt efficiently.

Method used

The cloud-edge collaboration method is adopted, and the training method of pre-training and transfer learning is used to train the exception detection model for the log sequence of the cloud-edge end system. The method includes feature extraction and model adaptation at the edge end, and pre-training and self-supervised training of the initial model in the cloud.

Benefits of technology

It realizes that the abnormal detection model with good results is trained in the case of only a small amount of labeled data or incomplete labeled data, and meets the abnormal detection accuracy and real-time requirements of the huge cloud edge collaborative system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766518B_ABST
    Figure CN115766518B_ABST
Patent Text Reader

Abstract

The present invention provides an anomaly detection model training method, an anomaly detection method and a system for a cloud-edge-terminal system. The model training method includes: each edge terminal extracts features from the first sample log sequence to obtain a first vector representation sequence, and sends the first vector representation sequence and the corresponding anomaly annotation information to the cloud; the cloud performs supervised training on the initial classification model based on the first vector representation sequences and the corresponding anomaly annotation information sent by each edge terminal, obtains a pre-trained classification model and sends it to each edge terminal; each edge terminal adds an adaptation layer to the pre-trained classification model as a transfer model, and extracts features from the second sample log sequence to obtain a second vector representation sequence; each edge terminal performs supervised training on the transfer model based on the second vector representation sequence and the corresponding anomaly annotation information to obtain an anomaly detection model corresponding to each edge terminal. The present invention realizes efficient training and adaptation of anomaly detection models for a large number of terminal applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing and intelligent operation and maintenance technology, and more specifically, to an anomaly detection model training, anomaly detection method and system for a cloud-edge system. Background Art

[0002] With the rapid development of 5G (The 5th Generation Mobile Communication Technology) and IoT (Internet of Things) networking devices, traditional cloud computing architecture cannot meet the computing needs of massive terminal devices. The cloud-edge-end collaborative system can give full play to the high efficiency of cloud computing and the low latency of edge computing, and is an important architecture for future digital transformation. As the system becomes larger and larger and the terminal devices become more and more complex, the requirements for intelligent operation and maintenance are getting higher and higher.

[0003] Traditional anomaly detection methods rely heavily on system rules and domain knowledge, which consumes a lot of manpower and has poor versatility. Using machine learning algorithms to automatically learn, refine and summarize rules from massive operation and maintenance data, and turning the past process of manually summarizing operation and maintenance rules into an automatic learning process, that is, intelligent operation and maintenance, is an inevitable trend in the development of operation and maintenance technology. However, unsupervised learning has poor accuracy; supervised learning methods require a large amount of labeled data when training models. In actual scenarios, the cost of obtaining annotations is often very high, and traditional cloud computing platforms cannot efficiently train and adapt models for massive terminal applications. Summary of the invention

[0004] In view of the defects of the prior art, the purpose of the present invention is to provide an anomaly detection model training, anomaly detection method and system for a cloud-edge system, aiming to solve the problem that the prior art cannot efficiently train anomaly detection models for massive terminal applications.

[0005] To achieve the above objectives, in a first aspect, the present invention provides a method for training an anomaly detection model of a cloud-edge-device system, comprising:

[0006] S101 Each edge terminal extracts features from a first sample log sequence to obtain a first vector representation sequence, and sends the first vector representation sequence and corresponding anomaly annotation information to the cloud;

[0007] S102 The cloud performs supervised training on the initial classification model based on the first vector representation sequence and the corresponding abnormal annotation information sent by each edge, obtains a pre-trained classification model and sends it to each edge;

[0008] At each edge device, an adaptation layer is added to the pre-trained classification model as a transfer model, and feature extraction is performed on the second sample log sequence to obtain a second vector representation sequence;

[0009] At each edge device, based on the second vector representation sequence and the corresponding anomaly annotation information, supervised training is performed on the transfer model to obtain an anomaly detection model corresponding to each edge device.

[0010] In an optional example, before step S101, it further includes:

[0011] The cloud performs self-supervised training on the pre-trained model based on the third sample log sequence, obtains a feature extraction model, and sends it to each edge device;

[0012] Each edge device performs feature extraction on the first sample log sequence to obtain a first vector representation sequence, including:

[0013] Each edge device performs feature extraction on the first sample log sequence based on the feature extraction model to obtain a first vector representation sequence.

[0014] In an optional example, each edge device performs feature extraction on the first sample log sequence to obtain a first vector representation sequence, including:

[0015] Each edge device tokenizes each log in the first sample log sequence to obtain a tokenization result, and converts the tokenization result into a sub-word sequence based on a preset vocabulary;

[0016] Each edge device encodes the sub-word sequence to obtain a vector sequence, and takes the average of each vector in the vector sequence to obtain a vector representation;

[0017] Each edge device forms the vector representations of the respective logs into a first vector representation sequence.

[0018] In an optional example, the adaptation layer includes a first projection layer, a second projection layer, and an activation layer between the first projection layer and the second projection layer.

[0019] In a second aspect, the present invention provides an anomaly detection method for a cloud-edge device system, which is applied to each edge device, and the method includes:

[0020] S601 Perform feature extraction on the log sequence reported by the terminal to obtain a vector representation sequence;

[0021] S602 Input the vector representation sequence into the anomaly detection model, obtain the anomaly detection result of the log sequence, and send it to the cloud;

[0022] Wherein, the anomaly detection model is trained based on the anomaly detection model training method of the cloud-edge device system as described in the first aspect.

[0023] In a third aspect, the present invention provides an abnormal detection model training system for a cloud-edge-terminal system, and the system includes each edge terminal and a cloud;

[0024] Each of the edge terminals is used to extract features from the first sample log sequence to obtain a first vector representation sequence, and send the first vector representation sequence and the corresponding abnormal annotation information to the cloud;

[0025] The cloud is used to perform supervised training on the initial classification model based on the first vector representation sequence and the corresponding abnormal annotation information sent by each edge terminal, obtain a pre-trained classification model, and send it to each edge terminal;

[0026] Each of the edge terminals is further used to add an adaptation layer to the pre-trained classification model as a transfer model, and extract features from the second sample log sequence to obtain a second vector representation sequence;

[0027] Each of the edge terminals is further used to perform supervised training on the transfer model based on the second vector representation sequence and the corresponding abnormal annotation information to obtain an abnormal detection model corresponding to each edge terminal.

[0028] In an optional example, the cloud is further used to perform self-supervised training on the pre-trained model based on a third sample log sequence, obtain a feature extraction model, and send it to each edge terminal;

[0029] Each of the edge terminals is specifically used to extract features from the first sample log sequence based on the feature extraction model to obtain a first vector representation sequence.

[0030] In an optional example, each of the edge terminals is specifically used to perform word segmentation on each log in the first sample log sequence to obtain a word segmentation result, and convert the word segmentation result into a sub-word sequence based on a preset vocabulary; encode the sub-word sequence to obtain a vector sequence, and take the average of each vector in the vector sequence to obtain a vector representation; and, form a first vector representation sequence with the vector representations of each log.

[0031] In an optional example, the adaptation layer added by each of the edge terminals includes a first projection layer, a second projection layer, and an activation layer between the first projection layer and the second projection layer.

[0032] In a fourth aspect, the present invention provides an abnormal detection system for a cloud-edge-terminal system, and the system is applied to each edge terminal, and the system includes:

[0033] A feature extraction module, configured to extract features from the log sequence reported by the terminal to obtain a vector representation sequence;

[0034] Anomaly detection module, configured to input the vector representation sequence into an anomaly detection model, obtain the anomaly detection result of the log sequence, and send it to the cloud;

[0035] Wherein, the anomaly detection model is trained based on the anomaly detection model training method of the cloud-edge-terminal system as described in the first aspect.

[0036] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects:

[0037] The present invention provides an anomaly detection model training, anomaly detection method and system for a cloud-edge-terminal system. For the log sequence of the cloud-edge-terminal system, through cloud-edge collaboration, a training method of pre-training plus transfer learning is adopted to finally obtain an anomaly detection model suitable for each edge terminal, which can efficiently train and adapt the model for a large number of terminal applications, and can also train an anomaly detection model with good training effect even when there is only a small amount of labeled data or the labeled data is imperfect. At the same time, it can meet the requirements of the accuracy and real-time performance of anomaly detection for the increasingly large-scale cloud-edge-terminal collaborative system. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flowchart of the anomaly detection model training method provided by an embodiment of the present invention;

[0039] Figure 2 is an architecture diagram of the cloud-edge-terminal system provided by an embodiment of the present invention;

[0040] Figure 3 is a schematic diagram of the vector representation of the log text provided by an embodiment of the present invention;

[0041] Figure 4 is an architecture diagram of the initial classification model provided by an embodiment of the present invention;

[0042] Figure 5 is an architecture diagram of the transfer model provided by an embodiment of the present invention;

[0043] Figure 6 is a flowchart of the anomaly detection method provided by an embodiment of the present invention;

[0044] Figure 7 is an architecture diagram of the anomaly detection model training system provided by an embodiment of the present invention;

[0045] Figure 8 is an architecture diagram of the anomaly detection system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] The cloud-edge-end system mainly includes: cloud platform layer, communication network layer, edge computing layer and terminal layer; the cloud platform layer includes cloud computing server center, database, and cloud file storage system, which has high computing efficiency; the network communication layer includes a variety of network communication methods, which is responsible for information transmission between the cloud platform layer, edge layer and terminal devices; the edge computing layer includes edge devices and local file storage system; the terminal layer includes a variety of access devices and services. As the system becomes larger and larger and the terminal devices become more and more complex, the requirements for intelligent operation and maintenance are getting higher and higher.

[0048] In recent years, many log-based automatic anomaly detection methods have been proposed, which are mainly divided into the steps of log parsing, feature extraction, and anomaly detection. Representative log parsing methods include Drain and Spell, which aim to convert unstructured log texts in various formats into structured log templates. The methods used for feature extraction and anomaly detection are mainly divided into unsupervised learning methods and supervised learning methods. Unsupervised learning methods usually use machine learning methods such as clustering and PCA (Principal Component Analysis), without the need for additional annotation of normal logs or abnormal logs; supervised learning methods generally learn the abnormal pattern of logs based on the annotation of abnormal logs, so as to achieve the purpose of anomaly detection, and usually use deep learning methods such as CNN (Convolutional Neural Networks) and LSTM (Long Short-Term Memory). In recent years, some methods based on natural language processing have also emerged.

[0049] The above methods usually cannot meet the efficiency and accuracy requirements of real-time processing of massive logs. Simply put: First, the accuracy of unsupervised learning is poor; supervised methods require a large amount of labeled data when training models. In actual scenarios, the cost of obtaining annotations is often very high, and traditional cloud computing platforms cannot efficiently train and adapt models for massive terminal applications. At the same time, traditional transfer learning methods do not perform well in the current complex types of IOT systems. Existing technologies cannot meet the accuracy and real-time requirements of anomaly detection in the increasingly large-scale cloud-edge-end collaborative system.

[0050] To overcome the deficiencies of the prior art, the present invention provides a method for training an anomaly detection model for a cloud-edge-end system. Through cloud-edge collaboration, a better-performing anomaly detection model can be trained even with only a small amount of labeled data or imperfect labeled data, so as to timely and proactively discover service anomalies in systems or applications from log information, and thus timely take countermeasures to improve the stability of the system.

[0051] Figure 1 It is a schematic flowchart of the anomaly detection model training method provided by an embodiment of the present invention. As Figure 1 shown, the method specifically includes:

[0052] Step S101, each edge-end extracts features from the first sample log sequence to obtain a first vector representation sequence, and sends the first vector representation sequence and the corresponding anomaly annotation information to the cloud;

[0053] Step S102, the cloud performs supervised training on the initial classification model based on the first vector representation sequence and the corresponding anomaly annotation information sent by each edge-end, obtains a pre-trained classification model, and sends it to each edge-end;

[0054] Step S103, each edge-end adds an adaptation layer to the pre-trained classification model as a transfer model, and extracts features from the second sample log sequence to obtain a second vector representation sequence;

[0055] Step S104, each edge-end performs supervised training on the transfer model based on the second vector representation sequence and the corresponding anomaly annotation information, and obtains an anomaly detection model corresponding to each edge-end.

[0056] Here, the sample log sequence is a sequence composed of multiple consecutive sample logs, and the sample log here is the historical log text used as the model training sample. After each log text undergoes feature extraction, a corresponding vector representation will be obtained, and thus a vector representation sequence can be formed. The first vector representation sequence is used for the pre-training of the anomaly detection classification model, and the second vector representation sequence is used for the transfer learning of the anomaly detection classification model.

[0057] Specifically, Figure 2 It is an architecture diagram of the cloud-edge-end system provided by an embodiment of the present invention. As Figure 2 shown, the architecture of the cloud-edge-end system includes a terminal, edge-ends 1-N, and a cloud that are sequentially communicatively connected. The terminal and the edge-ends are installed in the same local network segment, and the cloud is installed in a remote network segment. The terminal generates log texts and transfers them to the edge-ends; the edge-ends receive the log texts and transfer them to the cloud.

[0058] Each edge terminal uses a pre-built feature extraction model to extract features from the first sample log sequence composed of part of the local log text, obtaining a first vector representation sequence, and performs anomaly annotation on the first sample log sequence to obtain anomaly annotation information, and transmits the first vector representation sequence and the corresponding anomaly annotation information to the cloud. Subsequently, the cloud performs supervised pre-training on the initial classification model of the multi-layer transformer structure according to the first vector representation sequences and the corresponding anomaly annotation information sent by all edge terminals, obtains a pre-trained classification model, and distributes the pre-trained classification model to each edge terminal.

[0059] In the transfer learning stage of the system, each edge terminal receives the pre-trained classification model transmitted by the cloud, adds an adaptation layer with a small number of parameters to the model to obtain a transfer model, and performs anomaly annotation on the second sample log sequence composed of a small amount of local edge domain log text collected. The feature extraction model is used to extract features from the second sample log sequence, and model transfer training is performed according to the obtained second vector representation sequence and the corresponding anomaly annotation information to obtain an independent anomaly detection model for each edge terminal. Finally, each edge terminal can use the feature extraction model and the anomaly detection model to extract features and judge outliers for the real-time collected log sequence, and feedback the detection results to the cloud for management and storage.

[0060] It should be noted that since whether a log is determined to be abnormal is related to its own content and also related to its previous and subsequent logs, the embodiments of the present invention train the anomaly detection model based on the log sequence, which can improve the accuracy of the anomaly detection model, and thus can improve the anomaly detection accuracy of the cloud-edge-terminal system. In addition, when the edge terminal performs model transfer training, each edge domain is independently trained, and only the parameters of the adaptation layer are iterated during training, and other parameters of the model remain fixed.

[0061] For the cloud-edge-terminal system log sequence, first use the computing efficiency of the cloud to perform pre-training to obtain a pre-trained classification model for the log anomaly detection task that is common to each edge terminal, and then perform model transfer on this basis to obtain an independent anomaly detection model for each edge terminal, which can improve the training efficiency and reliability of the anomaly detection model. Moreover, by adopting the training method of pre-training plus transfer learning, it is possible to use the existing labels and only need to obtain a small number of additional labels to perform model transfer.

[0062] The method provided by the embodiments of the present invention, for the cloud-edge-end system log sequence, through cloud-edge collaboration, adopts a training method of pre-training plus transfer learning to finally obtain an anomaly detection model suitable for each edge-end, which can efficiently train and adapt the model for a large number of terminal applications, and can also train an anomaly detection model with good training effects in the case of only a small amount of labeled data or imperfect labeled data, and at the same time can meet the requirements of the accuracy and real-time performance of anomaly detection for the increasingly large-scale cloud-edge-end collaborative system.

[0063] Based on the above embodiments, before step S101, it further includes:

[0064] The cloud end performs self-supervised training on the pre-trained model based on the third sample log sequence, obtains a feature extraction model and sends it to each edge end;

[0065] In step S101, each edge end extracts features from the first sample log sequence to obtain a first vector representation sequence, including:

[0066] Each edge end extracts features from the first sample log sequence based on the feature extraction model to obtain a first vector representation sequence.

[0067] Specifically, in order to ensure the application effect of the feature extraction model in the log feature extraction task, in the embodiments of the present invention, the feature extraction model needs to be obtained by performing self-supervised training on the pre-trained model according to a large amount of diverse log text data, that is, the third sample log sequence. And because this training requires high computational efficiency and a large amount of diverse log data, preferably, this step is performed on the cloud end, that is, the cloud end receives a large amount of log text transmitted by each edge end, trains the feature extraction model, and distributes the feature extraction model to each edge end for edge domain log feature extraction.

[0068] Here, the pre-trained model can be a RoBERTa model. When performing self-supervised training on the pre-trained model, a dynamic masking mechanism can be used, that is, randomly mask some sub-words from the input text content, so that the model predicts the sub-word through the context content.

[0069] After each edge end receives the feature extraction model, it can use the feature extraction model to extract features from each log in the first sample log sequence, obtain the vector representation of each log and thus form a first vector representation sequence, and send it to the cloud end for subsequent pre-training of the classification model. Similarly, each edge end can also use the feature extraction model to extract features from each log in the second sample log sequence, obtain the vector representation of each log and thus form a second vector representation sequence for subsequent training of the transfer model.

[0070] The method provided by the embodiments of the present invention, for the cloud-edge-end system logs, first obtains a feature extraction model through first-level pre-training, and then based on the vector representation sequence obtained by the feature extraction model, completes the pre-training and transfer learning of the classification model, so as to use the existing unlabeled data and existing labels, and obtains an anomaly detection model by using a training method of two-level pre-training plus transfer learning, further ensuring that an anomaly detection model with good training effect can be obtained even when there is only a small amount of labeled data or the labeled data is imperfect.

[0071] Based on any of the above embodiments, each edge-end extracts features from the first sample log sequence to obtain a first vector representation sequence, including:

[0072] Each edge-end tokenizes each log in the first sample log sequence to obtain a tokenization result, and converts the tokenization result into a sub-word sequence based on a preset vocabulary;

[0073] Each edge-end encodes the sub-word sequence to obtain a vector sequence, and takes the average of each vector in the vector sequence to obtain a vector representation;

[0074] Each edge-end forms the vector representations of each log into a first vector representation sequence.

[0075] Specifically, Figure 3 is a schematic diagram of the vector representation of the log text provided by the embodiments of the present invention. The feature extraction model (i.e., the log language model in Figure 3 ) can be composed of a tokenizer and a multi-layer Transformer encoder to complete the feature extraction of each log text in the first sample log sequence. For a log text with a length of L, the tokenizer tokenizes the log text to obtain a tokenization result and converts the tokenization result into a sub-word sequence with a length of M (L is less than or equal to M) based on a pre-set vocabulary Z with sub-word granularity, and then the multi-layer Transformer encoder converts it into a vector sequence with M dimensions each with a length of D. Finally, each input log text can be converted into a multi-dimensional vector containing log semantic information, and at the same time, the context information of the log text content can be well retained and represented. These M vector sequences are context-related, that is, the value of each vector in the vector sequence is affected by the other vectors. Take the average of these M vectors as the vector representation x of this log text i .

[0076] Sub-words can better avoid the Out Of Vocabulary problem compared with words, balancing the vocabulary size and semantic independence. Further, to ensure that the vocabulary is applicable to the log field, in the embodiments of the present invention, the cloud constructs a log corpus from the log texts forwarded by each edge device from different domains and terminal devices, and uses the byte pair encoding method for the log corpus to generate a sub-word granularity vocabulary A, which is merged with the sub-word granularity vocabulary B used by the original RoBERTa model to form a log-specific vocabulary Z and transmitted to each edge device, where Z = A ∪ B.

[0077] It should be noted that existing deep learning methods usually first use the log parsing method, but the log parsing process itself will bring parsing errors, which will in turn affect the accuracy of subsequent anomaly detection. In this regard, the embodiments of the present invention do not parse the original log into a fixed template, but directly input the original log into the network, which is segmented into individual sub-word tokens by a tokenizer, and the Transformer converts each token into a context-related feature vector. Therefore, finally, each input log text can be converted into a multi-dimensional vector containing log semantic information, thereby avoiding the errors caused by parsing by adopting a set of non-parsing log analysis methods and further improving the accuracy of the anomaly detection model.

[0078] Based on any of the above embodiments, the cloud inputs the first vector representation sequence X and the corresponding anomaly annotation information into an initial classification model containing multiple layers of Transformer encoders for supervised learning training to obtain a pre-trained classification model.

[0079] Considering whether a log is determined to be abnormal is related to its own content and also related to the order of its previous and subsequent logs. Figure 4 is the architecture diagram of the initial classification model provided by the embodiments of the present invention. As Figure 4 shown, in the embodiments of the present invention, the cloud applies the sine encoder to generate a position encoding vector p for the received first vector representation sequence X according to the sine and cosine functions at the position i of the vector representation x of each log in X i in X, and then adds the corresponding position x i and p i and inputs the sum into the initial classification model for parameter training. i

[0080] The initial classification model consists of E Transformer encoders, and typically, E ranges from 2 to 20. For a single Transformer encoder, the multi-head attention layer calculates the attention score matrix for each log message with different attention patterns. The attention scores are calculated by training the Q and K matrices of the attention layer. Different attention patterns are obtained through multi-head attention, which enables the model to consider which attention scores are significant, similar to the multiple convolutional kernels used in each layer of a CNN. The typical value of the number of attention heads is 5 to 20. The multi-head attention layer is followed by a fully connected network, and the typical value of the network size is 1000 to 5000.

[0081] The output of the multi-layer Transformer encoder is fed into a pooling layer, a dropout layer, and a fully connected layer, and the softmax method is used to calculate the anomaly probability value, as Figure 4 shown. The initial learning rate can be set to 3e-4, the mini-batch size and the dropout rate can be set to 64 and 0.1 respectively; cross-entropy is used as the loss function.

[0082] Similarly, when each edge device performs model transfer training, it also needs to input the second vector representation sequence and the corresponding position encoding sequence into the transfer model for parameter training.

[0083] Based on any of the above embodiments, the adaptation layer includes a first projection layer, a second projection layer, and an activation layer between the first projection layer and the second projection layer.

[0084] Specifically, transfer learning of the anomaly detection classification model is performed during the deployment phase in the system edge domain. Figure 5 is the architecture diagram of the transfer model provided by the embodiments of the present invention. Each edge device receives the pre-trained classification model distributed by the cloud as the base model, keeps the network parameters in the model unchanged, and adds an adaptation layer (i.e., Figure 5 the adapter in Figure 5 as the transfer model in each Transformer encoder structure of the pre-trained classification model. Here, the adaptation layer includes a first projection layer (i.e., Figure 5 the upper projection layer in

[0085] ), a second projection layer (i.e., up the lower projection layer in down ), and an activation layer between the first projection layer and the second projection layer.

[0086] h' = W up tanh(W down h) + h

[0087] The edge terminals 1 - N respectively use a small amount of local edge domain log texts collected by themselves to form a second sample log sequence, perform anomaly annotation, and through the aforementioned feature extraction model, complete the feature extraction of the second sample log sequence and output a second vector representation sequence. Subsequently, using the method of supervised learning, the second vector representation sequence and the corresponding anomaly annotation information are input into the transfer model for training the parameters of the adaptation layer, and finally an anomaly detection model more suitable for each edge domain is obtained.

[0088] Based on any of the above embodiments, the present invention further provides an anomaly detection method for a cloud - edge - terminal system. Figure 6 It is a schematic flowchart of the anomaly detection method provided by the embodiment of the present invention, as Figure 6 shown. This method is applied to each edge terminal, and specifically includes:

[0089] Step S601, extract features from the log sequence reported by the terminal to obtain a vector representation sequence;

[0090] Step S602, input the vector representation sequence into the anomaly detection model, obtain the anomaly detection result of the log sequence and send it to the cloud;

[0091] Among them, the anomaly detection model is trained based on the anomaly detection model training method for the cloud - edge - terminal system provided in any of the above embodiments.

[0092] Specifically, the log sequence is a sequence composed of multiple consecutive log texts. The edge terminal collects the log data reported by the terminal in real - time. After each log text undergoes feature extraction, a corresponding vector representation can be obtained, and thus a vector representation sequence can be formed.

[0093] Immediately, according to the anomaly detection model trained by the anomaly detection model training method for the cloud - edge - terminal system provided in any of the above embodiments, the vector representation sequence is input into the anomaly detection model to obtain the anomaly detection result of the log sequence, and it is fed back to the cloud for management and storage.

[0094] The method provided by the embodiment of the present invention, through cloud - edge collaboration, adopts a training method of pre - training plus transfer learning to finally obtain an anomaly detection model suitable for each edge terminal, complete real - time and accurate anomaly detection of log data, so as to timely and actively discover service anomalies of systems or applications from log information, and thus take corresponding measures in time to improve the stability of the system.

[0095] Based on any of the above embodiments, step S601 can be completed by the edge side using the feature extraction model provided in the above embodiments. For each log text in the log sequence, the log text is tokenized by a tokenizer to obtain a tokenization result, and the tokenization result is converted into a sub-word sequence according to the pre-set vocabulary Z of sub-word granularity, and then converted into a fixed-dimensional vector sequence containing log semantic information by a multi-layer Transformer encoder. The average of the vectors in the vector sequence is taken as the vector representation of this log text. In this way, the vector representations of all log texts in the log sequence can be obtained, forming a vector representation sequence.

[0096] It should be noted that existing deep learning methods usually first use log parsing methods, but the log parsing process itself will bring parsing errors, which will affect the accuracy of subsequent anomaly detection. In this regard, the embodiments of the present invention do not parse the original log into a fixed template, but directly input the original log into the network, cut it into individual sub-word tokens by a tokenizer, and the Transformer converts each token into a context-related feature vector. Therefore, finally, each input log text can be converted into a multi-dimensional vector containing log semantic information, thereby avoiding the errors caused by parsing by adopting a set of non-parsing log analysis methods, and further improving the accuracy of subsequent anomaly detection.

[0097] Based on any of the above embodiments, each edge side can apply a sine encoder to perform positional encoding on each vector representation in the vector representation sequence, thereby obtaining each positional encoding vector and forming a positional encoding sequence. Then, the vector representation sequence and the corresponding positional encoding sequence are added and input into the anomaly detection model, output to a pooling layer, a dropout layer and a fully connected layer through a multi-layer Transformer encoder, and then the softmax layer is used to calculate the anomaly probability value, and finally the anomaly detection result of the log sequence is obtained and fed back to the cloud for management and storage.

[0098] Based on any of the above embodiments, the present invention provides a cloud-edge-end system fault diagnosis method based on log feature extraction, and the method includes:

[0099] (1) The pre-training part of the feature extraction model:

[0100] 1.1 The cloud receives a large number of existing log sequences collected from different edge sides as a log corpus. The log sequence is a sequence composed of single logs, for example:

[0101] “L1:081109 203521 146 INFO dfs.DataNode$PacketResponder:PacketResponder for block blk_7503483334202473044 terminating

[0102] L2:081109 203521 146 INFO dfs.DataNode$PacketResponder:Received blockblk_7503483334202473044 of size 233217 from / 10.251.71.16

[0103] L3:081109 203521 148 INFO dfs.DataNode$PacketResponder:PacketResponder 2 for block blk_7503483334202473044 terminating”

[0104] L4:1117838570 2005.06.03 R02-M1-N0-C:J12-U11 2005-06-03-15.42.50.363779 R02-M1-N0-C:J12-U11 RAS KERNEL INFO instruction cache parityerror corrected

[0105] L5:1117979336 2005.06.05 R02-M0-NC-C:J04-U01 2005-06-05-06.48.56.695380 R02-M0-NC-C:J04-U01 RAS KERNEL INFO generating core.1415

[0106] L6:1118538740 2005.06.11 R30-M0-N9-C:J16-U01 2005-06-11-18.12.20.931990 R30-M0-N9-C:J16-U01 RAS KERNEL FATAL data TLB error interrupt“

[0107] L7:1104566421 2005.01.01 sadmin1 Jan 1 00:00:21sadmin1 / sadmin1kernel:hda:drive not ready for command

[0108] L8:1104566421 2005.01.01sadmin1 Jan 1 00:00:21sadmin1 / sadmin1 kernel:hda:status error:status=0x00{}

[0109] L9:1104566423 2005.01.01 sn209 Jan 1 00:00:23sn209 / sn209 sendmail

[17795] :unable to qualify my own domain name(sn209)--using short name”

[0110] They are logs collected from three different edge devices. Among them, L1 - L3 are from edge device 1, L4 - L6 are from edge device 2, and L7 - L9 are from edge device 3. The log corpus is randomly combined after preprocessing the log text. The preprocessing mainly refers to making each log occupy a single line of text and converting all uppercase letters to lowercase letters, etc.

[0111] 1.2 Based on the log language corpus described in the previous section, the cloud platform server, i.e., the cloud, uses the byte - pair encoding method to generate a sub - word vocabulary from the log text. This method uses bytes as the basic sub - word units and merges the most frequently occurring sub - word pairs. The goal is to segment the input log text into individual sub - words, each of which has a relative linguistic meaning. Sub - words can better avoid the Out Of Vocabulary problem compared to words, balancing the vocabulary size and semantic independence.

[0112] 1.3 Denote the sub - word vocabulary described in Section 1.2 as set X, and the sub - word vocabulary used by the original RoBERTa model as set A. Merge the two to generate a sub - word vocabulary B dedicated to the log language, where Z = A ∪ B.

[0113] 1.4 Use the 1.1 festival log corpus to continue self-supervised pre-training on the original RoBERTa model to obtain a new RoBERTa model adapted to the log domain, that is, the feature extraction model. Compared with its predecessor, the BERT model, the main changes in the RoBERTa model are the use of the byte-level byte pair encoding method, the use of the dynamic MASK mechanism, and the removal of the NSP (Next-Sentence-Prediction) task, etc. The MASK mechanism means that when training the model, some subwords are randomly masked from the input text content, enabling the model to predict the subword through the context content.

[0114] It should be noted that the original Roberta model is an existing pre-trained model for natural language processing. Its corpus includes a large unlabeled text corpus and a book corpus. Its tokenizer, subword vocabulary, and Transformer encoder do not have the characteristics of log language. Log language has both commonalities and differences with natural language. The purpose of this step is to utilize the commonalities to reduce the pre-training workload, and at the same time incorporate the differences to train a Roberta model that is more adapted to the log domain. In addition, due to the need for high computational efficiency and a large amount of diverse log data for training based on the log corpus, this step is carried out in the cloud.

[0115] (2) Feature extraction part of the log text:

[0116] The cloud transmits the feature extraction model obtained in step (1) to each edge device 1-N respectively. The edge device cleans the single log text data and inputs it into the feature extraction model to obtain the vector representation of the single log text, as Figure 3 shown. Subsequently, the sliding window technique is used to crop the data to reduce the data complexity; and the feature extraction for the log sequence is realized.

[0117] The feature extraction model includes a tokenizer and multiple layers of Transformer. Among them, the tokenizer realizes the segmentation of the log text into tokens. Therefore, the method adopted in this method is a non-parsing method. This method does not parse the original log into a fixed template (there is a parsing error here), but directly inputs the original log into the network and cuts it into individual tokens through the tokenizer. The Transformer architecture adopts the attention mechanism and realizes the function of the encoder, that is, embedding (converting each token into a context-related feature vector). Therefore, ultimately, each input log text can be converted into a multi-dimensional vector containing log semantic information, and at the same time, it can well retain and represent the context information of the log text content.

[0118] For example, a log text with a length of L is transformed into a sub-word sequence with a length of M (L is less than or equal to M) through a tokenizer and a sub-word vocabulary Z, and then transformed into a sequence of M vectors with a dimension length of D by a multi-layer Transformer architecture. These M vector sequences are context-related, that is, the value of each vector in the vector sequence is affected by the other vectors. Take the average of these M vectors as the vector representation x of this log text i . The value of D depends on the feature extraction model and is usually 768 or 1024.

[0119] (3) Pre-training part of the classification model:

[0120] 3.1 Edge 1-N performs partial anomaly annotation on the historical log text data to obtain the source log data, and at the same time uses the above feature extraction model to extract features in two dimensions of time and status from the source log data to obtain the input vector representation sequence x 1 , x 2 , …, x n ; each log corresponds to a vector representation x i (i ∈ [1, n]); select n consecutive vector representations to form a vector representation sequence X = [x 1 , x 2 , …, x n , n is the selected window length, that is, the vector representations of n consecutive logs are a vector representation sequence, and the value range is generally 10-100, and the vector representation sequence and the corresponding anomaly annotation information are transmitted to the cloud for summary and model training.

[0121] 3.2 The cloud inputs the vector representation sequence X and the corresponding anomaly annotation information into the initial classification model containing multiple layers of Transformer encoders for supervised learning training to obtain the pre-trained classification model.

[0122] Considering whether a log is determined to be abnormal is related to its own content and also related to the order of its previous and subsequent logs. As Figure 3 shown, the cloud applies the sine encoder to generate the position encoding vector p i for the received vector representation sequence X according to the sine and cosine functions of the position i of each log's vector representation x i in the vector representation sequence X, and then adds the corresponding position's x i and p i and inputs them into the initial classification model for parameter training.

[0123] The initial classification model consists of E Transformer encoders, and the typical value range of E is 2 - 20; for a single Transformer encoder, the multi-head attention layer calculates the attention score matrix for each log message with different attention patterns. The attention scores are calculated by training the Q and K matrices of the attention layer. Different attention patterns are obtained through multi-head attention, which enables the model to consider which attention scores are significant, similar to the multiple convolutional kernels used in each layer of a CNN. The typical number of attention heads is 5 - 20. The multi-head attention layer is followed by a fully connected network, and the typical value of the network size is 1000 - 5000.

[0124] The outputs of multiple layers of Transformer encoders are sent to a pooling layer, a dropout layer, and a fully connected layer, and the softmax method is used to calculate the anomaly probability values, as Figure 4 shown. The calculated anomaly probability results are compared with the corresponding annotation information of the data, and the loss function is calculated, and the model parameters are trained iteratively; the pre-trained classification model obtained by the cloud is distributed to each edge device. The initial learning rate can be set to 3e-4, the mini-batch size and the dropout rate can be set to 64 and 0.1 respectively; the cross-entropy is used as the loss function.

[0125] It should be noted that the cloud does not have the professional domain knowledge to judge whether the log is abnormal. Therefore, the annotation of the log data is performed at the edge device.

[0126] (4) Transfer learning and anomaly detection part:

[0127] 4.1 The cloud transfers the pre-trained classification model in step (3) to each of the edge devices 1 - N respectively. At the same time, each edge device collects a small amount of log data of the terminal devices in its own edge domain and performs a small amount of anomaly annotation as the target database 1 - N. The target databases used by each edge device are independent of each other. The method in step (2) is used to extract the feature vectors of the log texts in the target database to obtain the vector representation sequence.

[0128] 4.2 Each edge device keeps the network parameters in the fixed model of the pre-trained classification model received in 4.1 unchanged, and adds an adaptation layer with a small number of parameters as a transfer model in the Figure 5 structure of the Transformer encoder shown. Here, the adaptation layer consists of a down-projection layer, an activation layer, and an up-projection layer. Among them, the dimensions of the up and down projection layers are the same as those of the multi-head attention layer in the Transformer encoder, denoted as d, and the dimension of the activation layer is denoted as m, and the value of m / d is between 0.5 - 8%; the output of the adaptation layer is shown as follows: h is the input of the adaptation layer, W up and W downAssign weight parameters to the upper and lower projection layers respectively; the adapted layer structure added to each edge remains unchanged.

[0129] h′ = W up tanh(W down h) + h

[0130] 4.3 The edge terminals 1 - N respectively use the target databases 1 - N obtained in 4.1, and adopt the method of supervised learning to input the vector representation sequence and annotation information into the transfer model for the training of the adapted layer parameters, and finally obtain independent anomaly detection models for each edge domain.

[0131] 4.4 Each edge terminal collects the log data reported by the terminal in real time, obtains the corresponding vector representation sequence by using the log feature extraction method in step (2), inputs the vector representation sequence into its own anomaly detection model to obtain its anomaly detection result, and feeds it back to the cloud for management and storage.

[0132] In the initial stage of the system of the present invention, the training of the feature extraction model and the pre - training of the anomaly detection classification model are first carried out. In the deployment stage of the system edge domain, the transfer learning of the anomaly detection classification model is carried out, and finally independent anomaly detection models for each edge terminal are obtained. On this basis, each edge terminal can use the feature extraction model and the anomaly detection model to complete anomaly detection for the real - time collected log data. Compared with the prior art, the present invention reduces the pressure on the network bandwidth and improves the system anomaly detection efficiency and reliability at the same time.

[0133] Based on any one of the above - mentioned embodiments, the embodiment of the present invention provides an anomaly detection model training system for a cloud - edge - terminal system. Figure 7 is the architecture diagram of the anomaly detection model training system provided by the embodiment of the present invention. As Figure 7 shown, the system includes each edge terminal 710 and the cloud 720;

[0134] Each edge terminal 710 is used to extract features from the first sample log sequence to obtain the first vector representation sequence, and send the first vector representation sequence and the corresponding anomaly annotation information to the cloud;

[0135] The cloud 720 is used to perform supervised training on the initial classification model based on the first vector representation sequence and the corresponding anomaly annotation information sent by each edge terminal, obtain the pre - trained classification model and send it to each edge terminal;

[0136] Each edge terminal 710 is further used to add an adapted layer to the pre - trained classification model as a transfer model, and extract features from the second sample log sequence to obtain the second vector representation sequence;

[0137] Each edge device 710 is also used to perform supervised training on the transfer model based on the second vector representation sequence and the corresponding anomaly annotation information, and obtain an anomaly detection model corresponding to each edge device.

[0138] The system provided by the embodiments of the present invention, for the cloud-edge-end system log sequence, through cloud-edge collaboration, adopts a training method of pre-training plus transfer learning to finally obtain an anomaly detection model suitable for each edge device, which can efficiently train and adapt the model for a large number of terminal applications, and can also train an anomaly detection model with good training effects even in the case of only a small amount of labeled data or imperfect labeled data. At the same time, it can meet the requirements of accuracy and real-time performance for anomaly detection in the increasingly large-scale cloud-edge-end collaborative system.

[0139] Based on any of the above embodiments, the embodiments of the present invention provide an anomaly detection system for a cloud-edge-end system. Figure 8 It is an architecture diagram of the anomaly detection system provided by the embodiments of the present invention, as Figure 8 shown. This system is applied to each edge device, and specifically includes:

[0140] A feature extraction module 810, configured to extract features from the log sequence reported by the terminal to obtain a vector representation sequence;

[0141] An anomaly detection module 820, configured to input the vector representation sequence into the anomaly detection model, obtain the anomaly detection result of the log sequence, and send it to the cloud;

[0142] Wherein, the anomaly detection model is trained based on the anomaly detection model training method of the cloud-edge-end system provided in any of the above embodiments.

[0143] The system provided by the embodiments of the present invention, through cloud-edge collaboration, adopts a training method of pre-training plus transfer learning to finally obtain an anomaly detection model suitable for each edge device, and completes real-time and accurate anomaly detection of log data, so as to be able to timely and actively discover service anomalies of systems or application programs from log information, so as to take corresponding measures in time and improve the stability of the system.

[0144] It can be understood that the detailed function implementation of each of the above modules can be referred to the introduction in the foregoing method embodiments, and will not be elaborated here.

[0145] In addition, the embodiments of the present invention provide another anomaly detection model training device for a cloud-edge-end system, which includes: a memory and a processor;

[0146] The memory is used to store a computer program;

[0147] The processor is configured to implement the anomaly detection model training method in the above embodiments when executing the computer program.

[0148] In addition, an embodiment of the present invention provides another anomaly detection device for a cloud-edge-end system, which includes: a memory and a processor;

[0149] The memory is used to store a computer program;

[0150] The processor is configured to implement the anomaly detection method in the above embodiment when executing the computer program.

[0151] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method in the above embodiment is implemented.

[0152] Based on the method in the above embodiment, an embodiment of the present invention provides a computer program product, and when the computer program product runs on a processor, the processor is caused to execute the method in the above embodiment.

[0153] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for training an anomaly detection model of a cloud-edge-terminal system, characterized in that, it includes: S101 Each edge terminal extracts features from the first sample log sequence to obtain a first vector representation sequence, and sends the first vector representation sequence and the corresponding anomaly annotation information to the cloud; S102 The cloud performs supervised training on the initial classification model based on the first vector representation sequence and the corresponding anomaly annotation information sent by each edge terminal, obtains a pre-trained classification model, and sends it to each edge terminal; S103 Each edge terminal adds an adaptation layer to the pre-trained classification model as a transfer model, and extracts features from the second sample log sequence to obtain a second vector representation sequence; S104 Each edge terminal performs supervised training on the transfer model based on the second vector representation sequence and the corresponding anomaly annotation information to obtain an anomaly detection model corresponding to each edge terminal; Before step S101, it further includes: The cloud performs self-supervised training on the pre-trained model based on the third sample log sequence, obtains a feature extraction model, and sends it to each edge terminal. The pre-trained model uses a preset vocabulary, and the preset vocabulary is obtained according to the vocabulary corresponding to the third sample log sequence and the vocabulary used by the pre-trained model; Each edge terminal extracts features from the first sample log sequence to obtain a first vector representation sequence, including: Each edge terminal tokenizes each log in the first sample log sequence based on the feature extraction model to obtain a tokenization result, and converts the tokenization result into a sub-word sequence based on a preset vocabulary; Each edge terminal encodes the sub-word sequence to obtain a vector sequence, and takes the average of each vector in the vector sequence to obtain a vector representation; Each edge terminal forms the vector representations of each log into a first vector representation sequence.

2. The anomaly detection model training method according to claim 1, characterized in that, The adaptation layer includes a first projection layer, a second projection layer, and an activation layer between the first projection layer and the second projection layer.

3. An anomaly detection method for a cloud-edge-terminal system, characterized in that, The method is applied to each edge terminal, and the method includes: S601 Extract features from the log sequence reported by the terminal to obtain a vector representation sequence; S602 Input the vector representation sequence into the anomaly detection model, obtain the anomaly detection result of the log sequence, and send it to the cloud; Wherein, the anomaly detection model is trained based on the anomaly detection model training method of the cloud-edge-terminal system according to any one of claims 1 to 2.

4. An anomaly detection model training system for a cloud-edge-terminal system, characterized in that, The system includes each edge terminal and the cloud; Each edge terminal is used to extract features from the first sample log sequence to obtain a first vector representation sequence, and send the first vector representation sequence and the corresponding anomaly annotation information to the cloud; The cloud is used to perform supervised training on the initial classification model based on the first vector representation sequence and the corresponding anomaly annotation information sent by each edge terminal, obtain a pre-trained classification model, and send it to each edge terminal; Each edge terminal is further configured to add an adaptation layer to the pre-trained classification model as a transfer model, and extract features from the second sample log sequence to obtain a second vector representation sequence; Each edge terminal is further configured to perform supervised training on the transfer model based on the second vector representation sequence and the corresponding anomaly annotation information to obtain an anomaly detection model corresponding to each edge terminal; The cloud performs self-supervised training on the pre-trained model based on the third sample log sequence to obtain a feature extraction model and sends it to each edge terminal. The pre-trained model uses a preset vocabulary, and the preset vocabulary is obtained according to the vocabulary corresponding to the third sample log sequence and the vocabulary used by the pre-trained model; Specifically, each edge terminal is configured to perform word segmentation on each log in the first sample log sequence based on the feature extraction model to obtain a word segmentation result, and convert the word segmentation result into a sub-word sequence based on the preset vocabulary; encode the sub-word sequence to obtain a vector sequence, and average each vector in the vector sequence to obtain a vector representation; and, form a first vector representation sequence with the vector representations of each log.

5. The anomaly detection model training system according to claim 4, wherein, The adaptation layer added by each edge terminal includes a first projection layer, a second projection layer, and an activation layer between the first projection layer and the second projection layer.

6. An anomaly detection system for a cloud-edge-terminal system, wherein, The system is applied to each edge terminal, and the system includes: A feature extraction module, configured to extract features from the log sequence reported by the terminal to obtain a vector representation sequence; An anomaly detection module, configured to input the vector representation sequence into an anomaly detection model, obtain an anomaly detection result of the log sequence, and send it to the cloud; wherein, the anomaly detection model is trained based on the anomaly detection model training method of the cloud-edge-terminal system according to any one of claims 1 to 2.

Citation Information

Patent Citations

  • Intrusion detection method and device for edge computing scene

    CN113037722A

  • Method and device for training model, electronic equipment and storage medium

    CN114756658A