A fault diagnosis method and device based on unsupervised learning

By processing the log data of the operation and maintenance tools through unsupervised learning methods, the problems of insufficient automation and accuracy in fault diagnosis in the existing technology are solved, and fast and accurate fault location and diagnosis are achieved, thereby improving the operation and maintenance efficiency and accuracy of the system.

CN116910657BActive Publication Date: 2025-09-16CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310054333.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-03
Publication Date
2025-09-16
Estimated Expiration
2043-02-03

AI Technical Summary

Technical Problem

Existing fault diagnosis methods based on log data have deficiencies in automation and accuracy, making it difficult to quickly and accurately discover system anomalies and root causes of faults.

Method used

An unsupervised learning method is used to process the log data of operation and maintenance tools. The log data is represented as a vector through natural language processing. The autoencoder and clustering algorithm are combined to learn the intrinsic pattern of the log. The decoder is used for fault diagnosis to achieve rapid classification and fault location of the current log.

Benefits of technology

It realizes the automatic and rapid location and diagnosis of system faults, improves the efficiency and accuracy of fault detection, reduces manual processing costs, and improves the processing efficiency and classification accuracy of log data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_17
    Figure SMS_17
  • Figure SMS_36
    Figure SMS_36
Patent Text Reader

Abstract

The present invention discloses a fault diagnosis method and device based on unsupervised learning. The method is used to diagnose the fault of the operation and maintenance object of the operation and maintenance tool, including the following steps: obtaining the historical log data of the operation and maintenance tool, expressing it as a log vector using a natural language processing method, and embedding time and location information; clustering all historical log vectors; using i historical log vectors under the same category to learn the intrinsic pattern of the corresponding category log; obtaining the current log data of the operation and maintenance tool, expressing it as a current log vector with location and time; using the intrinsic pattern of the category to which the current log vector belongs, according to the current log vector and the i-1 historical log vectors approximating the category to which it belongs, obtaining a reference log vector; comparing the current log vector with the reference log vector to determine whether the operation and maintenance object has a fault and the module where the fault is located. The present invention can quickly diagnose and locate the fault of the operation and maintenance object, and improve the fault detection efficiency and accuracy of the operation and maintenance object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent operation and maintenance technology, and specifically relates to a fault diagnosis method and equipment based on unsupervised learning. Background Art

[0002] With the development of artificial intelligence, the concept of intelligent operations and maintenance (IOM) was first proposed by Gartner in 2016. This concept involves using machine learning and other algorithms to analyze large amounts of data from a variety of operations and maintenance tools and equipment, automatically identifying and responding to system issues in real time, thereby improving information technology operations and automation capabilities. Under the AIOps trend, intelligent fault discovery and root cause diagnosis technologies driven by multi-source operations and maintenance data and centered around machine learning and other algorithms have attracted widespread attention. Among these technologies, log data-based fault diagnosis, which uses intelligent methods to analyze log data generated during system operation to automatically discover system anomalies, diagnose system failures, locate the root cause of failures, and facilitate self-healing, is becoming a research hotspot in both mathematics and industry. Summary of the Invention

[0003] The present invention provides a fault diagnosis method based on unsupervised learning, which quickly diagnoses and locates faults of operation and maintenance objects by analyzing log data of operation and maintenance tools, thereby improving the fault detection efficiency and accuracy of operation and maintenance objects.

[0004] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0005] A fault diagnosis method based on unsupervised learning is used to diagnose faults of an operation and maintenance object of an operation and maintenance tool, comprising the following steps:

[0006] Obtain a large amount of historical log data from the operation and maintenance tools, use natural language processing methods to represent each log data as a log vector, and embed the log data's time and corresponding module location into it;

[0007] The obtained historical log vectors are clustered, and the number of categories is the same as the number of modules of the operation and maintenance object; then, the i historical log vectors under the same category are used to learn the intrinsic pattern of the corresponding category log;

[0008] Obtain the current log data of the operation and maintenance tool and obtain the current log vector embedded with the location and time in the same way as the historical log data;

[0009] Determine the category to which the current log vector belongs by distance, and then use the intrinsic pattern of the category to obtain the reference log vector based on the current log vector and the i-1 approximate historical log vectors under its category;

[0010] Compare the current log vector with the reference log vector to determine whether the operation and maintenance object is faulty and the module where the fault is located.

[0011] Furthermore, an autoencoder is used to learn the intrinsic patterns of each type of log, specifically:

[0012] First, the i historical log vectors of the same category are concatenated into a matrix;

[0013] Then, some log vectors in the log matrix are masked and input into the encoder for encoding to obtain the encoding matrix;

[0014] Then, the encoding vector at the mask position in the encoding matrix is ​​input to the decoder for decoding; the decoding output is recorded as the predicted log vector;

[0015] Finally, the decoder network is trained using a loss function based on all predicted log vectors and the corresponding log vectors before mask processing. The trained decoder can represent the intrinsic log pattern under this category.

[0016] Furthermore, the intrinsic pattern of the category to which it belongs is used to obtain a reference log vector based on the current log vector and the i-1 approximate historical log vectors of the category to which it belongs, specifically:

[0017] First, the current log vector and its i-1 approximate historical log vectors under the category to which it belongs are concatenated into a matrix;

[0018] Then, the current log vector in the concatenated log matrix is ​​masked and input into the encoder for encoding to obtain the encoding matrix;

[0019] Finally, the encoding vector corresponding to the current log vector in the encoding matrix is ​​input to the trained decoder, and the reference log vector is output.

[0020] Furthermore, the loss function is expressed as:

[0021]

[0022] Where, is the weight factor, is the j-th predicted log vector output by the decoder, express The corresponding log vector before masking

[0023] Furthermore, in the process of learning the intrinsic pattern, a random masking method is used to mask half of the log vectors in the log matrix.

[0024] Furthermore, the K-means method is used to perform unsupervised clustering on historical log vectors.

[0025] Furthermore, the Word2Vec algorithm is used to represent the log data into log vectors.

[0026] Furthermore, before embedding the position and time into the log vector, principal component analysis is used to reduce the dimensionality of the log vector.

[0027] Furthermore, the fault location, ie, the module where the fault is located, is determined based on the location information embedded in the current log vector.

[0028] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor implements the fault diagnosis method based on unsupervised learning described in any of the above technical solutions.

[0029] Beneficial effects

[0030] This invention can automatically locate and diagnose system faults based on the massive amount of log data from operation and maintenance tools, significantly reducing manual processing costs while improving the efficiency and accuracy of system fault detection. Furthermore, when representing log data as vectors, dimensionality reduction is used to improve algorithm efficiency. An unsupervised method is used to automatically classify logs, and the classification outputs reference vectors for comparison, improving the accuracy of fault diagnosis. DETAILED DESCRIPTION

[0031] The following is a detailed description of an embodiment of the present invention. This embodiment is based on the technical solution of the present invention, provides a detailed implementation method and a specific operation process, and further explains the technical solution of the present invention.

[0032] Example 1

[0033] This embodiment provides a fault diagnosis method based on unsupervised learning, comprising the following steps:

[0034] Phase 1: Unsupervised Log Classifier

[0035] Step 1: Obtain a large amount of historical log data from the operation and maintenance tool, use natural language processing methods to represent each log data as a log vector, and embed the time and corresponding module position of the log data into it.

[0036] 1.1, Log information embedding

[0037] The first thing to do is to extract the potential information in the operation and maintenance tool log data and express it as a signal that can be processed by the computer. This embodiment uses a method based on natural language processing to express the log data into log vectors through Word2Vec. .

[0038] Considering the large scale of log data, in order to process log vectors more efficiently, in a more optimal embodiment, the current log vector can be reduced in dimension. The specific method is to reduce the dimension of the log vector through principal component analysis, that is:

[0039] ,in

[0040] In this implementation, the operation and maintenance tool makes two assumptions: ① Log data from the same module has the highest similarity; ② Log data generated within the same time window also has the highest similarity. Therefore, the above log vectors can be further embedded with location and time information to facilitate subsequent learning of the decoder network.

[0041] 1.2 Location Information Embedding

[0042] The log vector obtained by the above dimensionality reduction , embedding location information into it , including system location, module number, etc.:

[0043]

[0044] 1.3 Temporal Information Embedding

[0045] The log vector after the above embedding position , and then embed time information into it :

[0046]

[0047] Step 2: Cluster the obtained historical log vectors, and the number of categories is the same as the number of modules of the operation and maintenance object.

[0048] Based on the log vector obtained in step 1 , all log vectors are clustered unsupervisedly by the K-means method. Through clustering operation, massive log vectors are classified to achieve unsupervised distinction of log vectors of different modules and different time windows. This embodiment supervises the clustering process by adding weight factors to achieve the purpose of distinguishing logs of different modules. For two log vectors , and defines its similarity score as follows:

[0049]

[0050] in, Represents a weight matrix consisting of m log features and weight factors of location and time.

[0051] Phase 2: Unsupervised Training

[0052] Step 3, use the same category The historical log vectors learn the intrinsic patterns of the corresponding category logs.

[0053] In step 2, all log vectors are divided into Classes: Then, in step 3, the intrinsic pattern of each category is learned in an unsupervised manner.

[0054] In this embodiment, an autoencoder is used to learn the intrinsic patterns of each category. The intrinsic pattern learning method for each category includes:

[0055] First, take the log vector Splice into a matrix :

[0056]

[0057] Then, the log matrix The partial log vectors in are masked and input to the encoder for encoding to obtain the encoding matrix;

[0058] The mask can be expressed as:

[0059]

[0060]

[0061] In this embodiment, half of the log vectors in the log matrix are masked by random masking, and the resulting mask sequence is Input to the Transformer encoder. The Transformer encoder is obtained by migrating from the upstream task and does not need to be updated during network training. In addition, it can use the self-attention mechanism in the Transformer encoder to focus on the global information of the log matrix to achieve log vector encoding.

[0062] Alternatively, the encoding can be expressed as:

[0063]

[0064] Where, The encoding matrix obtained by the encoder includes the original The log vector corresponds to Encoded vectors.

[0065] Then, the encoding vector at the mask position in the encoding matrix is ​​input to the decoder for decoding; the decoding output is recorded as the predicted log vector; the decoder decodes the encoding vector at the mask position and can be expressed as:

[0066]

[0067] Where, is the predicted log vector of the decoded output.

[0068] Finally, based on all the predicted log vectors and the corresponding log vectors before masking, the decoder network is trained using the loss function. The trained decoder can represent the intrinsic log pattern under this category. The loss function is expressed as:

[0069]

[0070] Where, is the weight factor, The decoder output prediction log vectors, express The corresponding log vector before masking.

[0071] Phase 3: Abnormal log detection

[0072] Step 4: Obtain the current log data of the operation and maintenance tool, and obtain the current log vector embedded with position and time in the same way as the historical log data.

[0073] Step 5: Determine the category to which the current log vector belongs by distance, and then use the intrinsic pattern of the category to obtain the reference log vector based on the current log vector and the i-1 approximate historical log vectors under its category.

[0074] The category to which the current log vector belongs is determined by comparing the distance between the current log vector and each cluster center, selecting the category with the smallest distance, and determining it as the category to which the current log vector belongs.

[0075] As described above, this embodiment uses an autoencoder to learn the intrinsic pattern of each category. The trained decoder can represent the log intrinsic pattern under this category. Therefore, this embodiment uses the decoder to obtain the reference log vector, which specifically includes:

[0076] First, concatenate the current log vector and its i-1 approximate historical log vectors for its category into a matrix. The i-1 approximate historical log vectors are the i-1 historical log vectors closest to the current log vector among all the historical log vectors for its category. The specific matrix concatenation method is the same as described in step 3.

[0077] Then, the current log vector in the concatenated log matrix is ​​masked and input into the encoder for encoding to obtain the encoding matrix.

[0078] Finally, the encoding vector corresponding to the current log vector in the encoding matrix is ​​input to the trained decoder, and the reference log vector is output.

[0079] Step 6: Compare the current log vector with the reference log vector to determine whether the operation and maintenance object is faulty and the module where the fault is located, so as to achieve rapid fault diagnosis.

[0080] Example 2

[0081] This embodiment provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor implements the fault diagnosis method based on unsupervised learning described in Example 1.

[0082] The above embodiments are preferred embodiments of the present application. Ordinary technicians in this field can also make various changes or improvements on this basis. Without departing from the overall concept of the present application, these changes or improvements should fall within the scope of protection required by the present application.

Claims

1. A fault diagnosis method based on unsupervised learning, characterized in that: This procedure is used to diagnose faults on the operation and maintenance objects of the operation and maintenance tool, including the following steps: Obtain a large amount of historical log data from the operation and maintenance tools, use natural language processing methods to represent each log data as a log vector, and embed the log data's time and corresponding module location into it; Cluster the obtained historical log vectors, and the number of categories is the same as the number of modules of the operation and maintenance object; Use i historical log vectors under the same category to learn the intrinsic pattern of the corresponding category log; Obtain the current log data of the operation and maintenance tool and obtain the current log vector embedded with location and time in the same way as the historical log data; Determine the category to which the current log vector belongs by distance, and then use the intrinsic pattern of the category to obtain the reference log vector based on the current log vector and the i-1 approximate historical log vectors under its category; Compare the current log vector with the reference log vector to determine whether the operation and maintenance object is faulty and the module where the fault occurs; Among them, the autoencoder is used to learn the intrinsic pattern of each type of log, specifically: First, the i historical log vectors of the same category are concatenated into a matrix; Then, some log vectors in the log matrix are masked and input into the encoder for encoding to obtain the encoding matrix; Then, the encoding vector at the mask position in the encoding matrix is ​​input to the decoder for decoding; the decoding output is recorded as the predicted log vector; Finally, based on all the predicted log vectors and the corresponding log vectors before masking, the decoder network is trained using a loss function. The trained decoder can represent the intrinsic log pattern under this category. The intrinsic pattern of the category to which it belongs is used to obtain a reference log vector based on the current log vector and its i-1 approximate historical log vectors under the category to which it belongs, specifically: First, the current log vector and its i-1 approximate historical log vectors under the category to which it belongs are concatenated into a matrix; Then, the current log vector in the concatenated log matrix is ​​masked and input into the encoder for encoding to obtain the encoding matrix; Finally, the encoding vector corresponding to the current log vector in the encoding matrix is ​​input to the trained decoder, and the reference log vector is output.

2. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: The expression of the loss function is: ; Where, is the weight factor, The decoder output prediction log vectors, express The corresponding log vector before masking.

3. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: In the process of learning intrinsic patterns, a random masking method is used to mask half of the log vectors in the log matrix.

4. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: The K-means method is used to perform unsupervised clustering on historical log vectors.

5. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: The Word2Vec algorithm is used to represent log data into log vectors.

6. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: Before embedding the position and time into the log vector, principal component analysis is used to reduce the dimensionality of the log vector.

7. The fault diagnosis method based on unsupervised learning according to claim 1, characterized in that: The fault location, i.e., the module where the fault is located, is determined based on the location information embedded in the current log vector.

8. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the computer program is executed by the processor, the processor is caused to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Abnormal state detection method and system, storage medium, program and server

    CN111552609A

  • Tool-specific alerting rules based on abnormal and normal patterns obtained from history logs

    US20200160230A1