Log anomaly detection method and system based on multi-source feature fusion
By adopting the multi-source feature fusion method in log exception detection, the weighted hypergraph is established and fusion is solved, and the shortcomings in the existing technology in modeling complex log relationships and considering different impacts are achieved, and more efficient log exception detection is achieved.
Patent Information
- Application Number
- CN202510193504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-27
AI Technical Summary
The existing log-based exception detection methods have shortcomings in effectively modeling complex log relationships and considering different impacts between logs. They also have incomplete utilization of the semantic features of logs, and have failed to consider homophones or polysense issues.
The log exception detection method based on multi-source feature fusion is adopted to process the logs to be detected through a pre-trained log exception detection model, including parsing units, building units, fusion units and prediction units. This method uses multiple semantic embedding methods to extract multiple semantic features, establish a weighted hypergraph, and comprehensively consider the multiple features of the sample through the hypergraph fusion mechanism.
This method can more effectively model the higher-order relationships between logs, establish complex graph structures, and measure different effects through vertex weights, balance the impact between samples, and automatically measure the importance of each feature, thereby improving the accuracy and adaptability of anomaly detection.
Smart Images

Figure CN120216283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log data detection, and particularly to a log anomaly detection method and system based on multi-source feature fusion. Background Art
[0002] During the anomaly detection process, there are very complex relationships among logs, and the existing log-based anomaly detection methods do not model these relationships and do not consider the influence of different logs in the detection.
[0003] Most methods use semantic embedding techniques to simply convert events into numerical vectors. Different semantic embedding techniques will have different effects, such as homonymy and polysemy problems.
[0004] The early log anomaly detection algorithms used include support vector machine (SVM), logistic regression, convolutional neural network (CNN), etc. Support vector machine (SVM) is a supervised learning method widely used for classification tasks. Liang et al. developed three classifiers, including a rule-based classifier (RIPPER), SVM, and nearest neighbor method, for predicting fault events from log data. In addition, Fulp et al. applied the sliding window technique to analyze system logs and used SVM to predict system failures. Bod′ik et al. proposed a method using L1-regularized logistic regression to automatically select the most relevant metrics from large datasets collected from data centers. This method helps to identify anomalies by focusing on key information and ignoring irrelevant data, thus simplifying the detection process. By analyzing historical labeled data, it can predict potential anomalies and trigger an alarm when there is a deviation from the expected behavior.
[0005] In order to utilize the linear information contained in logs, many models are used to extract linear relationships, such as long short-term memory (LSTM), gated recurrent unit (GRU), etc. This greatly improves the effectiveness of log anomaly detection. LogRobust captures the semantic information of log events by converting log events into semantic vectors, and then uses an attention-based bidirectional LSTM model to detect anomalies while considering the context information and importance of different log events. The author evaluated LogRobust using public datasets and industrial datasets, and the results showed that it can effectively detect anomalies in log data. PLELog adopted a GRU neural network based on the attention mechanism, which eliminated the need for a large amount of manual labeling and only used normal log sequences to infer the labels of unlabeled log sequences, achieving excellent performance. DeepLog uses long short-term memory (LSTM) to model system logs as natural language sequences, and it can automatically learn log patterns from normal executions and detect anomalies when the log patterns deviate from the trained model.
[0006] In recent years, due to the fact that graph structures can represent other relationships including linear relationships, graph-based log anomaly detection has become increasingly popular. Compared with other log anomaly detection methods, graph anomaly detection can better utilize the spatial structure of log graphs, thereby improving detection performance and accuracy. Various graph neural networks (GNNs), such as DynAD, STGAN, and TopoMAD, have been effectively applied to graph anomalies. Lakha et al. used a GNN model to capture the proximity and structural relationships between nodes, and then encoded complex log features and generated robust feature representations. Graph convolutional neural networks (GCNs) have also achieved remarkable success in various graph data mining applications. LogGD utilizes a custom graph transformation neural network (TransformerConv) to convert log sequences into graph structures, effectively capturing the structural relationships between log events by detecting anomalies using node semantics and graph structure features. CoLA introduced a contrastive learning method for anomaly detection in large-scale attributed graphs by defining instance pairs based on the following, node neighbor relationships, which can effectively capture local structural information and enhance the adaptability of anomaly detection.
[0007] However, existing log-based anomaly detection methods are insufficient in effectively modeling complex log relationships, considering different impacts between logs, do not comprehensively utilize the semantic features of logs, and do not consider possible homonymy or polysemy problems. Summary of the Invention
[0008] The purpose of the present invention is to provide a log anomaly detection method and system based on multi-source feature fusion to solve at least one of the technical problems existing in the above background technology.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] In a first aspect, the present invention provides a log anomaly detection method based on multi-source feature fusion, including:
[0011] Obtain the log to be detected;
[0012] The obtained logs to be detected are processed using a pre-trained log anomaly detection model to obtain log anomaly detection results. Among them, the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit. The parsing unit is used to convert logs into structured data using a log parsing method. Different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of log events are determined by summarizing the labels of the logs. The construction unit is used to extract multiple semantic features using multiple semantic embedding methods, using Word2Vec to focus on the semantic meaning of context, using TF-IDF to focus on the frequency information of word occurrences, and using BERT to focus on the semantics of words. Then, according to the extracted multiple features, a weighted hypergraph is established respectively. The fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph. The prediction unit is used to predict the fused hypergraph to obtain log detection results.
[0013] As a further limitation of the first aspect of the present invention, when and only when all logs with the same log event are normal or abnormal, it is determined that the event is normal or abnormal, otherwise the log event is an event to be classified; the event to be classified is classified using the parameter information obtained by log parsing.
[0014] As a further limitation of the first aspect of the present invention, the construction of s hypergraphs includes: the s-th hypergraph is represented as where V s represents the set of points, ε s represents the set of hyperedges, represents the weight of the hyperedge; after log parsing, a total of N + M + 1 independent log events are obtained, which are regarded as the nodes of the hypergraph. The K-nearest neighbor algorithm is used to mine the high-order relationships between them and connect them with hyperedges, and these relationships are stored in the hypergraph relationship matrix H s ; according to the hyperedge weight the point weight u s and the relationship matrix H s , the degree δ s of the hyperedge and the degree d s of the point are calculated, and then according to δ s , d s , u s , diagonal matrices D e,s , D v,s , U s are generated respectively.
[0015] As a further limitation of the first aspect of the present invention, the learning result F of a single hypergraph s :
[0016] First, the hypergraph convolution Θ
[0017]
[0018] wherein represents the transpose of H s Then define the label matrix Y, with a size of (N + M + 1) * 2. For each row of data, the column where it is located is 1, and the other columns are 0;
[0019] The hypergraph learning result F s is:
[0020]
[0021] where I represents the identity matrix and τ is a parameter.
[0022] As a further limitation of the first aspect of the present invention, the hypergraph fusion objective formula is:
[0023]
[0024] where Δ s is the hypergraph Laplacian matrix, and y respectively represent the prediction result and the true result of the validation set, λ, μ, γ are parameters used to balance each part, and ζ is a variable.
[0025] As a further limitation of the first aspect of the present invention, after the hypergraph fusion is completed, the final result is obtained according to the following formula:
[0026]
[0027] Obtain the prediction result according to F:
[0028] In the second aspect, the present invention provides a log anomaly detection system based on multi-source feature fusion, including:
[0029] An acquisition module for acquiring the log to be detected;
[0030] A detection module is used to process the obtained logs to be detected by using a pre-trained log anomaly detection model, and obtain a log anomaly detection result. The log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit. The parsing unit is used to convert logs into structured data by using a log parsing method. Different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the logs. The construction unit is used to extract multiple semantic features by using multiple semantic embedding methods, use Word2Vec to focus on the semantic meaning of the context, use TF-IDF to focus on the frequency information of word occurrences, and use BERT to focus on the semantics of words. Then, according to the extracted multiple features, a weighted hypergraph is established respectively. The fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph. The prediction unit is used to predict the fused hypergraph to obtain a log detection result.
[0031] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the method for log anomaly detection based on multi-source feature fusion as described in the first aspect is implemented.
[0032] In a fourth aspect, the present invention provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the method for log anomaly detection based on multi-source feature fusion as described in the first aspect.
[0033] In a fifth aspect, the present invention provides an electronic device, including: a processor, a memory, and a computer program. The processor is connected to the memory, and the computer program is stored in the memory. When the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes the instructions for implementing the method for log anomaly detection based on multi-source feature fusion as described in the first aspect.
[0034] The beneficial effects of the present invention are as follows: It can model the high-order relationships between logs, and it is easier to establish a complex graph structure. In addition, weighting the vertices can measure the different influences of the vertices in the graph and balance the influences between samples. The multi-hypergraph fusion mechanism comprehensively considers the multiple features of the samples and automatically measures the importance of each feature during the learning process.
[0035] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or can be understood through the practice of the present invention. Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 It is a flowchart of the log anomaly detection method based on multi-source feature fusion according to the embodiments of the present invention.
[0038] Figure 2 It is a schematic diagram of the process for establishing log event tags according to the embodiments of the present invention. Detailed implementation manners
[0039] The following details the implementation manners of the present invention. The examples of the implementation manners are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described through the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.
[0040] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used here have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.
[0041] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0042] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the", and "said" used here may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or their groups.
[0043] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0044] To facilitate the understanding of the present invention, the following further explains the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation to the embodiments of the present invention.
[0045] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0046] The present invention provides a log anomaly detection method based on multi-source feature fusion, using a multi-hypergraph fusion strategy with vertex weights to solve the deficiencies of existing log-based anomaly detection methods in effectively modeling complex log relationships and considering different influences between logs, not comprehensively utilizing the semantic features of logs, and not considering possible homonym or polysemy problems. The hypergraph structure effectively solves the problem of complex relationship modeling, especially in establishing high-order associations between logs. Compared with traditional graphs, hypergraphs can connect multiple vertices through hyperedges and can better capture the association relationships between logs. In addition, some logs may have greater value in the anomaly detection process, and vertex weights can enable logs to play different roles in the hypergraph. Aiming at the deficiencies of single-semantic methods, multiple hypergraph structures are adopted to fully express semantic representations, and then complement each other in hypergraph learning, and multiple semantic features are comprehensively learned through the fusion strategy.
[0047] Embodiment 1
[0048] In this Embodiment 1, first, a log anomaly detection system based on multi-source feature fusion is provided, including: an acquisition module for acquiring logs to be detected; a detection module for processing the acquired logs to be detected by using a pre-trained log anomaly detection model to obtain a log anomaly detection result. The log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit. The parsing unit is used to convert logs into structured data by using a log parsing method. Different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of log events are determined by summarizing the labels of the logs. The construction unit is used to extract multiple semantic features by using multiple semantic embedding methods, use Word2Vec to focus on the semantic meaning of the context, use TF-IDF to focus on the frequency information of word occurrences, and use BERT to focus on the semantics of words. Then, according to the extracted multiple features, a weighted hypergraph is established respectively. The fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph. The prediction unit is used to predict the fused hypergraph to obtain a log detection result.
[0049] As Figure 1 shown, in this embodiment, the above system is used to implement a log anomaly detection method based on multi-source feature fusion, including:
[0050] First, use a log parsing method (such as Drain) to convert logs into structured data. Different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of log events are determined by summarizing the labels of the logs. Here, when and only when all logs with the same log event are normal or abnormal, we determine that the event is normal or abnormal. Otherwise, we regard the log event as an event to be classified. Next, use the parameter information obtained by log parsing to classify the event to be classified, and supplement it sequentially from the parameter and system feedback levels, as Figure 2 shown, where '-' 'APPOUT' and 'KERNPROG' represent the labels of the logs. Except for '-', they all represent anomalies. 'FATAL' and 'INFO' represent system feedback. If neither the parameter nor the system feedback can distinguish the event to be classified, we can only regard the event as an anomaly.
[0051] Next, we use multiple semantic embedding methods to extract multiple semantic features. We use Word2Vec to focus on the semantic meaning of the context, use TF-IDF to focus on the frequency information of word occurrences, and use BERT to focus on the semantics of words. Then, according to the extracted multiple features, a weighted hypergraph is established respectively.
[0052] The hypergraph can be represented as where V s represents the point set, ε s represents the set of hyperedges, Denote the weight of the hyperedge. We consider the weights of all hyperedges as 1. Let s represent the number of the hypergraph. We adopt three semantic embedding methods respectively, so s = 3. After log parsing, we obtain a total of N + M + 1 independent log events, which are regarded as the nodes of the hypergraph. We use the K-nearest neighbor algorithm to mine the high-order relationships among them and connect them with hyperedges, and store these relationships in the hypergraph relation matrix H s as follows
[0053]
[0054] This means that if the point v ∈ V s is connected to other points through the hyperedge e ∈ ε s , then H s (v, e) is 1, otherwise it is 0. The hypergraph with point weights introduces the point weight u s to make the samples play different roles in the hypergraph. Here, we calculate the average distance between each sample and other samples with the same label (normal / abnormal) to measure the role of the sample in the hypergraph. Since dense points exert influence more frequently, we assign smaller weights to them. For outliers, we assign higher weights so that they can have a greater impact even if the influence is not frequent. The formula is as follows
[0055]
[0056] represents the average distance between the i-th sample and other samples, m represents that there are m samples with the same label as i, represents the distance matrix constructed when building the s-th hypergraph, and t represents the total number of distance matrices constructed. For example, two distance matrices can be constructed according to the Euclidean distance and the Manhattan distance to measure the weights of the samples.
[0057] According to the hyperedge weight the point weight u s and the relation matrix H s (v, e), calculate the degree δ s of the hyperedge and the degree d s of the point. The formula is as follows
[0058]
[0059] Then, according to δ s , d s , u s generate the diagonal matrices D e,s , D v,s , U s respectively, and the construction of s hypergraphs is completed.
[0060] The multi-point-weighted hypergraph fusion includes:
[0061] Single hypergraph The learning result F s Can be calculated by the following formula. First, the hypergraph convolution Θ
[0062]
[0063] Where Represents the transpose of H s Then define the label matrix Y, with size (N + M + 1) * 2. For each row of data, the column corresponding to its label is 1, and other columns are 0. The hypergraph learning result F s Is obtained by the following formula
[0064]
[0065] Where I represents the identity matrix, and τ is a parameter.
[0066] Due to the differences in hypergraph structures, different hypergraphs may produce different learning results. It is neither wise nor realistic to select the best result as the common result. Here, we regard the learning results as pseudo-labels and allow different hypergraphs to complement each other. This means that although these learning results may be different, they still have reference value. In addition, we need to evaluate the reference value of each hypergraph. Here, the hypergraph weight α s Is used to evaluate the reference value of the hypergraph, and the validation set can calculate this metric more intuitively.
[0067] The hypergraph fusion objective formula is as follows
[0068]
[0069] Where Δ s Is the hypergraph Laplacian matrix, And y represent the predicted result and the true result of the validation set respectively. λ, μ, γ are parameters used to balance each part, and ζ is a variable, the reason for which will be explained later.
[0070] The objective formula is solved by two-step derivation. First, fix α s , and solve for F s , the formula is as follows
[0071]
[0072] Where F * Represents the predicted result of the previous round.
[0073] Then fix F s , and solve for α s , the formula is as follows
[0074]
[0075] To ensure that each hypergraph has reference value, i.e., minα s >0, ζ becomes a variable to control this goal, and the calculation formula of ζ is as follows
[0076] ζ = 10 x
[0077]
[0078] To intuitively measure the hypergraph weight α s , we optimize the non-validation set part of φ s , and the new calculation method is as follows
[0079]
[0080] where represents the hypergraph learning result of the validation set, and Y dev represents the label matrix of the validation set, and the generation method is the same as that of the label matrix Y.
[0081] After sequentially updating F s and α s until the target formula converges.
[0082] After the hypergraph fusion is completed, the final result is obtained according to the following formula
[0083]
[0084] The prediction result is obtained according to F In this way, the anomaly detection task is completed.
[0085] In summary, the log anomaly detection method described in this embodiment proposes a multi-hypergraph fusion framework for log anomaly detection. More specifically, by the fusion strategy, multiple features are concerned, and each feature is constructed as a vertex-weighted hypergraph to model the high-order correlation between logs and let the features complement each other through hypergraph learning. The proposed framework is used to solve the polysemy and homonym problems caused by a single semantic embedding method. The vertex weights are used to measure the different influences of vertices, and different weights are assigned to the samples in the hypergraph according to the distance between them. The closer the distance, the more similar the samples. Since the dense points exert influence more frequently, smaller weights are assigned to them. For the outliers, higher weights are assigned so that they can have a greater impact even if the influence is not frequent.
[0086] Embodiment 2
[0087] Embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the above-mentioned log anomaly detection method based on multi-source feature fusion is implemented. The method includes: obtaining logs to be detected; using a pre-trained log anomaly detection model to process the obtained logs to be detected to obtain a log anomaly detection result; wherein, the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit; the parsing unit is used to convert logs into structured data using a log parsing method; different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the logs; the construction unit is used to extract multiple semantic features using multiple semantic embedding methods, use Word2Vec to focus on the context meaning of words, use TF-IDF to focus on the frequency information of word occurrences, use BERT to focus on the semantics of words, and then establish a weighted hypergraph respectively according to the extracted multiple features; the fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph; the prediction unit is used to predict the fused hypergraph to obtain a log detection result.
[0088] Embodiment 3
[0089] Embodiment 3 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute the above-mentioned log anomaly detection method based on multi-source feature fusion. The method includes: obtaining logs to be detected; using a pre-trained log anomaly detection model to process the obtained logs to be detected to obtain a log anomaly detection result; wherein, the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit; the parsing unit is used to convert logs into structured data using a log parsing method; different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the logs; the construction unit is used to extract multiple semantic features using multiple semantic embedding methods, use Word2Vec to focus on the context meaning of words, use TF-IDF to focus on the frequency information of word occurrences, use BERT to focus on the semantics of words, and then establish a weighted hypergraph respectively according to the extracted multiple features; the fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph; the prediction unit is used to predict the fused hypergraph to obtain a log detection result.
[0090] Embodiment 4
[0091] Embodiment 4 of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory to enable the electronic device to execute instructions for implementing the above-mentioned log anomaly detection method based on multi-source feature fusion. The method includes: obtaining the log to be detected; using a pre-trained log anomaly detection model to process the obtained log to be detected to obtain a log anomaly detection result; wherein, the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit, and a prediction unit; the parsing unit is used to convert the log into structured data using a log parsing method; different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the log; the construction unit is used to extract multiple semantic features using multiple semantic embedding methods, use Word2Vec to focus on the context meaning of words, use TF-IDF to focus on the frequency information of word occurrences, use BERT to focus on the semantics of words, and then establish a weighted hypergraph according to the extracted multiple features; the fusion unit is used to fuse the constructed weighted hypergraph to obtain a fused hypergraph; the prediction unit is used to predict the fused hypergraph to obtain a log detection result.
[0092] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0093] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0094] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.
[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to perform a series of operational steps on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 in one or more of the processes and / or blocks Figure 1 specified in the function.
[0096] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative efforts should be covered by the scope of the present invention.
Claims
1. A log anomaly detection method based on multi-source feature fusion, characterized in that: include: Get the log to be tested; The obtained log to be detected is processed by using a pre-trained log anomaly detection model to obtain a log anomaly detection result; wherein the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit and a prediction unit; the parsing unit is used to convert the log into structured data using a log parsing method; different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the logs; the construction unit is used to extract multiple semantic features using multiple semantic embedding methods, using Word2Vec to focus on the meaning of the context, using TF-IDF to focus on the frequency information of the word occurrence, using BERT to focus on the semantics of the word, and then establishing weighted hypergraphs respectively according to the extracted multiple features; the fusion unit is used to fuse the constructed weighted hypergraphs to obtain a fused hypergraph; the prediction unit is used to predict the fused hypergraph to obtain a log detection result.
2. The log anomaly detection method based on multi-source feature fusion according to claim 1 is characterized in that: If and only if all logs with the same log event are normal or abnormal, the event is considered normal or abnormal, otherwise the log event is an event to be divided; the event to be divided is divided using the parameter information obtained from log parsing.
3. The log anomaly detection method based on multi-source feature fusion according to claim 1 is characterized in that: The construction of s hypergraphs includes: the sth hypergraph is represented as Where V s represents the point set, ε s represents the set of hyperedges, Represents the weight of the hyperedge; after log parsing, a total of N+M+1 independent log events are obtained, which are regarded as nodes of the hypergraph. The K nearest neighbor algorithm is used to mine the high-order relationships between them and use hyperedges to connect them. These relationships are stored in the hypergraph relationship matrix H s In; According to the hyperedge weight Point weight u s and the relationship matrix H s , calculate the degree δ of the hyperedge s and the degree of the point d s , then according to δ s , d s ,u s Generate diagonal matrices D respectively e,s , D v,s , U s .
4. The log anomaly detection method based on multi-source feature fusion according to claim 3 is characterized in that: Single Hypergraph The learning result F s : First, the hypergraph convolution Θ in Indicates H s Then define the label matrix Y, which is (N+M+1)*2 in size. For each row of data, its label column is 1 and other columns are 0. Hypergraph learning results F s for: Where I represents the identity matrix and τ is a parameter.
5. The log anomaly detection method based on multi-source feature fusion according to claim 4 is characterized in that: The hypergraph fusion objective formula is: Where Δ s is the hypergraph Laplacian matrix, and y represent the predicted results and true results of the validation set respectively, λ, μ, γ are parameters used to balance each part, and ζ is a variable.
6. The log anomaly detection method based on multi-source feature fusion according to claim 5 is characterized in that: After the hypergraph fusion is completed, the final result is obtained according to the following formula: The prediction results are obtained according to F:
7. A log anomaly detection system based on multi-source feature fusion, characterized in that: include: The acquisition module is used to obtain the logs to be detected; The detection module is used to process the acquired log to be detected by using a pre-trained log anomaly detection model to obtain the log anomaly detection result; wherein the log anomaly detection model includes a parsing unit, a construction unit, a fusion unit and a prediction unit; the parsing unit is used to convert the log into structured data using a log parsing method; different log events obtained by log parsing are regarded as nodes in a graph, and then the labels of the log events are determined by summarizing the labels of the logs; the construction unit is used to extract multiple semantic features using multiple semantic embedding methods, using Word2Vec to focus on the meaning of the context, using TF-IDF to focus on the frequency information of word occurrence, using BERT to focus on the semantics of the words, and then establishing weighted hypergraphs respectively according to the extracted multiple features; the fusion unit is used to fuse the constructed weighted hypergraphs to obtain a fused hypergraph; the prediction unit is used to predict the fused hypergraph to obtain the log detection result.
8. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the log anomaly detection method based on multi-source feature fusion as described in any one of claims 1 to 6 is implemented.
9. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the log anomaly detection method based on multi-source feature fusion as described in any one of claims 1-6.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the log anomaly detection method based on multi-source feature fusion as described in any one of claims 1-6.