A log partition knowledge graph construction and anomaly detection method

By establishing a knowledge graph for log templates, analyzing field name associations between log templates, clustering and analyzing log exceptions in the area, and performing timing analysis, the problem of difficult to clarify the log relationship is solved, and efficient log exception detection is achieved.

CN119337298BActive Publication Date: 2025-05-09INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411888385.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-09
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively clarify the relationship between the relationship between the internal relationship of the log and the key information of the log. When building a knowledge graph, the dispersed professional knowledge points lead to the huge graph and is not conducive to the discovery of the relationship.

Method used

By establishing a knowledge graph for each log template with the field name as node, analyzing the same field names between the log templates to establish undirected edges, clustering and analyzing logs in the region, and analyzing exceptions between regions through timing.

Benefits of technology

It achieves better analysis of the correlation between logs and discovers abnormal logs, which improves the efficiency and effectiveness of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119337298B_ABST
    Figure CN119337298B_ABST
Patent Text Reader

Abstract

The present invention discloses a log partition knowledge graph construction and anomaly detection method. On the basis of analyzing a log template, the field names in the log template are used as nodes, and directed edges are established in the order of the nodes, thereby establishing a knowledge graph of the log template. Logs belonging to the same log template are in the same area, and different areas are associated through the same field names between the log templates to establish wireless edges; in the same area, classification is performed according to the values ​​of key nodes, and cluster analysis of abnormal logs is performed according to other field names in the same classification; between different areas, on the basis of specifying key nodes, anomaly detection is performed by analyzing the time sequence formed by logs in different areas. The present invention can better analyze the association relationship between logs and find abnormal logs by establishing a log partition knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of log analysis and relates to a log partition knowledge graph construction and anomaly detection method. Background Art

[0002] For the log information of the information system, the information within each log constitutes an association relationship. In addition, different logs constitute an association relationship. The existing method of constructing logs cannot well define the internal association relationship of the log, nor can it track the association relationship between the key information of the log.

[0003] Knowledge graphs provide a way to associate information, which is more conducive to in-depth analysis of the relationship between information. Existing methods for establishing knowledge graphs mostly appear in the form of triples. For certain types of professional knowledge, because the knowledge points are scattered, the knowledge graphs established are not only huge, but also not conducive to discovering the relationship inside, and not conducive to the reasoning and discovery of new knowledge. Summary of the invention

[0004] The purpose of the present invention is to provide a log partition knowledge graph construction and anomaly detection method, which can better analyze the association between logs and discover abnormal logs by establishing a log partition knowledge graph.

[0005] The present invention is implemented by the following technical solution: A log partition knowledge graph construction and anomaly detection method, the steps are as follows:

[0006] Step 1: For each log template, a knowledge graph is established with the field name as the node. Each log template contains several nodes, and all the nodes of each log template constitute the knowledge graph of the region;

[0007] Step 2: Analyze the same field names between log templates, and establish undirected edges based on the same field names between log templates;

[0008] Step 3: Analyze log anomalies in the area through clustering;

[0009] Step 3.1: Extract logs in the region;

[0010] Step 3.2: Determine the key nodes in the region and classify them according to the key nodes;

[0011] According to the nature of the specified event, a specific field name is designated as a key node in the region, and logs with the same field value are classified according to the field value corresponding to the key node;

[0012] Step 3.3: Analyze abnormal behavior by key nodes;

[0013] For one category in step 3.2, cluster analysis is performed based on a field name or a vector consisting of multiple field names in the region to obtain the cluster center. Logs whose distance from the cluster center is greater than a threshold are considered abnormal logs.

[0014] Step 4: Analyze the abnormal logs between regions through time series analysis;

[0015] Step 4.1: Specify a specific field name as a key node;

[0016] According to the nature of the specified event, specific field names are designated as key nodes;

[0017] Step 4.2: Find the area related to the key node;

[0018] Step 4.3: Specify the field value of the key node;

[0019] Step 4.4: Analyze the paths formed by the logs between various regions, and regard the logs related to the abnormal paths as abnormal logs.

[0020] Further preferably, in step 1, by analyzing all system log information, a log analysis tool is used to analyze the log templates contained in the system log information to form a log template library , indicating that there are p different field name combinations. are the 1st, 2nd, …, pth log templates respectively. A log template can be formally expressed as , indicating that the i-th log template contains "field name-field value" format, Respectively represent the 1st, 2nd, ..., field names, Respectively represent the 1st, 2nd, ..., The field value corresponding to the field name.

[0021] Further preferably, in step 1, a knowledge graph is established for each log template with field name as a node. Log Templates Included nodes, forming the The knowledge graph of the region Respectively represent 1,2,…, nodes; for the , , Log templates ( ), respectively , , field name, corresponding to , , The knowledge graph of the region contains , , nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes; directed edges are established between nodes in the region in the order in which the field names appear, and directed edges indicate the adjacency relationship between two nodes.

[0022] Further preferably, in step 2, the same field names between log templates are obtained by string search and matching technology; Log Templates , for the The first of the log templates Field Name , , respectively match all field names in other log templates and record the same field names:

[0023] ;

[0024] In the formula, Indicates obtaining the same field name in the log template. Indicates that all log templates are related to Log Templates The Field Name The same field name, Indicates Log Templates The Field Name , Indicates Log Templates The Field Name ;

[0025] according to , in Nodes in a region , No. Nodes in a region , No. Nodes in a region Construct undirected edges in .

[0026] Further preferably, in step 4.2, the specified key node is found in all regions, and the region number containing the key node is recorded. ,in represents the number of the rth region containing the key node, , The total number of area numbers containing key nodes; record the area number containing the user name ,in is the number of the rth region containing the user name, , The total number of zones containing user names.

[0027] Further preferably, in step 4.3, the area numbering containing the key nodes is Find all the fields specified by the key nodes. logs, and sort them by log generation time in each area; specify the field value of the key node user name as "UserName1", In the example above, extract the logs with the username value "UserName1".

[0028] Further preferably, in step 4.4, the extracted For each log, record the number of the area where the log is located in sequence according to the log generation time to form a log number sequence ,in is the number of the region where the u-th log is located; the time series formed by the log number sequence is the path formed by the specified log corresponding to the keyword between different regions; the log number sequence is detected using the time series anomaly detection model Perform anomaly detection, that is, analyze the paths formed by logs between various areas, and regard logs related to abnormal paths as abnormal logs.

[0029] Based on the analysis of log templates, the present invention uses the field names in the log templates as nodes, establishes directed edges in the order of nodes, and thus establishes a knowledge graph of the log template. Logs belonging to the same log template are in the same area, and different areas are associated through the same field names between log templates to establish wireless edges; within the same area, classification is performed according to the values ​​of key nodes, and cluster analysis of abnormal logs is performed according to other field names in the same classification; between different areas, on the basis of specifying key nodes, anomaly detection is performed by analyzing the time series formed by logs between different areas. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 Schematic diagram of log partition knowledge graph.

[0031] Figure 2 Schematic diagram of the log anomaly detection process within the region.

[0032] Figure 3 Schematic diagram of the log anomaly detection process between regions. DETAILED DESCRIPTION

[0033] For the large number of logs generated by the information system, the log template is used to divide the area, and the abnormal logs are analyzed within and between areas. In the example process, the user name and file name are used as the key fields of the log to analyze the abnormal behavior of a user or the abnormal behavior associated with a file.

[0034] A log partition knowledge graph construction and anomaly detection method, the steps are as follows:

[0035] Step 1: Create a regional knowledge graph based on the log template;

[0036] Step 1.1: Analyze the log template;

[0037] Each log record has its own specific format, including multiple "field name-field value" forms, indicating that the log record of an event needs to include the field name used and the field value corresponding to the field name. By analyzing all system log information, log analysis tools are used to analyze the log templates contained in the system log information to form a log template library. , indicating that there are p different field name combinations. are the 1st, 2nd, …, pth log templates respectively. A log template can be formally expressed as , indicating that the i-th log template contains "field name-field value" format, Respectively represent the 1st, 2nd, ..., field names, Respectively represent the 1st, 2nd, ..., The field value corresponding to the field name.

[0038] Step 1.2: Log partition knowledge graph construction;

[0039] For each log template, a knowledge graph is created with field names as nodes (e.g. Figure 1 As shown), for Log Templates Included nodes, forming the The knowledge graph of the region Respectively represent 1,2,…, For the first , , Log templates ( ), respectively , , field name, corresponding to , , The knowledge graph of the region contains , , nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes; directed edges are established between nodes in the region in the order in which the field names appear, and directed edges indicate the adjacency relationship between two nodes.

[0040] Step 2: Analyze the correlation between the knowledge graphs of each region;

[0041] Step 2.1: Analyze the same field names between log templates;

[0042] Through string search and matching technology, the same field names between log templates are obtained. Log Templates , for the The first of the log templates ( ) field names , respectively, match all field names in other log templates and record the same field names:

[0043] ;

[0044] In the formula, Indicates obtaining the same field name in the log template. Indicates that all log templates are related to Log Templates The Field Name The same field name, Indicates Log Templates The Field Name , Indicates Log Templates The Field Name .

[0045] Step 2.2: Establish the correlation between the knowledge graphs of each region;

[0046] Undirected edges are established based on the same field names between log templates; , in Nodes in a region , No. Nodes in a region , No. Nodes in a region Undirected edges are constructed between nodes to indicate the association between these nodes. The association is that the nodes connected by the undirected edges describe the same log field.

[0047] Step 3: Cluster analysis of log anomalies in the region, such as Figure 2 As shown;

[0048] Step 3.1: Extract logs in the region;

[0049] Each log in all system log information belongs to a certain area; logs in the area are extracted from all system log information, that is, specified events using the same log template.

[0050] Step 3.2: Determine the key nodes in the region and classify them according to the key nodes;

[0051] According to the nature of the specified event, specify a specific field name as a key node in the area, such as specifying a user name or file name as a key node, and classify logs with the same field value according to the field value corresponding to the key node.

[0052] When the user name is used as the key node, the logs with the same user name field value are classified into one category, such as the logs with the user name "UserName1" are classified into one category, the logs with the user name "UserName2" are classified into one category, and so on. Logs related to the same user are classified into one category.

[0053] Similarly, when the file name is used as the key node, logs with the same file name field value are grouped into one category, such as grouping logs with the file name "FileName1" into one category, and so on, grouping logs related to the same file into one category.

[0054] Step 3.3: Analyze abnormal behavior by key nodes;

[0055] For one category in step 3.2, cluster analysis is performed based on a field name or a vector consisting of multiple field names in the area to obtain the cluster center, and logs whose distance from the cluster center is greater than a threshold are considered abnormal logs.

[0056] For example, for the classification of the user name "UserName1", clustering is performed according to the field name such as access time to obtain the cluster center , and set the distance threshold , calculate the log to cluster center Distance , the distance Greater than distance threshold Logs that deviate from regular access time are considered abnormal logs, that is, logs that deviate from regular access time are considered abnormal user behavior logs.

[0057] For example, for the classification of the file name "FileName1", clustering is performed according to the field name such as the access user, and the cluster center is obtained. , and set the distance threshold , calculate the log to cluster center Distance , the distance Greater than distance threshold The logs that deviate from the normal access users are regarded as abnormal logs, that is, the logs that deviate from the normal access users are regarded as abnormal logs for file access.

[0058] Step 4: Analyze the abnormal logs between regions through time series, such as Figure 3 As shown;

[0059] Step 4.1: Specify a specific field name as a key node;

[0060] According to the nature of the specified event, a specific field name is designated as a key node, such as specifying a user name or a file name as a key node.

[0061] Step 4.2: Find the area related to the key node;

[0062] Find the specified key node in all regions and record the region number containing the key node ,in represents the number of the rth region containing the key node, , is the total number of area numbers that contain key nodes. Record the area number that contains the user name ,in is the number of the rth region containing the user name, , The total number of zones containing user names.

[0063] Step 4.3: Specify the field value of the key node;

[0064] Number the area containing the key nodes Find all the fields specified by the key nodes. logs and sort them by log generation time in each area. Specify the key node user name field value as "UserName1" and In the example above, extract the logs with the username value "UserName1".

[0065] Step 4.4: Analyze the paths formed by the logs between various areas, and regard the logs related to the abnormal paths as abnormal logs;

[0066] Extracted in step 4.3 For each log, record the number of the area where the log is located in sequence according to the log generation time to form a log number sequence ,in is the number of the region where the uth log is located. The time series formed by the log number sequence is the path formed by the specified log corresponding to the keyword between different regions. Use the time series anomaly detection model to detect the log number sequence Perform anomaly detection, that is, analyze the paths formed by logs between various areas, and regard logs related to abnormal paths as abnormal logs.

[0067] This embodiment provides a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the method for constructing a log partition knowledge graph and detecting anomalies is implemented.

[0068] The invention described above only expresses the implementation methods of the embodiments of the present invention, and cannot be understood as limiting the scope of the invention patent, nor does it impose any form of limitation on the structure of the embodiments of the present invention. It should be pointed out that for ordinary technicians in this field, several changes and improvements can be made without departing from the concept of the embodiments of the present invention, which all belong to the protection scope of the embodiments of the present invention.

Claims

1. A log partition knowledge graph construction and anomaly detection method, characterized in that: Here are the steps: Step 1: For each log template, a knowledge graph is established with the field name as the node. Each log template contains several nodes, and all the nodes of each log template constitute the knowledge graph of the region; Step 2: Analyze the same field names between log templates, and establish undirected edges based on the same field names between log templates; Step 3: Analyze log anomalies in the area through clustering; Step 3.1: Extract logs in the region; Step 3.2: Determine the key nodes in the region and classify them according to the key nodes; According to the nature of the specified event, a specific field name is designated as a key node in the region, and logs with the same field value are classified according to the field value corresponding to the key node; Step 3.3: Analyze abnormal behavior by key nodes; For one category in step 3.2, cluster analysis is performed based on a field name or a vector consisting of multiple field names in the region to obtain the cluster center. Logs whose distance from the cluster center is greater than a threshold are considered abnormal logs. Step 4: Analyze the abnormal logs between regions through time series analysis; Step 4.1: Specify a specific field name as a key node; According to the nature of the specified event, specific field names are designated as key nodes; Step 4.2: Find the area related to the key node; Step 4.3: Specify the field value of the key node; find all the fields specified by the key node. Logs are sorted by log generation time in each area, and the field value uses the user name value; Step 4.4: Analyze the paths formed by the logs between various regions, and regard the logs related to the abnormal paths as abnormal logs.

2. The log partition knowledge graph construction and anomaly detection method according to claim 1 is characterized in that: In step 1, by analyzing all system log information, log analysis tools are used to analyze the log templates contained in the system log information to form a log template library. , indicating that there are p different field name combinations. are the 1st, 2nd, …, pth log templates respectively; a log template is formally expressed as , indicating that the i-th log template contains "field name-field value" format, Respectively represent the 1st, 2nd, ..., field names, Respectively represent the 1st, 2nd, ..., The field value corresponding to the field name.

3. The log partition knowledge graph construction and anomaly detection method according to claim 2 is characterized in that: In step 1, a knowledge graph is created for each log template with field name as node. Log Templates Included nodes, forming the The knowledge graph of the region Respectively represent 1,2,…, nodes; for the , , Log templates ( ), respectively , , field name, corresponding to , , The knowledge graph of the region contains , , nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes, with Respectively represent 1,2,…, nodes; directed edges are established between nodes in the region in the order in which the field names appear, and directed edges indicate the adjacency relationship between two nodes.

4. The log partition knowledge graph construction and anomaly detection method according to claim 3 is characterized in that: In step 2, the same field names between log templates are obtained through string search and matching technology; Log Templates , for the The first of the log templates Field Name , , respectively match all field names in other log templates and record the same field names: ; In the formula, Indicates obtaining the same field name in the log template. Indicates that all log templates are related to Log Templates The Field Name The same field name, Indicates Log Templates The Field Name , Indicates Log Templates The Field Name ; according to , in Nodes in a region , No. Nodes in a region , No. Nodes in a region Construct undirected edges in .

5. The log partition knowledge graph construction and anomaly detection method according to claim 4 is characterized in that: In step 4.2, the specified key node is found in all regions, and the region number containing the key node is recorded. ,in represents the number of the rth region containing the key node, , The total number of area numbers containing key nodes; record the area number containing the user name ,in is the number of the rth region containing the user name, , The total number of zones containing user names.

6. The log partition knowledge graph construction and anomaly detection method according to claim 5 is characterized in that: In step 4.3, the area number containing the key node is In the example, specify the field value of the key node User Name as "UserName1", respectively in In the example above, extract the logs with the username value "UserName1".

7. The log partition knowledge graph construction and anomaly detection method according to claim 6 is characterized in that: In step 4.4, the extracted For each log, record the number of the area where the log is located in sequence according to the log generation time to form a log number sequence ,in is the number of the region where the u-th log is located; the time series formed by the log number sequence is the path formed by the specified log corresponding to the keyword between different regions; the log number sequence is detected using the time series anomaly detection model Perform anomaly detection, that is, analyze the paths formed by logs between various areas, and regard logs related to abnormal paths as abnormal logs.

Citation Information

Patent Citations

  • Web log abnormal behavior identification method based on knowledge graph

    CN114328962A

  • Multi-source heterogeneous log anomaly detection method and system based on hybrid drive

    CN118296532A