Log analysis method, system and device and storage medium
By constructing a dynamic topological network and multi-scale anomaly detection, combined with causal analysis, the shortcomings of existing log analysis methods in real-time, correlation and abnormal detection are solved, and efficient processing of complex log data and accurate identification of exception patterns are achieved.
Patent Information
- Application Number
- CN202510093992.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Existing log analysis methods lack real-time, correlation and abnormal detection capabilities when processing complex log data, making it difficult to adapt to the needs of large-scale log data, resulting in low troubleshooting efficiency and system stability affected.
By building a dynamic topological network, topological characteristics are extracted, combined with multi-scale anomaly detection and causal analysis, the system abnormal patterns are identified and the abnormal root cause devices or events are located.
It realizes efficient processing of multi-source log data, comprehensively identify exception patterns, accurately locate the root cause of abnormalities, and improves the efficiency and accuracy of system troubleshooting.
Smart Images

Figure CN120066833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of log data processing and analysis, and specifically to a log analysis method, system, device, and storage medium. Background Art
[0002] With the rapid development of information technology, large-scale distributed systems and complex network environments have become important components of modern data centers and enterprise IT architectures. In these environments, the massive log data generated by system operations has become a key data source for monitoring, operation and maintenance, and fault troubleshooting. However, existing log analysis methods and tools face many challenges when processing these complex log data.
[0003] Existing technologies usually rely on simple keyword search and time series analysis to process log data. These methods can meet certain basic needs, but it is difficult to adapt to the real-time nature and complexity of large-scale log data. On the one hand, traditional methods often lack the ability to model the complex interaction relationships between devices and cannot comprehensively depict the dynamic behavior of the system; on the other hand, in the face of complex abnormal scenarios caused by multi-device collaboration or linkage in a distributed system, the correlation analysis ability of existing technologies is limited, and it is difficult to identify potential abnormal propagation paths. In addition, existing log analysis tools have low automation levels and mostly rely on manual analysis, resulting in low efficiency of fault troubleshooting and problem location, seriously affecting the stability and operation and maintenance efficiency of the system.
[0004] In practical applications, due to the heterogeneity of log data sources and the diversity of formats, existing technologies also lack unified preprocessing capabilities in the data processing stage and cannot effectively integrate log data from different devices. At the same time, the abnormal detection methods of current log analysis systems are mostly limited to the identification of abnormalities at a single scale, and the comprehensive analysis of short-term sudden abnormalities and long-term trend abnormalities cannot be realized, resulting in the omission of some abnormal patterns. In addition, existing systems generally lack the ability to analyze abnormal propagation paths and root cause devices. During fault troubleshooting, operation and maintenance personnel can only rely on experience for inference, further increasing the difficulty and cost of system maintenance.
[0005] Therefore, existing technologies have significant deficiencies in terms of real-time nature, accuracy, intelligence, and comprehensiveness, and there is an urgent need for a log analysis method that can efficiently process log data, comprehensively identify abnormal patterns, and accurately locate the root cause of abnormalities. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention provides a log analysis method, system, device, and storage medium, which efficiently constructs a dynamic topology network from multi-source log data, and accurately identifies system abnormal patterns and locates the root cause device or event of the abnormality through topology feature extraction, multi-scale abnormal detection, and causal analysis.
[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A log analysis method, comprising the following steps: Collect log data from multi-source devices, clean and structure the log data, and extract the timestamp, event type, device identifier, and description information; Based on the interaction relationships between devices in the log data, generate a dynamic topological network with nodes representing devices and edges representing device interactions, where the network evolves dynamically over time windows; Extract topological characteristics based on the dynamic topological network, including the connectivity of the topological structure, loop structure, and complexity characteristics of the network; Analyze the changes in network topological characteristics on different time scales, and combine local and global anomaly detection to identify abnormal patterns in the log data; Construct a dynamic causal network between devices, extract the abnormal propagation chain, and locate the root cause device or event that triggers the anomaly.
[0008] Preferably, the steps for constructing the dynamic topological network specifically include: Map the source device of each log data to a node in the dynamic network; Map the interaction relationships between devices in the log data to the edges of the network, and the weight of the edge is determined by the number of associated events or interaction intensity in the log; Segment the log data based on the set time window to generate a sequence of dynamic topological networks that evolve over time.
[0009] Preferably, the steps for extracting the topological characteristics specifically include: Assign weights to the edges of the dynamic topological network and construct a filtering sequence to generate a subgraph of the network; Perform persistent homology analysis on the network subgraph to extract the topological characteristics of the network, including features such as connectivity, loop structure, and holes; Calculate topological invariants based on the topological characteristics, including Betti numbers and network topological complexity indicators.
[0010] Preferably, the steps for identifying abnormal patterns in the log data specifically include: Calculate the topological characteristic vector within each time window, and calculate the local anomaly difference degree based on the changes in topological characteristics between time windows; Calculate the global anomaly difference degree based on the change trend of the overall topological characteristic vector in the dynamic network; Integrate the local and global anomaly difference degrees, and determine the abnormal patterns in the log data by setting a threshold.
[0011] Preferably, the steps for locating the root cause device or event that triggers the anomaly specifically include: Construct a dynamic causal network among devices based on the interaction relationships and dynamic topology network among devices in the log data; Extract the abnormal propagation chain in the causal network, and locate the abnormal propagation path through the causal relationship among devices; Combine the abnormal scores of the devices in the propagation chain to determine the root cause device or event that triggers the abnormality.
[0012] Preferably, the determination of the root cause device or event is based on the following conditions: In the abnormal propagation chain, the device or event that has the earliest abnormality and the highest relevance is the root cause; The association among devices in the dynamic causal network is estimated through conditional probability, and the root cause priority of the devices is determined by maximum likelihood estimation.
[0013] Preferably, the method further includes: Generate a visual display based on the log analysis results, including a system health dashboard, an abnormal trend chart, and a root cause analysis report; The content of the visual display includes the changes in network topology characteristics, abnormal detection results, and device status statistics that are updated in real time.
[0014] The present invention also provides a log analysis system, including: A data collection module for collecting log data from multiple devices and preprocessing the log data; A dynamic network construction module for generating a dynamic topology network that evolves over time; A topology characteristic extraction module for extracting the topology characteristics of the dynamic network; An abnormal detection module for detecting local abnormalities and global abnormalities in the log data at different time scales; A root cause analysis module for locating the root cause device or event of the abnormality based on the abnormal propagation path; A visualization module for displaying the log analysis results, including network topology characteristics, abnormal trends, and root cause analysis information.
[0015] The present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method as described above is implemented.
[0016] The present invention also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method as described above is implemented.
[0017] The present invention provides a log analysis method, system, device, and storage medium. It has the following beneficial effects: 1. The present invention can effectively process log data from various devices and in various formats. Through structured preprocessing, it converts different types of log data into standardized inputs, adapts to the heterogeneous data environment in complex systems, and enhances the generality and applicability of log analysis.
[0018] 2. By constructing a dynamic topology network, the present invention can clearly describe the interaction relationships between devices and their dynamic changes over time, comprehensively depict the state and correlation characteristics of system operation, and provide structured basic data support for anomaly detection and analysis.
[0019] 3. By combining local anomaly detection and global anomaly detection methods, the present invention can capture different patterns of short-term sudden anomalies and long-term trend anomalies, provide comprehensive anomaly recognition capabilities, and effectively improve the coverage and detection accuracy of system anomalies.
[0020] 4. Through the construction of a dynamic causal network and the extraction of propagation chains, the present invention can reveal the propagation mechanism of abnormal events between devices, accurately locate the root cause device or event that triggers the anomaly, provide efficient support for system fault troubleshooting and problem location, and reduce the time and cost of fault handling.
[0021] 5. The present invention supports the visual display of analysis results, including topological structures, anomaly trends, root cause devices, etc., provides an intuitive and easy-to-understand overview of the system health status and problem location basis for operation and maintenance personnel, and thus improves the efficiency and accuracy of operation and maintenance work. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a schematic flowchart of the method of the present invention; Figure 2 is a schematic structural diagram of the system of the present invention; Figure 3 is a schematic structural diagram of the computer device of the present invention.
[0023] Among them, 100, data acquisition module; 200, dynamic network construction module; 300, topological feature extraction module; 400, anomaly detection module; 500, root cause analysis module; 600, visualization module; 40, computer device; 41, processor; 42, memory; 43, storage medium. DETAILED DESCRIPTION OF THE INVENTION
[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0025] Please refer to the appendix Figure 1 , the present invention provides a log analysis method. By introducing technical means of dynamic topology network and high-order topology analysis, combined with multi-scale anomaly detection and causal network analysis, it solves the deficiencies in real-time performance, relevance, and anomaly detection in traditional log analysis methods, and can achieve accurate identification of abnormal patterns and root cause location in log data. The following details the specific implementation steps of the present invention.
[0026] As Figure 1 shown, the log analysis method may include the following steps: S1. Collect log data from multi-source devices and perform preprocessing; S2. Build a dynamic topology network based on the log data; S3. Extract topological features based on the dynamic topology network; S4. Analyze the changes in network topological features to identify abnormal patterns in the log data; S5. Build a dynamic causal network between devices to locate the anomaly source.
[0027] For step S1, in this embodiment, step S1 mainly involves the collection, cleaning, and structured processing of log data. The present invention provides basic data support for subsequent analysis steps through the standardized processing of log data from multi-source devices. The following details the specific implementation manner of this step.
[0028] In this embodiment, the log data mainly comes from a variety of different devices, including but not limited to servers, switches, routers, application logs, virtualization platform logs, and container logs, etc. The format of the log data may be structured, semi-structured, or unstructured, and common formats include plain text logs, JSON files, XML files, and system-specific binary log formats.
[0029] As an option, the collection of log data can be achieved through distributed log collection tools. For example, using tools such as Fluentd and Logstash, log data can be collected from multiple sources and uniformly formatted for processing.
[0030] It should be noted that during the log data collection process, the real-time collection mode or batch collection mode can be set according to the generation speed and storage capacity of the logs. For example: The real-time collection mode is suitable for scenarios with high-frequency log generation, and the log data is directly passed to subsequent modules through streaming processing.
[0031] The batch collection mode is suitable for logs of non-critical systems, which are collected regularly and then uploaded for unified processing.
[0032] In a possible implementation, the cleaning of log data includes deduplication, denoising, and filtering of invalid information from the original log data. For example, for log entries with duplicate timestamps and content, redundant records can be identified and deleted through a hashing algorithm or direct string comparison.
[0033] Specifically, the cleaning of log data may include the following: Invalid log filtering: For example, removing log records containing null values or meaningless fields.
[0034] Log time alignment: For device logs with clock deviations, a time synchronization algorithm (such as timestamp calibration based on the NTP protocol) can be used to align the timestamps uniformly.
[0035] Abnormal character handling: Some logs may contain garbled characters or special characters, and these abnormal characters can be cleaned or replaced through regular expressions.
[0036] In some embodiments, the description fields of logs can be cleaned through natural language processing (NLP) techniques. Specifically, stop words (such as "is", "a", "the", etc.) in the log description can be removed and the case can be unified for subsequent structured extraction of the description information.
[0037] Exemplarily, the structured processing of log data can be parsed based on the following fields: Timestamp ( ): By parsing the time information in the log record, it is converted into a unified time format, such as the ISO 8601 standard (YYYY-MM-DDTHH:MM:SSZ). It should be noted that if the log is missing a timestamp, it can be supplemented with its file generation time or other associated information.
[0038] Event type ( ): The event type of the log is extracted through keyword matching or regular expressions. For example, matching "ERROR" to mark error-type logs and matching "WARNING" to mark warning-type logs.
[0039] Device identifier ( ): For logs in a multi-device scenario, the device identifier is extracted by parsing the log source information (such as IP address, hostname, or device ID).
[0040] Description information ( ): The detailed description field in the log is extracted, and the complete information content is retained for subsequent analysis.
[0041] As an option, the extraction of descriptive information can further adopt word segmentation technology to split long texts into key phrases for subsequent pattern matching or classification tasks. For example, for the log "Memory usage exceeds 90%", the word segmentation results may be "Memory", "usage", "exceeds", "90%".
[0042] It should be noted that for the structured storage of log data, a relational database (such as MySQL, PostgreSQL) or a non-relational database (such as Elasticsearch, MongoDB) can be selected. In a possible implementation, using an index structure based on Elasticsearch to store log data can support efficient full-text retrieval and complex queries.
[0043] In some embodiments, to further improve the processing efficiency, a data streaming processing framework (such as Apache Kafka) can be used to transmit structured log data to the storage module in real time.
[0044] To better illustrate, the specific operations of the present invention in an actual application scenario are as follows: Suppose the following logs are collected from two servers (Server_A and Server_B): [20xx-12-30 10:23:45] [ERROR][Server_A] CPU utilization exceeds 95%. [20xx-12-30 10:24:05] [INFO][Server_B] Memory utilization stable. Through cleaning and structuring, the standardized log data obtained is: [ { "timestamp": "2024-12-30T10:23:45Z", "event_type": "ERROR", "source_device": "Server_A", "description": "CPU utilization exceeds 95%"},{ "timestamp": "2024-12-30T10:24:05Z", "event_type": "INFO", "source_device":"Server_B", "description": "Memory utilization stable"}] This structured data provides a standardized input for subsequent steps.
[0045] It is understandable that the quality of log data collection and cleaning directly affects the accuracy and timeliness of subsequent analysis results. In this embodiment, each step in the preprocessing stage can be implemented in a programmed manner and supports distributed deployment to meet the processing requirements of large-scale log data.
[0046] It should be noted that the log data preprocessing method in the present invention can be adapted to various log sources and formats, providing strong generality and flexibility for log analysis.
[0047] For step S2, in this embodiment, through the construction of a dynamic topology network, the association relationship between devices in the log data is abstracted into a network structure that evolves over time. This network structure can accurately describe the interaction behavior between devices and its dynamic changes, providing data support for subsequent anomaly detection and analysis.
[0048] In this embodiment, the construction of the dynamic topology network first maps devices to nodes in the network and the interaction relationship between devices to edges in the network based on the device identifiers, timestamps, and interaction information extracted from the log data. It should be noted that the dynamic topology network not only depicts the static relationship between devices but also introduces the time dimension, enabling the network to evolve dynamically with the division of time windows.
[0049] As an option, the node set of the dynamic topology network can be generated from the set of device identifiers that can be extracted from the log data. For example, in the log records, the source device of each log (such as a server, switch, etc.) can be directly mapped to a node . The generation of nodes does not depend on the device type or log format and can adapt to the data input of various heterogeneous devices.
[0050] Specifically, the edge set in the dynamic topology network is established through the interaction events between devices. The existence of an edge indicates that there is an interaction relationship between devices and within a certain time range. Exemplarily, the interaction relationship can be determined in the following ways: If there is a direct communication event between two devices (such as data transfer logs, request and response logs), an edge can be established between the devices.
[0051] If a collaboration event between devices appears in the log records (such as an alarm event triggered jointly by multiple devices), an edge can be established for these devices.
[0052] In a possible implementation, the weight of the edge For quantifying the intensity of interactions between devices. For example, the edge weight can be defined as the number of interaction events between two devices in the log: where is an indicator function that takes the value 1 when the condition is satisfied, and is the time window. It should be noted that by setting the weight threshold
[0053] edges with low interaction intensity can be filtered out, thereby reducing the complexity of the network. In terms of dynamics, in this embodiment, the log data is segmented according to timestamps, and a fixed time window method is used to divide multiple time slices, and a static topology network is generated within each time slice. Exemplarily, the width of the time window can be set to one minute, five minutes or other time scales, and the specific selection can be adjusted according to the generation frequency of the log data and system requirements. By combining the static networks of multiple time slices in time series, a complete dynamic topology network sequence is formed: where each time slice corresponds to a static network and the overall network evolves over time, depicting the dynamic interaction relationship between system devices.
[0054] In some embodiments, in order to capture the complex linkage behavior between devices, the dynamic topology network can also introduce the concept of higher-order simplicial complexes. Higher-order simplicial complexes can not only describe pairwise interactions between devices, but also describe the joint behavior between multiple devices. For example: when the log record shows that three devices , , are simultaneously involved in an event, it can be represented as a three-dimensional simplex in the network.
[0055] The higher-order simplicial complex can be constructed in the following way: where is the threshold for filtering weak interactions represents the dimension of the simplex.
[0056] It should be noted that the introduction of higher-order simplicial complexes can more comprehensively capture the complex correlation characteristics between devices, and is applicable to scenarios with frequent multi-device collaboration, such as cluster management in data centers or multi-device linkage analysis in network attack events.
[0057] In a possible implementation, the construction of the dynamic topology network can further consider the attribute characteristics of the relationships between devices. For example, by attaching interaction type tags (such as data transmission, alarm triggering, etc.) to each edge, more rich semantic information can be provided for subsequent analysis processes. The extraction of these attribute characteristics can be parsed from the log description fields. For example: For the log "Server_A sends data to Server_B", the edge can be added with the attribute "data transmission".
[0058] For the log "Switch_1 triggers alarm affecting Switch_2", the edge can be added with the attribute "alarm triggering".
[0059] It can be understood that the construction of the dynamic topology network provides a structured input for the extraction of topological characteristics and anomaly detection in subsequent steps. The dynamic topology network construction method of the present invention has high flexibility and adaptability, and can adapt to a variety of log data sources and complex interaction patterns. At the same time, the constructed dynamic network can intuitively present the association relationships between devices in the system, providing important basic data support for anomaly analysis.
[0060] It should be noted that by adjusting the size of the time window, the threshold of the edge weight, and the dimension of the simplex, the dynamic topology network construction method of the present invention can flexibly adapt to the requirements in different scenarios, such as a smaller time window in the high-frequency log scenario or a higher-dimensional simplex modeling in the multi-device linkage scenario. The network designed in this way can comprehensively reflect the state and interaction pattern of the system, providing support for log analysis in complex systems.
[0061] For step S3, in this embodiment, by extracting the topological characteristics of the dynamic topology network, the quantitative description of the global and local characteristics of the network is realized, providing basic data support for subsequent anomaly detection and analysis.
[0062] In this embodiment, the topological characteristic extraction takes the dynamic topology network as the input, and comprehensively depicts the dynamic behavior of the system by calculating the topological structure characteristics of the network, including connectivity, loop structure, and network complexity, etc. It should be noted that the method of topological characteristic extraction is based on persistent homology analysis, which can effectively capture the dynamic changes of topological features.
[0063] As an option, the topological characteristic extraction first constructs a filtration function and generates a sequence of multi-scale network subgraphs based on the weights of nodes and edges in the dynamic topology network. Specifically, the filtration function sorts the edges in the network and gradually introduces the edges in ascending order of weight, thereby gradually constructing the subgraphs of the network: Among them, is a filtering parameter, representing the filtering threshold of the current subgraph.
[0064] It can be understood that by gradually increasing the filtering threshold, a network sequence from sparse to dense can be generated, thereby capturing the topological characteristics of the network at different scales.
[0065] In a possible implementation, by performing persistent homology analysis on the filtered network sequence, the following two types of key topological characteristics can be extracted: 0-dimensional characteristic (connectivity): Represents the number of connected components in the network, that is, the number of unconnected subgraphs in the current network.
[0066] 1-dimensional characteristic (loop structure): Represents the number of closed paths in the network, that is, the loops formed in the network.
[0067] Specifically, persistent homology forms a persistence barcode by calculating the birth and death times of topological features for quantifying the stability of features. For example: For a connected component, the birth time represents the filtering threshold at which it first appears, and the death time represents the filtering threshold at which it is incorporated into other components.
[0068] For a loop structure, the birth time represents the filtering threshold at which it first appears, and the death time represents the filtering threshold at which the loop is filled or broken.
[0069] As an option, topological feature extraction can also further quantify the features of the network based on topological invariants. Specifically, the following topological invariants are calculated in this embodiment: Betti number: Used to quantify the number of topological features in different dimensions. For k-dimensional topological features, its Betti number calculation formula is: Among them, is the boundary operator of the k-dimensional chain complex, representing the boundary information of topological features. For example: represents the number of connected components in the network.
[0070] represents the number of loops in the network.
[0071] Topological entropy: Used to measure the topological complexity of the network, defined as: Among them, represents the proportion of topological features of each dimension. It should be noted that the higher the topological entropy, the more complex the topological structure of the network.
[0072] In some embodiments, the results of topological feature extraction can be represented by a feature vector, for example: Among them, represents at time The number of connected components represents the number of loops and represents the topological complexity. These feature vectors can comprehensively reflect the changes of the dynamic network in the time dimension and provide quantitative inputs for anomaly detection.
[0073] Exemplarily, for a dynamic topological network , if it is found during the filtering process that the number of connected components decreases significantly and the number of loops increases significantly, it may indicate that there is an anomaly in the cooperation of multiple devices in the system. For example, in a distributed system, the cooperation of multiple devices may lead to high load or network congestion, and this phenomenon is manifested as a significant change in the network structure in the topological characteristics.
[0074] It should be noted that in order to further improve the computational efficiency of topological feature extraction, in this embodiment, parallel processing can be performed on persistent homology analysis. For example: In a multi-core computing environment, persistent barcodes can be calculated in parallel for different time slices of the network.
[0075] In a large-scale network scenario, topological features can be calculated in parallel for different regions of the network.
[0076] In the above manner, topological feature extraction can not only accurately reflect the global and local characteristics of the dynamic network, but also maintain high computational efficiency in large-scale data scenarios.
[0077] It can be understood that topological feature extraction is an important step of the present invention. Through the calculation of persistent homology and topological invariants, it can accurately capture the complex correlation information in the dynamic network. These topological features provide quantitative support for subsequent anomaly detection, and at the same time, through the comparison of global and local characteristics, they can reveal the dynamic evolution law of the system.
[0078] For step S4, in this embodiment, by analyzing the changes of dynamic topological features within different time scales and combining the dual methods of local anomaly detection and global anomaly detection, the anomaly patterns in the log data are accurately identified. Multi-scale anomaly detection can comprehensively monitor the operating state of the system from the micro and macro levels.
[0079] In this embodiment, multi-scale anomaly detection takes the topological features of the dynamic topological network as input, analyzes the changes of topological features within each time window, determines the occurrence of local anomalies and global anomalies, and calculates the comprehensive anomaly index to provide accurate anomaly identification results for subsequent anomaly propagation path analysis.
[0080] As an option, the time scale for anomaly detection can be adjusted according to the specific application scenario. Specifically, in a high-frequency log environment, the time window can be set to the second level, while in a low-frequency log environment, the time window can be set to the minute level or a larger time interval. By flexibly adjusting the time scale, this embodiment can adapt to different scales of log analysis scenarios.
[0081] Specifically, the multi-scale anomaly detection method in this embodiment includes the following core steps: Local anomaly detection In a possible implementation, local anomaly detection analyzes the topological property changes between adjacent time windows. By comparing the topological property vectors of adjacent windows, the local difference degree is calculated to identify abnormal fluctuations within a short period of time.
[0082] The calculation formula for the local difference degree is: Where, is the topological property vector of the time window , represents the Euclidean distance.
[0083] As an option, the determination criterion for local anomaly detection can be achieved by setting a threshold . When the local difference degree exceeds the threshold, it is determined that there is a local anomaly in the time window .
[0084] Exemplarily, in some embodiments, if the topological properties of the system change significantly within a short period of time, such as a sudden decrease in the number of connected components ( decrease) or a sudden increase in the number of loop structures ( increase), these changes can be captured through local anomaly detection.
[0085] Global anomaly detection In a possible implementation, global anomaly detection identifies abnormal trends over a longer period of time by analyzing the differences between the current time window and the global average characteristics. Specifically, by comparing the topological property vector of the time window with the global average topological property vector , the global difference degree is calculated.
[0086] The calculation formula for the global difference degree is: Where, is the average value of the topological properties of all time windows, is the total number of time windows.
[0087] It should be noted that the determination threshold It can be set according to the tolerance range of the system. When the global difference degree exceeds the threshold, it is determined that the time window has a global anomaly.
[0088] Abnormality index calculation In this embodiment, in order to comprehensively combine the detection results of local anomalies and global anomalies, an abnormality index is defined as a metric for overall anomaly detection. The calculation formula of the abnormality index is: Where and are the standard deviations of the local difference degree and the global difference degree respectively, and are used for normalization processing.
[0089] It can be understood that the higher the abnormality index, the more serious the overall anomaly degree within the current time window. By analyzing the dynamic changes of the abnormality index, potential problems that may exist in the system can be effectively captured.
[0090] In some embodiments, the change of the topological characteristics of the dynamic topological network can reflect the abnormality of the system operation state. For example: In a distributed system, if some nodes lose connection, it may cause a significant increase in the number of connected components ( increase), which can be captured by local anomaly detection.
[0091] In a system with collaborative operations, if multiple nodes form an abnormal collaboration loop, it may cause an increase in the number of loop structures ( increase), which can be captured by global anomaly detection.
[0092] It should be noted that in order to improve the calculation efficiency of multi-scale anomaly detection, the following optimization measures can be introduced in this embodiment: In local anomaly detection, the sliding window technique is adopted to calculate only the differences between adjacent time windows, reducing redundant calculations.
[0093] In global anomaly detection, the incremental update technique is adopted to update only the global average feature vector without having to recalculate the means of all windows.
[0094] Through the above optimization measures, this embodiment can maintain a high detection efficiency in the scenario of large-scale log data.
[0095] It can be understood that multi-scale anomaly detection realizes the accurate identification of abnormal patterns in log data by combining the dual analysis methods of local anomalies and global anomalies. Based on the dynamic topological network, the multi-scale anomaly detection method of the present invention can effectively capture the short-term fluctuations and long-term trends of the system, providing important inputs for subsequent abnormal propagation path analysis and root cause location.
[0096] For step S5, in this embodiment, by constructing a dynamic causal network between devices, the abnormal propagation chain is analyzed and the root cause device or event that caused the abnormality is located. Based on the abnormality detection results of the aforementioned step S4, this step further reveals the cause and propagation path of the abnormality, providing precise guidance for system operation and maintenance and problem repair.
[0097] In this embodiment, the core of abnormal propagation path analysis is to extract the causal relationship between devices through log data and dynamic topology network. The construction of dynamic causal network can reveal the dependency relationship between devices and the transmission mechanism of abnormal events. As an option, the construction of causal relationship can be based on conditional probability inference, by analyzing the interaction mode between devices, to identify the causal chain between abnormal events.
[0098] Specifically, in this embodiment, the nodes of the dynamic causal network correspond to devices or events, and the edges represent causal relationships. For example, if the abnormality of device A triggers the abnormality of device B, a directed edge from device A to device B is established in the causal network.
[0099] Construction of causal networks In one possible implementation, the causal network is constructed through conditional probabilities Indicates the causal relationship between devices. The calculation of conditional probability can be based on the time and event association information in the log data. For example: in, Indicates the device Following the abnormality, the device The number of unusual events, Indicates the device The total number of anomaly events.
[0100] As an option, weights can be assigned to the edges in the causal network to quantify the strength of the causal relationship. For example, the edge weight can be defined as a normalized form of the conditional probability value to screen the paths with strong causal relationships.
[0101] Abnormal propagation chain extraction Specifically, the extraction of the abnormal propagation chain in this embodiment is based on the directed path in the causal network. By performing a depth-first search or a breadth-first search on the abnormal event nodes and their causal relationships, a complete abnormal propagation chain can be extracted.
[0102] It should be noted that the result of the abnormal propagation chain extraction is in the form of: in, It is the strength threshold of the causal relationship, which is used to filter the paths of weak correlation. It can be understood that by setting a suitable threshold, the most important abnormal propagation path can be focused on to avoid the causal network being too complicated.
[0103] In some embodiments, if there are multiple branches in the propagation chain of the causal network, the main path of abnormal propagation can be identified by accumulating metrics such as edge weights or path lengths. For example, if the accumulated edge weight of a certain propagation chain is the largest, it can be determined as the main path of abnormal propagation.
[0104] Root cause localization In this embodiment, root cause localization identifies the root cause device or event that triggers the abnormality based on the starting node in the abnormal propagation chain. Specifically, the criteria for determining the root cause device are as follows: If the device is located at the starting point of the abnormal propagation chain, and the weights of its related edges are all higher than the set threshold , then it is determined to be the root cause device.
[0105] If there are multiple starting devices in the abnormal propagation chain, the device with the highest conditional probability is selected as the root cause device.
[0106] Exemplarily, the calculation formula for root cause localization is: It should be noted that this formula determines the most likely root cause device by maximizing the influence of the device on subsequent abnormal events.
[0107] In some embodiments, the application scenarios of the dynamic causal network and the abnormal propagation chain include: Tracing the source of device failures in a distributed system: If the CPU overload of server A causes an abnormality in the load balancer, which further leads to a connection timeout of server B, the causal network can identify the propagation chain from server A to server B and locate server A as the root cause device.
[0108] Analysis of the attack path in a network attack event: If the firewall record shows that the attack traffic from IP1 causes an abnormality in the router, which further affects the internal server, the causal network can identify the attack path and lock IP1 as the root cause of the abnormality.
[0109] In a possible implementation, to improve the efficiency of abnormal propagation path analysis, the following optimization measures can be adopted in this embodiment: Incrementally construct the causal network: For newly generated log data, only update the affected part of the causal network to avoid recalculating the relationships between all devices.
[0110] Parallelize path search: For a large-scale causal network, the propagation paths starting from different starting nodes can be searched in parallel, significantly shortening the calculation time.
[0111] Through the above optimization measures, this embodiment can adapt to the scenario of high-frequency log data and complete the analysis of abnormal propagation paths in a short time.
[0112] In this embodiment, by constructing a dynamic causal network and combining the propagation chain extraction and root cause localization methods, the abnormal propagation mechanism and root cause devices in the system can be accurately identified. It can be understood that the abnormal propagation path analysis method of the present invention can not only locate the cause of abnormal problems, but also provide a clear path guidance for the fault handling of the system.
[0113] The method for log data analysis provided by the present invention realizes the monitoring of the system operation state and the identification and traceability of potential problems by collecting, processing, analyzing and detecting anomalies in log data. The following is the overall work flow of the method of the present invention: 1. Collection and preprocessing of log data The first step of the present invention is to collect raw log data from multiple data sources (such as devices like servers, switches, routers, etc.). These data may come from various formats and types of logs, including plain text logs, JSON format logs or XML logs, etc.
[0114] Preprocess the collected log data, including: Clean invalid log records (such as meaningless debug information or redundant records).
[0115] Extract key fields in the logs, such as timestamps, event types, device identifiers and event description information.
[0116] Unify log data in different formats into a structured format to provide a standardized input for subsequent analysis.
[0117] 2. Construction of dynamic topology network By analyzing the device interaction relationships in the log data, construct a dynamic topology network, specifically including: Map the devices extracted from the logs to nodes in the network.
[0118] Map the interaction behaviors between devices to edges in the network and assign weights to the edges according to the interaction intensity.
[0119] Divide the log data into multiple time slices according to time windows, construct a static topology network for each time slice, and form a time - serialized dynamic topology network.
[0120] The dynamic topology network can intuitively present the association relationships between devices in the system and their changes over time, laying a foundation for subsequent feature extraction and anomaly detection.
[0121] 3. Extraction of topology features Based on the constructed dynamic topology network, extract the topology features of the network, mainly including: Features describing the network structure, such as connectivity and loop structure, used to depict the basic topology form of the network.
[0122] By quantifying the indicators of network complexity, the overall state change of the system is further reflected.
[0123] The extracted topological features are updated over time with the dynamic changes of the network, providing a quantitative basis for subsequent anomaly detection.
[0124] 4. Multi-scale Anomaly Detection By analyzing the changes in network topological features at different time scales, anomaly patterns in log data are detected, specifically including: Local Anomaly Detection: Analyze the changes in topological features within adjacent time windows to identify sudden anomalies in the short term.
[0125] Global Anomaly Detection: Analyze the differences between the current time window and the historical state of the overall system to identify long-term trend anomalies.
[0126] Combining the local and global detection results, calculate the comprehensive anomaly index and quantify the degree of anomaly, providing anomaly nodes and time window information for subsequent analysis of the anomaly propagation path.
[0127] 5. Anomaly Propagation Path Analysis Based on the results of anomaly detection, construct a dynamic causal network between devices to analyze the propagation mechanism of anomalies, specifically including: Extract the anomaly propagation chain and identify the causal relationships between devices or events.
[0128] By tracing the starting point of the propagation chain, locate the root cause device or event that triggers the anomaly.
[0129] Finally, generate the anomaly propagation path, providing a clear logical basis for system fault troubleshooting and problem location.
[0130] 6. Output and Application Output the analysis results, including: Visual display of the system operation status, such as device topology, health status, and anomaly trends.
[0131] Reports on the anomaly propagation chain and root cause devices, providing decision-making support for operation and maintenance personnel.
[0132] The method of the present invention realizes the intelligent analysis and problem location of the system operation status through the whole process operation from the collection, processing of log data to anomaly detection and root cause location. The whole process has clear steps and close logical relationships, and is applicable to the log data analysis requirements in various system scenarios.
[0133] Generally speaking, through the collection, cleaning, and structured processing of multi-source device log data, the present invention constructs a topology network that dynamically changes over time, extracts the topological characteristics of the network, and performs multi-scale anomaly detection. It identifies local and global anomaly patterns in the system and analyzes the anomaly propagation path based on a dynamic causal network to accurately locate the root cause device or event that triggers the anomaly. This method can adapt to various log formats and multi-source data inputs of complex systems, support real-time analysis and dynamic monitoring, and provide efficient and accurate technical support for system fault troubleshooting and operation status optimization.
[0134] The log analysis system described below can be correspondingly referred to the log analysis method described above.
[0135] Please refer to the attached Figure 2 , the present invention also provides a log analysis system, including: A data acquisition module 100, configured to collect log data from multiple devices and preprocess the log data; A dynamic network construction module 200, configured to generate a dynamic topology network that evolves over time; A topological characteristic extraction module 300, configured to extract the topological characteristics of the dynamic network; An anomaly detection module 400, configured to detect local anomalies and global anomalies in the log data at different time scales; A root cause analysis module 500, configured to locate the root cause device or event of the anomaly based on the anomaly propagation path; A visualization module 600, configured to display the log analysis results, including network topological characteristics, anomaly trends, and root cause analysis information.
[0136] The system of this embodiment can be used to execute the method embodiment above. The principle and technical effect are similar and will not be elaborated here.
[0137] Please refer to the attached Figure 3 , the present invention also provides a computer device 40, including: a processor 41 and a memory 42. The memory 42 stores a computer program executable by the processor. When the computer program is executed by the processor, it executes the method as described above.
[0138] The present invention also provides a storage medium 43. A computer program is stored on the storage medium 43. When the computer program is run by the processor 41, it executes the method as described above.
[0139] Among them, the storage medium 43 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc.
[0140] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A log analysis method, characterized in that: The following steps are involved: Collect log data from multiple source devices, clean and structure the log data, and extract timestamps, event types, device identifiers, and description information; Based on the interaction relationship between devices in log data, a dynamic topology network is generated, where nodes represent devices and edges represent device interactions. The network evolves dynamically as the time window changes. Extract topological characteristics based on dynamic topological networks, including the connectivity of topological structures, ring structures, and network complexity characteristics; Analyze changes in network topology characteristics over different time scales, and identify abnormal patterns in log data by combining local and global anomaly detection; Build a dynamic causal network between devices, extract the anomaly propagation chain, and locate the root cause device or event that causes the anomaly.
2. The log analysis method according to claim 1, characterized in that: The steps of constructing the dynamic topology network specifically include: Map the source device of each log data to a node of the dynamic network; Map the interaction relationship between devices in the log data to the edge of the network. The weight of the edge is determined by the number of associated events or the interaction intensity in the log. The log data is segmented based on the set time window to generate a dynamic topology network sequence that evolves over time.
3. The log analysis method according to claim 1, characterized in that: The step of extracting the topological characteristics specifically includes: Assign weights to the edges of a dynamic topology network and construct a filtering sequence to generate a subgraph of the network; Perform persistent homology analysis on network subgraphs to extract network topological characteristics, including connectivity, ring structure, and holes; Topological invariants are calculated based on topological properties, including Betti numbers and network topological complexity indicators.
4. The log analysis method according to claim 1, characterized in that: The step of identifying abnormal patterns in log data specifically includes: Calculate the topological characteristic vector in each time window, and calculate the local anomaly difference based on the change of topological characteristics between time windows; Based on the changing trend of the overall topological characteristic vector in the dynamic network, the global anomaly difference is calculated; The difference between local anomalies and global anomalies is comprehensively considered, and the abnormal patterns in the log data are determined by setting thresholds.
5. The log analysis method according to claim 1, characterized in that: The step of locating the root cause device or event causing the abnormality specifically includes: Based on the interaction relationship and dynamic topology network between devices in the log data, a dynamic causal network between devices is constructed; Extract the abnormal propagation chain in the causal network and locate the abnormal propagation path through the causal relationship between devices; Combine the anomaly scores of the devices in the propagation chain to determine the root cause device or event that caused the anomaly.
6. The log analysis method according to claim 5, characterized in that: The root cause device or event is determined based on the following conditions: In the abnormality propagation chain, the device or event that first occurs abnormality and has the highest correlation is the root cause; The inter-device associations in the dynamic causal network are estimated by conditional probability, and the root cause priorities of the devices are determined by maximum likelihood estimation.
7. The log analysis method according to claim 1, characterized in that: The method further comprises: Generate visual displays based on log analysis results, including system health dashboards, abnormal trend graphs, and root cause analysis reports; The content of the visual display includes real-time updated network topology characteristic changes, anomaly detection results and device status statistics.
8. A log analysis system, used to execute the log analysis method according to any one of claims 1 to 7, characterized in that: include: A data collection module is used to collect log data from multiple devices and pre-process the log data; Dynamic network building module, used to generate dynamic topological networks that evolve over time; A topological feature extraction module is used to extract the topological features of the dynamic network; Anomaly detection module, used to detect local and global anomalies in log data at different time scales; The root cause analysis module is used to locate the root cause device or event of the anomaly based on the anomaly propagation path; The visualization module is used to display log analysis results, including network topology characteristics, abnormal trends, and root cause analysis information.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Monitoring data acquisition and analysis method based on Internet of Things
CN117076934A
Computer network anomaly detection method
CN118784364A
Marine ecology-oriented time-space diagram neural network anomaly detection method and system
CN119312267A
Cited By
Variable information sign safety processing method and device based on multi-dimensional anomaly detection
CN120429799A
Multi-modal project data analysis method
CN120822146A