Abnormal behavior detection method and system based on multi-source log and electronic equipment
By performing alignment processing, attribute filtering and association mode mining on multi-source logs, combined with adaptive time windows and LSTM neural networks, the problems of insufficient analysis complexity and association analysis capabilities of multi-source logs are solved, and efficient abnormal behavior detection is achieved.
Patent Information
- Application Number
- CN202510245419.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is difficult to effectively fuse heterogeneous data sources when processing multi-source logs, and cannot capture potential associations and complex patterns in cross-origin data. Inconsistent log formats lead to complex analysis, weak correlation analysis capabilities of multi-source logs, and immature abnormal identification.
A method of abnormal behavior detection based on multi-source logs is proposed, including log alignment processing, information gain ratio algorithm attribute filtering, DBSCAN algorithm clustering, improved Apriori algorithm mining association mode, adaptive time window construction time series, and abnormal behavior detection is performed using LSTM neural network.
Effectively process multi-source logs, improve log processing efficiency, enhance multi-source log analysis association analysis capabilities, improve the accuracy and comprehensiveness of abnormal identification, and ensure the scalability and adaptability of the system.
Smart Images

Figure CN120050108A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to an abnormal behavior detection method, system and electronic device based on multi-source logs. Background Art
[0002] With the rapid development of information technology, especially the wide application of cloud computing, Internet of Things (IoT) and big data technologies, companies, organizations, school official agencies, etc. all adopt computer systems and cloud servers to process business, and a large amount of data and information are generated and recorded, including important security events and operation records. As an important part of information system security management, log auditing can discover and prevent potential security threats and risks through the analysis and monitoring of these records, ensuring the security and integrity of information assets.
[0003] Logs are the key data sources for recording information such as system operations, user behaviors, error events, and system statuses. By analyzing logs, anomalies that can be identified include: system anomalies, performance anomalies, security events, etc. Traditional anomaly recognition methods mainly use machine learning to analyze log data. However, since a large amount of time is spent on feature engineering during the training of the model, and the feature extraction process highly depends on the knowledge and experience of relevant personnel in a specific industry. When using traditional methods to process multi-source logs, it is often impossible to effectively integrate these diverse and heterogeneous data sources, without considering the correlation between different source logs, and unable to capture potential associations and complex patterns in cross-source data. In recent years, many new solutions have been proposed, from methods for identifying anomalies from multi-source logs based on association rules to semi-supervised log analysis methods, which can all well improve the recognition accuracy and help solve the problem of slow anomaly recognition speed.
[0004] Existing solutions provide more measures for anomaly recognition. However, in the actual application process, they still face many problems, resulting in the inability to meet the requirements of anomaly recognition, such as:
[0005] 1. The non-uniform log data format makes log analysis complex and difficult. The diversity and complexity of log data are a challenge. Log data generated by different applications, systems, and devices have different formats and structures, and flexible technologies and algorithms are required to process and parse these diverse log data. The differences between various types of logs are large, causing great trouble for analysts to compare, and the marking of logs is a huge task, which will make log analysis more complex.
[0006] 2. The ability to analyze multi-source logs in relation is relatively weak. Traditional correlation algorithms and technologies may not be able to effectively handle complex multi-log correlation problems. Multi-log correlation involves considering multiple variables and relationships and requires the use of more advanced correlation analysis techniques. In practical applications, when dealing with log data from multiple different sources, effectively correlating and analyzing this log data is a rather complex task. Moreover, nowadays, the steps of network attacks are complex and attackers' operations are more concealed. Correlating and analyzing the logs of relevant devices, using technologies such as deep learning to dig out hidden attack paths and discover abnormal behaviors will become new challenges.
[0007] 3. Abnormality recognition is not yet mature. Currently, in the field of logs for anomaly detection and threat recognition, there are still some deficiencies and challenges. The main challenges faced in this field include how to represent heterogeneous events and unstructured messages; when multiple processes are running simultaneously or collecting distributed logs centrally, relevant log events may be intertwined, and it is not easy to retrieve the original events if the events lack session identifiers. An abnormality usually refers to unexpected behavior, and unknown instances cannot be trained, which limits semi-supervised learning and unsupervised learning. This means that anomaly detection and threat recognition systems cannot promptly identify and respond to unknown security threats, and the algorithms and models need to be continuously updated and improved.
[0008] In summary, despite the continuous development of abnormality recognition technologies, existing solutions still do not achieve the ideal systematicness and simplicity. For this reason, researchers are committed to innovating and improving abnormality recognition technologies to find truly effective solutions and better meet the actual needs of the security field. Summary of the Invention
[0009] To solve the problems of the existing technology, the present invention proposes an abnormal behavior detection method and system based on multi-source logs, including:
[0010] In the first aspect, the present invention proposes an abnormal behavior detection method based on multi-source logs, and this method includes:
[0011] S1: Collect log data from various network devices and applications, perform alignment processing and data preprocessing on multi-source log events, and generate normalized log data.
[0012] S2: Use the information gain ratio algorithm to screen the attributes of the normalized log data set, and encapsulate the log data after attribute screening into JSON format.
[0013] S3: Extract features from the log data encapsulated in JSON format, convert them into data points, use the DBSCAN algorithm to cluster each data point, and delete irrelevant log data according to the clustering results, retaining the log clusters.
[0014] S4: Use the improved Apriori algorithm to mine the hidden association patterns among the log attribute items in the log clusters, generate association rules, and construct a knowledge graph based on the association rules.
[0015] S5: Use an adaptive time window to construct the time series of the log data, combine the context information in the knowledge graph and the label sequence generated by the association rules, and input the time series of the log data and the label sequence into the LSTM neural network to obtain the abnormal behavior detection result.
[0016] In a second aspect, the present invention proposes an abnormal behavior detection system based on multi-source logs, and the system includes:
[0017] A collection module, which includes each collection component or collection tool, installs each collection component or collection tool to each network device, and collects log data in real time;
[0018] A data storage module for storing the collected log data;
[0019] A data preprocessing module for preprocessing the collected log data;
[0020] A feature extraction module for extracting features from the preprocessed log data;
[0021] An abnormal behavior detection module, which is provided with a trained LSTM neural network, a predefined association rule library and a knowledge graph, and inputs the extracted features into the trained LSTM neural network, the predefined association rule library and the knowledge graph to obtain the abnormal behavior detection result;
[0022] An output module for outputting the abnormal behavior detection result.
[0023] In a third aspect, the present invention proposes an electronic device, including: a memory and at least one processor, and computer-readable instructions are stored in the memory;
[0024] The at least one processor calls the computer-readable instructions in the memory to execute the above-mentioned abnormal behavior detection method based on multi-source logs.
[0025] The present invention can effectively process log data from different sources, mine the potential relationships among the logs through association analysis, and generate corresponding association features; using the LSTM network to model and learn these association features can discover the feature patterns in the multi-source logs and identify abnormal behaviors.
[0026] Compared with the existing technologies, the present invention has the following advantages:
[0027] (1) Unify the log data format and improve the log processing ability. The present invention designs an efficient log unification preprocessing mechanism, effectively solving the heterogeneity problems of multi-source logs in terms of format, semantics, and structure. Through a standardized data processing flow, the processing efficiency of the system for logs of various devices such as servers, switches, and routers is significantly improved.
[0028] (2) The correlation analysis ability of multi-source log analysis is significantly enhanced. An association rule mining model is constructed based on the improved Apriori algorithm to achieve automatic identification of complex association relationships between log events. This model can not only accurately capture known abnormal patterns but also effectively mine potential abnormal behavior characteristics, capable of dealing with the emergence of unknown anomalies.
[0029] (3) Significantly improve the accuracy and comprehensiveness of multi-source log anomaly recognition. The present invention introduces an optimized LSTM deep learning model for time series feature analysis, establishing a deep understanding of the time-dependent relationship of log events. Through the accurate modeling of complex time series patterns, the false alarm rate of the system is significantly reduced, and the accuracy and reliability of anomaly detection are improved.
[0030] (4) Ensure the scalability and adaptability of the system. The present invention adopts a distributed architecture design. By deploying acquisition components on each terminal device and uniformly managing them by the central server, high scalability of the system is achieved. At the same time, the system supports multiple log acquisition protocols such as SNMP, Syslog, and IPMI, and can flexibly adjust the monitoring strategy according to the actual application scenario, with good practicality. Description of the Drawings
[0031] Figure 1 is the step flow chart of the embodiment of the present invention;
[0032] Figure 2 is the structural schematic diagram of the anomaly behavior detection system based on multi-source logs in the embodiment of the present invention. Detailed Embodiments
[0033] Terms such as "first", "second", "third", "fourth", etc. in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application.
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0035] An embodiment of the present invention proposes an abnormal behavior detection method based on multi-source logs. Referring to Figure 1 as shown, the method includes:
[0036] S1: Collect log data from various network devices and applications, perform alignment processing and data preprocessing on multi-source log events, and generate normalized log data.
[0037] The log data comes from various network devices, and the collection methods include SNMP, Syslog, IPMI, Agent, etc. In the embodiments of the present invention, the types of logs include system logs and traffic logs.
[0038] The network devices include devices such as servers, switches, and routers. The main ways to collect logs are to collect relevant log information through the Simple Network Management Protocol (SNMP), Intelligent Platform Management Interface Protocol (IPMI), Syslog, Agent, etc. For server devices, the system obtains hardware status log data through the IPMI protocol, including system-level information such as CPU usage, memory usage, disk status, and fan operation; obtains system logs through the Syslog protocol, recording system startup information, process status, error information, etc.; for network devices such as switches and routers, collects operation logs through the SNMP protocol and Syslog protocol respectively, including information such as port status, traffic statistics, interface status, and security events.
[0039] It should be noted that the Simple Network Management Protocol SNMP is a protocol widely used in network device management and monitoring, especially widely used in network devices such as routers, switches, and firewalls. SNMP can help managers collect status information, performance data, and fault alarms of network devices, so as to realize remote management and monitoring of network devices. The log data collected from network devices using SNMP is shown in Table 1.
[0040] Table 1 Log data collected from network devices using SNMP
[0041]
[0042] It should be noted that the Intelligent Platform Management Interface (IPMI) is a standard interface defined for monitoring the physical characteristics of peripheral devices. At the same time, it provides a function of log management, supporting the monitoring and recording of hardware and system-level events. IPMI by default uses port 623 of the UDP protocol and has good autonomy characteristics. It is an independent board card located on the motherboard and is not affected by whether the server operating system, processor, and BIOS are running. It is equivalent to a subsystem that can run independently within the same operating system. The log data is collected from network devices using IPMI as shown in Table 2.
[0043] Table 2 Log data collected from network devices using IPMI
[0044] Field Name Field Type Field Comment timestamp float UNIX timestamp of hardware monitoring data, accurate to seconds device_id string Unique identification code of the server device host_ip string IP address of the monitored server cpu_temp float CPU temperature sensor reading, in degrees Celsius system_temp float System temperature sensor reading, in degrees Celsius fan_speed int Fan speed, in revolutions per minute (RPM) power_consumption float Total power consumption of the machine, in watts (W) event_description string Detailed description information of the event voltage float Motherboard voltage value, in volts
[0045] It should be noted that Syslog is a long-term used log tool, usually used to configure the logs of Unix / Linux systems. During the operation of the system, devices such as the kernel and application programs will require errors, warnings, and other information to be recorded, and corresponding records will be generated in the log files. Through these log messages, the running state of the system can be well analyzed and it is convenient for subsequent analysis. The log data is collected from network devices using Syslog as shown in Table 3.
[0046] Table 3 Log data collected from network devices using Syslog
[0047] Field Name Field Type Field Comment timestamp float UNIX timestamp of monitoring data, accurate to seconds device_id string Unique identification code of the server device host_ip string IP address of the monitored server severity int Severity level hostname string Host name message string Content of the log message
[0048] It should be noted that the Agent is a proxy program deployed locally on servers and network devices, used to collect and send log data. The Agent can work in coordination with protocols such as IPMI, Syslog, and SNMP to ensure the integrity and accuracy of log data. During the process of collecting log data, the Agent sends heartbeat messages to the detection system at regular intervals to let the monitoring system determine the running state of the Agent. When the log data is updated, the Agent will automatically collect the updated data without the need to set a timing mechanism to collect it on time, ensuring that the collected log data has high real-time performance. The Agent has autonomy. When it senses a change in the surrounding environment, it will respond to relevant events. For example, when the Agent is collecting log data and detects that the CPU occupancy of the server is too high, it will automatically slow down the collection speed to ensure the stable operation of the server. When it detects that the CPU resources are abundant, it will increase the log collection speed within a certain threshold range.
[0049] Introduce an event alignment mechanism (Log Synchronization, logSyc) to process the collected log data, and achieve the alignment of multi-source log events through time, space, and attribute context analysis.
[0050] Perform alignment processing on the collected log data to achieve the alignment of multi-source log events. The process includes:
[0051] S101: Calibrate the time of the log data, and group the log events calibrated in the time window into event sets.
[0052] Eliminate the clock drift between network devices to ensure the accurate alignment of events on a unified timeline. Time reference selection: Select the time of the central log management server as the reference time, or obtain the global synchronized time through NTP. Offset calculation: For the original timestamp of each device, calculate its offset from the reference time:
[0053] Δt offset,i =t base -t raw,i
[0054] where Δt offset,i represents the time offset, which is the difference between the log time of the i-th device and the reference time, in seconds, and is used to correct the clock drift. Δt offset,i can be estimated through the device heartbeat signal or the time synchronization record in the log metadata; t base represents the reference time, usually the time of the central log management server or the NTP synchronized time, in seconds, serving as a unified time reference; t raw,i represents the original timestamp of the log of the i-th device, in seconds, and is directly extracted from the log data.
[0055] Time calibration: Adjust the timestamp of each log to the calibrated time:
[0056] t aligned,i =t raw,i +Δt offset,i
[0057] In the formula, t aligned,i represents the calibrated time of each log, the calibrated timestamp of the log of the i-th device, in seconds, representing the corrected unified time.
[0058] Time window division: Define a fixed time window length W (for example, 5 seconds), and group the calibrated log events into event sets:
[0059] E W ={e j ∣t aligned,j ∈[t start ,tstart +W]}
[0060] In the formula, E W represents the set of events within the time window, and e j represents the j-th log event, which includes fields such as timestamp, device ID, description, etc.; j represents the index of the j-th event in the log dataset, used to identify and traverse all events; t aligned,j represents the calibrated timestamp of the j-th log event, with the unit of seconds; t start represents the start time of the window, and W represents the fixed window length, with the unit of seconds. For example, W = 5 seconds, which defines the size of the time window, and W can be adjusted according to the application scenario.
[0061] S102: Generate a network device topology relationship diagram, and identify the event correlation between physically or logically adjacent devices within the time window. The event correlation includes the topological distance and log attribute similarity between network devices.
[0062] Topological matching of spatial context:
[0063] Extract the network device topology: Obtain the nodes N i (devices) and edges E ij (connection relationships) from the generated network device topology relationship diagram G.
[0064] Define the neighborhood: For the device N W in the event e i in the window E i , its neighborhood is:
[0065] Neighbor(N i ) = {N j ∣E ij ∈ Topology}
[0066] In the formula, Neighbor(N i ) represents the neighborhood set of the device N i related to the i-th event e i in the time window, N i represents the i-th device node, N i represents the device where the log event occurs, N j represents the j-th device node, N j may be adjacent to N i , E ij represents the edge in the topology graph, indicating the direct connection relationship between devices N i and N j , and Topology represents the network device topology relationship diagram, which is a set containing all device nodes and connection edges.
[0067] Calculate the topological distance between network devices:
[0068]
[0069] Dist(N i ,N j ) = min(hops between N i and N j )
[0070] In the formula, S space (e i ,e j ) represents the spatial similarity score, whose value range is [0, 1], indicating the topological proximity between two event devices; Dist(N i ,N j ) represents the topological distance, whose unit is hops, indicating the shortest path length between device N i and N j ; N i represents the device node where event e i occurs, and N j represents the device node where e j occurs; hops represents the number of hops in the network topology, that is, the direct or indirect connection quantity between two devices. If Dist ≤ 1, then S space = 1.0.
[0071] Feature matching of attribute context: Identify associated events through the similarity of log attributes within the time window. Extract log attributes: Extract key attributes from the event set E W within the time window, such as source IP S IP、 target IP D IP、 port Port.
[0072] Calculate attribute similarity: Calculate the attribute similarity for each pair of events (e i ,e j ) within the window:
[0073] S attr (e i ,e j ) = w 1 ·δ(S IP,i ,S IP,j ) + w 2 ·δ(D IP,i ,D IP,j ) + w 3 ·δ(Port i ,Port j )
[0074] In the formula, S attr (ei , e j ) represents the attribute similarity score of each pair of events (e i , e j ) within the time window. Its value range is [0, 1], indicating the similarity degree of two events in key attributes; w 1 represents the weight of the source IP of the first weight coefficient, indicating the contribution of the source IP to the similarity; δ(S IP,i , S IP,j ) represents the source IP matching function, w 2 represents the weight of the destination IP of the second weight coefficient, indicating the contribution of the destination IP; δ(D IP,i , D IP,j ) represents the destination IP matching function, w 3 represents the weight of the port of the third weight coefficient, indicating the contribution of the port; δ(Port i , Port j ) represents the port matching function, where δ(x, y) = 1 (if x = y), otherwise 0.
[0075] For example, the weights w 1 = 0.5, w 2 = 0.3, w 3 = 0.2. w 1 , w 2 , and w 3 are not fixed standards, but are set according to the experience of general network security scenarios.
[0076] S103: Event integration and output description: Calculate the comprehensive scores of each event in the log in terms of time, space, attribute, and semantic context, and filter out the events with comprehensive scores greater than the threshold and merge them into an aligned event sequence.
[0077] For each pair of events (e W in the event set E i , e j ) within the time window, calculate the comprehensive score:
[0078] S total (e i , e j ) = w t ·S time + w s ·S space + w a ·S attr
[0079] In the formula, S total (e i , e j ) represents the comprehensive score of each pair of events (e W ) in the event set Ei , e j ) Calculate the comprehensive score, whose value range is [0, 1], S total (e i , e j ) represents the overall relevance of the two events and is used to determine whether to merge; w t represents the fourth weight coefficient, the time context weight, indicating the contribution of temporal proximity; S time represents the time similarity score. If e i , e j are in the same time window, then S time is 1, otherwise S time is 0; w s represents the spatial context weight, indicating the contribution of topological proximity; S space represents the spatial similarity score, see the topological distance formula; w a represents the attribute context weight, indicating the contribution of attribute similarity. S attr represents the attribute similarity score, see the attribute similarity formula.
[0080] For example, S time = 1.0 (in the same window), the weight w t = 0.3, w s = 0.2, w a = 0.3, w m = 0.2.
[0081] Filter out the events whose comprehensive score is greater than the threshold ψ, that is, S total > ψ, then merge them into an event sequence. The threshold ψ can be adjusted according to the actual situation. For example, ψ = 0.8.
[0082] Perform data preprocessing on the aligned event sequence, including: data cleaning, missing value handling, standardizing the log format, and log compression and encryption. Specifically:
[0083] (a) Data cleaning: Sort the logs according to the "occurrence time" to ensure that similar events can be arranged adjacent to each other; set a fixed-size time window T (for example, T = 60 seconds), and slide this window on the sorted log dataset; within each window, check the exact similarity between records. For the exactly identical records detected within the window, keep one record and delete the rest of the duplicates. For example, within a time window, multiple log records show that the same source IP sent the same request to the same destination IP, and these records will be regarded as duplicate events and only one record will be kept to reduce redundancy.
[0084] (b) Handling missing values: Conduct a completeness check on the key attributes in the log records, including event type, timestamp, source IP address, user ID, etc. If more than three key attributes are missing in a log record, it is considered that the record information is insufficient, and the record will be filtered out from the dataset; for records with fewer missing attributes, each missing attribute will be filled with the value "Null" to maintain the integrity of the dataset. For example, in a device log, if a record only lacks the source IP address and event type, these two attributes of the record will be marked as "Null" instead of deleting the entire record; in addition to filling the missing attributes with the value "Null", filling can also be performed based on context information. For example, if the source IP address of a log record is missing, but the source IP address of this record is the same as that of the previous and next few records, the missing source IP address can be inferred based on the context information.
[0085] (c) Standardizing the log format: The timestamps in the log records may exist in multiple formats, and they need to be converted to the unified Unix timestamp format; for data with a standard format such as a database, the data fields can be directly obtained; for unstructured log records, regular expressions can be used for field identification and extraction.
[0086] (d) Log compression and encryption: For historical log data, compression algorithms can be used for compression to reduce storage space occupancy, and the compressed log data can still be accessed and analyzed through decompression; for sensitive log data, encryption algorithms can be used for encrypted storage to prevent data leakage, and the encrypted log data needs to be decrypted during access.
[0087] The normalized log is the result of accurate processing of the fields in the content of Table 1-3. The field content is arranged in the order of the fields in the table and separated by commas.
[0088] Example: 1696515840.0,server2,192.168.1.2,4,server2,"Failed password for root from 192.168.1.200 port 22 ssh2"
[0089] After preprocessing the multi-source heterogeneous log data, normalized log data can be obtained, effectively solving the heterogeneity problems of multi-source logs in terms of format, semantics, and structure, and significantly improving the processing efficiency of log data.
[0090] Generate a network device topology relationship diagram based on the standardized log data. Specifically: collect and preprocess the SNMP logs, routing table information, ARP table logs of network devices, and network connection data obtained by actively scanning through the network discovery protocol. Extract the neighbor information of the devices from the SNMP logs, obtain the next-hop information of the devices from the routing table, extract the IP-MAC address mapping relationship of the devices from the ARP table logs, and obtain the connection relationship between the devices through the network discovery protocol. Abstract each network device as a node in the topology relationship diagram. The nodes include network devices such as servers, switches, and routers, and set device attributes for each node. The device attributes at least include device type, device name, IP address, and MAC address. At the same time, abstract the physical or logical connections between the devices as edges in the topology relationship diagram; use a graph database to store and represent the device nodes and connection edges in the form of a graph structure to generate a complete network device topology relationship diagram.
[0091] As can be seen from Tables 1-3 above, in the standardized log data set, there are still problems such as field redundancy, inconsistent field definitions, high field correlation, and excessive field noise, and attribute screening is required.
[0092] S2: Use the information gain ratio algorithm to perform attribute screening on the standardized log data set, and encapsulate the log data after attribute screening into JSON format data.
[0093] It should be noted that the information gain ratio algorithm is a measurement method used to select the best splitting attribute in the decision tree algorithm, and it also belongs to the attribute screening method based on information theory. The information gain ratio normalizes the information gain by considering the intrinsic value of the attribute (that is, the number of all possible values of the attribute), thereby overcoming the problem that the information gain tends to select attributes with a large number of possible values. In the embodiments of the present invention, the information gain ratio algorithm is used to evaluate the contribution degree of each attribute in the standardized log data to the detection of abnormal logs. Its core idea is to calculate the magnitude of the information gain ratio by quantifying the discrimination ability of the attribute to the classification target, and screen out the attributes that are most valuable for anomaly detection. The standardized log data set D includes multiple attributes A = {A 1 , A 2 ,..., A n}, such as device type, IP address, etc. The information entropy is used to measure the uncertainty of the data set.
[0094] Use the information gain ratio algorithm to perform attribute screening on the standardized log data set to obtain the log data after attribute screening. The specific process includes:
[0095] S201: Calculate the information entropy H(D) of the standardized log data set. The calculation formula is:
[0096]
[0097] Wherein, H(D) represents the information entropy of the normalized log dataset D, and p k represents the proportion of the k-th type of samples (such as normal logs or abnormal logs) in the log dataset D.
[0098] The larger the information entropy H(D), the higher the uncertainty of the normalized log dataset D.
[0099] S202: Calculate the conditional entropy H(D|A i ) of each attribute A in the normalized log dataset D under given conditions, and its calculation formula is: i )
[0100]
[0101] Wherein, D v represents the subset where the value of attribute A i is v, for example, all logs with the device type being "router"; v ∈ Values(A i ) is the value set of attribute A i , v takes values from this set, |D v | represents the number of elements (log data records) in the subset D i where the value of attribute A v is v, that is, the number of log records that satisfy the value of attribute A i being v; |D| represents the total number of elements (log data records) in the normalized log dataset D, which is the scale of the entire dataset.
[0102] The conditional entropy reflects the uncertainty of the log dataset D under the given condition that the attribute A i is known.
[0103] S203: Calculate the information gain IG(D,A i ) of the normalized log dataset D, and its calculation formula is:
[0104] IG(D,A i ) = H(D) - H(D|A i )
[0105] The information gain is used to measure the contribution degree of the attribute A i to the attribute classification target in the log dataset D. The larger the information gain, the greater the contribution of the attribute to distinguishing normal logs and abnormal logs.
[0106] S204: Calculate each type of attribute A in the normalized log dataset D iThe eigenvalue IV(A i ), and its calculation formula is:
[0107]
[0108] The eigenvalue IV(A i ) is used to measure the value distribution of each type of attribute A in the normalized log dataset D i . The larger the eigenvalue, the more dispersed the value distribution of the attribute.
[0109] S205: Calculate the information gain ratio in the normalized log dataset D, and its calculation formula is:
[0110]
[0111] The information gain ratio is the ratio of the information gain to the eigenvalue, which is used to eliminate the influence of the attribute value distribution on the information gain. The larger the information gain ratio, the higher the contribution degree of the attribute to the abnormal log detection.
[0112] S206: Retain the attributes in the normalized log dataset D whose information gain ratio IGR(D, A i ) is greater than the threshold, and delete the remaining attributes to obtain the log data after attribute screening.
[0113] The threshold can be obtained through repeated experiments according to user requirements.
[0114] The attributes of the log data after attribute screening include but are not limited to timestamp, network device ID, device type, IP address, log level, and event description, etc.
[0115] It should be noted that JSON is a lightweight data exchange format, which has good readability and is easy to write quickly. It can be used for data exchange between different platforms. Therefore, the collected log data is encapsulated in JSON format. Logstash is used as the component for data collection and conversion. Logstash is one of the core tools of the Elastic Stack and provides a flexible and extensible data processing pipeline. With the help of Logstash, the aggregation of multi-source heterogeneous logs and JSON format encapsulation can be realized, and the extracted information is shown in Table 4.
[0116] Table 4 JSON data after formatting processing
[0117]
[0118]
[0119] After this step, the preference of information gain for attributes with more values is eliminated, ensuring that the attribute selection process is more fair and reasonable; it can avoid selecting redundant or noisy features, thus constructing a more representative attribute set; by reducing the selection of redundant attributes, the information gain ratio can reduce the dimension of attributes and improve the overall efficiency.
[0120] After step S2, the log data (i.e., JSON-formatted data) still includes a large amount of duplicate and irrelevant log data, which needs to be processed. Otherwise, it will increase the unnecessary computational amount and affect anomaly detection.
[0121] In the embodiments of the present invention, the fields selected from the JSON-formatted data include: log_type, message, timestamp, host_ip, port, severity. Among them, log_type represents the type of log event, message represents the content of the log event, timestamp represents the timestamp of the log event, host_ip represents the host IP address, port represents the port, and severity represents the severity level of the log. Each log event can describe its detailed information through these feature fields, and the remaining fields in the JSON-formatted data are irrelevant fields and should be deleted.
[0122] S3: Extract features from the log data encapsulated in JSON format, convert it into data points, and use the DBSCAN algorithm to cluster each data point. Delete the irrelevant log data according to the clustering results and retain the log clusters.
[0123] Encode, filter, and extract multi-dimensional features for the fields of the JSON-formatted data respectively. Specifically:
[0124] Perform one-hot encoding (i.e., One-Hot encoding) on the type data in the JSON-formatted data. The type data in the JSON-formatted data are fields such as log_type. The One-Hot encoding refers to a coding method that converts categorical data into binary vectors. It converts each category into a vector, in which only one element is 1 and the rest are 0. Each category corresponds to a unique index, indicating its position in the vector. By using One-Hot encoding to process the log_type data, it can be converted into multiple binary value vectors, thus generating a new feature matrix.
[0125] Use the TF-IDF algorithm to extract features from the text data in the JSON-formatted data to obtain word vectors.
[0126] It should be noted that the Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is a text feature extraction algorithm commonly used in text analysis and information retrieval tasks. It weights each term (such as a word) in the text by evaluating the combination of term frequency (TF) and inverse document frequency (IDF), thereby effectively screening out the most informative words in the document. The message field of each log is regarded as an independent "document", and each term is weighted through TF-IDF calculation to obtain the feature representation of this field, namely the word vector.
[0127] The TF represents the frequency of a term in a single document, which reflects the relative importance of a certain word in this document. Intuitively, the higher the frequency of a word in a document, the greater its importance in this document. The TF calculation formula is as follows:
[0128]
[0129] where t is the term and d is the document, representing the frequency of term t in document d.
[0130] The IDF represents a measure of the distribution of a term in the entire document collection, reflecting the ability of a certain word to distinguish documents. If a word appears frequently in all documents, then it has no distinguishability, so the IDF value will be low; conversely, if a word appears in a few documents, then it has high distinguishability and its IDF value is high. The IDF calculation formula is as follows:
[0131]
[0132] where N is the total number of documents, and the numerator is the number of documents containing term t. IDF measures the rarity of a term in the document set, and frequently occurring terms will get a lower IDF value.
[0133] It should be noted that the core idea of TF-IDF is to evaluate the relative importance of a term in a certain document by multiplying the term frequency (TF) and the inverse document frequency (IDF). A high TF-IDF value of a term in a certain document indicates that this word has high importance in the document. The TF-IDF calculation formula:
[0134] TF-IDF(t,d) = TF(t,d) × IDF(t)
[0135] By combining the term frequency (TF) and the inverse document frequency (IDF), the TF-IDF value of term t in document d is obtained.
[0136] The above-mentioned one-hot encoding and TF-IDF algorithm are used to complete the extraction of the attributes (i.e., numerical features) of the JSON-formatted data, and convert them into data points to form a data point set. The DBSCAN algorithm is used to cluster the data point set.
[0137] It should be noted that DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density-based clustering algorithm suitable for clustering analysis of spatial data. It can identify clusters of any shape and can label noise points (i.e., data points that do not belong to any cluster).
[0138] The DBSCAN algorithm clusters the data point set, and its specific process includes:
[0139] S301: Define the neighborhood radius ∈ and MinPts, and arbitrarily select a data point p from the data point set;
[0140] S302: If the data point p satisfies the parameters ∈ and MinPts, that is, the number of points included in the neighborhood with radius ∈ is greater than or equal to MinPts, then the data point p is regarded as a core point, and all data points that are density-reachable from p are found to form a cluster;
[0141] S303: If the data point p does not satisfy the parameters ∈ and MinPts, that is, it does not meet the core point condition and belongs to the boundary of a certain cluster, then it is determined that the data point p is an edge point (also regarded as an abnormal data point, which is noise);
[0142] S304: Repeat steps S302 and S303 until all data points are processed to obtain the clustering result.
[0143] The neighborhood refers to all points within a given radius ∈. For a point p, its neighborhood refers to the set of all points whose distance from point p is less than or equal to ∈. The points within its neighborhood refer to all points whose Euclidean distance from point p is less than or equal to ∈. The formula is:
[0144] N(p,∈)={q∣distance(p,q)≤∈}
[0145] Among them, distance(p,q) is the distance between point p and q, and ∈ is the preset radius.
[0146] The formula for Euclidean distance is:
[0147]
[0148] where p=(p 1 ,p 2 ,…,p n) and q = (q 1 , q 2 , …, q n ) are two points in an n-dimensional space. d(p, q) is the Euclidean distance between point p and point q. This formula represents the square root of the sum of the squared differences between two points.
[0149] The cluster is a set composed of core points and their neighborhood points. The points in the cluster have a high density with each other and can be connected through the core points.
[0150] Direct density reachability is: If sample point q is within the ∈-neighborhood of p and p is a core point, then p directly reaches q, which can be expressed as: and p is a core point directly reaches q if there exists a series of sample points p 1 , p 2 , …, p n , where p i to p i+1 is directly density reachable, then it is said that p 1 to p n is density reachable. p 1 to p n is density reachable if p i directly density reaches p i +1 A core point is the "center" of the cluster, and there are enough points in its neighborhood. The core point determination formula is as follows:
[0151]
[0152] According to the clustering result, delete the irrelevant log data (i.e., delete the abnormal data points) and retain the log clusters (i.e., the clustering clusters).
[0153] S4: Use the improved Apriori algorithm to mine the hidden association patterns between the log attribute items in the log clusters, generate association rules, and construct a knowledge graph based on the association rules.
[0154] It should be noted that the Apriori algorithm is an association rule mining algorithm widely used in the field of data mining. The Apriori algorithm aims to discover the frequently occurring item sets in the dataset and generate association rules based on these frequent item sets. The algorithm reduces the computational amount by gradually expanding the size of the frequent item sets and using the "prior knowledge". Specifically, it assumes that all subsets of a frequent item set are also frequent.
[0155] Use the improved Apriori algorithm to mine the hidden association patterns between the log attribute items in the log clusters, generate association rules, and its specific steps include:
[0156] (1) Initialization: Define items as the specific values of various features or events in the log cluster. Treat each item as an independent candidate item set. Define an item set as a set containing a group of items in the log event. Construct a hash tree to store and retrieve item set information.
[0157] The item set refers to a set of items contained in the log event. An "item" is the specific value of a feature or event in the log, such as device type, IP address, log level, etc. Each log event can be regarded as a set composed of several items. Item sets are divided into single-item sets and multi-item sets. A single-item set is an item set that contains only one item, such as {Router} or {WARNING}. A multi-item set is an item set that contains multiple items, such as {Router, 192.168.1.1} or {WARNING, Connection Timeout}.
[0158] (2) Traverse the log cluster data set, use the hash tree to quickly count the occurrence times of each single-item set, calculate its support degree, and regard the item set with a support degree greater than the minimum support degree threshold as a frequent item set.
[0159] The frequent item set refers to an item set with a support degree greater than the preset threshold. That is to say, if an item set appears frequently enough in the log data, then it is considered a frequent item set. The support degree reflects the universality of the item set's appearance. If the support degree is high, it means that this combination of item sets is relatively common in the data.
[0160] The calculation formula for the support degree (X) is:
[0161]
[0162] Among them, the number of transactions containing the item set (X) is the number of transactions containing the item set X, and the total number of transactions is the total number of all transactions in the data set.
[0163] (3) In each round of iteration, by combining the frequent item sets with other frequent item sets, generate new candidate item sets, and use the hash function to map the new candidate item sets to the corresponding nodes of the hash tree; for each new candidate item set, check whether all its n-1 dimensional subsets are frequent item sets. If so, retain the new candidate item set; otherwise, skip it.
[0164] The candidate item set refers to the item set generated in each round of iteration and waiting to be evaluated for frequency. The candidate item set can be generated by combining the frequent item sets generated in the previous round. Suppose the frequent item sets in the previous round are {Router}, {192.168.1.1}, {WARNING}, then the new candidate item sets can be {Router, 192.168.1.1}, {192.168.1.1, WARNING}, etc.
[0165] (4) Traverse the log cluster dataset again. When encountering a record containing a new candidate itemset, quickly locate the position of the corresponding new candidate itemset through the hash tree, calculate its support degree, and retain the itemset with a support degree greater than the minimum support degree threshold as the new frequent itemset. Continuously iterate until no new frequent itemset can be generated.
[0166] In the frequent itemset generation stage, for each frequent itemset, by checking all its subsets, eliminate the candidate itemset whose subsets are not frequent, ensuring that only potential frequent itemsets are generated, thereby improving the calculation efficiency.
[0167] (5) Generate association rules from the frequent itemset, calculate the confidence, lift, differential lift, and weighted differential lift of each association rule, and filter out the association rules with a confidence greater than the minimum confidence threshold, a lift greater than 1, and a weighted differential lift greater than the set threshold.
[0168] The association rule refers to the rule generated from the frequent itemset and is used to describe the association relationship between item sets in the log data. In the Apriori algorithm, the association rule is usually expressed as A→B, which means that if item set A appears, then item set B may also occur. Confidence is used to measure the probability of another event occurring given that a certain event has occurred. It is usually used to generate association rules and represents the strong association relationship between the condition and the conclusion.
[0169] The specific formula for calculating the confidence is:
[0170]
[0171] Among them, A→B represents the association rule from item set A to item set B, support degree(A∪B) represents the support degree that simultaneously contains item set A and item set B, and support degree(A) represents the support degree of item set A.
[0172] The specific formula for calculating the lift is:
[0173]
[0174] In the formula, lift(A→B) represents the lifting effect of the association rule A→B on the occurrence of item set B.
[0175] The differential lift is used to eliminate the absolute value influence of the lift and directly measure the influence degree of the antecedent on the consequent. A positive differential lift indicates a positive correlation between the antecedent and the consequent, and a negative differential lift indicates a negative correlation between the antecedent and the consequent.
[0176] The specific formula for calculating the differential lift is:
[0177] Differential lift(A→B) = lift(A→B) - 1
[0178] The weighted differential lift is a metric that combines support and differential lift, and is used to more accurately measure the practical significance of rules. The weighted differential lift not only considers the correlation, but also takes into account the universality of association rules in log data.
[0179] The specific calculation formula of the weighted differential lift is as follows:
[0180] Weighted differential lift(A→B) = Support(A∪B) × Differential lift(A→B)
[0181] The association rules reflect the potential dependencies between log events, and these association rules can be used as the basis for abnormal behavior pattern recognition in subsequent steps. Based on the generated association rules, a knowledge graph is constructed to represent the complex relationships between log attribute items.
[0182] The construction of the knowledge graph is based on the results of association rule mining, and defines nodes and edges to represent the relationships between log attribute items. The nodes in the knowledge graph represent each attribute item in the log data, each node corresponds to an entity or attribute, and attributes can be attached to describe its characteristics; the edges in the knowledge graph represent the dependencies between nodes, and each edge can be attached with attributes to quantify the strength or characteristics of the relationship. The knowledge graph is represented in the formalized form of a graph structure G=(V, E), where V is the set of nodes and E is the set of edges, and this graph structure is stored through a graph database, intuitively showing the complex relationship network.
[0183] S5: Use an adaptive time window to construct the time series of log data, combine the context information in the knowledge graph and the label sequence generated by the association rules, and input the time series of log data and the label sequence into the LSTM neural network to obtain the abnormal behavior detection result.
[0184] The time series refers to: the log data is segmented according to the time window, and the log events and their associated rule trigger situations within each time window together constitute a time series input. Within each time window, the features included are not only the attributes of the log events, but also the status of the association rules triggered by the events.
[0185] Define the length of each time window, which can be a fixed time interval or based on the frequency of event triggers.
[0186] T w ={e 1 ,e 2 ,…,e n}
[0187] In the formula, T w represents the set of log events within the time window w, and e iDenote the i-th log time within the time window, where i ∈ {1, 2, …, n}, and n represents the number of log events.
[0188] In the traditional fixed-time window method, a constant window length (such as 60 seconds) is used when constructing the time series, which may not be able to flexibly adapt to the event distribution characteristics of the log data. In the embodiments of the present invention, an adaptive time window mechanism is introduced to adaptively adjust the time window and construct the time series of the log data.
[0189] The adaptive time window mechanism can enhance the detection of sudden events and optimize the event context coverage by introducing event density and severity as dynamic adjustment factors. When the event density is high (such as a traffic surge) or the severity is high (such as an intrusion attempt), the window length is shortened to capture fine-grained anomalies. When the event density is low but the severity is high (such as a privilege escalation), the window length is extended to ensure that the complete context is included in the analysis. The specific calculation steps for adaptively adjusting the time window are as follows:
[0190] Calculate the event density:
[0191]
[0192] In the formula, density represents the event density within a unit time period, N represents the number of events within a unit time period, and T unit represents the unit time length. For example, T unit = 60 seconds.
[0193] Normalize the event density, and its calculation formula is:
[0194]
[0195] In the formula, density norm represents the normalized event density, with a range of [0, 1], which is used for window adjustment; density max represents the preset maximum density. The event density density within a unit time reflects the event occurrence frequency.
[0196] Calculate the event severity:
[0197]
[0198] In the formula, severity represents the event severity within a unit time period, s i represents the i-th event e iThe severity value. For example, the range of severity is 1 - 5, obtained from the log field. The event severity within a unit of time reflects the importance of the event occurrence.
[0199] Normalize the event severity, and its calculation formula is:
[0200]
[0201] In the formula, severity norm represents the event severity value after normalization, and its value range is [0, 1], which is used for window adjustment; s min represents the minimum value of the event severity, s max represents the maximum value of the event severity.
[0202] For example, the range of severity is 1 - 5. After normalization, the range of severity is mapped to 0 - 1.
[0203] Adaptively adjust the time window length according to the event density and event severity, and its calculation formula is:
[0204] W adaptive = W base ×(1 + α·density + β·severity)
[0205] In the formula, W adaptive represents the dynamic length of the time window, in seconds, indicating the dynamically adjusted window size, W base represents the base length of the time window, in seconds, serving as the benchmark for adjustment; α represents the density adjustment coefficient, controlling the impact of density on the window size; density represents the event density within a unit of time; β represents the severity adjustment coefficient, controlling the impact of severity on the window size; severity represents the event severity within a unit of time.
[0206] For example, W base is initially set to 60 seconds as the benchmark value; set α = 0.5, β = 0.5, which are the initial adjustment coefficients and can be intelligently optimized according to the actual scenario.
[0207] Convert the discrete sequences of logs and events into time series suitable for the LSTM model, generate label sequences by combining context information and association rules in the knowledge graph, and use the optimized LSTM neural network model for training and anomaly prediction.
[0208] The construction of the time series refers to arranging the extracted features in chronological order to generate a time series. The shape of the time series is (batch_size, sequence_length, feature_dim), where:
[0209] batch_size represents the number of samples, and sequence_length represents the number of time windows;
[0210] feature_dim represents the feature dimension (including log event attributes and association rule status).
[0211] The generation of the label sequence refers to generating specific label values for the log events within each time window according to the analysis results of association rules and knowledge graphs. The label values reflect different behavior patterns or anomaly types. For example, 0 represents normal behavior, 1 represents abnormal IP address access, 2 represents executing sensitive commands after successful login, etc. Arrange the generated labels in chronological order to generate a label sequence. The shape of the label sequence is (batch_size, sequence_length), where:
[0212] batch_size represents the number of samples, and sequence_length represents the number of time windows.
[0213] The LSTM neural network is a special recurrent neural network used to process and predict time series data. Different from traditional RNNs, LSTM solves the long-term dependence problem through its internal memory cells, that is, it can retain and update information over a long time span, avoiding the problem of gradient vanishing or gradient explosion that occurs in traditional RNNs. The core of LSTM is its internal memory cell, which is controlled by three main gating mechanisms (forget gate, input gate, output gate) to control the flow of information. Its main function is to determine which information should be retained, which should be forgotten, and which information should be output. This enables LSTM to effectively learn the long-term and short-term dependence relationships in time series data.
[0214] The forget gate refers to determining how much information in the memory information of the previous moment (i.e., the cell state of the previous moment) needs to be forgotten. Its calculation formula is:
[0215] f t =σ(W f ·[h t-1 ,x t +b f )
[0216] where σ represents the sigmoid activation function, W f represents the weight matrix of the forget gate, h t-1 is the hidden state at the previous moment, x t is the current input, b f is the bias term.
[0217] The input gate refers to determining how much of the current input information needs to be saved into the cell state. Its calculation formula is: i t = σ(W i ·[h t-1 , x t +b i )
[0218]
[0219] where, i t represents the activation value of the input gate, represents the current candidate cell state, and tanh is the hyperbolic tangent activation function. The cell state update refers to updating the current cell state according to the outputs of the forget gate and the input gate. Its calculation formula is:
[0220]
[0221] where, C t-1 is the cell state at the previous moment, f t is the output of the forget gate, i t is the output of the input gate, is the candidate cell state.
[0222] The output gate refers to determining the output at the current moment according to the cell state. Its calculation formula is:
[0223] o t = σ(W o ·[h t-1 , x t +b o )
[0224] h t = o t ·tanh(C t )
[0225] where, o t is the activation value of the output gate, C t is the current cell state, and h t is the current hidden state.
[0226] The optimization of the LSTM refers to adding a Dropout layer and a fully connected layer. To prevent the LSTM model from overfitting, Dropout can be used during training. Dropout randomly discards some neuron connections, forcing the network not to overly rely on certain specific neurons during training, thereby improving the generalization ability of the model. The calculation formula of Dropout is:
[0227]
[0228] where is the output after Dropout processing, y is the original output, and p is the retention probability.
[0229] The fully connected layer refers to adding a fully connected layer at the last layer of the LSTM network. Usually, the fully connected layer is added to map the output of the LSTM to the target space. This helps to further transform the learned temporal features into the output of classification or regression tasks. The calculation formula of the fully connected layer is:
[0230]
[0231] where is the output of the fully connected layer, W is the weight matrix, h is the output of the hidden layer of the LSTM, and b is the bias term.
[0232] An embodiment of the present invention proposes an abnormal behavior detection system based on multi-source logs. Referring to Figure 2 as shown, the system includes:
[0233] A collection module, which includes various collection components or tools, and installs each collection component or tool on each network device to collect log data in real time.
[0234] Each network device includes servers, switches, routers, etc. Each device will deploy appropriate log collection tools, such as Syslog clients, SNMP collection tools, or specialized IPMI log modules, to collect the running status information of various devices and form a log data set.
[0235] A data storage module for saving the collected log data.
[0236] A data preprocessing module for preprocessing the collected log data.
[0237] A feature extraction module for extracting features from the preprocessed log data.
[0238] An abnormal behavior detection module, which is equipped with a trained LSTM neural network, a predefined association rule library, and a knowledge graph. The extracted features are input into the trained LSTM neural network, the predefined association rule library, and the knowledge graph to obtain the abnormal behavior detection result;
[0239] An output module for outputting the abnormal behavior detection result.
[0240] The data preprocessing module, the feature extraction module, and the abnormal behavior detection module are all set on the central log management server.
[0241] The abnormal behavior detection result is a detailed abnormal report, which includes the Apriori analysis result, the frequent patterns and abnormal rules mined from the log data, and the LSTM analysis result, such as time-dependent anomalies and prediction anomalies based on time series analysis. In addition, the report also includes the specific description of the event, the occurrence time, the possible impact, and the anomaly severity. All detected abnormal results will be recorded in the log database.
[0242] An embodiment of the present invention provides an electronic device, including: a memory and at least one processor, and computer-readable instructions are stored in the memory;
[0243] The at least one processor invokes the computer-readable instructions in the memory to execute the above-mentioned abnormal behavior detection method based on multi-source logs.
[0244] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: ROM, RAM, disk, or optical disc, etc.
[0245] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for detecting abnormal behavior based on multi-source logs, characterized in that: include: Collect log data from a variety of network devices and applications, perform alignment and data preprocessing of multi-source log events, and generate standardized log data; The information gain ratio algorithm is used to perform attribute screening on the normalized log data set, and the attribute-screened log data is encapsulated into JSON format; Extract features from log data encapsulated in JSON format and convert them into data points. Use the DBSCAN algorithm to cluster each data point. Delete irrelevant log data based on the clustering results and retain the log clusters. The improved Apriori algorithm is used to mine the hidden association patterns between log attribute items in the log cluster, generate association rules, and build a knowledge graph based on the association rules; The time series of log data is constructed using an adaptive time window. The context information in the knowledge graph and the label sequence generated by the association rules are combined. The time series and label sequence of log data are input into the LSTM neural network to obtain the abnormal behavior detection results.
2. According to claim 1, the abnormal behavior detection method based on multi-source logs is characterized in that: The collected log data is aligned, and the specific process includes: Perform time calibration on the log data and group the calibrated log events in the time window into event sets; Generate a network device topology diagram to identify the event correlation of physically or logically adjacent devices within a time window. Event correlation includes topological distance between network devices and log attribute similarity. Calculate the comprehensive scores of each event in the log in terms of time, space, attributes, and semantic context, and filter out events with comprehensive scores greater than the threshold and merge them into aligned event sequences.
3. According to claim 1, the abnormal behavior detection method based on multi-source logs is characterized in that: The attributes of the log data after attribute screening include timestamp, network device ID, device type, IP address, log level and event description; the fields in the JSON format data include: log_type, message, timestamp, host_ip, port, severity, wherein log_type indicates the type of log event, message indicates the content of the log event, timestamp indicates the timestamp of the log event, host_ip indicates the host IP address, port indicates the port, and severity indicates the severity level of the log.
4. The abnormal behavior detection method based on multi-source logs according to claim 1 or 3 is characterized in that: The information gain ratio algorithm is used to perform attribute screening on the normalized log data set, and the specific process includes: calculating the information entropy of the normalized log data set; According to the information entropy, calculate each attribute A in the log data set i Conditional entropy under given conditions; Calculate the information gain of the normalized log data set D according to the information entropy and conditional entropy; Calculate the intrinsic value of each attribute in the normalized log dataset; Calculate the information gain ratio in the normalized log data set D according to the information gain and the intrinsic value; The attributes whose information gain ratio is greater than the threshold are retained, and the remaining attributes are deleted to obtain log data after attribute screening.
5. According to claim 1, the abnormal behavior detection method based on multi-source logs is characterized in that: The DBSCAN algorithm is used to cluster each data point. The clustering process includes: Define the neighborhood radius ∈ and MinPts, and arbitrarily select a data point p from the data point set; If the data point p satisfies the parameters ∈ and MinPts, that is, the number of points contained in the neighborhood of radius ∈ is greater than or equal to MinPts, then the data point p is regarded as a core point, and all data points that are density-reachable from p are found to form a cluster; If the data point p does not satisfy the parameters ∈ and MinPts, that is, it does not meet the core point condition and belongs to the boundary of a cluster, then the data point p is determined to be an edge point; The process is iterated until all data points are processed and the clustering results are obtained.
6. The abnormal behavior detection method based on multi-source logs according to claim 1 is characterized in that: The improved Apriori algorithm is used to mine the hidden association patterns between the log attribute items in the log cluster and generate association rules. The specific steps include: Define an item as the specific value of each feature or event in the log cluster, regard each item as an independent candidate item set, define an item set as a set of items in the log event, and build a hash tree; Traverse the log cluster data set, use the hash tree to count the number of occurrences of each single item set, calculate its support, and regard the item sets with support greater than the minimum support threshold as frequent item sets; In each iteration, a new candidate item set is generated by combining frequent item sets with other frequent item sets, and a hash function is used to map the new candidate item set to the corresponding node of the hash tree; for each new candidate item set, check whether all its n-1-dimensional subsets are frequent item sets. If so, keep the new candidate item set; otherwise, skip it; Traverse the log cluster data set again. When encountering a record containing a new candidate item set, quickly locate the corresponding new candidate item set through the hash tree, calculate its support, and retain the item set with support greater than the minimum support threshold as the new frequent item set. Continue iterating until no new frequent item set can be generated. Generate association rules from frequent item sets, calculate the confidence, lift, differential lift and weighted differential lift of each association rule, and filter out association rules whose confidence is greater than the minimum confidence threshold, lift is greater than 1 and weighted differential lift is greater than the set threshold.
7. The abnormal behavior detection method based on multi-source logs according to claim 1 is characterized in that: The event density and event severity within a unit time period are calculated respectively, and the length of the time window is adaptively calculated according to the event density and event severity.
8. An abnormal behavior detection system based on multi-source logs, characterized in that: include: A collection module, which includes various collection components or collection tools, and each collection component or collection tool is installed on each network device to collect log data in real time; A data storage module, used to store the collected log data; A data preprocessing module, used for preprocessing the collected log data; A feature extraction module, used to extract features from the preprocessed log data; An abnormal behavior detection module, which is equipped with a trained LSTM neural network, a predefined association rule base and a knowledge graph, inputs the extracted features into the trained LSTM neural network, the predefined association rule base and the knowledge graph to obtain abnormal behavior detection results; The output module is used to output abnormal behavior detection results.
9. The abnormal behavior detection system based on multi-source logs according to claim 8 is characterized in that: The data preprocessing module, feature extraction module and abnormal behavior detection module are all arranged on the central log management server.
10. An electronic device, characterized in that: include: a memory having computer-readable instructions stored therein and at least one processor; The at least one processor calls the computer-readable instructions in the memory to execute the abnormal behavior detection method based on multi-source logs as described in any one of claims 1-7.
Citation Information
Cited By
Abnormal behavior pattern recognition method and system applied to network security
CN120567572A
Abnormal behavior pattern recognition method and system applied to network security
CN120567572B
Audit method and system based on time sequence characteristic monitoring
CN120611172A
Automatic log analysis method for operating system repair
CN120892242A
Construction method of trusted industrial data space
CN120934911A