Security log research and judgment method and system based on threat intelligence
By constructing a log semantic inconsistency index and an intelligence log time matching failure index, the problems of semantic mapping inconsistency and time synchronization failure of multi-source heterogeneous log data are solved, realizing the accuracy and completeness of security log analysis based on threat intelligence and ensuring the security of power networks.
Patent Information
- Application Number
- CN202511109442.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-11
AI Technical Summary
Existing threat intelligence-based security log analysis methods lack stable assessment and adaptive adjustment mechanisms for semantic mapping quality when processing multi-source heterogeneous log data, leading to semantic ambiguity and time synchronization failure, which affects the accuracy of threat identification and response efficiency.
By constructing a log semantic inconsistency index and an intelligence log time matching failure index, and jointly calculating the unavailability index of the dynamic association mechanism, dynamic reconstruction is performed to ensure the consistency of semantic and intelligence matching and time synchronization, and the final judgment result is output.
This improves the accuracy and completeness of security log analysis, avoids delays, misjudgments, and omissions in identifying potential threats, and ensures the reliability of power network security information.
Smart Images

Figure CN120934826A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid data security technology, specifically to a security log analysis method and system based on threat intelligence. Background Technology
[0002] As next-generation power networks continue to evolve towards intelligence, distributed computing, and deep interconnection, network boundaries are becoming increasingly blurred, and attack surfaces are rapidly expanding. Traditional network security protection methods relying on static rules and fixed characteristics are no longer effective in addressing complex and ever-changing threat scenarios. Therefore, threat intelligence-driven security log analysis technology is becoming one of the core and key tools in the power network security protection system. This technology continuously acquires threat intelligence data from external or internal sources (such as attack indicators, attacker profiles, and malicious behavior patterns), and performs multi-dimensional comparison, analysis, and cross-validation with real-time generated security logs, thereby enabling rapid discovery, tracing, and response to unknown threats.
[0003] In existing threat intelligence-based security log analysis methods, the dynamic correlation mechanism between semantics and intelligence is one of the core technologies for improving the intelligence and accuracy of analysis. This mechanism works by semantically parsing and structurally modeling behavioral information recorded in multi-source security logs, enabling dynamic matching at the semantic level with attack characteristics, attack chain steps, and behavioral intent contained in threat intelligence. This achieves "semantic cascading" and "dynamic fusion" between cross-modal data. This mechanism significantly enhances the contextual understanding capability of analysis, allowing the system to not only identify surface-level abnormal behavior but also uncover potential complex threats based on semantic understanding.
[0004] However, most current security log analysis methods based on threat intelligence still suffer from key technical bottlenecks. On the one hand, the dynamic correlation mechanism between semantics and intelligence lacks stability assessment and adaptive adjustment methods for semantic mapping quality when processing multi-source heterogeneous log data. This leads to the system's inability to effectively identify semantic inconsistencies when semantic ambiguity, structural deviations, or mapping mismatches occur during log semantic mapping, affecting the accuracy of subsequent intelligence comparisons. On the other hand, due to the inherent delays and timeliness fluctuations in threat intelligence acquisition, current systems lack mechanisms for analyzing the time-dimensional matching deviations between log events and intelligence information. This makes it difficult to accurately determine whether logs and intelligence are in the same attack period or behavioral stage in a dynamic environment, potentially causing delays in threat identification and failures in coordinated responses.
[0005] Therefore, it is urgent to construct a security log analysis method based on threat intelligence that has the ability to self-check the semantic consistency of multi-source logs and the ability to perceive the temporal matching of intelligence logs, so as to solve the problems of delayed identification, misjudgment and missed judgment of potential threats to power network security caused by the instability of semantic matching and the failure of time synchronization in the above-mentioned dynamic association mechanism. Summary of the Invention
[0006] The purpose of this invention is to solve the problems mentioned above and provide a method and system for security log analysis based on threat intelligence.
[0007] In a first aspect of this invention, a security log analysis method based on threat intelligence is first proposed, the method comprising:
[0008] S1: Obtain security log data from multiple log sources in the target power network, perform semantic parsing, and form a semantic mapping sequence;
[0009] S2: Construct a log semantic inconsistency index based on the semantic mapping sequence to measure the degree of deviation in semantic mapping of multi-source logs;
[0010] S3: Obtain threat intelligence data related to security log data, extract time-series information of log data and related threat intelligence data, and construct an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension;
[0011] S4: Calculate the unavailability index of the dynamic association mechanism by combining the log semantic inconsistency index and the intelligence log time matching failure index, and compare the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, reconstruct the dynamic reconstruction mechanism. Continue until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and output the final judgment result based on the reconstructed dynamic reconstruction mechanism.
[0012] Optionally, the steps of acquiring security log data from multiple log sources in the target power network and performing semantic parsing to form a semantic mapping sequence are as follows:
[0013] Security log data is collected in real time from multiple log sources in the target power network and recorded as the raw log set;
[0014] For each raw log in the raw log set, execute the field extraction function to extract the key fields of the raw log and structure them into a standard form as the standard log;
[0015] Each standard log is mapped to a final semantic vector, and all semantic vectors are used as a sequence of semantic mappings for the logs.
[0016] Optionally, the calculation steps and formula for the log semantic inconsistency index are as follows:
[0017] Based on the semantic mapping sequence of the logs , semantic mapping sequence The final semantic vector for each log source It is represented as a concatenated vector of semantic embeddings containing three semantic dimensions: action verb, operation object, and action result. ; , and These represent the semantic vectors of the action verb, the semantic vector of the object being operated on, and the semantic vector of the action result, respectively. This represents the total number of log sources and is a positive integer.
[0018] For any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix :
[0019] in These represent the action verb, the object of the operation, and the result of the action, respectively.
[0020] Based on the semantic consistency matrix of action verbs Semantic consistency matrix of the operation object Semantic consistency matrix of action verbs Construct a directed graph with semantically conflicting edges. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, that is, when there exists Make When the conflict occurs, construct an edge pointing to the conflict.
[0021] Calculation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Minimum number of cut edges required to cut a subgraph into which all nodes have completely identical internal semantics. ;
[0022] The log semantic inconsistency index is calculated using the following formula: In the formula, This is the log semantic inconsistency index.
[0023] Optionally, the steps for constructing the intelligence log time-matching failure index based on time-series information are as follows:
[0024] Obtain security log data generated in the power grid and form a log time series. , ,in, Indicates the first The timestamp of the log entry Indicates the total number of log entries;
[0025] Simultaneously, acquire external or internal threat intelligence data related to log behavior to form an intelligence time series. ,in, Indicates the first The timestamp of the intelligence message Indicates the total number of intelligence reports;
[0026] Calculate the set of adjacent time differences for both log time series and intelligence time series. and :
[0027] Find the sets respectively and median interval and ;
[0028] Build a time envelope interval for each log. , ;
[0029] For each piece of intelligence, construct a time envelope interval. , ;
[0030] Construct the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ;
[0031] For matrix each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows.
[0032] The coverage quantity is obtained by counting the number of log entries that intersect with the time of any intelligence event. , ; Calculate matching coverage , ;
[0033] Combining the maximum consecutive unmatched length P and the matching coverage The intelligence log time matching failure index is calculated as follows: In the formula, The expiration index is matched with the time of the intelligence log.
[0034] Optionally, the step of jointly calculating the unavailability index of the dynamic association mechanism based on the log semantic inconsistency index and the intelligence log time matching failure index, and comparing the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable is as follows:
[0035] The log semantic inconsistency index and the intelligence log time matching failure index are normalized, and the normalized log semantic inconsistency index and the intelligence log time matching failure index are weighted and summed to obtain the unavailability index of the dynamic association mechanism.
[0036] The unavailability index of the dynamic association mechanism is compared with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available. The final judgment result of the current dynamic association mechanism is then output.
[0037] If the unavailability index is not less than the preset threshold, it means that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
[0038] In a second aspect of this invention, a security log analysis system based on threat intelligence is proposed, the system comprising:
[0039] Sequence module: Acquires security log data from multiple log sources in the target power network, performs semantic parsing, and forms a semantic mapping sequence;
[0040] Semantic mapping deviation module: Constructs a log semantic inconsistency index based on the semantic mapping sequence to measure the degree of semantic mapping deviation among multi-source logs;
[0041] Matching Failure Module: Acquires threat intelligence data related to security log data, extracts time-series information of log data and related threat intelligence data, and constructs an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension;
[0042] The analysis module calculates the unavailability index of the dynamic association mechanism by jointly calculating the log semantic inconsistency index and the intelligence log time matching failure index, and compares the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, the dynamic reconstruction mechanism is reconstructed. This process continues until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and the final analysis result is output based on the reconstructed dynamic reconstruction mechanism.
[0043] Optionally, the sequence module includes:
[0044] Raw log module: Collects security log data in real time from multiple log sources in the target power network, and records it as the raw log set;
[0045] Standard Log Module: Executes a field extraction function on each raw log in the raw log set, extracts the key fields of the raw log, and structures them into a standard form as the standard log;
[0046] Semantic Mapping Sequence Module: Maps each standard log to a final semantic vector, and uses all semantic vectors as a semantic mapping sequence of logs.
[0047] Optionally, the semantic mapping deviation module includes:
[0048] Vector module: Sequences of semantic mapping from logs , semantic mapping sequence The final semantic vector for each log source It is represented as a concatenated vector of semantic embeddings containing three semantic dimensions: action verb, operation object, and action result. ; , and These represent the semantic vectors of the action verb, the semantic vector of the object being operated on, and the semantic vector of the action result, respectively. This represents the total number of log sources and is a positive integer.
[0049] Equivalence determination module: for any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix :
[0050] in These represent the action verb, the object of the operation, and the result of the action, respectively.
[0051] Edge construction module: based on three semantic consistency matrices , , Construct a directed graph with semantically conflicting edges. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, that is, when there exists Make When the conflict occurs, construct an edge pointing to the conflict.
[0052] Minimum number of cut edges module: computation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Minimum number of cut edges required to cut a subgraph into which all nodes have completely identical internal semantics. ;
[0053] Log semantic inconsistency index module: Calculates the log semantic inconsistency index using the following formula: In the formula, This is the log semantic inconsistency index.
[0054] Optionally, the matching failure module includes:
[0055] Log Time Module: Acquires security log data generated in the power network and forms a log time series. , ,in, Indicates the first The timestamp of the log entry Indicates the total number of log entries;
[0056] Intelligence Time Module: Simultaneously acquires external or internal threat intelligence data related to log behavior, forming an intelligence time series. ,in, Indicates the first The timestamp of the intelligence message Indicates the total number of intelligence reports;
[0057] Time Difference Set Module: Calculates adjacent time difference sets for both log time series and intelligence time series. and :
[0058] The median interval module obtains the set respectively. and median interval and ;
[0059] First Time Envelope Interval Module: Constructs a time envelope interval for each log file. , ;
[0060] Second time envelope interval module: Constructs a time envelope interval for each piece of intelligence. , ;
[0061] Matrix construction module: Constructs the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ;
[0062] Unmatched length module: for matrices each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows.
[0063] The coverage module counts the number of log entries that intersect with the time of any intelligence event, thus obtaining the coverage quantity. , ; Calculate matching coverage , ;
[0064] Intelligence log time matching failure index module: combines the maximum consecutive unmatched length P and the matching coverage. The intelligence log time matching failure index is calculated as follows: In the formula, The expiration index is matched with the time of the intelligence log.
[0065] Optionally, the analysis module includes:
[0066] Unavailability Index Module: Normalize the log semantic inconsistency index and the intelligence log time matching failure index, and then sum the normalized log semantic inconsistency index and the intelligence log time matching failure index by weight to obtain the unavailability index of the dynamic association mechanism.
[0067] The first comparison module compares the unavailability index of the dynamic association mechanism with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available, and the final judgment result of the current dynamic association mechanism is output.
[0068] The second comparison module: If the unavailability index is not less than the preset threshold, it means that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
[0069] The beneficial effects of this invention are:
[0070] This invention proposes a security log analysis method and system based on threat intelligence. It acquires security log data from multiple log sources in a target power network and performs semantic parsing to form a semantic mapping sequence. A log semantic inconsistency index is constructed based on the semantic mapping sequence to measure the degree of deviation in the semantic mapping of multi-source logs. Threat intelligence data related to the security log data is acquired, and the time-series information of the log data and related threat intelligence data is extracted. An intelligence log time-matching failure index is constructed based on the time-series information to measure the effectiveness of the matching between intelligence data and log data in the time dimension. An unavailability index of the dynamic association mechanism is jointly calculated based on the log semantic inconsistency index and the intelligence log time-matching failure index. The unavailability index is compared with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If unavailable, the dynamic reconstruction mechanism is reconstructed. This process continues until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and the final analysis result is output based on the reconstructed dynamic reconstruction mechanism. The above methods can be used to assess the dynamic correlation mechanism between semantics and intelligence in existing threat intelligence-based security log analysis, determine whether the semantic mapping quality of multi-source heterogeneous security log data is consistent, and whether the threat intelligence data related to the security log data is effectively matched in the time dimension. Furthermore, the dynamic correlation mechanism between semantics and intelligence in threat intelligence-based security log analysis can be reconstructed to ensure the accuracy and completeness of the security log analysis results output by the dynamic correlation mechanism between semantics and intelligence, and to prevent problems such as delayed identification, misjudgment, and omission of potential threats to power network security information. Attached Figure Description
[0071] The invention will now be further described with reference to the accompanying drawings.
[0072] Figure 1 A flowchart of a security log analysis method based on threat intelligence;
[0073] Figure 2 This is a framework diagram of a security log analysis system based on threat intelligence. Detailed Implementation
[0074] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0075] This invention provides a method for security log analysis based on threat intelligence. See also... Figure 1 , Figure 1A flowchart illustrating a security log analysis method based on threat intelligence, provided as an embodiment of the present invention. The method includes the following steps:
[0076] S1: Obtain security log data from multiple log sources in the target power network, perform semantic parsing, and form a semantic mapping sequence;
[0077] S2: Construct a log semantic inconsistency index based on the semantic mapping sequence to measure the degree of deviation in semantic mapping of multi-source logs;
[0078] S3: Obtain threat intelligence data related to security log data, extract time-series information of log data and related threat intelligence data, and construct an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension;
[0079] S4: Calculate the unavailability index of the dynamic association mechanism by combining the log semantic inconsistency index and the intelligence log time matching failure index, and compare the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, reconstruct the dynamic reconstruction mechanism. Continue until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and output the final judgment result based on the reconstructed dynamic reconstruction mechanism.
[0080] Based on the security log analysis method based on threat intelligence provided by this invention, the dynamic correlation mechanism between semantics and intelligence in existing security log analysis based on threat intelligence can be judged. This method can determine whether the semantic mapping quality of multi-source heterogeneous security log data is consistent, whether the threat intelligence data related to the security log data is effectively matched in the time dimension, and reconstruct the dynamic correlation mechanism between semantics and intelligence in security log analysis based on threat intelligence. This ensures the accuracy and completeness of the security log analysis results output by the dynamic correlation mechanism between semantics and intelligence, and avoids problems such as delayed identification, misjudgment, and omission of potential threats in power network security information.
[0081] It should be noted that in step S1, the security log data from multiple log sources in the target power network is typically collected in real time by various security components deployed at key nodes of the power information system, such as security monitoring equipment, host protection systems, boundary firewalls, network intrusion detection systems (NIDS), intrusion prevention systems (IPS), industrial control security audit platforms, dispatch automation systems, and smart substation equipment. These log sources cover multiple business levels and areas, including power dispatch master stations, secondary equipment, substation monitoring systems, energy management systems (EMS), distribution automation systems (DAS), substation integrated automation platforms, office information networks, and operation and maintenance terminals. Generally, the types of security log data acquired include, but are not limited to: identity authentication logs (such as login / logout information, account lockout records), access control logs (such as resource call traces, instruction execution status), system event logs (such as abnormal restarts, system crashes), network traffic logs (such as port scans, communication connection status), attack alarm logs (such as malicious IP connections, abnormal traffic detection), equipment operation logs (such as SCADA instruction execution anomalies, IED device reconfiguration status), and industrial control instruction operation audit logs, etc. These logs reflect the dynamic changes in user behavior, equipment status, communication processes, and potential threat activities in the power network from multiple perspectives, and serve as the basic data support for subsequent semantic parsing and threat intelligence correlation analysis.
[0082] In one embodiment, the steps of acquiring security log data from multiple log sources in the target power network and performing semantic parsing to form a semantic mapping sequence are as follows:
[0083] Security log data is collected in real time from multiple log sources (such as firewalls, hosts, control systems, dispatching platforms, etc.) in the target power network and recorded as the original log set. , , Indicates the first One original log entry, This indicates the total number of log sources. The original logs may be unstructured (such as syslog format, plain text lines) or semi-structured (such as JSON, XML, CSV, etc.). It should be noted that the original logs may have problems such as inconsistent field naming, redundant information, lengthy content, or semantic ambiguity, making it difficult to achieve efficient semantic analysis by direct judgment.
[0084] For the original log set Each raw log Execute field extraction function Extract the original logs Key fields are structured into a standard format for use in standard logs. , In the formula, The first The original log entries include the device identifier, action verbs (e.g., "login", "send command"), operation object (e.g., "device X", "port Y"), and action result (e.g., "success", "failure"). This step standardizes the log format, facilitating semantic modeling in the next step. Field extraction functions are also included. Key fields are typically extracted based on regular expression matching rules, log template recognition algorithms (such as Drain and Spell), or natural language parsing methods (such as dependency parsing) to form a data format with a unified structure.
[0085] Each standard log Mapped to the final semantic vector , indicating the first The meaning of each log entry in the semantic space, defining the semantic mapping function. The details are as follows: In the formula, This is a semantic mapping function that maps fields to vector form. Indicates terms Semantic embedding vectors (which can be obtained through word vectors, entity embeddings, rule mappings, etc.): , indicating the first The final semantic vector of each log. Dimensionality (e.g., d=3 represents a 3D representation of behavior + object + result); semantic mapping function This is typically achieved using word embedding models (such as Word2Vec, GloVe, BERT) or knowledge graph entity encoding (such as TransE, ComplEx), i.e., for behavioral fields. Mapped to vector form;
[0086] All semantic vectors As a semantic mapping sequence of logs , .
[0087] It should be noted that the above steps are illustrated with an example, such as a real-world scenario of a power dispatch center. Assume its network security system obtains security log data from three sources: firewall logs, protected host logs, and dispatch control system logs. Firewall logs are unstructured text, such as "2025-07-01 10:21:32 blocked IP 192.168.0.1 from accessing port 8080"; host logs are in JSON structure, such as {"timestamp":"2025-07-01T10:21:35","event":"login","user":"admin","result":"fail"}; and dispatch system logs are in XML format, such as... <log> <time>2025-07-01T10:21:40< / time> <action> sendCommand< / action> <device> UnitX< / device> <result> success< / result> < / log> Because the log formats differ and the meanings of fields are inconsistent, a field extraction function is required. Standard fields are extracted and standardized into structured quintuples: timestamp (e.g., 2025-07-01 10:21:32), device identifier (e.g., Firewall / Host / UnitX), action verb (e.g., "block", "login", "send command"), operation object (e.g., IP address, username, device name), and action result (e.g., "success", "failure"). Then, a semantic mapping function is used to convert fields such as "login", "send command", and "failure" into corresponding semantic vectors. For example, using the Word2Vec model, "send command" is mapped to [0.45, 0.12, -0.33], and "failure" is mapped to [-0.25, 0.91, 0.60]. Each log entry is ultimately represented as a multi-dimensional semantic vector Si, and a semantic mapping sequence is formed based on the final semantic vectors of all log sources for subsequent threat behavior chain matching and dynamic analysis. This process not only achieves format standardization and semantic unification but also provides a solid data foundation for semantic association between cross-source logs.
[0088] It should be noted that the semantic mapping sequence obtained above The system retains the occurrence of events in the logs in real time and encodes the evolution of behavioral chains at the semantic level, providing a unified data foundation for dynamic semantic association mechanisms. In subsequent analysis, this sequence can be used for intelligent analysis scenarios such as behavioral chain identification, context-based attack tracing, and attack chain detection. The entire process ensures a clear structure, semantic accuracy, and coherent processing from the raw logs to the semantic space representation.
[0089] In one embodiment, a log semantic inconsistency index is constructed based on the semantic mapping sequence to measure the degree of deviation in the semantic mapping of multi-source logs;
[0090] In one implementation, the calculation steps and formula for the log semantic inconsistency index are as follows:
[0091] Based on the semantic mapping sequence of the logs , semantic mapping sequence The final semantic vector for each log source It is represented as a concatenated vector of semantic embeddings containing three semantic dimensions: action verb, operation object, and action result. ; , and These represent the semantic vectors of the action verb, the semantic vector of the object being operated on, and the semantic vector of the action result, respectively. This represents the total number of log sources and is a positive integer.
[0092] For any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix :
[0093] in These represent the action verb, the object of the operation, and the result of the action, respectively.
[0094] Based on the semantic consistency matrix of action verbs Semantic consistency matrix of the operation object Semantic consistency matrix of action verbs These three semantic consistency matrices are used to construct a directed graph with semantically conflicting edges. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, that is, when there exists Make When the conflict occurs, construct an edge pointing to the conflict.
[0095] Calculation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Cut into subgraphs where all nodes have completely identical internal semantics (i.e., any two nodes within a subgraph have completely identical internal semantics). of Minimum number of cut edges required ;
[0096] The log semantic inconsistency index is calculated using the following formula: In the formula, This is the log semantic inconsistency index.
[0097] It's important to note that in semantic embedding space, "structural equivalence" refers to two vectors not only being numerically similar but also highly consistent in the semantic relationships and contextual structures they represent. Specifically, structural equivalence means that two semantic vectors have similar distributed expressions in the semantic features of the action verbs, objects of operation, or results of actions, reflecting their homogeneity or equivalence in actual semantic meaning and function. For example, in the security logs of a power network, if the semantic embedding vectors of the "login successful" actions described in different log sources are structurally equivalent, it indicates that regardless of the differences in the log's expression or details, they all point to the same semantic event, thus being considered semantically consistent. The determination of structural equivalence considers not only the distance between vectors (such as Euclidean distance or cosine similarity) but also the internal relationships of the semantic context and the association structure in the semantic graph, ensuring that the judgment reflects a true semantic match rather than simple numerical similarity, thereby improving the accuracy and stability of multi-source log semantic mapping. A specific method for "structural equivalence" is the discrete semantic center projection re-method: that is, each... The semantic clustering center of the field is mapped (through clustering or rule mapping). By comparing the projection relationship between the corresponding semantic vectors of the field and the main cluster center in different log sources, it is determined whether they belong to the same semantic category. If the projection distance between two semantic vectors on the same semantic main cluster center is small enough, they are considered structurally equivalent, indicating that the semantic expression of the two logs in this field is highly consistent. This method effectively alleviates the semantic offset problem caused by differences in expression, details, or noise in multi-source logs, ensuring stable alignment of semantic mapping across log sources. It is particularly suitable for unified semantic parsing and consistency verification of logs from heterogeneous devices in power networks.
[0098] It's important to note that the Log Semantic Inconsistency Index (LSI) quantifies and reflects the degree of semantic discrepancies and conflicts in security logs from multiple sources. Specifically, this index measures the inconsistency and mismatch between core semantic dimensions such as action verbs, operation objects, and action results when describing the same security event from different log sources. It is used to assess the semantic fusion quality of multi-source log data and is a crucial indicator for judging whether the semantic expression of log data is unified and accurate. A higher SSI indicates significant semantic discrepancies and confusion in the security logs recorded by different security devices (such as firewalls, hosts, control systems, and dispatch platforms) in the power network. These discrepancies may be caused by differences in log formats, non-standard field naming, diverse description methods, or poor semantic mapping algorithms. In such cases, during threat intelligence-based security log analysis, the dynamic association mechanism between semantics and intelligence cannot effectively align the semantic information of different log sources, leading to a significant decrease in the usability and accuracy of the dynamic association mechanism. This semantic inconsistency can result in the fragmentation of semantic information for important security events, making it difficult for the system to correctly identify continuous behaviors in the attack chain and affecting the overall judgment of complex attacks. For example, if an attack targeting a substation control system is described as "abnormal login failure" in one log source and "remote access denied" in another, the system may misjudge these behaviors as unrelated isolated events if the semantics are not correctly mapped and unified, thus missing real intrusion threats. Conversely, some unrelated events may be misjudged as attacks due to semantic mismatch, generating a large number of false alarms and causing great trouble for maintenance personnel. Therefore, when the log semantic inconsistency index is high, the probability of the dynamic correlation mechanism between semantics and intelligence in security log analysis based on threat intelligence becoming unusable increases significantly. If this dynamic correlation mechanism is still used for analysis, it may lead to the missed detection of serious security threats, rampant false alarms, and delays and failures in security response, ultimately affecting the overall security situation awareness and risk prevention capabilities of the power network, and even triggering more widespread system failures and power supply interruptions, causing serious social and economic losses.
[0099] It should be noted that calculating the log semantic inconsistency index using the above method can accurately quantify the semantic discrepancies and conflicts between multi-source logs, avoiding misjudgments or omissions caused by simple similarity calculations. This method not only considers the multidimensional semantic relationships of action verbs, operation objects, and action results in the logs, but also identifies and isolates semantic conflict areas through graph cut algorithms, thus more comprehensively reflecting the overall consistency of semantic mapping. This fine-grained and structured analysis approach helps to promptly detect semantic anomalies in log data, improving the accuracy and robustness of the judgment system in fusion of heterogeneous logs. The log semantic inconsistency index calculated in this way can provide a scientific basis for judging the availability of the dynamic association mechanism between semantics and intelligence in security log judgment based on threat intelligence, ensuring that the mechanism is adjusted or reconstructed when semantic conflicts are significant, avoiding false alarms or missed alarms caused by erroneous associations, and greatly improving the intelligence level and early warning reliability of power network security protection.
[0100] In one embodiment, S3: acquire threat intelligence data related to security log data, extract time-series information of log data and related threat intelligence data, and construct an intelligence log time matching failure index based on the time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension.
[0101] In one implementation, the steps of acquiring threat intelligence data related to security log data, extracting time-series information from the log data and related threat intelligence data, and constructing an intelligence log time-matching failure index based on the time-series information are as follows:
[0102] Obtain security log data generated in the power grid and form a log time series. , ,in, Indicates the first The timestamp of the log entry Indicates the total number of log entries;
[0103] Simultaneously, acquire external or internal threat intelligence data related to log behavior to form an intelligence time series. ,in, Indicates the first The timestamp of the intelligence message Indicates the total number of intelligence reports;
[0104] Calculate the set of adjacent time differences for both log time series and intelligence time series. and :
[0105] ;
[0106] ;
[0107] Find the sets respectively and median interval and , , ;
[0108] Build a time envelope interval for each log. , ;
[0109] For each piece of intelligence, construct a time envelope interval. , ;
[0110] Construct the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ;
[0111] For matrix each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows.
[0112] The coverage quantity is obtained by counting the number of log entries that intersect with the time of any intelligence event. , ; Calculate matching coverage , ;
[0113] Combining the maximum consecutive unmatched length P and the matching coverage The intelligence log time matching failure index is calculated as follows: In the formula, The expiration index is matched with the time of the intelligence log.
[0114] It should be noted that in the process of constructing the intelligence log time-matching failure index described above, the acquisition of log data and threat intelligence data is typically achieved through an automated collection platform. Specifically, security log data generated in the power network can be collected from multiple log sources through an integrated Security Information and Event Management System (SIEM) or log management platform. These log sources include firewalls, host operating systems, secure endpoint agents, dispatch control systems, intrusion detection systems (IDS), and authentication services. For example, logs of "denying external IP access to the control system" are collected from the firewall equipment in the main control center, recording the precise event timestamp and source target address; while the dispatch control platform may generate operation logs of "illegal command attempts." After cleaning and extraction, these logs form a time-ordered log time series. Threat intelligence data can be obtained by accessing dedicated intelligence sharing platforms in the power industry (such as power grid intelligence exchange centers, industry threat analysis platforms, CTI services, etc.), or by combining them with the organization's self-built threat intelligence collection mechanisms, such as buzzer feedback, honeypot system monitoring results, or APT group activity fingerprints. For example, if an intelligence platform releases information at 07:35 stating that "an attack group has attempted to attack the SCADA control channel in the South China region," with a clear timestamp, behavioral tag, and target characteristics, the system can record this data as a node in the intelligence time series. By extracting the time fields from the logs and intelligence separately and unifying the time format (e.g., UTC time), their respective time series can be constructed, providing a foundation for subsequent matching analysis. This method features strong real-time performance, diverse sources, and high time accuracy, effectively supporting the construction of a highly reliable time matching failure index.
[0115] It's important to note that the intelligence log time-match failure index is a key indicator used to measure the effectiveness of the temporal correspondence between security logs and threat intelligence data in power networks. Essentially, it reflects whether there is a reasonable temporal linkage between logs and intelligence. A smaller index value indicates good temporal synchronization and overlap between logs and intelligence data, high matching coverage, and short consecutive mismatch intervals, suggesting an effective temporal correlation that facilitates the construction of dynamic and reliable semantic association paths. Conversely, a larger intelligence log time-match failure index indicates a low temporal overlap between logs and intelligence, severe gaps in matching intervals, and long periods of log activity without effective intelligence coverage. This often means that current intelligence data cannot timely cover ongoing suspicious behavior, or that the behavior recorded in the logs and the attack chain described by the intelligence cannot be closed temporally. This makes it impossible for dynamic analysis mechanisms based on "temporal semantic coupling" to continuously track attack phases or identify latent behavioral chains. This situation is particularly dangerous in power networks. For example, the dispatch host might continuously record multiple device configuration change operations at 00:13, but the currently accessed threat intelligence only covers the activity trajectory of a certain attack organization before 00:10. This lack of time support during semantic matching causes the system to mistakenly classify these log behaviors as "low-risk operations," thus missing the identification of real threat behaviors in the later stages of an attack. This could lead to serious consequences such as dispatch configuration tampering and abnormal power grid operation. Therefore, when this failure index rises, it often indicates that the dynamic correlation mechanism between semantics and intelligence is facing the risk of "disconnection." If the system continues to rely on the original mechanism for log correlation analysis, it may cause problems such as intelligence mismatch, broken threat chains, and misleading analysis, ultimately reducing the overall situational awareness capability and affecting the timeliness and accuracy of power network security response and coordinated defense.
[0116] It's important to note that the core advantage of calculating the intelligence log time-matching failure index using the above method lies in its comprehensive measurement of the continuity and coverage integrity of threat intelligence and security logs across the time dimension. It doesn't just focus on single-point time overlap, but constructs a complete assessment system for time-series correlation from multiple dimensions, including time interval patterns, median envelope synchronization, continuous unmatched areas, and overall coverage. This approach avoids the misjudgments or omissions caused by traditional simple window overlap judgments, making it particularly suitable for complex scenarios in power networks where periodic scheduling behaviors and sudden attacks coexist. The greatest advantage of analyzing based on this index is its ability to objectively determine whether the current dynamic semantic and intelligence correlation mechanism still possesses a temporally reasonable basis. A low failure index indicates that semantic inference has time support and the judgment path is reliable; a high failure index can trigger timely mechanism reconstruction, avoiding judgment failures or security blind spots caused by "time misalignment" between intelligence and logs. This enables the dynamic correlation mechanism to achieve self-diagnosis and flexible adjustment capabilities, ensuring that the security situation awareness of the power network remains effective, synchronized, and controllable.
[0117] In one embodiment, S4: The unavailability index of the dynamic association mechanism is jointly calculated based on the log semantic inconsistency index and the intelligence log time matching failure index, and the unavailability index is compared with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, the dynamic reconstruction mechanism is triggered.
[0118] In one implementation, the steps for jointly calculating the unavailability index of the dynamic association mechanism based on the log semantic inconsistency index and the intelligence log time matching failure index are as follows:
[0119] The log semantic inconsistency index and the intelligence log time matching failure index are normalized, and then the weighted sum of the normalized log semantic inconsistency index and the intelligence log time matching failure index is used to obtain the unavailability index of the dynamic association mechanism; the calculation formula is as follows: In the formula This is an index indicating the unavailability of the dynamic association mechanism. and These are the log semantic inconsistency index after normalization and the intelligence log time matching failure index, respectively. These represent the preset deweighting coefficients for the log semantic inconsistency index after multi-normalization processing and the intelligence log time matching failure index, respectively. All are greater than 0;
[0120] It should be noted that the above formulas are all dimensionless calculations. Commonly used methods for removing dimensions include Min-Max normalization and Z-Score standardization, which will not be elaborated here. Settings should be set according to the actual situation, generally They are equal and their sum is 1, for example, It can be 0.5 or 0.5.
[0121] In one embodiment, the step of comparing the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable, and triggering the dynamic reconstruction mechanism if it is unavailable, is as follows:
[0122] The unavailability index of the dynamic association mechanism is compared with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available. The final judgment result of the current dynamic association mechanism is then output.
[0123] If the unavailability index is not less than the preset threshold, it means that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
[0124] It is important to note that during the security log analysis process based on threat intelligence, to ensure the effectiveness and stability of the dynamic correlation mechanism between semantics and intelligence, its availability status needs to be assessed in real time. Specifically, the unavailability index of the dynamic correlation mechanism, calculated above, is compared with a system-preset threshold. When the unavailability index is less than the threshold, it indicates good consistency between the current semantic correlation and temporal matching, with log semantic fusion and intelligence time coverage within acceptable ranges. The dynamic correlation mechanism still possesses good reasoning accuracy and context awareness. In this case, the threat analysis result under the current dynamic correlation mechanism is directly output. For example, confirming that a certain IP address is highly correlated with a specific intrusion behavior and associated with an APT attack intelligence, the system generates an alert accordingly. However, when the unavailability index is greater than or equal to the threshold, it means there is significant semantic conflict between multiple source logs, or a large time drift between intelligence and logs. The semantic reasoning under the current mechanism has lost consistency and synchronization, which is highly likely to lead to biased analysis results or incorrect correlations. For example, an attack behavior log might actually be associated with intelligence from last week, but due to time mismatch, it is incorrectly included in the current attack chain, thus misleading subsequent responses. At this point, the system will immediately trigger a dynamic reconstruction mechanism, re-execute semantic mapping correction, time window adjustment, and matching strategy update to ensure that the dynamic mechanism returns to a reasonable state and the analysis path converges again to the actual attack context. This enables the mechanism to have self-healing capabilities and improves the adaptability of power networks in intelligent protection scenarios facing complex threats.
[0125] In one embodiment, when the dynamic association mechanism is determined to be unavailable, a dynamic reconstruction mechanism will be triggered to correct the judgment failure caused by semantic bias or temporal drift. The reconstruction process mainly includes two core components: semantic mapping correction and time window reconfiguration, which work together to reconstruct the semantic and temporal association path between logs and threat intelligence.
[0126] First, during the semantic mapping correction phase, the semantic mapping process is traced back to analyze fields in the log source that exhibit semantic conflicts or ambiguities. By clustering log vectors with semantic anomalies and comparing them with preset semantic master class centers, key semantic fragments that may cause inconsistencies are identified (e.g., the action verb "upload" may represent data submission on some devices, but remote command push on others). Subsequently, by enabling supplementary semantic ontology libraries, device context rules, or context association models based on graph neural networks, the semantic vectors of this portion of the logs are regenerated, making them structurally closer to the semantic master class consistency standard.
[0127] Secondly, during the time window reconfiguration phase, volatility analysis is performed on the intelligence time series and log time series. Based on historical statistical information (such as anomaly frequency, attack duration, and communication cycle), the upper and lower limits of the time envelope are adaptively adjusted, dynamically widening or narrowing the boundary for determining time intersection. For example, if a certain type of malicious behavior exhibits short-term, high-frequency outbreaks, the time envelope is narrowed to a finer granularity to avoid misjudgments; while for APT intelligence with a long incubation period, the time window is expanded to ensure matching coverage. This adjustment allows more effective log behaviors to overlap with intelligence information, improving the time synchronization of the dynamic mechanism.
[0128] The entire reconstruction process is an iterative optimization process. After each reconstruction, the updated unavailability index is recalculated. If the unavailability index is less than a preset threshold, it indicates that the repair is effective, and the current semantic-intelligence association path is restored to a usable state. At this point, based on the repaired dynamic path, joint reasoning and matching of logs and intelligence are performed again to output the final judgment result, such as locating the attack path, identifying the attacker's TTP (tactics, techniques, and processes), generating threat alerts, or triggering adaptive response strategies. If the unavailability index is not less than the preset threshold, iterative reconstruction continues until the unavailability index meets the usability threshold, ensuring that the final output result is based on a robust and highly consistent data foundation. This dynamic reconstruction mechanism greatly enhances the resilience and intelligence level of power networks in the face of complex environments such as fuzzy semantics, multi-source data drift, and evolving attack methods.
[0129] Based on the same inventive concept, this invention also provides a security log analysis system based on threat intelligence. See also Figure 2 , Figure 2 A framework diagram of a security log analysis system based on threat intelligence provided in this embodiment of the invention. The system includes:
[0130] Sequence module: Acquires security log data from multiple log sources in the target power network, performs semantic parsing, and forms a semantic mapping sequence;
[0131] Semantic mapping deviation module: Constructs a log semantic inconsistency index based on the semantic mapping sequence to measure the degree of semantic mapping deviation among multi-source logs;
[0132] Matching Failure Module: Acquires threat intelligence data related to security log data, extracts time-series information of log data and related threat intelligence data, and constructs an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension;
[0133] The analysis module calculates the unavailability index of the dynamic association mechanism by jointly calculating the log semantic inconsistency index and the intelligence log time matching failure index, and compares the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, the dynamic reconstruction mechanism is reconstructed. This process continues until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and the final analysis result is output based on the reconstructed dynamic reconstruction mechanism.
[0134] Based on the security log analysis system based on threat intelligence provided in this embodiment of the invention, the existing dynamic correlation mechanism between semantics and intelligence in security log analysis based on threat intelligence can be judged in the above manner. It can determine whether the semantic mapping quality of multi-source heterogeneous security log data is consistent, whether the threat intelligence data related to security log data is effective in the time dimension matching, and reconstruct the dynamic correlation mechanism between semantics and intelligence in security log analysis based on threat intelligence. This ensures the accuracy and completeness of the security log analysis results output by the dynamic correlation mechanism between semantics and intelligence, and avoids problems such as delayed identification, misjudgment, and omission of potential threats in power network security information.
[0135] The sequence module includes:
[0136] Raw log module: Collects security log data in real time from multiple log sources in the target power network, and records it as the raw log set;
[0137] Standard Log Module: Executes a field extraction function on each raw log in the raw log set, extracts the key fields of the raw log, and structures them into a standard form as the standard log;
[0138] Semantic Mapping Sequence Module: Maps each standard log to a final semantic vector, and uses all semantic vectors as a semantic mapping sequence of logs.
[0139] In one embodiment, the semantic mapping deviation module includes:
[0140] Vector module: Sequences of semantic mapping from logs , semantic mapping sequence The final semantic vector for each log source It is represented as a concatenated vector of semantic embeddings containing three semantic dimensions: action verb, operation object, and action result. ; , and These represent the semantic vectors of the action verb, the semantic vector of the object being operated on, and the semantic vector of the action result, respectively. This represents the total number of log sources and is a positive integer.
[0141] Equivalence determination module: for any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix :
[0142] in These represent the action verb, the object of the operation, and the result of the action, respectively.
[0143] Edge construction module: based on three semantic consistency matrices , , Construct a directed graph with semantically conflicting edges. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, that is, when there exists Make When the conflict occurs, construct an edge pointing to the conflict.
[0144] Minimum number of cut edges module: computation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Minimum number of cut edges required to cut a subgraph into which all nodes have completely identical internal semantics. ;
[0145] Log semantic inconsistency index module: Calculates the log semantic inconsistency index using the following formula: In the formula, This is the log semantic inconsistency index.
[0146] In one embodiment, the matching failure module includes:
[0147] Log Time Module: Acquires security log data generated in the power network and forms a log time series. , ,in, Indicates the first The timestamp of the log entry Indicates the total number of log entries;
[0148] Intelligence Time Module: Simultaneously acquires external or internal threat intelligence data related to log behavior, forming an intelligence time series. ,in, Indicates the first The timestamp of the intelligence message Indicates the total number of intelligence reports;
[0149] Time Difference Set Module: Calculates adjacent time difference sets for both log time series and intelligence time series. and :
[0150] The median interval module obtains the set respectively. and median interval and ;
[0151] First Time Envelope Interval Module: Constructs a time envelope interval for each log file. , ;
[0152] Second time envelope interval module: Constructs a time envelope interval for each piece of intelligence. , ;
[0153] Matrix construction module: Constructs the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ;
[0154] Unmatched length module: for matrices each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows.
[0155] The coverage module counts the number of log entries that intersect with the time of any intelligence event, thus obtaining the coverage quantity. , ; Calculate matching coverage , ;
[0156] Intelligence log time matching failure index module: combines the maximum consecutive unmatched length P and the matching coverage. The intelligence log time matching failure index is calculated as follows: In the formula, The expiration index is matched with the time of the intelligence log.
[0157] In one embodiment, the analysis module includes:
[0158] Unavailability Index Module: Normalize the log semantic inconsistency index and the intelligence log time matching failure index, and then sum the normalized log semantic inconsistency index and the intelligence log time matching failure index by weight to obtain the unavailability index of the dynamic association mechanism.
[0159] The first comparison module compares the unavailability index of the dynamic association mechanism with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available, and the final judgment result of the current dynamic association mechanism is output.
[0160] The second comparison module: If the unavailability index is not less than a preset threshold, it indicates that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
[0161] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A security log analysis method based on threat intelligence, characterized in that, Includes the following steps: Acquire security log data from multiple log sources in the target power network, perform semantic parsing, and form a semantic mapping sequence; A log semantic inconsistency index is constructed based on the semantic mapping sequence to measure the degree of deviation in semantic mapping among multi-source logs; Acquire threat intelligence data related to security log data, extract time-series information of log data and related threat intelligence data, and construct an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension; The unavailability index of the dynamic association mechanism is calculated by jointly using the log semantic inconsistency index and the intelligence log time matching failure index. The unavailability index is then compared with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, the dynamic reconstruction mechanism is reconstructed. This process continues until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold. Finally, the analysis result is output based on the reconstructed dynamic reconstruction mechanism.
2. The security log analysis method based on threat intelligence according to claim 1, characterized in that, The steps for acquiring security log data from multiple log sources in the target power network and performing semantic parsing to form a semantic mapping sequence are as follows: Security log data is collected in real time from multiple log sources in the target power network and recorded as the raw log set; For each raw log in the raw log set, execute the field extraction function to extract the key fields of the raw log and structure them into a standard form as the standard log; Each standard log is mapped to a final semantic vector, and all semantic vectors are used as a sequence of semantic mappings for the logs.
3. The security log analysis method based on threat intelligence according to claim 1, characterized in that, The calculation steps and formula for the log semantic inconsistency index are as follows: The final semantic vector of each log source in the semantic mapping sequence is represented as a concatenated vector containing semantic embeddings of three semantic dimensions: action verb, operation object, and action result. These are the semantic vectors of the action verb. semantic vectors of the objects being operated on Semantic vector of behavioral outcome ; For any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix : ; in These represent the action verb, the object of the operation, and the result of the action, respectively. Construct a directed graph with semantic conflict edges based on the semantic consistency matrix of the action verb, the semantic consistency matrix of the operation object, and the semantic consistency matrix of the action verb. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, i.e. When the conflict occurs, construct an edge pointing to the conflict. Calculation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Minimum number of cut edges required to cut a subgraph into which all nodes have completely identical internal semantics. ; The log semantic inconsistency index is calculated using the following formula: In the formula, This is the log semantic inconsistency index.
4. The security log analysis method based on threat intelligence according to claim 1, characterized in that, The steps for constructing an intelligence log time-matching failure index based on time-series information are as follows: Acquire security log data generated in the power network and form a log time series; Simultaneously, acquire external or internal threat intelligence data related to log behavior to form an intelligence time series; Calculate the set of adjacent time differences for both log time series and intelligence time series. and : Find the sets respectively and median interval and ; Build a time envelope interval for each log. , ; Indicates the first The timestamp of each log entry; For each piece of intelligence, construct a time envelope interval. , ; Indicates the first The timestamp of the intelligence report; Construct the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ; For matrix each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows. The coverage quantity is obtained by counting the number of log entries that intersect with the time of any intelligence event. , ; Calculate matching coverage , ; Indicates the total number of log entries; Combining the maximum consecutive unmatched length P and the matching coverage The intelligence log time matching failure index is calculated as follows: In the formula, The failure index is matched with the time of the intelligence log.
5. The security log analysis method based on threat intelligence according to claim 1, characterized in that, The steps for determining whether the dynamic association mechanism is currently unusable are as follows: First, the unusability index of the dynamic association mechanism is calculated by jointly using the log semantic inconsistency index and the intelligence log time matching failure index. Then, the unusability index is compared with a preset threshold. The log semantic inconsistency index and the intelligence log time matching failure index are normalized, and the normalized log semantic inconsistency index and the intelligence log time matching failure index are weighted and summed to obtain the unavailability index of the dynamic association mechanism. The unavailability index of the dynamic association mechanism is compared with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available. The final judgment result of the current dynamic association mechanism is then output. If the unavailability index is not less than the preset threshold, it means that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
6. A security log analysis system based on threat intelligence, characterized in that, The system includes: Sequence module: Acquires security log data from multiple log sources in the target power network, performs semantic parsing, and forms a semantic mapping sequence; Semantic mapping deviation module: Constructs a log semantic inconsistency index based on the semantic mapping sequence to measure the degree of semantic mapping deviation among multi-source logs; Matching Failure Module: Acquires threat intelligence data related to security log data, extracts time-series information of log data and related threat intelligence data, and constructs an intelligence log time matching failure index based on time-series information to measure the matching effectiveness of intelligence data and log data in the time dimension; The analysis module calculates the unavailability index of the dynamic association mechanism by jointly calculating the log semantic inconsistency index and the intelligence log time matching failure index, and compares the unavailability index with a preset threshold to determine whether the current dynamic association mechanism is unavailable. If it is unavailable, the dynamic reconstruction mechanism is reconstructed. This process continues until the unavailability index of the reconstructed dynamic reconstruction mechanism is less than the threshold, and the final analysis result is output based on the reconstructed dynamic reconstruction mechanism.
7. A security log analysis system based on threat intelligence according to claim 6, characterized in that, The sequence module includes: Raw log module: Collects security log data in real time from multiple log sources in the target power network, and records it as the raw log set; Standard Log Module: Executes a field extraction function on each raw log in the raw log set, extracts the key fields of the raw log, and structures them into a standard form as the standard log; Semantic Mapping Sequence Module: Maps each standard log to a final semantic vector, and uses all semantic vectors as a semantic mapping sequence of logs.
8. A security log analysis system based on threat intelligence according to claim 6, characterized in that, The semantic mapping deviation module includes: Vector module: Represents the final semantic vector of each log source in the semantic mapping sequence as a concatenated vector containing semantic embeddings of three semantic dimensions: action verb, operation object, and action result. Semantic vectors of action verbs semantic vectors of the objects being operated on Semantic vector of behavioral outcome , Equivalence determination module: for any two log sources Construct a cross-source semantic structure equivalence decision function for the semantic vector of the action verb, the object of operation, and the result of the action. Calculate the corresponding semantic consistency matrix : ; in These represent the action verb, the object of the operation, and the result of the action, respectively. Edge construction module: Constructs a directed graph with semantically conflicting edges based on the semantic consistency matrix of the action verb, the semantic consistency matrix of the operation object, and the semantic consistency matrix of the action verb. , where vertex set Corresponding log source number, edge set It consists of any pair of log source nodes that are inconsistent in any dimension, i.e. When the conflict occurs, construct an edge pointing to the conflict. Minimum number of cut edges module: computation graph Total number of conflict edges And calculate the graph using the graph cut approximation algorithm. Minimum number of cut edges required to cut a subgraph into which all nodes have completely identical internal semantics. ; Log semantic inconsistency index module: Calculates the log semantic inconsistency index using the following formula: In the formula, This is the log semantic inconsistency index.
9. A security log analysis system based on threat intelligence according to claim 6, characterized in that, The matching failure module includes: Log time module: Acquires security log data generated in the power network and forms a log time series; Intelligence Time Module: Simultaneously acquires external or internal threat intelligence data related to log behavior, forming an intelligence time series; Time Difference Set Module: Calculates adjacent time difference sets for both log time series and intelligence time series. and : The median interval module obtains the set respectively. and median interval and ; First Time Envelope Interval Module: Constructs a time envelope interval for each log file. , ; Indicates the first The timestamp of each log entry; Second time envelope interval module: Constructs a time envelope interval for each piece of intelligence. , ; Indicates the first The timestamp of the intelligence report; Matrix construction module: Constructs the time intersection matrix Determine any log With intelligence Determine if there is time overlap, and define the overlap value. , ; Unmatched length module: for matrices each line Calculate the first The maximum length of consecutive zeros in a line is denoted as . This yields the maximum consecutive unmatched length P across all rows. The coverage module counts the number of log entries that intersect with the time of any intelligence event, thus obtaining the coverage quantity. , ; Calculate matching coverage , ; Indicates the total number of log entries; Intelligence log time matching failure index module: combines the maximum consecutive unmatched length P and the matching coverage. The intelligence log time matching failure index is calculated as follows: In the formula, The failure index is matched with the time of the intelligence log.
10. A security log analysis system based on threat intelligence according to claim 6, characterized in that, The analysis module includes: Unavailability Index Module: Normalize the log semantic inconsistency index and the intelligence log time matching failure index, and then sum the normalized log semantic inconsistency index and the intelligence log time matching failure index by weight to obtain the unavailability index of the dynamic association mechanism. The first comparison module compares the unavailability index of the dynamic association mechanism with a preset threshold. If the unavailability index is less than the preset threshold, it means that the current dynamic association mechanism is available, and the final judgment result of the current dynamic association mechanism is output. The second comparison module: If the unavailability index is not less than the preset threshold, it means that the current dynamic association mechanism is unavailable, and the dynamic reconstruction mechanism will be reconstructed immediately.
Citation Information
Cited By
Semantic fusion and cross-view self-supervision-based log intention recognition method and system
CN122113003A