Log data processing method and device, equipment, medium and product

By identifying, mapping, segmenting, and completing missing features in security log data from multiple vendors, and combining word vector algorithms and AI technology, the analysis challenges caused by differences in the formats of log data from multiple vendors have been solved. This has achieved standardization and quality improvement of security log data, and improved the accuracy of threat identification and the efficiency of analysis.

CN120973624APending Publication Date: 2025-11-18CHINA RESOURCES POWER TECH RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511074322.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Security logs from different vendors have varying formats, different logical content, and different interpretations of security events, making log data analysis difficult. Traditional methods of interfacing with security systems rely on manual parsing, which is costly and inefficient, resulting in low-quality security data and affecting the accuracy of threat characterization.

Method used

By acquiring log data to be processed, identifying key fields and mapping them to the target log format, performing word segmentation and missing feature completion, and using word vector algorithms and AI technology for data standardization and fusion noise reduction, the standardization and quality improvement of security log data are achieved.

Benefits of technology

It has achieved standardization of security log data, improved the efficiency of data quality assessment, reduced labor costs, and improved the accuracy of threat identification and data parsing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973624A_ABST
    Figure CN120973624A_ABST
Patent Text Reader

Abstract

The invention discloses a log data processing method and device, equipment, a medium and a product. The method comprises the steps of obtaining to-be-processed log data, performing key field identification on the to-be-processed log data, and mapping the obtained key field to a target log format to obtain log data in the target log format; performing word segmentation processing on the log data in the target log format, converting an obtained word into a word vector represented by a vector, and performing missing feature completion processing on a missing word in the word vector to obtain a log in the target format; and performing fusion noise reduction processing on the similar target format logs to obtain target log data. According to the technical scheme, standardization of the security log data can be achieved, and the data quality evaluation efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of network security and information security, and particularly relate to a log data processing method, device, equipment, medium and product. BACKGROUND

[0002] The log data formats of multiple vendors are different, the data content logics are different, the security events are understood differently by each security vendor, the development modes are different, and meanwhile, there is a lack of unified standard in the industry, which leads to an increased difficulty in analyzing the log data.

[0003] The traditional connection adopts a manual analysis mode, and the connection work is complex and high in cost. The traditional connection relies on manual writing of analysis rules to convert log fields. According to research, 5-10 people / day of customization workload is needed to customize the analysis and mapping of key fields of a security device log in the industry, and alarm repetition can only be found after analysis, and there is a problem of a large increase in workload.

[0004] In addition, the security data quality is not high, and after the log formats of various security devices in the present network are initially unified, key features may be missing, which seriously affects the accuracy of subsequent threat qualification. SUMMARY

[0005] Embodiments of the present application provide a log data processing method, device, equipment, medium and product to realize standardization of security log data and improve data quality evaluation efficiency.

[0006] According to an aspect of the present application, a log data processing method is provided, comprising:

[0007] obtaining to-be-processed log data, identifying key fields of the to-be-processed log data, mapping the obtained key fields to a target log format, and obtaining log data in the target log format;

[0008] performing word segmentation processing on the log data in the target log format, converting the obtained words into word vectors in vector representation, performing missing feature completion processing on missing words in the word vectors, and obtaining a target format log;

[0009] performing fusion and noise reduction processing on similar target format logs, and obtaining target log data.

[0010] According to another aspect of the present application, a log data processing device is provided, comprising:

[0011] an obtaining module, configured to obtain to-be-processed log data, identify key fields of the to-be-processed log data, map the obtained key fields to a target log format, and obtain log data in the target log format;

[0012] The first processing module is configured to perform word segmentation processing on the log data in the target log format, convert the obtained words into word vectors in vector representation, perform missing feature completion processing on missing words in the word vectors, and obtain target format logs.

[0013] The second processing module is configured to perform fusion and noise reduction processing on similar target format logs, and obtain target log data.

[0014] According to another aspect of the present application, an electronic device is provided, which comprises:

[0015] at least one processor; and

[0016] a memory in communication with the at least one processor; wherein

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the log data processing method according to any one of the embodiments of the present application.

[0018] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the log data processing method according to any one of the embodiments of the present application when executed by the processor.

[0019] According to another aspect of the present application, the embodiments of the present application further provide a computer program product, which comprises a computer program, and the computer program implements the log data processing method according to any one of the embodiments of the present application when executed by a processor.

[0020] The embodiments of the present application obtain the log data to be processed, first perform key field identification on the log data to be processed, map the obtained key fields to a target log format, obtain log data in the target log format, then perform word segmentation processing on the log data in the target log format, convert the obtained words into word vectors in vector representation, perform missing feature completion processing on missing words in the word vectors, obtain target format logs, and finally perform fusion and noise reduction processing on similar target format logs, and obtain target log data. Through the technical solution of the present application, the safety log data can be standardized and the data quality evaluation efficiency can be improved.

[0021] It should be understood that the contents described in this part are not intended to identify the key or important features of the embodiments of the present application, nor are they used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be regarded as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0023] Figure 1 is a flow chart of a log data processing method in the embodiments of the present application;

[0024] Figure 2 is a structural schematic diagram of a log data processing device in the embodiments of the present application;

[0025] Figure 3 is a structural schematic diagram of an electronic device for implementing the log data processing method in the embodiments of the present application. DETAILED DESCRIPTION

[0026] In order to make the person skilled in the art better understand the present application, the following will combine the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of the present application.

[0027] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and the like are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario and the like of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.

[0029] Embodiment one

[0030] Figure 1is a flowchart of a log data processing method in an embodiment of the present application, the embodiment can be applicable to the case of log data processing, and the method can be executed by a log data processing device in the embodiment of the present application, which can be realized in the form of software and / or hardware, as shown in Figure 1 The method specifically includes the following steps:

[0031] S101, obtaining log data to be processed, identifying a key field of the log data to be processed, mapping the obtained key field to a target log format, and obtaining log data in the target log format.

[0032] It should be noted that the log data to be processed can be security device log data to be processed. The target log format can be a standard log format.

[0033] In the embodiment, the collected log data to be processed can be obtained from, for example, a network side flow collection node and a server / PC side data collection node.

[0034] Through years of security investment and construction, the network side flow collection probe has been basically covered. Through the deployed situation awareness system, threat detection system, firewall system, intrusion prevention system, and WEB application firewall system, the network side can collect related malicious threat data such as vulnerability, service detection, host detection, website attack, backdoor communication, account cracking, attack exploitation, mail attack, DOS attack, vulnerability exploitation, hacker tool, abnormal traffic protocol, NetFlow, Payload, and the like, as shown in Table 1.

[0035] The server has deployed host security software, anti-virus software, and the PC has deployed a unified security platform. These security software can collect related malicious threat data on the host side, such as malicious virus information, account abnormalities, malicious processes, exposed ports, vulnerabilities, reverse Shell (reverse Shell is a network attack technology. Attackers use this technology to create a reverse connection from the victim's computer back to the attacker's computer. This connection allows the attacker to connect to the attacker's computer through the victim's network, thereby bypassing traditional firewalls and network security devices, as these devices usually only check outgoing connections and do not monitor incoming connections), Web Rce (Web Shell is a backdoor technology that provides remote access to an attacked computer through a web interface. Attackers can control the victim's computer through the Web Shell in the browser, just like sitting in front of the victim's machine), security baseline, etc., as shown in Table 2.

[0036] Table 1

[0037] Serial number Vendor Device type Data type Number 1 A Situation awareness Security log 1 set of primary platform 2 B Threat detection Security log 2 3 A Firewall Security log 10 4 B Web application firewall Security log 4 5 B Intrusion prevention system Security log 4

[0038] Table 2

[0039] Serial number Vendor Device type Data type Number 1 C PC antivirus Security log 1 set 2 A Host protection Security log, behavior log 1 set 3 D Host protection Security log 1 set 4 E Host antivirus Security log 1 set

[0040] Specifically, the log data of each security device in the live network is accessed, and is converted into a unified format through XStream (extreme stream processing, a stream processing technology used to process and analyze large-scale data streams. It can help developers build applications that can process and analyze data streams in real time) technology. Through key field identification, the key fields are mapped to the standard log format, and the log format is preliminarily unified.

[0041] S102, the log data in the target log format is processed for word segmentation, and the obtained words are converted into word vectors represented by vectors. The missing words in the word vectors are processed for missing feature completion to obtain the target format log.

[0042] It needs to be explained that the missing words can be unknown words in the log stream that need to be inferred according to the context, mainly including: unregistered words (new terms that do not appear in the training stage, such as the error code ERR NEW SLA FAILURE introduced by the new version), dynamic parameters (variable placeholders in log templates: Failed to connect to <server> : <port> → <server>and <port>is missing and the key information is masked (replaced for security de-sensitization: User******attempted admin access).

[0043] In actual operation, due to the differences in source security logs of various manufacturers, the mapped standard log format may have missing key features, i.e., "threatClass" (major category), "threatType" (minor category), "threatSubType" (sub-category), and the like (which can be understood as title classification, the major category corresponds to a first-level title, the minor category corresponds to a second-level title under the first-level title, and the sub-category corresponds to a third-level title under the second-level title. Among them, there are five or six minor categories in the major category, and multiple sub-categories in the minor category, and each sub-category corresponds to corresponding log data), the main function of these key fields is to present the core identifiers of the threat category and the threat type (by extracting all the attributes in the log data into key fields, the key fields are associated with threats and alarms), and subsequent alarm aggregation and threat qualification need classification identifiers as key features to improve the accuracy of subsequent threat qualification scene classification. For example, a SQL (Structured Query Language, structured query language) injection attack, threatClass is "vulnerability risk", threatType is "vulnerability risk", and threatSubType is "SQL injection vulnerability".

[0044] To improve data quality, missing data needs to be supplemented. The missing key features can be supplemented by word vector algorithm processing to form complete logs.

[0045] Specifically, a CRF (Causal Random Field, causal random field) or bidirectional LSTM (Long Short-Term Memory networks, long short-term memory network) model can be used to finely segment the log data, for example, "user login failure" is segmented into "user / login / failure" three semantic units. A domain-specific dictionary (such as containing "CPU occupancy rate is too high", "disk I / O is abnormal" and the like operation and maintenance terms) is constructed, and a mixed word table is formed in combination with a general corpus. According to the log characteristics, Word2Vec (CBOW / Skip-Gram) (suitable for 90% of log analysis scenes, especially in tasks with high sparsity and fine-grained relationship) or GloVe algorithm (only used when the log size is extremely large, such as more than 1 million, and focuses on high-frequency patterns) is selected. For short text logs, Skip-Gram performs better; and CBOW is suitable for long text scenarios.

[0046] Among them, Word2Vec (Word Vectorization) is a technology that converts words into vectors, which can be used for various NLP (Natural Language Processing) tasks. Word2Vec usually uses two main models: CBOW (Continuous Bag of Words) and Skip-Gram (Skip-Gram model). The CBOW model works by representing the vector of a word as the average of the vector of its context words, which assumes that the meaning of a word can be predicted by its surrounding words. The Skip-Gram model extends the CBOW model by considering not only the context words but also the intervals between words, and the Skip-Gram model generates the vector representation of a word by considering the vector product of the word and its context words, which can capture longer distance word relationships. GloVe (Global Vectors for Word Representation) is an unsupervised learning method that uses global matrix factorization to learn dense vector representations of words, and the GloVe model learns word vectors by maximizing the log probability of word co-occurrence in the entire corpus, which can capture global semantic information of words and can be extended to other languages.

[0047] In the implementation process, the vector dimension can be set to 100-300 dimensions, the window size can be set to 5-10, and the minimum word frequency threshold can be set to 2-5 times. Negative sampling technology is used to accelerate training, and the proportion of negative samples (positive samples with threat properties, negative samples are log data without threat properties) is controlled at 5-20.

[0048] Next, according to the context, the missing words can be filled by weighted average vector or similarity matching, and then multi-feature fusion is performed. Time series features (such as periodic fluctuations) and statistical features (such as abnormal word frequency) are introduced, and the LSTM network is used for context-aware completion decision. The adversarial training method is used to distinguish normal logs and completed logs through the discriminator, to prevent the model from deviating from the original semantic space. A log template library is constructed, and the structured log is filled in the template to complete the missing fields.

[0049] In addition, the semantic similarity between the completed content and the original log can also be measured by indicators. A sampling review mechanism is established to manually review high-risk operation logs (such as permission changes).

[0050] For the consideration of the word vector model iteration problem, generally, the log format of the online security device will not change, and the model does not need to be updated. If the log format of the online security device changes greatly after software upgrade, the word vector model can be upgraded and trained by collecting the original logs of the related devices to update the training set, and then the word vector model is updated to the basic platform.

[0051] S103, the similar target format log is fused and denoised to obtain target log data.

[0052] The target log data can be unified standardized security data logs.

[0053] In actual operation, there can be similar security logs from multiple manufacturers. After comparison, the similar standard format logs are fused and denoised, and finally the unified standardized security data logs are formed.

[0054] In the embodiment of the application, the log data to be processed is obtained, the key fields of the log data to be processed are identified, the obtained key fields are mapped to the target log format, the log data in the target log format is obtained, the log data in the target log format is segmented, the obtained words are converted into word vectors, the missing features in the word vectors are completed, the target format log is obtained, and the similar target format log is fused and denoised to obtain the target log data. Through the technical scheme of the application, the security log data standardization can be realized and the data quality evaluation efficiency can be improved.

[0055] Optionally, the obtained key fields are mapped to the target log format, including:

[0056] The target field set in the target log format is defined, and a mapping table between the key fields of the log data and the target field set is established.

[0057] In this embodiment, the target field set can be a standard field set, such as timestamp (Timestamp), source IP (SourceIP), event type (EventType), severity (Severity), etc.

[0058] Specifically, the standard field set (Timestamp, SourceIP, EventType, Severity, etc.) is defined, and a mapping table of the manufacturer-specific field to the standard field (such as "% FIREWALL-6-106023" mapped to "RuleMatch" event type) is established.

[0059] In network devices, the "%FIREWALL-6-106023" log message usually represents a specific warning or error event. IOS (Internetwork Operating System) software uses this format to record system events to facilitate network administrators in troubleshooting and monitoring. The format of this log message usually contains the following parts: timestamp: the time when the event occurred; device name: the name of the device that records the log; module: the IOS module that generates the log message; process ID: the process identifier that generates the log message; message ID: a unique identifier that identifies a specific message; severity level: a level that indicates the severity of the event (e.g., notification, warning, error, critical, etc.); description: a brief description of the event. In the above example "%FIREWALL-6-106023", FIREWALL can refer to a "Firewall" related warning or error; 6 can refer to the severity level, usually 6 indicates "Notification"; 106023 can be the message ID that uniquely identifies this message. Mapping to the "RuleMatch" event type can mean that this log message is related to a security rule match, such as an access control list rule, intrusion prevention system rule, or other type of security policy rule. This can indicate that the traffic matches a predefined security rule and takes corresponding actions (such as allowing, denying, logging, etc.) according to the rule.

[0060] Based on the mapping table, the obtained key field is mapped to a field in the target field set to obtain a target log format.

[0061] Specifically, by key field recognition, the key field is mapped to the standard log format, realizing the preliminary unification of the log format.

[0062] In addition, for the log mode not covered, the Prodigal online learning algorithm can be used to automatically generate parsing rules, and the artificial review mechanism is used to ensure the accuracy of the new rules.

[0063] The technical scheme of the embodiment of the application researches network security alarm and event data, performs security data standardization, defines a rich security log format specification, and can adapt to the format of the existing network and mainstream security manufacturers in the industry. At the same time, it supports the access of log data of various security devices in the existing network and converts it into a unified format through XStream, and performs key field recognition and maps the feature field to the standard log format, realizing the preliminary unification of the log format.

[0064] Optionally, the word vector is subjected to missing feature completion processing, including:

[0065] For each missing word, the target number of word vectors adjacent to the missing word are obtained based on a sliding window to obtain a word vector set.

[0066] The target number can be represented by N. Specifically, the target number can be determined according to the window size of the sliding window. For example, the window size is 5, that is, 2 words are taken before and after the center word.

[0067] It should be noted that the word vector set can be a set composed of the target number of word vectors adjacent to the missing word.

[0068] Specifically, a sliding window can be set. For each missing word, N words (i.e., context words) before and after the missing word are obtained to form a word vector set.

[0069] The weighted average vector corresponding to the word vector set is used as the vector representation of the missing word.

[0070] Specifically, the weighted average of the word vectors of the context words (the weights can be assigned according to the distance from the center position, such as the position weight of the center word being 1, the weight of the adjacent word being 0.8, and the weight of the farther word being 0.5, etc., or the average can be directly taken) is calculated, and the obtained weighted average vector is used as the vector representation of the missing word.

[0071] Optionally, the word vector is subjected to missing feature completion processing, including:

[0072] The cosine similarity between the word vector and the missing word is obtained.

[0073] Specifically, the cosine similarity between the missing word and each word vector is calculated.

[0074] The candidate word corresponding to the missing word is determined based on the cosine similarity.

[0075] Specifically, the cosine similarity is used to search for Top-K (K can be set according to actual conditions, and the present embodiment does not limit this) candidate words in the word vector space, and the candidate words are filtered according to the business rules (such as excluding entity words such as device models). The preliminary result obtained through "cosine similarity calculation in the word vector space" and "Top-K search" is the candidate word of the missing word.

[0076] The word vector space: converting words into vectors (such as through Word2Vec, BERT (Bidirectional Encoder Representations from Transformers, bidirectional encoder representations from transformers), etc.), the direction and distance of the vector reflect the semantic relationship of the words (such as "apple" and "fruit" vectors are close, and "apple" and "computer" vectors also have some association).

[0077] Cosine similarity: measures the cosine value of the angle between two word vectors (range [-1, 1]), the closer the value is to 1, the more similar the semantics (such as the cosine similarity of "happy" and "happy" is close to 1).

[0078] Top-K search: select the top K words with the highest cosine similarity to the target word (such as K = 3), these words are candidate words, which are the most semantically similar words to the target word.

[0079] The technical scheme of the embodiment of the application realizes key feature completion through a word vector algorithm, automatically fills in key fields for improving threat qualification accuracy, and improves the completeness and availability of access data through the combination of automatic access analysis and log information fusion of multi-source telemetry data, greatly improves the log analysis efficiency, and reduces the labor cost.

[0080] Optionally, the similar target format log is fused and denoised, including:

[0081] If the target separator does not exist in the target format log, it is determined that the target format log is an unstructured log.

[0082] For example, the target separator can be, for example, JSON (JavaScript Object Notation or JavaScript Object Tag), XML (eXtensible Markup Language), CSV (Comma-Separated Values or Comma-Separated Variables), or a specific separator (such as the <pri>(uniform delimiters such as

[0083] In the implementation process, when performing similar log fusion noise reduction, first, structured conversion is performed in the log analysis layer.

[0084] Specifically, a regular expression library is constructed: a log analysis rule library covering mainstream manufacturers is constructed to support Syslog (System Log), SNMP Trap (Simple Network Management Protocol Trap), and other protocol analysis. In the construction of the log analysis rule library covering mainstream manufacturers, regular expressions are core tools for extracting structured information from unstructured log text generated by Syslog, SNMP Trap, and other protocols. The following is a detailed description of regular expressions in log analysis: the role of regular expressions: log analysis: extract key information (such as timestamp, device name, event type, IP address, etc.) from unstructured log text; multi-manufacturer adaptation: different manufacturers have different log formats, and regular expressions need to be customized to match specific formats; event classification: identify alarm levels, fault types, or security events by matching specific patterns; data standardization: convert extracted information into a unified format for subsequent analysis. Common log formats and regular expression examples: Syslog example: log format: Mar 10 14:23:45 router1 %LINK-3-UPDOWN: Interface GigabitEthernet0 / 1, changed state to up. Regular expressions can be expressed as follows, for example: ^(\w+\s+\d+\s+\d+:\d+:\d+)\s+(\S+)\s+%(\S+)-(\d+)-(\S+):\s+(.*)$. Extracted groups: timestamp: Mar 10 14:23:45; device name: router1; facility: LINK; severity: 3; message type: UPDOWN; detailed information: Interface GigabitEthernet0 / 1, changed state to up.

[0085] Specifically, after the regular expression library is constructed, natural language processing is performed: for unstructured logs (such as firewall policy change descriptions), a BERT (Bidirectional Encoder Representations from Transformers)-BiLSTM (Bidirectional Long Short-Term Memory)-CRF (Conditional Random Field) model is used for entity recognition to extract key fields such as time, source IP, event type, etc.

[0086] In this embodiment, the structured log features can include: fixed format: such as JSON, XML, CSV or specific delimiters (such as Syslog <pri>). Machine readability: Fields and values are one-to-one, easy for programs to parse. Processing method: Direct parsing, e.g. JSON / XML can use built-in parsing libraries (like Python's json.loads), CSV / Tab can split fields by delimiter, Syslog can split by <pri>Timestamp device name message format parsing. In addition, the processing method can also include tool assistance: Logstash: define patterns (such as %{TIMESTAMP_ISO8601:timestamp}) through the grok plug-in. Regular expression: match fixed format parts (such as extracting the priority of Syslog). Finally, standardize the data: map the parsing results to a unified data model (such as uniformly naming policy_change events from different manufacturers).

[0087] Unstructured features include: no fixed format: log content is free text, no uniform field separator (such as comma, tab); strong human readability: usually descriptive text generated automatically by the system, suitable for manual reading; key information is scattered: time, IP, event, etc. Information may appear in different order or form.

[0088] Specifically, whether it is a structured log can be determined by checking whether there is a uniform separator (such as {} for JSON,, for CSV).

[0089] If the target format log is an unstructured log, extract the key information from the target format log.

[0090] For example, key information may be, for example, timestamp, device name, event type, IP address, etc.

[0091] Specifically, structured logs can be directly parsed, and unstructured logs (such as firewall policy change descriptions) can use the BERT-BiLSTM-CRF model for entity recognition to extract time, source IP, event type, and other key fields.

[0092] In the specific implementation process, when performing similar log fusion and noise reduction, first, structured conversion is performed in the log parsing layer, and then semantic alignment is performed in the feature engineering layer. Specifically, word vector enhancement and semantic fingerprint extraction are required. Among them, word vector enhancement: inject security domain knowledge (such as CVE (Common Vulnerabilities and Exposures, Common Vulnerabilities and Exposures) vulnerability number, attack method term) on the basis of general word vectors, and construct a security special word vector space. Semantic fingerprint extraction: generate a semantic fingerprint (such as "user [admin] at [2025-06-16, 10:23] executes [high-risk command deletes system log]") for each log event, and weight the key fields through TF-IDF (Term Frequency-Inverse Document Frequency, Term Frequency-Inverse Document Frequency).

[0093] Next, intelligent deduplication can be performed through the fusion and noise reduction layer, and the specific process is as follows:

[0094] According to the key information, determine the first similarity and the second similarity corresponding to the target format log.

[0095] It should be noted that the first similarity can be Jaccard similarity. Jaccard similarity is an index for measuring the similarity of two sets, widely used in data mining, text similarity calculation, machine learning and other fields. It measures the similarity of two sets by calculating the ratio of their intersection and union.

[0096] It should be noted that the second similarity can be semantic similarity.

[0097] Specifically, calculate the Jaccard similarity of the time, IP, port and other precise matching fields, and use Sentence-BERT to calculate the semantic similarity of the event description.

[0098] If the gap between the first similarity and the second similarity corresponding to the target format log exceeds the preset threshold, determine the target similarity according to the conflict resolution strategy.

[0099] Among them, the target similarity can be the priority similarity, that is, when there is a conflict between the two ways (the gap between the first similarity and the second similarity exceeds the preset threshold), the priority similarity is determined.

[0100] Specifically, if there is a conflict between the two ways, the Jaccard similarity is high but the semantic similarity is low, or the semantic similarity is high but the Jaccard similarity is low, the core principle of conflict processing is: conflict processing needs to determine the priority around the business target (such as "prioritize identifying similar attack events" and "prioritize associating operations of the same device"), the core principle is:

[0101] Business scenario priority: Different scenarios have different dependencies on "hard matching" and "soft matching" (such as network attack analysis, where the accuracy of IP / time may be higher than the semantic; while in audit logs, the semantic consistency of operation description may be more important).

[0102] Minimum misjudgment principle: Avoid misjudgment due to a single indicator (such as Jaccard high but semantic low, which needs to be alert to "misassociation of same IP different events").

[0103] Interpretability: The processing logic needs to be traceable (such as "IP belongs to the same network segment, even if Jaccard is low, the association is retained" needs to be clear rules).

[0104] Specific conflict resolution strategy:

[0105] 1. Rule filtering based on business priority:

[0106] Scenario A: Network security events (IP / time as the core):

[0107] Priority: Jaccard similarity (structured field) > semantic similarity.

[0108] Processing logic: If Jaccard >= 0.6 (high match in structured field), even if semantic similarity < 0.85, it is still considered as an associated event (possibly different stages of the same event); if Jaccard < 0.3 (large difference in structured field), even if semantic similarity is high, further verification is required (such as whether the IP belongs to the same network segment, whether the time is within the attack window period). The above thresholds can be set by the user according to actual situation or experience value.

[0109] Scenario B: Operation audit event (description semantic is the core):

[0110] Priority: semantic similarity > Jaccard similarity.

[0111] Processing logic: If semantic >= 0.85 (consistent description), even if the IP / port is different (such as an administrator performing the same operation on different devices), it is still considered as an associated event; if semantic < 0.7, even if Jaccard is high (such as different operations on the same IP), it is not associated.

[0112] 2. Weight fusion: comprehensive calculation of "mixed similarity":

[0113] Weighted fusion of the two similarities into a comprehensive score, or add additional structured fields or rules to alleviate conflicts:

[0114] Auxiliary fields: device ID, event type (such as "attack", "configuration change"), user ID, etc. For example, if Jaccard is low but semantic similarity is high, if "device ID belongs to the same cluster" or "event type is 'SQL injection'", the association is retained.

[0115] Domain rules: whether the IP belongs to the same network segment (such as 192.168.1.1 and 192.168.1.2 belong to / 24 network segment, which can increase the equivalent score of Jaccard); whether the time is within the preset window period (such as within 2 hours is considered as the same event).

[0116] According to the target similarity, the similar target format log is fused and processed.

[0117] The technical scheme of the embodiment of the application automatically understands the log content through NLP technology, automatically analyzes the security log based on AI technology, and maps the fragmented alarms of the same attack scattered on different devices to the same alarm ID, that is, the alarms are fused into one, and then accurately sent to the corresponding detection engine. Based on AI technology, log understanding is enabled, automatic log access and analysis of security devices are realized, and the subsequent construction cost is greatly reduced.

[0118] Optionally, the similar target format logs are fused and denoised according to the target similarity, including:

[0119] The log event graph is constructed according to the target similarity.

[0120] In the log event graph, the nodes are the target format logs, and the edge weights are the target similarities.

[0121] Specifically, the log event graph is constructed, the nodes are the log entries, and the edge weights are the similarity scores.

[0122] The similar target format logs are clustered based on a preset community discovery algorithm.

[0123] In this embodiment, the preset community discovery algorithm can be a Louvain community discovery algorithm. The Louvain community discovery algorithm is a high-efficiency algorithm for detecting community structures in complex networks. It identifies communities in networks by optimizing the modularity of the network and is widely used in social network analysis, bioinformatics, information science, etc.

[0124] Specifically, the Louvain community discovery algorithm is used for clustering, and events in the same community are determined as repeated or associated events.

[0125] Violent cracking noise is filtered based on a frequency threshold.

[0126] For example, the frequency threshold can be, for example, > 100 times of failed login per second for a single IP.

[0127] Specifically, the noise filtering rule can be, for example, filtering violent cracking noise based on a frequency threshold (e.g., > 100 times of failed login per second for a single IP).

[0128] An abnormal log pattern is detected based on an isolation forest algorithm.

[0129] It can be known that the isolation forest algorithm is an unsupervised learning algorithm for anomaly detection, which is particularly suitable for high-dimensional data sets. It constructs isolated trees by randomly selecting features and randomly selecting feature split points, thereby isolating data points. Anomaly points are usually more easily isolated, so the depth of the isolated tree can be used to determine whether a data point is abnormal.

[0130] The technical scheme of the embodiment of the application studies a large amount of network security device alarm metadata, and through natural semantic analysis technology, a multi-source network security device telemetry data intelligent management scheme is formed. Through the combination of multi-source telemetry data automatic access, threat understanding, intelligent verification, attack information fusion, etc., data standardization is realized, and data quality evaluation efficiency is improved.

[0131] Example 2

[0132] Figure 2 This is a schematic diagram of a log data processing device according to an embodiment of the present invention. This embodiment is applicable to log data processing applications. The device can be implemented using software and / or hardware, and can be integrated into any device that provides log data processing functionality, such as… Figure 2 As shown, the log data processing device specifically includes: an acquisition module 201, a first processing module 202, and a second processing module 203.

[0133] The acquisition module 201 is used to acquire log data to be processed, identify key fields in the log data to be processed, and map the obtained key fields to the target log format to obtain log data under the target log format.

[0134] The first processing module 202 is used to perform word segmentation on the log data in the target log format, convert the obtained words into word vectors, and perform missing feature completion on the missing words in the word vectors to obtain the target format log;

[0135] The second processing module 203 is used to perform fusion and noise reduction processing on logs of similar target formats to obtain target log data.

[0136] Optionally, the first processing module 202 is specifically used for:

[0137] For each missing word, the target number of word vectors adjacent to the missing word are obtained based on a sliding window to obtain a set of word vectors;

[0138] The weighted average vector corresponding to the word vector set is used as the vector representation of the missing word.

[0139] Optionally, the first processing module 202 is specifically used for:

[0140] Obtain the cosine similarity between the word vectors other than the missing word and the missing word;

[0141] Candidate words corresponding to the missing word are determined based on the cosine similarity.

[0142] Optionally, the second processing module 203 includes:

[0143] The first determining unit is configured to determine that the target format log is an unstructured log if no target delimiter exists in the target format log.

[0144] An extraction unit is used to extract key information from the target format log if the target format log is an unstructured log.

[0145] a second determining unit, configured to determine a first similarity and a second similarity corresponding to the target format log according to the key information;

[0146] a third determining unit, configured to determine a target similarity according to a conflict resolution strategy if a gap between the first similarity and the second similarity corresponding to the target format log exceeds a preset threshold value;

[0147] a processing unit, configured to perform fusion and noise reduction processing on similar target format logs according to the target similarity.

[0148] Optionally, the processing unit is specifically configured to:

[0149] construct a log event graph according to the target similarity; in the log event graph, a node is a target format log, and an edge weight is the target similarity;

[0150] cluster similar target format logs based on a preset community discovery algorithm;

[0151] filter brute force cracking noise based on a frequency threshold value;

[0152] detect an abnormal log mode based on an isolation forest algorithm.

[0153] Optionally, the acquisition module 201 is specifically configured to:

[0154] define a target field set under a target log format, and establish a mapping table between a key field corresponding to log data and the target field set;

[0155] map the obtained key field to a field in the target field set based on the mapping table to obtain the target log format.

[0156] The product can execute the log data processing method provided in any embodiment of the application, and has the corresponding functional modules and beneficial effects of the execution method.

[0157] Embodiment Three

[0158] Figure 3 A structural diagram of an electronic device 30 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0159] As shown in Figure 3 The electronic device 30 includes at least one processor 31, and memory, such as read-only memory (ROM) 32, random access memory (RAM) 33, etc., communicatively connected to the at least one processor 31, where the memory stores computer programs executable by the at least one processor. The processor 31 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 32 or loaded from the storage unit 38 into the random access memory (RAM) 33. In the RAM 33, various programs and data required for the operation of the electronic device 30 can also be stored. The processor 31, the ROM 32, and the RAM 33 are connected to each other through a bus 34. An input / output (I / O) interface 35 is also connected to the bus 34.

[0160] Various components in the electronic device 30 are connected to the I / O interface 35, including an input unit 36, such as a keyboard, a mouse, etc., an output unit 37, such as various types of displays, speakers, etc., a storage unit 38, such as a magnetic disk, an optical disk, etc., and a communication unit 39, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 39 allows the electronic device 30 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0161] The processor 31 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 31 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 31 performs various methods and processes described above, such as the log data processing method:

[0162] Obtaining to-be-processed log data, performing key field identification on the to-be-processed log data, and mapping the obtained key field to a target log format to obtain log data in the target log format;

[0163] Performing word segmentation processing on the log data in the target log format, converting the obtained words into word vectors in vector representation, performing missing feature completion processing on missing words in the word vectors, and obtaining a target format log;

[0164] Performing fusion and noise reduction processing on similar target format logs to obtain target log data.

[0165] In some embodiments, the log data processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as storage unit 38. In some embodiments, part or all of the computer program can be loaded and / or installed onto electronic device 30 via ROM 32 and / or communication unit 39. When the computer program is loaded onto RAM 33 and executed by processor 31, one or more steps of the log data processing method described above can be performed. Alternatively, in other embodiments, processor 31 can be configured, by any other suitable means (e.g., by means of firmware), to perform the log data processing method.

[0166] The various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0167] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flow diagrams and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or a remote machine or a server.

[0168] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0169] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0170] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.

[0171] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are a host product in the cloud computing service system to solve the defects of great management difficulty and weak business scalability in traditional physical hosts and VPS services.

[0172] In an embodiment, the present embodiment further includes a computer program product, which includes a computer program, the computer program, when executed by a processor, implements the log data processing method of any embodiment of the present application.

[0173] The computer program product, in implementation, can be written in one or more programming languages or combinations thereof to implement computer program code for performing the operations of the present application, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. The program code can be executed entirely on a user computer, partially on a user computer, as a separate software package, partially on a user computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, through the Internet using an Internet service provider).

[0174] It should be understood that the various forms of flow shown above can be reordered, added or deleted steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0175] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.< / pri> < / pri> < / pri> < / port> < / server> < / port> < / server>

Claims

1. A log data processing method, characterized in that, include: Obtain log data to be processed, identify key fields in the log data to be processed, and map the obtained key fields to the target log format to obtain log data under the target log format; The log data in the target log format is segmented into words, and the resulting words are converted into word vectors. Missing words in the word vectors are filled in with missing features to obtain the target format log. Similar target format logs are fused and denoised to obtain the target log data.

2. The method according to claim 1, characterized in that, The word vectors are subjected to missing feature completion processing, including: For each missing word, the target number of word vectors adjacent to the missing word are obtained based on a sliding window to obtain a set of word vectors; The weighted average vector corresponding to the word vector set is used as the vector representation of the missing word.

3. The method according to claim 1, characterized in that, The word vectors are subjected to missing feature completion processing, including: Obtain the cosine similarity between the word vectors other than the missing word and the missing word; Candidate words corresponding to the missing word are determined based on the cosine similarity.

4. The method according to claim 1, characterized in that, Perform fusion and noise reduction processing on logs of similar target formats, including: If the target format log does not contain a target delimiter, then the target format log is determined to be an unstructured log. If the target format log is an unstructured log, then extract key information from the target format log; Based on the key information, determine the first similarity and the second similarity corresponding to the target format log; If the difference between the first similarity and the second similarity corresponding to the target format log exceeds a preset threshold, the target similarity is determined according to the conflict resolution strategy. Based on the target similarity, similar target format logs are fused and denoised.

5. The method according to claim 4, characterized in that, Based on the target similarity, similar target format logs are fused and denoised, including: A log event graph is constructed based on the target similarity; in the log event graph, the nodes are logs in the target format, and the edge weights are the target similarity. Clustering of logs with similar target formats based on a pre-defined community detection algorithm; Brute-force noise filtering based on frequency thresholds; Detecting abnormal log patterns based on the isolated forest algorithm.

6. The method according to claim 1, characterized in that, Map the obtained key fields to the target log format, including: Define the target field set under the target log format, and establish a mapping table between the key fields corresponding to the log data and the target field set; Based on the mapping table, the obtained key fields are mapped to fields in the target field set to obtain the target log format.

7. A log data processing device, characterized in that, include: The acquisition module is used to acquire log data to be processed, identify key fields in the log data to be processed, and map the obtained key fields to the target log format to obtain log data under the target log format. The first processing module is used to perform word segmentation on the log data in the target log format, convert the obtained words into word vectors, and perform missing feature completion on the missing words in the word vectors to obtain the target format log. The second processing module is used to perform fusion and noise reduction processing on logs with similar target formats to obtain target log data.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the log data processing method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the log data processing method according to any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the log data processing method according to any one of claims 1-6.

Citation Information

Cited By

  • Multi-cloud unified management platform oriented middleware log auditing method and device

    CN121841961A