SYSTEM AND METHOD FOR ANALYSIS OF DATA LOGS

The AI-powered system addresses the challenge of processing diverse data protocols by using NLP and neural networks to convert and analyze data formats, enhancing data security by detecting events in real-time and reducing security gaps.

DE102021212380B4Active Publication Date: 2025-09-25NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
DE102021212380
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-04
Filing Date
2021-11-03
Publication Date
2025-09-25
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

Conventional data protocol processing systems are unable to handle the large volumes of heterogeneous data generated by companies, leading to security gaps and inefficiencies in data analysis, particularly in cybersecurity protocols, due to their inability to scale and process data of varying formats without significant human intervention.

Method used

A flexible artificial intelligence (AI)-enabled system utilizing Natural Language Processing (NLP) and neural networks to analyze data protocols of known or unknown formats, including partial, incomplete, and degraded protocols, enabling real-time processing and detection of triggerable events.

Benefits of technology

The system efficiently processes large sets of data protocols in real-time, identifying triggerable events and reducing security gaps by converting diverse data formats into a common format for comprehensive analysis and alerting system administrators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for processing data logs, the method comprising: Receiving a data log from a data source, wherein the data log is received in a format native to a machine that generated the data log; Providing the data protocol for a neural network trained to process natural language inputs; Analyzing the data log with the neural network; Receiving an output from the neural network, wherein the output is generated by the neural network in response to the neural network analyzing the data protocol; and Storing the output of the neural network in a data log memory.
Need to check novelty before this filing date? Find Prior Art

Description

AREA OF REVELATION

[0001] The present disclosure relates generally to data protocols and, more particularly, to the analysis of data protocols of known or unknown formats. BACKGROUND

[0002] Data logs were originally developed as a mechanism to preserve historical information about important events. For example, bank transactions needed to be recorded for verification and audit purposes. With the development of technology and the proliferation of the internet, data logs have become increasingly common, and all data generated by a connected device is often stored in some form of data log.

[0003] For example, cybersecurity logs created for an organization can include data generated by endpoints, network devices, and perimeter devices. Even small businesses can expect to generate hundreds of gigabytes of data in log traffic. Even a small data loss can lead to security vulnerabilities within the organization. SHORT SUMMARY

[0004] Traditional systems designed to ingest data logs are unable to handle the current data volumes of most organizations. Furthermore, these traditional systems are not scalable to support a significant increase in data log traffic, frequently resulting in missing or lost data. In the context of cybersecurity logs, any amount of lost or missing data can lead to security vulnerabilities. Today, organizations are collecting, storing, and attempting to analyze more data than ever before. The data logs are heterogeneous in terms of source, format, and time. To make matters worse, the types and formats of data logs are constantly changing, meaning new types of data logs are being introduced into systems, and many of these systems are not designed to handle such changes without significant human intervention.In summary, traditional data log processing systems are unable to properly handle the volumes of data generated in many organizations.

[0005] EP 3 468 140 B1 relates to an artificial intelligence network for natural language processing and a data security system, hereinafter referred to as a “system,” that determines an emotion model for one or more users from the users’ electronic natural language interactions. The system comprises a natural language processing decoder for determining text features that may indicate users’ emotional states from the electronic natural language interactions. Examples of electronic natural language interactions, also simply referred to as “interactions,” may include an interaction between two or more users. The interactions may be contained in emails, documents, natural language such as recorded audio files, recorded videos, recorded videos with natural language and non-verbal communication, images, text messages, social media messages or posts, etc.The interactions may take place over a period of, for example, one day, a few minutes, or a few hours. An interaction capture system may capture the electronic interactions in natural language and store them in a data store for the natural language processing decoder. The system may use the stored interactions as a data set.

[0006] Li, Weixiang. "Automatic Log Analysis using Machine Learning: Awesome Automatic Log Analysis Version 2.0." (2013) focuses on automatic log analysis using machine learning techniques, as they are effective and efficient for automated log analysis and are well-suited to big data problems. Features are extracted from the logs and clustering algorithms are used to detect abnormal logs.

[0007] Embodiments of the present disclosure aim to address the above-mentioned deficiencies and other problems associated with data log processing. The embodiments described herein provide a flexible, artificial intelligence (AI)-enabled system configured to process large volumes of data logs in known or unknown formats.

[0008] The invention is defined by the claims. To illustrate the invention, aspects and embodiments are described herein, which may or may not fall within the scope of the claims.

[0009] In some embodiments, the AI-enabled system may utilize natural language processing (NLP) as a technique for processing data logs. NLP has traditionally been used for applications such as text translation, interactive chatbots, and virtual assistants. Using NLP to process data logs generated by machines may not appear immediately practical. However, the present disclosure recognizes the unique ability of NLP or other natural language neural networks to analyze data logs of known or unknown formats when properly trained. Embodiments of the present disclosure also enable a natural language neural network to analyze partial data logs, incomplete data logs, degraded data logs, and data logs of various sizes.

[0010] In one illustrative example, a method for processing data logs is disclosed, comprising: receiving a data log from a data source, wherein the data log is received in a format native to a machine that generated the data log; providing the data log to a neural network trained to process natural language-based inputs; analyzing the data log with the neural network; receiving an output from the neural network, wherein the output is generated by the neural network in response to the neural network analyzing the data log; and storing the output from the neural network in a data log repository.

[0011] The method may further comprise: receiving an additional data log from an additional data source, wherein the additional data source is different from the data source and the additional data log is received in a second format native to the additional data source; providing the additional data log to the neural network; analyzing the additional data log with the neural network; receiving an additional output from the neural network, wherein the additional output is generated by the neural network in response to the neural network analyzing the additional data log; and storing the additional output from the neural network in the data log storage. The output and the additional output may be stored in the data log storage in a common data format as part of a combined data log.The additional data protocol may be received as a data stream directly from the additional data source. In one embodiment, the machine that generated the data protocol may comprise a first device type, while the additional data source may comprise a second device type, and the first device type and the second device type may belong to a common network infrastructure.

[0012] The machine that generated the data log may include at least one communication endpoint, a network device, a network boundary device, a security device, or a sensor. The data log may include security data transmitted from the machine to another machine, and the neural network may include a machine learning model for natural language processing.

[0013] The method may further comprise splitting the data log into a plurality of data log chunks and providing the plurality of data log chunks to the neural network, wherein the neural network is trained with training data comprising log chunks, and wherein a size of one log chunk in the plurality of data log chunks differs from a size of another log chunk in the plurality of data log chunks.

[0014] The method may further comprise analyzing the data log memory, detecting a triggerable data event based on the analysis of the data log memory, and providing an alert to a communication device, the alert including information describing the triggerable data event.

[0015] The data log may include at least one of a file path name, an Internet Protocol (IP) address, a Media Access Control (MAC) address, a timestamp, a hexadecimal value, a sensor reading, a user name, an account name, a domain name, a hyperlink, host system metadata, a connection duration, a communication protocol, a communication port, and raw payload data. The data log may include at least one of a degraded log or an incomplete log.

[0016] In another example, a system for processing data logs is disclosed, comprising: a processor and a memory coupled to the processor, the memory storing data that, when executed by the processor, enables the processor to: receive a data log from a data source, wherein the data log is received in a format native to a machine that generated the data log; analyze the data log with a neural network trained to process natural language-based inputs; and store an output of the neural network in a data log memory, wherein the output of the neural network is generated in response to the neural network analyzing the data log.

[0017] The data stored in memory may further enable the processor to tokenize the data log prior to analyzing the data log with the neural network. The data stored in memory may further enable the processor to receive an additional data log from an additional data source, wherein the additional data source is different from the data source and wherein the additional data log is received in a second format native to the additional data source; analyze the additional data log with the neural network; and store an additional output from the neural network in the data log memory, wherein the additional output is generated by the neural network in response to the neural network analyzing the additional data log.The output and the additional output may be stored in the data log storage in a common data format as part of a combined data log. The additional data log may be received as a data stream directly from the additional data source. The machine that generated the data log may comprise a first device type, the additional data source may comprise a second device type, and the first device type and the second device type may belong to a common network infrastructure. The data log may comprise security data communicated from the machine to another machine, and the neural network may comprise a machine learning model for natural language processing.

[0018] The data stored in memory also enables the processor to: analyze the data log memory; detect a triggerable data event based on the analysis of the data log memory; and send an alert to a communication device, the alert comprising information describing the triggerable data event.

[0019] At least one of the processor and the memory may be housed in a graphics processing unit (GPU).

[0020] The data log may include at least one of a degraded log and an incomplete log.

[0021] In another example, a method for training a system for processing data logs is disclosed, comprising: providing a neural network with first training data, the neural network including a machine learning model for natural language processing, and the first training data including a first data log generated by a first type of machine; providing the neural network with second training data, the second training data including a second data log generated by a second type of machine; determining that the neural network has trained on the first training data and the second training data for at least a predetermined time; and storing the neural network in a computer memory such that the neural network is made available for processing additional data logs.

[0022] The first data log may comprise at least one of a raw data log and an analyzed data log, wherein the first data log is tokenized with at least one of word embedding, split words, and position encoding, and the method may further comprise adapting training of the neural network by at least one of: (i) shuffling the first training data and the second training data; (ii) varying a starting point of the first training data; (iii) varying a starting point of the second training data; and (iv) degrading at least one of the first training data and the second training data.

[0023] In another example, a processor is provided that includes one or more circuits that use one or more natural language-based neural networks to analyze one or more machine-generated data logs. The one or more circuits may correspond to logic circuits interconnected in a graphics processing unit. The one or more circuits may be configured to receive the one or more machine-generated data logs from a data source and generate an output in response to analyzing the one or more machine-generated data logs, the output configured to be stored as part of a data log store.In some examples, the one or more machine-generated data logs are received as part of a data stream, and at least one of the machine-generated data logs may include a degraded log and an incomplete log.

[0024] The one or more circuits may be configured to receive the one or more machine-generated data logs from a data source and generate an output in response to analyzing the one or more machine-generated data logs, the output configured to be stored as part of a data log store. The one or more machine-generated data logs may be received as part of a data stream. The one or more machine-generated data logs may include at least one of a degraded log, an incomplete log, a deduplicated log, a log summary, a log in which sensitive information is obfuscated, and a partial data log.

[0025] Further features and advantages are described here and can be seen from the following description and figures.

[0026] Any feature of one aspect or embodiment may be applied to other aspects or embodiments, in any suitable combination. In particular, any feature of a method aspect or method embodiment may be applied to a device aspect or device embodiment, and vice versa. BRIEF DESCRIPTION OF THE DIFFERENT VIEWS OF THE DRAWINGS

[0027] The present disclosure is described in conjunction with the accompanying figures, which are not necessarily drawn to scale: Fig. 1 is a block diagram illustrating a computer system in accordance with at least some embodiments of the present disclosure; Fig. 2 is a block diagram illustrating a training architecture for a neural network according to at least some embodiments of the present disclosure; Fig. 3 is a flowchart illustrating a method for training a neural network in accordance with at least some embodiments of the present disclosure; Fig. 4 is a block diagram illustrating an operational architecture of a neural network according to at least some embodiments of the present disclosure; Fig. 5 is a flowchart illustrating a method for processing data logs in accordance with at least some embodiments of the present disclosure; and Fig. 6 is a flowchart illustrating a method for preprocessing data logs in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTION

[0028] The following description contains only exemplary embodiments and is not intended to limit the scope, applicability, or embodiment of the claims. Rather, the following description is intended to provide a guide to the person skilled in the art for implementing the described embodiments. It is understood that various changes in the function and arrangement of the elements may be made without departing from the spirit and scope of the appended claims.

[0029] From the following description and for reasons of computational efficiency, it is clear that the components of the system can be arranged at any suitable location within a distributed network of components without affecting the operation of the system.

[0030] Furthermore, the various connections connecting the elements may be wired, cabled, or wireless, or any combination thereof, or any other suitable means known or later developed capable of delivering and / or transmitting data to and from the connected elements. Transmission media may include, for example, any suitable carrier for electrical signals, including coaxial cable, copper wire, and fiber optic cables, electrical traces on a printed circuit board, or the like.

[0031] As used herein, the terms "at least one," "one or more," "or," and "and / or" are indefinite terms that can be used both conjunctively and disjunctively. For example, each of the terms "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," "A, B, and / or C," and "A, B, or C" means: A alone; B alone; C alone; A and B together; A and C together; B and C together; or A, B, and C together.

[0032] The term "automatic" and variations thereof, as used herein, refers to any suitable process or operation that is performed without substantial human input when the process or operation is performed. However, a process or operation can be automatic even if performance of the process or operation requires tangible or intangible human input if the input is received before the process or operation is performed. Human input is considered substantial if it influences the performance of the process or operation. Human input that consents to the performance of the process or operation is not considered "substantial."

[0033] The terms “determine,” “calculate,” and “compute,” and their variations, are used interchangeably herein and include any appropriate methodology, process, procedure, or technique.

[0034] Various aspects of the present disclosure are described herein with reference to drawings that are schematic representations of idealized configurations.

[0035] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. It is further understood that terms as defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant prior art and this disclosure.

[0036] As used herein, the singular forms "a," "an," and "the" include the plural forms unless the context clearly indicates otherwise. It is further understood that the terms "comprise," "comprises," and / or "including," when used in this specification, specify the presence of certain features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "and / or" includes all combinations of one or more of the listed items.

[0037] With reference to Fig. 1-6, various systems and methods for analyzing data logs are now described. While various embodiments are described in the context of using AI, machine learning (ML), and similar techniques, it should be understood that embodiments of the present disclosure are not limited to the use of AI, ML, or other machine learning techniques, which may or may not include the use of one or more neural networks. Moreover, embodiments of the present disclosure contemplate the mixed use of neural networks for certain tasks, while algorithmic or predefined computer programs may be used for certain other tasks.In other words, the methods and systems described or claimed herein may be performed using conventional executable instruction sets that are finite and act on a fixed set of inputs to provide one or more defined outputs. Alternatively or additionally, the methods and systems described or claimed herein may be performed using AI, ML, neural networks, or the like. In other words, a system or components of a system as described herein may include finite instruction sets and / or AI-based models / neural networks to perform some or all of the processes or steps described herein.

[0038] In some embodiments, a natural language neural network is used to analyze machine-generated data logs. The data logs may be received directly from the machine that generated the data log, in which case the machine itself may be considered a data source. The data logs may be received from a memory area used to temporarily store data logs from one or more machines; in this case, the memory area may be considered a data source. In some embodiments, data logs may be received in real time, as part of a data stream transmitted directly from a data source to the natural language neural network. In some embodiments, the data logs may be received at a time after they have been generated by a machine.

[0039] Certain embodiments described herein contemplate the use of a natural language neural network. An example of a natural language neural network or an approach using a natural language neural network is NLP. Certain types of word representations in neural networks, such as Word2vec, are context-free. Embodiments of the present disclosure contemplate the use of such context-free neural networks that are capable of creating a single word embedding for each word in the vocabulary and are unable to distinguish between words with multiple meanings (e.g., the file on disk versus a single file line). Newer models (e.g., ULMFit and ELMo) have multiple representations for words based on context.These models achieve an understanding of the context by using the word and the previous words in the sentence to create the representations. Embodiments of the present disclosure also contemplate the use of context-based neural networks. A more specific, but non-limiting, example of a neural network type that may be used without departing from the scope of the present disclosure is a BERT (Bidirectional Encoder Representations from Transformers) model. A BERT model is capable of generating contextual representations but also taking into account the surrounding context in both directions—before and after a word. While the following describes embodiments that use a natural language-based neural network built on a data corpus containing English words, sentences, etc.trained, it should be considered that the natural language-based neural network can be trained on any data, including any human language (e.g., Japanese, Chinese, Latin, Greek, Arabic, etc.) or a collection of human languages.

[0040] Encoding context information (before and after a word) can be useful for understanding cyberprotocols and other types of machine-generated data protocols due to their ordered nature. For example, in several types of data protocols, a source address precedes a destination address. BERT and other context-aware NLP models can take this context-dependent / ordered information into account.

[0041] An additional challenge in applying a natural language model to cyberlogs and other types of machine-generated data logs is that many "words" in a cyberlog are not English words, but include things like file paths, hexadecimal values, and IP addresses. Other language models return an "out-of-dictionary" entry for an unknown word, but BERT and similar neural networks are configured to break down the words in cyberlogs into dictionary word pieces. For example, ProcessID becomes two in-dictionary WordPieces—Process and ##ID.

[0042] Different sets of data logs can be used to train one or more of the language-based neural networks described here. For example, data logs such as Windows event logs and Apache web logs can be used as training data. The language of the cyberlogs is not the same as the English language corpus on which the BERT tokenizer and neural network were trained.

[0043] The speed and accuracy of a model can be further improved by using a tokenizer and a representation trained from scratch on a large corpus of data logs. For example, a BERT word chunk tokenizer can decompose AccountDomain into A ##cco ##unt ##D ##oma ##in, which is considered more detailed than the meaningful word chunks of AccountDomain in the data log language. The use of a tokenizer is also conceivable without departing from the scope of the present disclosure.

[0044] It may also be possible to configure a parser to move at network speed to keep up with the high volume of generated data logs. In some embodiments, preprocessing, tokenization, and / or postprocessing may be performed on a graphics processing unit to achieve faster parsing without requiring back-and-forth communication with host memory. However, it should be appreciated that a central processing unit (CPU) or other type of processing architecture may also be used without departing from the scope of the present disclosure.

[0045] With reference to the Fig. 1-6, an illustrative computer system 100 is described in accordance with at least some embodiments of the present disclosure. A computer system 100 may include a communications network 104 configured to facilitate communication between machines. In some embodiments, the communications network 104 may enable communication between different types of machines, which may also be referred to herein as data sources 112. One or more of the data sources 112 may be provided as part of a common network infrastructure, meaning that the data sources 112 may be owned and / or operated by a common entity. In such a situation, the company that owns and / or operates the network containing the data sources 112 may be interested in obtaining data logs from the various data sources 112.

[0046] Non-limiting examples of data sources 112 may include communication endpoints (e.g., user devices, personal computers (PCs), computing devices, communication devices, point of service (PoS) devices, laptops, phones, smartphones, tablets, wearables, etc.), network devices (e.g., routers, switches, servers, network access points, etc.), network border devices (e.g., firewalls, session border controllers (SBCs), network address translators (NATs), etc.), security devices (access control devices, card readers, biometric readers, locks, doors, etc.), and sensors (e.g., proximity sensors, motion sensors, light sensors, sound sensors, biometric sensors, etc.). A data source 112 may alternatively or additionally include a data storage area for storing data logs generated by various other devices connected to the communication network 104.The data storage area may correspond to a location or device type used to temporarily store data logs until a processing system 108 is ready to retrieve and process the data logs.

[0047] In some embodiments, a processing system 108 is provided that receives data logs from the data sources 112 and analyzes the data logs for the purpose of analyzing the content contained in the data logs. The processing system 108 may be executed on one or more servers that are also connected to the communications network 104. The processing system 108 may be configured to analyze data logs and then evaluate / analyze the analyzed data logs to determine whether the information contained in the data logs includes triggerable data events. The processing system 108 is depicted as a single component in the system 100 for ease of discussion and understanding. It should be understood that the processing system 108 and its components (e.g.,Processor 116, circuitry 124, and / or memory 128) may be deployed in any number of computer architectures. For example, processing system 108 may be deployed as a server, as a collection of servers, as a collection of (server) blades within a single server, on bare metal, in the same premises as data sources 112, in a cloud architecture (enterprise cloud or public cloud), and / or via one or more virtual machines.

[0048] Non-limiting examples of a communication network 104 include an Internet Protocol (IP) network, an Ethernet network, an InfiniBand (IB) network, a FibreChannel network, the Internet, a cellular telephone communication network, a wireless communication network, combinations thereof (e.g., Fibre Channel over Ethernet), variations thereof, and the like.

[0049] As previously mentioned, data sources 112 may be considered host devices, servers, network devices, data storage devices, security devices, sensors, or combinations thereof. Note that data source(s) 112 may be assigned at least one network address, and the format of the assigned network address may depend on the type of network 104.

[0050] The processing system 108 is illustrated as including a processor 116 and a memory 128. Although the processing system 108 includes only a processor 116 and a memory 128, the processing system 108 may include one or more processing devices and / or one or more storage devices. The processor 116 may be configured to execute instructions stored in the memory 128 and / or the neural network 132 stored in the memory 128. As some non-limiting examples, the memory 128 may correspond to any suitable type of storage device or collection of storage devices configured to store instructions and / or commands. Non-limiting examples of suitable storage devices that may be used for the memory 128 include flash memory, random access memory (RAM), read-only memory (ROM), variations thereof, combinations thereof, or the like.In some embodiments, memory 116 and processor 128 may be integrated into a common device (e.g., a microprocessor may include integrated memory).

[0051] In some embodiments, processing system 108 may have processor 116 and memory 128 configured as a GPU. Processor 116 may include one or more circuits 124 configured to execute a neural network 132 stored in memory 128. Alternatively or additionally, processor 116 and memory 128 may be configured as a CPU. A GPU configuration may enable parallel operations on multiple data sets, which may facilitate real-time processing of one or more data logs from one or more data sources 112. When configured as a GPU, circuits 124 may be designed with thousands of concurrently running processor cores, each core dedicated to performing efficient computations.Further details of a suitable, but non-limiting, example of a GPU architecture that may be used to execute the neural network(s) 132 are described in U.S. Patent Application No. 16 / 596,755 to Patterson et al., entitled "GRAPHICS PROCESSING UNIT SYSTEMS FOR PERFORMING DATA ANALYTICS OPERATIONS IN DATA SCIENCE," the entire contents of which are hereby incorporated by reference.

[0052] Regardless of whether configured as a GPU and / or CPU, the circuitry 124 of the processor 116 can be configured to execute the neural network(s) 132 in a highly efficient manner, thereby enabling real-time processing of data logs received from various data sources 112. As the data logs are processed / analyzed by the processor 116 executing the neural network(s) 132, the outputs of the neural networks 132 can be provided to a data log memory 140. In some embodiments, as various data logs in different data formats and structures are processed by the processor 116 executing the neural network(s) 132, the outputs of the neural network(s) 132 can be stored in the data log memory 140 as a combined data log 144.The combined data log 144 may be stored in any format suitable for storing data logs or information from data logs. Non-limiting examples of formats used to store a combined data log 144 include spreadsheets, tables, delimited files, text files, and the like.

[0053] The processing system 108 may also be configured to analyze the data log(s) stored in the data log storage 140 (e.g., after the data logs received directly from the data sources 112 have been processed / analyzed by the neural network(s) 132). The processing system 108 may be configured to analyze the data log(s) individually or as part of the combined data log 144 by performing a data log evaluation 136 with the processor 116. In some embodiments, the data log evaluation 136 may be performed by a different processor 116 than the one used to execute the neural networks 132. Likewise, the storage devices used to store the neural network(s) 132 may, but need not, be the same storage devices used to store the data log evaluation 136 instructions.In some embodiments, the data log analysis 136 is stored in a different storage device 128 than the neural network(s) 132 and may be executed using a CPU architecture as compared to using a GPU architecture to execute the neural networks 132.

[0054] In some embodiments, when performing data log analysis 136, processor 116 may be configured to analyze combined data log 144, detect a triggerable event based on the analysis of combined data log 144, and communicate the triggerable event to a communication device 148 of system administrator 152. In some embodiments, the triggerable event may correspond to the detection of a network threat (e.g., an attack on computer system 100, the presence of malicious code in computer system 100, a phishing attempt in computer system 100, a data breach in computer system 100, etc.), a data anomaly, a user behavior anomaly in computer system 100, a application behavior anomaly in computer system 100, a device behavior anomaly in computer system 100, etc.

[0055] If the processor 116 detects a triggerable data event while executing the data log analysis 136, a report or alert may be sent to the communication device 148 operated by a system administrator 152. The report or alert for the communication device 148 may include identification of the machine / data source 112 that resulted in the triggering data event. The report or alert may alternatively or additionally include information about the time at which the data log was created by the data source 112 that resulted in the vulnerable data event. The report or alert may be delivered to the communication device 148 in the form of one or more electronic messages, an email, a Short Message Service (SMS) message, an audible alert, a visible alert, or the like. The communication device 148 may be any type of network-connected device (e.g.,PC, laptop, smartphone, mobile phone, portable device, PoS device, etc.) configured to receive electronic messages from the processing system 108 and to reproduce information from the electronic messages for a system administrator 152.

[0056] In some embodiments, the data log analysis 136 may be provided as a set of alert analysis instructions stored in memory 128 and executable by the processor 116. A non-limiting example of the data log analysis 136 is shown below:

[0057] The data log 136 evaluation code shown above, when executed by processor 116, may enable processor 116 to read cyber alerts, aggregate cyber alerts by day, and calculate the rolling Z-score value over multiple days to look for outliers in the set of alerts.

[0058] With reference to Fig. 2 and Fig. 3, further details of an architecture and method for training a neural network will now be described in accordance with at least some embodiments of the present disclosure. A neural network under training 224 may be trained by a training engine 220. After sufficient training, the training engine 220 may ultimately generate a trained neural network 132, which may be stored in memory 128 of the processing system 108 and used by the processor 116 to process / analyze data logs from data sources 112.

[0059] In some embodiments, the training engine 220 may receive tokenized inputs 216 from a tokenizer 212. The tokenizer 212 may be configured to receive training data 208a-N from a variety of different machine types 204a-N. In some embodiments, each machine type 204a-N may be configured to generate a different type of training data 208a-N, which may be in the form of a raw data log, a parsed data log, a partial data log, a degraded data log, a portion of a data log, or a data log that has been split into multiple parts.In some embodiments, each machine 204a-N may correspond to a different data source 112, and one or more of the different types of training data 208a-N may be in the form of a raw data log from a data source 112, an analyzed data log from a data source 112, or a partial data log. While some training data 208a-N is received as a raw data log, other training data 208a-N may be received as an analyzed data log.

[0060] In some embodiments, tokenizer 212 and training engine 220 may be configured to jointly process the training data 208a-N received from the various machine types 204a-N. Tokenizer 212 may correspond to a partial word tokenizer that supports non-truncation of logs / sentences. Tokenizer 212 may be configured to return an encoded tensor, an attention mask, and metadata to reform corrupted data logs. Alternatively or additionally, tokenizer 212 may correspond to a partial word tokenizer, a partial sentence tokenizer, a character-based tokenizer, or any other suitable tokenizer capable of tokenizing data logs into tokenized inputs 216 for training engine 220.

[0061] As a non-limiting example, the tokenizer 212 and the training function unit 220 may be configured such that neural networks are trained and tested in training 224 on entire data logs that are all small enough to fit into an input sequence and achieve a micro-F1 score of 0.9995. However, a model trained in this manner may not be able to analyze data logs larger than the maximum model input sequence, and the model's performance may suffer if the data logs from the same test set have been altered to have variable starting positions (e.g., micro-F1: 0.9634) or have been cut into smaller pieces (e.g., micro-F1: 0.9456). To prevent the neural network from learning the absolute positions of the fields during model training 224, it may be possible to train the neural network on pieces of data logs during training 224.It may also be desirable to train the neural network model during training 224 on variable starting points in data logs, degraded data logs, and data logs or log chunks of variable length. In some embodiments, training function unit 220 may include functions that allow training function unit 220 to adjust one, some, or all of these properties of training data 208a-N (or tokenized input 216) to improve the training of the neural network during training 224 of the model. In particular, but without limitation, training function unit 220 may include one or more components that allow for mixing of training data 228, starting point variation 232, degradation of training data 236, and / or length variation 240.Adjustments to the training data can lead to similar accuracy as with the fixed starting positions, and the resulting trained neural network(s) 132 can perform well on log pieces with variable starting positions (e.g., micro-F1: 0.9938).

[0062] A robust and effective trained neural network 132 can be achieved if the training function unit 220 trains the neural network on data log pieces during model training 224. The test accuracy of a trained neural network 132 can be measured by splitting each data log into overlapping data log pieces before inference, then reassembling them, and taking predictions from the middle half of each data log piece. This allows the model to have the most context in both directions for inference. With proper training, the trained neural network 132 may exhibit the ability to analyze data log types outside the training set (e.g., data log types that differ from the types of training data 208a-N used to train the neural network 132).When trained with only 1000 examples of each of the nine different Windows event log types, a trained neural network 132 can be configured to accurately analyze a never-before-seen Windows event log type or a data log from a non-Windows data source 112 (e.g., micro-F1: 0.9645).

[0063] Fig. 3 shows an illustrative but non-limiting method 300 for training a neural network, which may correspond to a language-based neural network. The method 300 may be used to train an NLP machine learning model, which is an example of a neural network in training 224. The method 300 may be used to start with a pre-trained NLP model that was originally trained on a data corpus in a particular language (e.g., English, Japanese, German, etc.). When training a pre-trained NLP model (sometimes referred to as fine-tuning), the training engine 220 may update the internal weights and / or layers of the neural network in training 224. The training engine 220 may also be configured to add a classification layer to the trained neural network 132.Alternatively, the method 300 can also be used to train a model from scratch. Training a model from scratch can benefit from using multiple data sources 112 and many different machine types 204a-N, each of which provides different types of training data 208a-N.

[0064] Regardless of whether the method 300 begins with fine-tuning an already trained model or with a fresh start, it may begin by obtaining initial training data 208a-N (step 304). The training data 208a-N may originate from one or more machines 204a-N of different types. Although in Fig. 2, more than three different types of machines 204a-N are shown, the training data 208a-N may originate from a greater or fewer number of different types of machines 204a-N. In some embodiments, the number N of different machine types may correspond to an integer value greater than or equal to one. Furthermore, the number of types of training data does not necessarily have to correspond to the number N of different machine types. For example, two different machine types may be configured to produce the same or similar types of training data.

[0065] The method 300 may continue by determining whether additional training data or other types of training data 208a-N are desired for the neural network during training 224 (step 308). If this question is answered affirmatively, the additional training data 208a-N is obtained from the appropriate data source 112, which may correspond to a different machine type 204a-N than the original training data.

[0066] Thereafter, or if the query in step 308 is answered negatively, the method 300 continues with the tokenizer 212, which tokenizes the training data and generates a tokenized input 216 for the training function unit 220 (step 316). It should be understood that the tokenization step may correspond to an optional step and is not required to sufficiently train a neural network in training 224. In some embodiments, the tokenizer 212 may be configured to provide a tokenized input 216 that tokenizes the training data through embedding, word separation, and / or positional encoding.

[0067] The method 300 may also include an optional step of dividing the training data into data log chunks (step 320). The size of the data log chunks may be selected based on a maximum size of the memory 128 ultimately used in the processing system 108. The optional dividing step may be performed before or after the training data is tokenized by the tokenizer 212. For example, the tokenizer 212 may receive training data 208a-N that has already been divided into data log chunks of a suitable size. In some embodiments, it may be possible to provide the training engine 220 with log chunks of different sizes.

[0068] In addition to optionally adjusting the size of the data log chunks used to train the neural network in training 224, the method 300 may also provide the ability to adjust other training parameters. Thus, the method 300 may proceed to determine whether or not to use other adjustments for training the neural network in training 224 (step 324). Such adjustments may include, without limitation, adjusting training by: (i) shuffling training data 228; (ii) varying a starting point of the training data 232; (iii) degrading at least some of the training data 236 (e.g., introducing errors into the training data or deleting some portions of the training data); and / or (iv) varying the length of the training data 240 or portions thereof (step 328).

[0069] The training function unit 220 may train the neural network in training 224 on the various types of training data 208a-N until it is determined that the neural network in training 224 is sufficiently trained (step 332). The determination of whether or not the training is sufficient / complete may be based on a time component (e.g., whether or not the neural network in training 224 has trained on the training data 208a-N for at least a predetermined time). Alternatively or additionally, determining whether or not the training is sufficient / complete may include analyzing a performance of the neural network in training 224 with a new data log that was not included in the training data 208a-N to determine whether the neural network in training 224 is capable of analyzing the new data log with at least a minimum required accuracy.Alternatively or additionally, determining whether training is sufficient / complete may involve requesting and receiving human input indicating that training is complete. If the request is answered negatively in step 332, method 300 continues training (step 336) and returns to step 324.

[0070] If the query is answered affirmatively in step 332, the neural network in training 224 may be output from the training engine 220 as a trained neural network 132 and stored in memory 128 for subsequent processing of data logs from data sources 112 (step 340). In some embodiments, additional feedback (human feedback or automated feedback) may be received based on the processing / analysis (parsing) of actual data logs by the neural network 132. This additional feedback may be used to further train or fine-tune the neural network 132 outside of a formal training process (step 344).

[0071] In the Fig. 4-6, further details of using one or more trained neural networks 132 to process or analyze data logs from data sources 112 are described in accordance with at least some embodiments of the present disclosure. Fig. 4 shows an illustrative architecture in which the trained neural network(s) 132 may be deployed. In the illustrated example, a plurality of different device types 404a-M provide data protocols 408a-M to the trained neural network(s) 132. The different device types 404a-M may, but do not have to, correspond to different data sources 112. In some embodiments, the first device type 404a may be different from the second device type 404b, and each device may be configured to provide data protocols 408a and 408b, respectively, to the trained neural network(s) 132. As described above, the neural network (or networks) 132 may be trained to process speech-based inputs and, in some embodiments, may include an NLP machine learning model.

[0072] One, some, or all of the data logs 408a-M may be received in a format native to the device type 404a-M that generated the data logs 408a-M. For example, the first data log 408a may be received in a format native to the first device type 404a (e.g., raw data format), the second data log 408b may be received in a format native to the second device type 404b, the third data log 408c may be received in a format native to the third device type 404c, ..., and the Mth data log 408M may be received in a format native to the Mth device type 404M, where M is an integer value greater than or equal to one. The data logs 408a-M do not necessarily have to be provided in the same format. Rather, one or more of the 408a-M data protocols may be provided in a different format than the other 408a-M data protocols.

[0073] The data logs 408a-M may correspond to full data logs, partial data logs, degraded data logs, raw data logs, or combinations thereof. In some embodiments, one or more of the data logs 408a-M may correspond to alternative representations or structured transformations of a raw data log. For example, one or more of the data logs 408aM provided to the neural network(s) 132 may include deduplicated data logs, summaries of data logs, sanitized data logs (e.g., data logs from which sensitive / personally identifiable information (PII) has been removed or obfuscated), combinations thereof, and the like. In some embodiments, one or more of the data logs 408a-M are received in a data stream directly from the data source 112 that generated the data log.For example, the first device type 404a may correspond to a data source 112 that transmits the first data protocol 408a as a data stream using any communication protocol suitable for transmitting data protocols over the communication network 104. As a more specific but non-limiting example, one or more of the data protocols 408a-M may correspond to a cyber protocol containing security data transmitted from one machine to another machine over the communication network 104.

[0074] Because the data log(s) 408a-M may be provided to the neural network 132 in a native format, the data log(s) 408a-M may include various types of data or data fields generated by a machine communicating over the communication network 104. One or more of the data logs 408a-M may include, for example, a file path name, an Internet Protocol (IP) address, a Media Access Control (MAC) address, a timestamp, a hexadecimal value, a sensor reading, a user name, an account name, a domain name, a hyperlink, host system metadata, connection duration information, communication protocol information, communication port identification, and / or a raw data payload. The type of data contained in the data log(s) 408a-M may depend on the type of device 404a-M that generates the data log(s) 408a-M.For example, a data source 112 corresponding to a communication endpoint may include application information, user behavior information, network connection information, etc. in a data log 408, while a data source 112 corresponding to a network device or network edge device may include network connectivity information, network behavior information, quality of service (QoS) information, connection times, port usage, etc.

[0075] In some embodiments, the data logs 408a-M may first be fed to a preprocessing stage 412. The preprocessing stage 412 may be configured to tokenize one or more of the data logs 408a-M before forwarding the data logs to the neural network 132. The preprocessing stage 412 may include a tokenizer, similar to tokenizer 212, that enables the preprocessing stage 412 to tokenize the data log(s) 408a-M using word embedding, split words, and / or positional encoding.

[0076] The preprocessing stage 412 may also be configured to perform other preprocessing tasks, such as dividing a data log 408 into a plurality of data log chunks and then providing the data log chunks to the neural network 132. The data log chunks may vary in size and may or may not overlap. For example, one data log chunk may have some overlap or common content with another data log chunk. The maximum size of the data log chunks may be determined based on the limitations of the memory 128 and / or the processor 116. Alternatively or additionally, the size of the data log chunks may be determined based on the size of the training data 232 used during training of the neural network 132.The preprocessing stage 412 may alternatively or additionally be configured to perform preprocessing techniques including deduplication processing, summarization processing, scrubbing / obfuscation of sensitive data, etc.

[0077] It should be noted that the data logs 408a-M are not necessarily complete or error-free. In other words, if the neural network 132 has been adequately trained, it may be possible for the neural network 132 to successfully analyze incomplete data logs 408a-M and / or degraded data logs 408a-M that are missing at least some information that was included when the data logs 408a-M were generated at the data source 112. Such losses may occur due to network connectivity issues (e.g., lost packets, delay, noise, etc.), and therefore, it may be desirable to train the neural network 132 to account for the possibility of incomplete data logs 408a-M.

[0078] The neural network 132 may be configured to analyze the data log(s) 408a-M and produce an output 416 that may be stored in the data log storage 140. For example, the neural network 132 may provide an output 416 that includes reconstituted complete key / value values ​​of the various data logs 408a-M that were analyzed. In some embodiments, the neural network 132 may analyze data logs 408a-M with different formats, whether these formats are known or unknown to the neural network 132, and produce an output 416 that represents a combination of the various data logs 408a-M.When the neural network 132 analyzes different data logs 408a-M, the output generated by the neural network 132 based on the analysis of the individual data logs 408a-M may be stored in a common data format as part of the combined data log 144.

[0079] In some embodiments, the output 416 of the neural network 132 may correspond to an entry for the combined data log 144, a set of entries for the combined data log 144, or new data to be referenced by the combined data log 144. The output 416 may be stored in the combined data log 144 to allow the processor 116 to perform the data log analysis 136 and search the combined data log 144 for actionable events.

[0080] With reference to the Fig. 4 and Fig. 5, a method 500 for processing data logs 408a-M will now be described in accordance with at least some embodiments of the present disclosure. The method 500 may begin with receiving data logs 408a-M from various data sources 112 (step 504). One or more of the data sources 112 may correspond to a first device type 404a, others of the data sources 112 may correspond to a second device type 404b, others of the data sources 112 may correspond to a third device type 404c, ..., while still others of the data sources 112 may correspond to an Mth device type 404M. The various data sources 112 may provide data logs 408a-M of different types and / or formats, which may be known or unknown to the neural network 132.

[0081] The method 500 may continue with preprocessing the data log(s) 408a-M in the preprocessing stage 412 (step 508). The preprocessing may include tokenizing one or more of the data logs 408a-M and / or dividing one or more of the data logs 408a-M into smaller data log pieces. The preprocessed data logs 408a-M may then be fed to the neural network 132 (step 512), where the data logs 408a-M are analyzed (step 516).

[0082] Based on the analysis step, the neural network 132 may generate an output 416 (step 520). The output 416 may be provided in the form of a combined data log 144, which may be stored in the data log memory 140 (step 524).

[0083] The method 500 may continue by enabling the processor 116 to analyze the data log memory 140 and the data contained therein (e.g., the combined data log 144) (step 528). The processor 116 may analyze the data log memory 140 by executing the data log evaluation 136 stored in the memory 128. Based on the analysis of the data log memory 140 and the data contained therein, the method 500 may continue by determining whether a triggerable data event has been detected (step 532). If the query is answered affirmatively, the processor 116 may be configured to generate an alert message that is transmitted to a communication device 148 operated by a system administrator 152 (step 536).The alert message may include information describing the triggerable data event, possibly including the data log 408 that triggered the triggerable data event, the data source 112 that generated the data log 408 that triggered the triggerable data event, and / or whether other data anomalies with some relationship to the triggerable data event were detected.

[0084] Thereafter, or if the query in step 532 is answered negatively, the method 500 may continue with the processor 116 waiting for another change in the data log memory 140 (step 540), which may or may not be based on the receipt of a new data log in step 504. In some embodiments, the method may return to step 504 or to step 528.

[0085] In Fig.6, a method 600 for preprocessing data logs 408 will now be described in accordance with at least some embodiments of the present disclosure. The method 600 may begin when one or more data logs 408a-M are received in the preprocessing stage 412 (step 604). The data logs 408a-M may correspond to raw data logs, parsed data logs, degraded data logs, lossy data logs, incomplete data logs, or the like. In some embodiments, the data logs 408a-M received in step 604 may be received as part of a data stream (e.g., an IP data stream).

[0086] The method 600 may continue with the preprocessing stage 412 determining that at least one data log 408 should be split into log chunks (step 608). Following this determination, the preprocessing stage 412 may split the data log 408 into log chunks of appropriate size (step 612). The data log 408 may be split into equal-sized log chunks, or the data log 408 may be split into log chunks of different sizes.

[0087] Thereafter, the preprocessing stage 412 may provide the data log chunks to the neural network 132 for analysis (step 616). In some embodiments, the size and variability of the data log chunks may be selected based on the characteristics of the training data 208a-N used to train the neural network 132.

[0088] In the description, specific details have been set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to one skilled in the art that the embodiments may be practiced without these specific details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be set forth without unnecessary detail in order not to obscure the embodiments.

[0089] While illustrative embodiments of the disclosure have been described in detail herein, it is to be understood that the inventive concepts may be embodied and utilized in other ways, and the appended claims are intended to be construed to include such variations unless limited by the prior art.

[0090] It is to be understood that the aspects and embodiments described above are exemplary only and that changes in detail may be made within the scope of the claims.

[0091] Each device, method, and feature disclosed in the description and (where appropriate) in the claims and drawings may be provided independently or in any suitable combination.

[0092] The reference signs contained in the claims are for illustrative purposes only and do not limit the scope of the claims.

Claims

[1] A method for processing data logs, the method comprising: Receiving a data log from a data source, wherein the data log is received in a format native to a machine that generated the data log; Providing the data protocol for a neural network trained to process natural language inputs; Analyzing the data log with the neural network; Receiving an output from the neural network, wherein the output is generated by the neural network in response to the neural network analyzing the data protocol; and Storing the output of the neural network in a data log memory. [2] The method of claim 1, further comprising: Receiving an additional data protocol from an additional data source, wherein the additional data source is different from the data source and wherein the additional data protocol is received in a second format that is native to the additional data source; Providing the additional data protocol for the neural network; Analyzing the additional data protocol with the neural network; Receiving an additional output from the neural network, wherein the additional output is generated by the neural network in response to the neural network analyzing the additional data protocol; and Storing the additional output of the neural network in the data log memory. [3] The method of claim 2, wherein the output and the additional output are stored in the data log memory in a common data format as part of a combined data log. [4] A method according to claim 2 or 3, wherein the additional data protocol is received as a data stream directly from the additional data source. [5] The method of claim 2, 3 or 4, wherein the machine that generated the data log comprises a first device type, wherein the additional data source comprises a second device type, and wherein the first device type and the second device type belong to a common network infrastructure. [6] A method according to any preceding claim, wherein the machine that generated the data log comprises at least one of a communications endpoint, a network device, a network boundary device, a security device, or a sensor. [7] A method according to any one of the preceding claims, wherein the data protocol comprises security data transmitted from the machine to another machine, and wherein the neural network comprises a machine learning model for natural language processing. [8] Method according to one of the preceding claims, further comprising: Dividing the data log into a plurality of data log pieces; and Providing the plurality of data log chunks to the neural network, wherein the neural network is trained with training data comprising log chunks, and wherein a size of one log chunk in the plurality of data log chunks is different from a size of another log chunk in the plurality of data log chunks. [9] Method according to one of the preceding claims, further comprising: Analyzing the data log storage; based on the analysis of the data log memory, detecting a triggerable data event; and Providing an alert message to a communications device, the alert message containing information describing the triggerable data event. [10] The method of any preceding claim, wherein the data log comprises at least one of a file path name, an Internet Protocol (IP) address, a Media Access Control (MAC) address, a timestamp, a hexadecimal value, a sensor reading, a user name, an account name, a domain name, a hyperlink, metadata of a host system, a duration of the connection, a communication protocol, a communication port, and raw payload data. [11] A method according to any one of the preceding claims, wherein the data log comprises at least one of a degraded log and an incomplete log. [12] Data log processing system comprising: a processor; and Memory coupled to the processor, the memory storing data that, when executed by the processor, enables the processor to: receive a data log from a data source, wherein the data log is received in a format that is native to a machine that generated the data log; analyze the data log using a neural network trained to process natural language inputs; and storing an output from the neural network in a data log memory, wherein the output from the neural network is generated in response to the neural network analyzing the data log. [13] The system of claim 12, wherein the data stored in the memory further enables the processor to tokenize the data log prior to analyzing the data log with the neural network. [14] The system of claim 12 or 13, wherein the data stored in the memory further enables the processor to: receive an additional data protocol from an additional data source, wherein the additional data source is different from the data source and wherein the additional data protocol is received in a second format that is native to the additional data source; to analyze the additional data protocol with the neural network; and storing an additional output from the neural network in the data log memory, the additional output from the neural network being generated in response to the neural network analyzing the additional data log. [15] The system of claim 14, wherein the output and the additional output are stored in the data log memory in a common data format as part of a combined data log. [16] A system according to claim 14 or 15, wherein the additional data protocol is received as a data stream directly from the additional data source. [17] The system of claim 14, 15 or 16, wherein the machine that generated the data log comprises a first device type, wherein the additional data source comprises a second device type, and wherein the first device type and the second device type belong to a common network infrastructure. [18] The system of any one of claims 12 to 17, wherein the data protocol comprises security data communicated from the machine to another machine, and wherein the neural network comprises a machine learning model for natural language processing. [19] The system of any one of claims 12 to 17, wherein the data stored in the memory further enables the processor to: analyze the data log storage; to detect a triggerable data event based on the analysis of the data log memory; and provide an alert message to a communications device, the alert message containing information describing the triggerable data event. [20] The system of any of claims 12 to 19, wherein at least one of the processor and the memory is provided in a graphics processing unit. [21] The system of any of claims 12 to 20, wherein the data log comprises at least one of a degraded log and an incomplete log. [22] A method for training a system for processing data logs, the method comprising: Providing a neural network with first training data, wherein the neural network comprises a machine learning model for natural language processing, and wherein the first training data comprises a first data log generated by a first type of machine; Providing second training data to the neural network, the second training data comprising a second data log generated by a second type of machine; Determining that the neural network has trained with the first training data and the second training data for at least a predetermined time; and Storing the neural network in computer memory so that the neural network is available to process further data protocols. [23] The method of claim 22, wherein the first data log comprises at least one of a raw data log and an analyzed data log, wherein the first data log is tokenized with at least one of a word embedding, separated words, and position encoding, and wherein the method further comprises: Adjusting a training of the neural network by at least one of: (i) mixing the first training data and the second training data; (ii) varying a starting point of the first training data; (iii) varying a starting point of the second training data; and (iv) degrading at least one of the first training data and the second training data. [24] Processor comprising: one or more circuits for using one or more natural language-based neural networks to analyze one or more machine-generated data logs. [25] The processor of claim 24, wherein the one or more circuits are configured to: receive the one or more machine-generated data protocols from a data source; and generate an output in response to analyzing the one or more machine-generated data logs, wherein the output is configured to be stored as part of a data log store. [26] A processor according to claim 24 or 25, wherein the one or more machine-generated data logs are received as part of a data stream. [27] The processor of claim 24, 25 or 26, wherein the one or more machine-generated data logs comprise at least one of a degraded log, an incomplete log, a deduplicated log, a log summary, a log with obfuscated sensitive information, and a partial data log.

Citation Information

Patent Citations

  • SYSTEMS AND METHODS FOR RISK ANALYSIS

    DE112018004325T5

  • ANOMALITY DETECTION USING COGNITIVE COMPUTING

    DE112018005462T5

  • Natural language processing artificial intelligence network and data security system

    EP3468140B1