Host-level data analysis for cyber-attack detection

JP2024521739A5Pending Publication Date: 2025-05-27MANDEX INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023572099
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-05-21
Filing Date
2022-05-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional cyber-attack detection methods focus on network-level analysis, missing sophisticated attacks that occur internally within host computer systems, allowing malicious actors to remain undetected and compromise systems for extended periods.

Method used

An internal system for host computer systems analyzes performance patterns to detect cyber-attacks by establishing a baseline of normal behavior, collecting and comparing system performance data from compromised and non-compromised systems, using statistical analysis and logistic regression to identify cyber-attack signatures.

Benefits of technology

Enhances the detection of cyber-attacks by identifying internal system indicators, allowing for timely intervention and prevention of data theft or system disruption, with high accuracy and reduced false positives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A host computer system may be monitored to track its system performance with respect to internal system parameters, and this monitoring may be performed when the host computer system is known to be under a cyber attack and when the host computer system is known not to be under a cyber attack. The host's system performance data in these conditions may be compared and analyzed by host level data analysis to find a subset of the internal system parameters and their corresponding data values ​​that discernibly correlate to a cyber attack. From this information, a cyber attack signature may be generated. The host system may then be monitored based on its system performance data to determine whether the system performance data matches the cyber attack signature, which may assist in determining whether the host is under a cyber attack.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE AND PRIORITY CLAIM FOR RELATED PATENT APPLICATIONS This patent application claims priority to U.S. Provisional Patent Application No. 63 / 191,464, entitled “Host Level Data Analytics for Cyberattack Detection,” filed May 21, 2021, the entire disclosure of which is incorporated herein by reference. [Background technology]

[0002] introduction: Conventional approaches to detecting cyber attacks on computer systems using data analysis focus on the network footprint of the cyber attack. Thus, conventional approaches to data analysis of cyber attacks perform pattern analysis at the network level of the computer system by focusing the analysis on computer network traffic to and from the host computer system, as shown at 150 in FIG. 1. This analysis often involves deep packet inspection that examines the sources and destinations of packet traffic to and from the host computer system 100, and possibly the contents of the payload of such packet traffic. Thus, conventional approaches to cyber attack detection observe activity that occurs at the edge of the computer system, such as network traffic, network traffic logs, and system events, for network and border devices to discover anomalies.

[0003] However, the inventor believes that improvements in the art are needed to enable systems to better detect cyber attacks. Many sophisticated cyber attacks allow malicious actors to surreptitiously slip through traditional network-level cyber defenses and remain undetected on host computer systems for extended periods of time while collecting and stealing valuable confidential information or data from the host computer systems.

[0004] Summary of the Invention To satisfy the need in the art for new approaches to cyber-attack detection, the present inventors focus on the internal performance of a host computer system, as shown at 152 in FIG. 1, to detect system performance patterns indicative of a cyber-attack. As used herein, a "cyber-attack" refers to an attempt to damage, disrupt, and / or gain unauthorized access to a computer, computer system, or electronic communications network. Examples of cyber-attacks include the deployment or installation of malware on a computer system that operates to extract information, disable access to information, and / or modify system applications without authorization, whether for financial, destructive, or other purposes.

[0005] The inventor has developed a technique that, rather than observing the external interactions and characteristics of a host computer system, observes changes in the host computer system internally and analyzes the internal performance of the host computer system to determine whether the host computer system is under a cyber attack.

[0006] In this approach, quantitative and exploratory analysis of datasets representative of the performance of host computer systems can facilitate the identification of resource usage indicators of cyber attacks.

[0007] As part of this approach, a baseline of the host computer system's behavior may be established by collecting system performance data for host computer systems that are known not to have been compromised by cyber attacks. Such host computer systems may be referred to as normal-state host computer systems. System performance data collected from normal-state host computer systems may be referred to as normal-state system performance data and may serve as a baseline control for host-level data analysis. The normal-state system performance data may include data values ​​over time for a number of different parameters, each of which represents a different aspect of the host computer system during operation.

[0008] Also, system performance data may be collected from a host computer system known to have been compromised by malware as a result of a cyber-attack. Such a host computer system may be referred to as an attacked host computer system, and system performance data collected therefrom may be referred to as attacked system performance data. The attacked system performance data may include data values ​​over time for the same parameters used for the normal system performance data.

[0009] Statistical analysis may be performed on the normal and attacked system performance data to identify system performance parameters and parameter values ​​that are discriminatively correlated to cyber attacks. Values ​​of various parameters in the normal and attacked system performance data may be evaluated as variables for positive and negative cyber attack theories. Logistic regression may then identify system indicators that indicate the presence of a cyber attack. These system indicators may then serve as cyber attack signatures for the host computer system.

[0010] A host computer system in an unknown cyber-attack state may then have its internal system performance parameters tested against the cyber-attack signature to determine whether the host computer system is in a cyber-attack compromised state. The host computer system may be referred to as a test host computer system. To accomplish this testing, system performance data may be collected from the test host computer system, which may be referred to as test system performance data. The test system performance data may include data values ​​over time of the same system parameters in the attacked state system performance data (or may include data values ​​over time of at least enough of those system parameters to determine whether they match the cyber-attack signature).

[0011] The test system parameter data may then be compared to the attack signature to determine if there is a pattern match between the two. The existence of a match allows the system to determine that the test host computer system should be reported as positive for the cyber attack. If no match is found, the test host computer system may be reported as negative for the cyber attack.

[0012] Through such host level data analysis, the inventor believes that more options are available so that cyber attacks can be easily detected and countered. For example, enumeration scans such as Nmap and Nessus scans, which often form the preparation phase of a cyber attack, can be detected through the use of such host level data analysis. Timely detection of such enumeration cyber attacks can help prevent loss and damage that may occur in subsequent phases of a cyber attack if the enumeration phase is not detected.

[0013] Additionally, these techniques may be used with multiple different types of cyber attacks to develop a library of cyber attack signatures corresponding to the different types of cyber attacks. Test system performance data may then be compared to any or all of these cyber attack signatures to determine whether a test host computer system has been compromised by a cyber attack.

[0014] These and other features and advantages of the present invention are described in more detail below. [Brief description of the drawings]

[0015] [Figure 1] 1 illustrates an exemplary host computer system. [Diagram 2] 1 illustrates an exemplary process flow for performing host-level data analysis to detect cyber-attacks. [Diagram 3] 1 illustrates an example process flow for performing host-level data analysis to detect any of several different types of cyber-attacks. [Figure 4] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Diagram 5] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 6] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 7] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 8A] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 8B] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 9] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 10] A further exemplary process flow for performing host level data analysis is shown in connection with a host computer system of an exemplary embodiment. [Figure 11] 1 illustrates an example plot of collected system performance data that may indicate the presence of an Nmap enumeration cyber-attack and a Nessus enumeration cyber-attack. [Figure 12] 1 illustrates an example plot of collected system performance data that may indicate the presence of an Nmap enumeration cyber-attack and a Nessus enumeration cyber-attack. [Figure 13] 1 illustrates an example plot of collected system performance data that may indicate the presence of an Nmap enumeration cyber-attack and a Nessus enumeration cyber-attack. [Figure 14] 1 illustrates an example plot of collected system performance data that may indicate the presence of an Nmap enumeration cyber-attack and a Nessus enumeration cyber-attack. [Figure 15] 1 illustrates an example plot of collected system performance data that may indicate the presence of an Nmap enumeration cyber-attack and a Nessus enumeration cyber-attack. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016] FIG. 1 illustrates an exemplary host computer system 100 that may be used in practicing exemplary embodiments of the present invention. The host computer system 100 includes one or more processors 102, which may take the form of one or more CPUs or other suitable computer resources. The host computer system 100 also includes one or more memories 104. The memory 104 may store code executed by the processor(s), such as an operating system (OS), software applications, etc. The memory 104 may also store data, such as files, records, etc., that are processed by the processor(s) 102 during system operation. The memory(s) 104 may include random access memory (RAM) devices, disk drives, or other suitable data storage devices. The host computer system 100 may also include one or more network interface cards 106, via which it may communicate with an external network 120, such as the Internet. The host computer system 100 may also include one or more peripherals 108 or other components, such as I / O devices (keyboards, monitors, printers, etc.), among others. These components of the host computer system 100 may be interconnected and communicate via a system bus 110 or other suitable interconnection technology.

[0017] 1 depicts a highly simplified version of host computer system 100 for purposes of elaboration, it should be understood that more complex system architectures may be deployed. For example, host computer system 100 may be an IT computer system that hosts a wide variety of services for a number of machines in a distributed network, such services may include email applications, word processing applications, and / or other business support or technical support applications.

[0018] Figure 2 illustrates an example process flow for performing host level data analysis to detect cyber attacks against a host computer system such as that shown in Figure 1. The process flow of Figure 2 may be embodied by machine readable code resident in a non-transitory machine readable storage medium, such as memory 104. The code may take the form of software and / or firmware that defines the processing operations discussed herein and is executed by one or more processors 102.

[0019] The process flows of FIG. 2 include an attack signature generation process flow 250 and a cyber-attack detection process flow 252. In an exemplary embodiment, these process flows 250 and 252 can be combined in a software application or other processor-executable code. However, it should be understood that these process flows 250 and 252 can also be implemented by different code, such as different software applications, as desired by the implementer. Furthermore, it should be understood that the process flows 250 and 252 do not have to be executed in conjunction with each other, or even close in time to each other. For example, the process flow 250 may be executed once (or relatively infrequently), while the process flow 252 may be executed multiple times (relatively frequently).

[0020] Generating attack signatures via host-level data analysis: Steps 200 - 206 of the attack signature generation process flow 250 operate to apply host level data analysis to one or more host computer systems 100 to generate an attack signature that can be used by a cyber-attack detection process flow 252 .

[0021] In step 200, the normal / control host runs a performance monitoring application to measure its performance over time across a number of different host system parameters. This measurement data may be referred to as normal system performance data as discussed above. Examples of host system parameters that may be monitored in step 200 may include any combination of the following: CPU Usage. Examples of CPU usage parameters may include measurements of processor switches (e.g. processor switches per unit of time such as seconds), and processor interrupts (e.g. interrupts per unit of time such as seconds), measurements of CPU usage by the operating system (e.g. system percentage (Sys%)), and / or measurements of CPU usage by user applications (e.g. user percentage (User%)). Power Consumption. Examples of power consumption parameters may include measurements of power consumption per unit time of host level system peripherals and / or system resources. Random Access Memory (RAM) Usage: An example of a RAM usage parameter may be a measure of RAM allocation (eg, used memory vs. free memory, etc.). Network card usage. Examples of network card usage parameters may include measurements of incoming network packets (e.g., count, size, etc.) and / or measurements of network traffic (e.g., amount of traffic per unit time (e.g., KB in / sec or KB out / sec, etc.)). User process elapsed time: An example of a measurement of a user process elapsed time parameter may be a measurement of the number of user processes per unit of time, such as seconds. CPU behavior. CPU behavior parameters may include any pattern over time in one or more CPU usage parameters. For example, CPU behavior parameters may be measurements of deviations in processor switches per unit time, process interrupts per unit time, system percentage, and / or user percentage from their respective baseline values. Significant and / or sudden increases in CPU usage reflected in CPU behavior may be used as a factor in detecting a cyber attack. Power behavior. A power behavior parameter may include any pattern over time in one or more power consumption parameters. For example, a power behavior parameter may be a measurement of deviation in power consumption per unit time relative to its baseline value. Significant and / or sudden increases in power usage reflected in the power behavior may be used as a factor in detecting a cyber attack. RAM behavior. The RAM behavior parameter may include any pattern over time in one or more RAM usage parameters. For example, the RAM behavior parameter may be a measure of deviation in memory allocation relative to its baseline value. Significant and / or sudden increases in memory allocation reflected in the RAM behavior may be used as a factor in detecting a cyber attack. Network card behavior. The network behavior parameters may include any pattern over time in one or more network card usage parameters. For example, the network card behavior parameters may be a measurement of deviation in incoming network packets and / or inbound / outbound network packets relative to respective baseline values. Significant and / or sudden increases in network card usage reflected in the network card behavior may be used as a factor in detecting a cyber attack. User behavior. A user behavior parameter may include any pattern over time in one or more user process elapsed time parameters. For example, a user behavior parameter may be a measurement of deviation in user processes per unit time relative to its baseline value. A significant and / or sudden increase in user processes or applications executed on a system by users (which may include unknown users and / or unknown user processes) as reflected in user behavior may be used as a factor in detecting a cyber attack.

[0022] It should be understood that these are merely examples of host system parameters that can be collected, and a practitioner may choose to use more, fewer, and / or different host system parameters when generating an attack signature. Additionally, a practitioner may select an appropriate duration for measuring such host system parameters based on experience and need to detect a particular type of cyber attack. These host system parameters may then serve as features, whose feature values ​​are evaluated to identify an appropriate feature set having feature values ​​that correlate with a dependent outcome (i.e., the presence of a cyber attack). Using logistic regression and model fitting, feature coefficients for the identified host system parameter set may be identified, which are used in a model to model the detection of a cyber attack with respect to the host system parameters of the feature set of interest.

[0023] Based on an estimate of the time span necessary to be able to distinguish between normal and abnormal operating system behavior, the implementer should select the length of time duration that the normal state system performance data should cover. A period such as 10 minutes or more may be used. However, it is understood that some implementers may feel that a longer or shorter period is desirable and / or sufficient.

[0024] In step 202, the under-attack host runs a performance monitoring application to measure its performance over time across a number of different host system parameters. This measurement data may be referred to as under-attack system performance data as discussed above. The host system parameters collected from the under-attack host in step 202 may include the same system parameters as those discussed above with respect to step 200, since the purpose of the two collections in steps 200 and 202 is to compare the internal operating behavior of the under-attack host with the internal operating behavior of the normal host to detect attack indicators based on differences in data sets that are discriminatively correlated to cyber attacks.

[0025] In step 202, the host computer system 100 may be subjected to any of a number of different types of cyber attacks. For example, enumeration scans are often used by malicious actors as a preparatory phase of a cyber attack, where they attempt to monitor the host and learn its structure so that the host's security flaws or weaknesses can be probed. Enumeration scans involve a process of extracting user names, machine names, network resources, and other services present on the host computer system. This information can then be leveraged by the malicious actor in executing later phases of a cyber attack. Examples of enumeration scan tools that may be used in such an enumeration cyber attack include Network Mapper (Nmap) and Nessus. Nmap is an open source network scanner used to discover hosts and services on a computer network by sending packets and analyzing the responses. Nessus is a proprietary scanner that operates in a similar manner. Thus, in an exemplary embodiment, the host computer system 100 may be attacked by an Nmap scanner and / or a Nessus scanner, and step 202 operates to monitor system performance while the Nmap scanner and / or Nessus scanner are operating within the host system.

[0026] However, it will be appreciated that cyber attacks other than enumeration attacks may be used in step 202. The inventors anticipate that the process flow of Figure 2 may be operational with respect to any type of cyber attack and generate attacked system performance data that is sufficiently different from normal system performance data such that the cyber attack can be detected using the techniques discussed below.

[0027] Any of a number of different performance monitoring applications may be used to perform steps 200 and 202. An example of a suitable performance monitoring application is Nmon. Nmon is an open source monitoring application that collects system performance data from a computer system for every second of a specified duration regarding system parameters such as those discussed above. For example, Nmon may record data specific to system parameters such as CPU performance, internal system processes, memory, disk usage, system resources, file systems, and network card activity from a Linux host system. Nmon data files may be collected from the host computer system 100 and imported into a suitable application for analysis (such as the IBM Nmon spreadsheet and analyzer).

[0028] Another example of a suitable performance monitoring application is Collectl. Collectl is a lightweight command line utility that collects system hardware and software data from a computer system for system parameters such as those discussed above, every second for a specified duration. For example, Collectl may record data from a host system specific to system parameters such as CPU performance, internal system processes, disk usage, file system, memory, network card activity, and network protocols. Collectl data files may be collected from the host computer system 100 and imported into a suitable application for analysis (e.g., imported as text files into Microsoft Excel or other spreadsheet program, where the data may be graphed and evaluated for patterns and changes).

[0029] Yet another example of a suitable performance monitoring application is Monitorix. Monitorix is ​​a software application designed to monitor system resources and services within a Linux operating system and can display performance output in a web browser. Monitorix operates to monitor and record specific data over time on CPU usage, power, memory, network cards, network traffic, internal system processes, and system users. Monitorix data files can also be imported into a suitable software application for analysis.

[0030] It should be appreciated that, if desired, steps 200 and 202 may be performed by running multiple performance monitoring applications on the host computer system 100. For example, a practitioner may find it useful to run both Nmap and Collectl on the host computer system 100 to collect normal and under attack system performance data.

[0031] It should further be appreciated that steps 200 and 202 may operate on a clone of host computer system 100 rather than directly on the host computer system itself. Thus, step 202 (and step 200, if the practitioner desires) may involve cloning host computer system 100 and then running the performance monitoring application(s) on the clone host. Such cloning allows the practitioner to avoid having to attack host computer system 100 itself.

[0032] In step 204, the system performs a statistical analysis on the normal state system performance data and the attacked state system performance data. Based on the statistical analysis, system performance parameters and parameter values ​​that are correlated with a cyber-attack can be identified. These system parameters and parameter values ​​can serve as system indicators of a cyber-attack.

[0033] This statistical analysis allows positive and negative infection theories to be tested for different parameters and parameter values ​​of normal and attacked system performance data. In positive predictive value theory, probability statistics may be used to confirm positive indications of a cyber-attack. In negative predictive value theory, a conclusion may be reached that a cyber-attack has not occurred because the system indicators did not reach a predefined threshold that would positively conclude that a cyber-attack has occurred. System parameters and parameter values ​​that serve as indicators that a cyber-attack has occurred or will occur may be identified using logistic regression analysis.

[0034] In this manner, a logistic regression model can be developed that models the probability of a cyber-attack in terms of the feature set and corresponding feature coefficients. The features can be host system parameters that correlate with the presence of a cyber-attack based on comparative statistical analysis of known normal and known attacked system performance data.

[0035] In positive predictive value theory and negative predictive value theory, a practitioner may use logistic regression analysis to test the probability that a cyber attack is present (positive) and the probability that a cyber attack is not present (negative). Using positive predictive value theory, host system parameters and parameter values ​​indicative of the presence of a cyber attack may be identified, and using negative predictive value theory, host system parameters and parameter values ​​indicative of the absence of a cyber attack may be identified.

[0036] Using positive and negative predictive value theory, practitioners may test the probability of a cyber attack on a host computer system by analyzing a combination of host system parameters for symptoms or indicators of a cyber attack (e.g., sudden fluctuations, spikes, or anomalies in performance that have been found to be highly correlated with a cyber attack). Positive and negative predictive value theory also allows host level system data analysis to test whether the system provides positive or negative indicators of a system compromise. By applying positive and negative predictive value theory to evaluate system performance data, performance monitoring results are provided to practitioners to ascertain the positive or negative probability of a possible cyber attack.

[0037] The use of positive predictive value identifies positive indications of a cyber attack by providing probability statistics for the host level system. The use of negative predictive value demonstrates that a cyber attack is not occurring because the host level system indicators do not reach a positive result threshold on the host level system (e.g., a sudden change, spike, or anomaly in performance that is sufficiently correlated with a cyber attack), thereby confirming that the system is not under attack. The use of logistic regression testing for positive predictive value theory and negative predictive value theory derives the probability that the host system is under cyber attack by performing calculations on the positive or negative values ​​and providing a percentage and indicator that the system is compromised, while reducing the amount of potential false positive cyber attack identifications.

[0038] A tool such as IBM SPSS may be used to provide statistical analysis of the normal and under attack system performance data sets, although it should be understood that this is not required and other tools for statistical analysis of the data sets may be used as desired by the practitioner.

[0039] It should be appreciated that process flow 250 may require repeating steps 200, 202, and 204 multiple times to reliably identify and validate system indicators of a cyber attack.

[0040] Next, using the cyber-attack system indicators identified in step 204, a cyber-attack signature may be created in step 206. The cyber-attack signature includes a number of system indicators, which may be expressed in terms of host system parameters and corresponding parameter values ​​(which may include ranges of parameter values), and serves to characterize the presence of a cyber-attack on the host system. In this manner, the cyber-attack signature serves to represent the cyber-attack in terms of its measurable and quantifiable impact on various host system parameters.

[0041] The cyber-attack signatures may be stored in memory 104 and later accessed when testing the host computer system 100 to determine whether a cyber-attack has occurred.

[0042] Detecting cyber attacks using host-level cyber attack signatures: Steps 210-218 of the cyber-attack detection process flow 252 operate to apply host level data analysis to test the host computer system 100 to determine whether the host computer system 100 has been compromised by a cyber-attack that corresponds to the cyber-attack signature created in step 206.

[0043] In step 210, the system triggers the cyber-attack detection process flow 252 to run on the host computer system 100. This host computer system 100 may be referred to as a test host computer system. This trigger may be configured to run on a period or event basis as the implementer may desire. For example, the implementer may select to run the cyber-attack detection process flow 252 every X minutes (e.g., every 10 minutes), or on other time bases (hourly, daily, weekly, etc.). Furthermore, the periods included in the cyber-attack detection process 252 may overlap when the cyber-attack detection process 252 is repeated. The degree of overlap may depend on the time duration characteristics of the cyber-attack signature. For example, if the cyber-attack signature requires a two-minute data value window to detect a cyber-attack from the host's system parameters, the implementer may desire to use two-minute (or slightly more than two-minute) overlapping periods when repeating the detection process 252. This may help to ensure that cyber-attacks that may occur near the time end of the detection process 252 are not missed. As another example, step 210 may trigger the cyber-attack detection process flow 252 upon user request or on some other event-driven basis.

[0044] In another exemplary embodiment, the system may run the cyber-attack detection process 252 continuously, in which case the trigger step 210 may not be required. In a continuous mode of operation, the system essentially constantly examines a sliding window of system performance data from a test host computer system to determine whether a cyber-attack is indicated.

[0045] In step 212, the system runs the performance monitoring application(s) used in steps 200 and 202 on a test host computer system and measures its performance over time across a number of different host system parameters. The system performance data generated in step 212 may be referred to as test system performance data. As discussed above with respect to steps 200 and 202, the system parameters for which data is collected in step 212 may include system parameter measurements indicative of any of the following: ·CPU usage ·Power consumption · Random Access Memory (RAM) used -Network card used -Time spent on each process CPU Behavior Power Behavior RAM behavior Network Card Behavior User Behavior

[0046] Thus, the test system performance data may include a number of host system parameters during operation of the test host computer system and their corresponding values ​​over time. If desired, the practitioner may limit the monitoring and collection in step 212 to only those system parameters necessary to assess whether a cyber-attack signature is present therein.

[0047] The period for collection may be of sufficient duration to allow detection of a cyber-attack by the cyber-attack signature, and the implementer may wish to set the collection period in step 212 in conjunction with the trigger frequency of step 210 so that the detection process 252 can operate on the host at all periods (thus avoiding the risk of a cyber-attack occurring and going undetected during omitted periods).

[0048] As discussed above, examples of performance monitoring applications that may be used in step 212 include Nmon, Collectl, and / or Monitorix.

[0049] In step 214, the system compares the test system performance data to the cyber-attack signature to determine whether a pattern match exists. This comparison may include comparing characteristics of the cyber-attack signature against a sliding window of the test system performance data to determine whether there is any portion of the test system performance data that matches the cyber-attack signature.

[0050] If step 214 results in a match being found between the portion of the test system performance data and the cyber-attack signature, process flow may proceed to step 216, where the system reports the test host computer system as positive for a cyber-attack. This report may trigger an alert on a user interface for a system administrator or other user responsible for securing the host computer system 100. The system may then provide the user with access to a log that provides data describing the detected cyber-attack, including an identification of the time the cyber-attack was detected and the portion of the test system performance data that triggered the match. This may enable the user to take appropriate corrective action if the positive report is deemed accurate.

[0051] If step 214 results in no match being found between the test system performance data and the cyber-attack signature, process flow may proceed to step 218, where the system reports that the test host computer system is negative for a cyber-attack. Negative results may be logged by the system, allowing a system administrator or other user to audit the test results and review these relevant data characteristics as desired.

[0052] It should therefore be appreciated that FIG. 2 describes a technique for detecting a cyber-attack using host-level data analysis in terms of the impact of the cyber-attack on a host's operational performance as compared to a control baseline of the host's normal state operational performance when not under cyber-attack.

[0053] Detection of multiple types of cyber attacks via a cyber attack signature library: In another exemplary embodiment, the process flow of Figure 2 may be expanded to provide the ability to test a host computer system for any of several different types of cyber-attacks. Figure 3 illustrates an exemplary process flow for this. The process flow of Figure 3 may be embodied by machine-readable code resident in a non-transitory machine-readable storage medium, such as memory 104. The code may take the form of software and / or firmware that defines the processing operations discussed herein and is executed by one or more processors 102.

[0054] The process flow of FIG. 3 includes a process flow 350 for generating a plurality of cyber-attack signatures and a cyber-attack detection process flow 352 that operates in conjunction with the plurality of attack signatures. In an exemplary embodiment, these process flows 350 and 352 can be combined in a software application or other processor-executable code. However, it should be understood that these process flows 350 and 352 can also be implemented by different code, such as different software applications, as desired by the implementer. Furthermore, it should be understood that the process flows 350 and 352 do not have to be executed in conjunction with each other or even in close temporal proximity to each other. For example, the process flow 350 may be executed once (or relatively infrequently), while the process flow 352 may be executed multiple times (relatively frequently).

[0055] Process flow 350 includes step 300, which includes performing steps 200-206 of FIG. 2 for a plurality of different types of cyber-attacks. As a result, a plurality of different cyber-attack signatures are created, each having a corresponding cyber-attack type. These cyber-attack signatures may then be stored in memory 104 as a library of cyber-attack signatures 310. For example, for the enumeration cyber-attack discussed above, step 300 may include (1) performing steps 200-206 for an Nmap enumeration cyber-attack to generate a cyber-attack signature for the Nmap enumeration cyber-attack, and (2) performing steps 200-206 for a Nessus enumeration cyber-attack to generate a cyber-attack signature for the Nessus enumeration cyber-attack. Adding the Nmap and Nessus cyber-attack signatures to library 310 may enable the system to detect the presence of either of these types of cyber-attacks via process flow 352.

[0056] Process flow 352 includes steps 310, 312, 314, 316, and 318, which are substantially similar to corresponding steps 210, 212, 214, 216, and 218 of FIG. 2. Thus, step 310 may operate similarly to step 210, and step 312 may operate similarly to step 212. However, it should be understood that the system parameters for which data is collected from the test host computer system in step 312 should be at least a superset of all system parameters necessary to evaluate all cyber-attack signatures in library 310. Step 314 may operate similarly to step 214, except that a matching process may be performed on the test system performance data with respect to each of the multiple cyber-attack signatures in library 310. Thus, if any matches are found from step 314, step 316 may report that the test host computer system is positive for a cyber-attack. Additionally, based on knowledge of which cyber-attack signature triggered the match, the system may also report the type of cyber-attack detected. As discussed above, a user interface may be provided that allows a system administrator or other user to evaluate the positive hits. Additionally, if a match is found in multiple cyber-attack signatures in step 314, each of these positive matches may be reported in step 316. If no match is found in any of the cyber-attack signatures in library 310 in step 314, step 318 may report a negative result, as discussed above.

[0057] Thus, FIG. 3 illustrates a technique for detecting multiple types of cyber-attacks using host-level data analysis in terms of the impact of multiple types of cyber-attacks on a host's operational performance as compared to a control baseline of the host's normal state operational performance when not under cyber-attack.

[0058] The inventors contemplate that the process flows of Figures 2 and / or 3 may be used by themselves as cybersecurity applications for computer systems, or in conjunction with other cybersecurity applications, such as the network level data analysis discussed above, that are well known to those of skill in the art. In this manner, a cybersecurity dashboard may be provided that evaluates a number of different aspects of a host system, including the host system's external traffic characteristics as well as the host system's internal operating characteristics, to detect anomalies that may indicate the presence of a cyber attack.

[0059] Exemplary embodiment - Detecting Nmap and Nessus enumeration attacks on Red Hat Enterprise Linux systems: In an exemplary embodiment, the host computer system 100 may be a Red Hat Enterprise Linux (RHEL) system, which is a common host system used in commercial and government fields, and the cyber-attack may be an Nmap enumeration cyber-attack. The inventors have experimentally tested the cyber-attack detection techniques described herein on an RHEL system with an Nmap enumeration cyber-attack and found that the host-level data analysis described herein can accurately detect an Nmap enumeration cyber-attack against an RHEL system.

[0060] In this embodiment, the performance monitoring application that may be used in steps 200, 202, and 212 may be the Nmon performance monitoring application and / or the Collectl performance monitoring application. Appendix A included herein describes an exemplary procedure for performing collection on a RHEL host system using Nmon and Collectl to collect normal and under attack system performance data, and then evaluating the resulting data to find anomalies that may be correlated with cyber attacks and used as cyber attack signatures.

[0061] By running process flows 250 for Nmon system collection and Nmap enumeration cyber attacks across 20 instances of RHEL system virtual machines, it was revealed that 15 of the RHEL systems had increased system activity and resource usage during the enumeration scan time. All 15 test positive virtual machines logged increased resource usage for incoming network packets to the network interface card, central processor usage, process switches per second, and processor interrupts specific to the time the Nmap scan occurred. Graphed results in both Nmon and Collectl displayed these increases in system activity and resource usage across the test positive virtual machine systems specific to the data captured for incoming network packets to the network interface card, central processor usage, process switches per second, and processor interrupts. These graphed metrics correlated with the Nmap scan time and were logged in both Nmon and Collectl. The documented increase in activity and graphed results confirmed the test positive designation of these virtual machines.

[0062] Indicators of Nmap enumeration scans recorded by both Nmon and Collectl were incoming network packets, process switches per second, a 5-8 second isolated increase in processor utilization activity, and an increase in processor interrupts. These characteristics therefore correlate appreciably with Nmap enumeration cyber attacks against RHEL systems running Nmon and / or Collectl to gather relevant system performance data.

[0063] In another exemplary embodiment, a RHEL system may be subject to a Nessus enumeration cyber attack. The inventors have experimentally tested the cyber attack detection techniques described herein on a RHEL system for Nessus enumeration cyber attacks and found that the host-level data analysis described herein can accurately detect Nessus enumeration cyber attacks on a RHEL system.

[0064] In this example, the performance monitoring application that may be used in steps 200, 202, and 212 may be a Nmon performance monitoring application and / or a Collectl performance monitoring application. Executing process flow 250 for Nmon system collection and Nessus enumeration cyber attacks across 20 instances of virtual machines of RHEL systems (along with retesting an additional 5 Nmon collections and 5 additional Collectl collections) revealed that 15 of the RHEL systems had increased system activity and resource usage during the enumeration scan time. All 15 test positive virtual machines logged increased resource usage for incoming network packets to the network interface card, central processor usage, process switches per second, and processor interrupts specific to the time that the Nessus scan occurred. Graphed Nessus scan attack metrics associated with the test positive systems were also recorded in both the Nmon and Collectl data. These graphed metrics were correlated with the Nessus scan time and recorded in both the Nmon and Collectl datasets. The documented increase in activity and graphed results confirmed the test positive designation of these virtual machines.

[0065] Indicators of a Nessus enumeration scan recorded in both Nmon and Collectl were a 6 second period of increased activity alone for incoming network packets, process switches per second, and processor usage, followed by an 8 second time frame of normal activity, followed by another 6 seconds of increased activity alone, and an increase in processor interrupts. These characteristics therefore correlate appreciably with Nessus enumeration cyber attacks against RHEL systems running Nmon and / or Collectl to gather relevant system performance data.

[0066] These experiments revealed similar system metrics recorded in both Nmon and Collectl data files that correlate with Nmap enumeration and Nessus enumeration attacks. Specifically, these experiments demonstrated the following performance metrics for detecting Nmap and Nessus enumeration scans using the Nmon and Collectl data sets: Sensitivity = 75%. Sensitivity is the percentage of hosts that test positive when an attack is present. Thus, the test sensitivity provides the percentage of hosts that showed measurable impact from Nmap and Nessus enumeration attacks that were discernible through host level data analysis. Specificity = 100%. Specificity is the percentage of machines that have not been attacked and test negative. The specificity percentage therefore represents the probability of a true negative and allows to distinguish between false positives. · Positive result predictive value = 100% and negative result predictive value = 100%. Predictive value is a measure of the probability that a positive or negative result is considered true. Thus, predictive value is the level of accuracy with which the result could be predicted and the percentage of predictive value that allows to exclude false negatives in order to correctly predict a test positive or test negative. · Test Efficiency = 87.5%. Test Efficiency is the percentage of hosts that produced correct results out of the total number of hosts tested. Test Efficiency provides a measure of the number of correct results out of the total number of tests performed, and the percentage of tested hosts that demonstrated a measurable ability to detect cyber attacks.

[0067] FIG. 4 illustrates the process flow for baseline data capture using the Nmon and Collectl applications associated with this test of a RHEL host system. This process flow generates the normal state system performance data discussed above in connection with FIGs. 2 and 3. In step 402, a user logs into the baseline RHEL host system. In step 404, a terminal window is opened. In step 406, directories for Nmon and Collectl logs are created. Nmon and Collectl can then be started (steps 408 and 410, with step 412 showing the end of the Collectl application). The directory with the baseline capture logs is then accessed (step 414) and copied to an external drive (or another location) for processing / analysis (step 416).

[0068] Figure 5 shows a process flow where collected baseline metrics (which may include metrics and graphed RHEL baseline system results from the collection process of Figure 4 - see 502) include time data (see 504), NIC activity data (see 506), and CPU activity data (see 508). This baseline data is then processed and analyzed to establish consistent / regular patterns of resource usage (step 510). From this baseline data, the system may document baseline NIC activity for each application on the baseline RHEL system (step 512) and baseline CPU activity for each application on the baseline RHEL system (step 514).

[0069] FIG. 6 illustrates a process flow for data capture for Nmap and Nessus cyber attacks against a RHEL host system running Nmon and Collectl applications. This process flow generates the attacked system performance data discussed above in connection with FIG. 2 and FIG. 3. In step 602, a full clone of the baseline RHEL system is created, which can then serve as the target RHEL system. In step 604, the attacking system (e.g., a Kali Linux attacking system) is booted and logged in. In step 606, a terminal window is opened on the attacking system and Nessus is started. In step 608, the target RHEL system is logged in. Next, the IP address of the attacking system and the IP address of the target RHEL system are identified (see steps 610 and 612), and a ping is completed between the attacking system and the target RHEL system to ensure they can communicate with each other (see steps 614 and 616). In step 618, Nmon and Collectl directories are created for log files from the Nmon and Collectl collections. Next, the Nmon and Collectl applications are started on the target RHEL system (step 620). The attacking system may then launch an attack against the target RHEL system (step 622), which may include performing an Nmap scan of the target RHEL system (step 624) and a Nessus scan of the target RHEL system (step 626). It should be appreciated that these attacks may be performed at different times to provide different collections, as desired by the practitioner. These attacks occur while the target RHEL system is being monitored by the running Nmon and Collectl applications, and step 628 depicts terminating the Collectl application.The directory with the capture log files is then accessed (step 630) and copied to an external drive (or another location) for processing and analysis to develop attack signatures (step 632). The target and attacking systems may then be logged off and powered down (steps 634 and 636).

[0070] FIG. 7 illustrates a process flow in which collected attack metrics 702 include time data (see 704), NIC activity data (see 706), and CPU activity data (see 708). The time information 704 allows the timing of the cyber attack to be verified (see step 710) and used to correlate with measurements of NIC and CPU activity. In step 712, the system may determine whether activity across the NIC (e.g., incoming network packets, etc.) has increased in correlation with the verified attack time. The correlated NIC activity in the attacked system performance data may be compared to the normal / baseline NIC activity of step 512 (see step 714), and any changes for each application and each attack (e.g., increased NIC activity resulting from the attack) may be documented (see step 716). The system may perform similar operations with respect to the CPU activity data. In step 718, the system may determine whether CPU usage has increased in correlation with the verified attack time. As shown at 720, this CPU usage may be measured by an increase in activity in at least one of the CPUs, where the increase in activity correlates to the time of Nmap enumeration scans and Nessus enumeration scans, if applicable (derived from Nmon collected data). As shown at 722, this CPU usage may also be measured by an increase in process switching activity per second, where the increase in activity correlates to the time of Nmap enumeration scans and Nessus enumeration scans, if applicable (derived from Nmon collected data). As shown at 724, this CPU usage may also be measured by an increase in processor interrupt activity per second, where the increase in activity correlates to the time of Nmap enumeration scans and Nessus enumeration scans, if applicable (derived from Collectl collected data).Correlated CPU usage activity in the attacked system performance data may be compared to normal / baseline NIC activity from step 514 (see step 726), and any changes for each application and each attack (e.g., increased NIC activity resulting from an attack) may be documented (see step 728). These changes in NIC and CPU activity may be recorded / documented for use in generating attack signatures (see 730).

[0071] 8A and 8B show a process flow for generating cyber attack signatures corresponding to Nmap and Nessus scans based on Nmon and / or Collectl datasets collected from a RHEL host system. For example, in step 802, the system may collect, graph, and analyze resource usage increase documents of the system under attack. In this embodiment, the system parameters used for the determination are NIC activity (e.g., a measurement of incoming packet activity) and CPU activity (e.g., processor switches (processor interrupts) per second), and the process flow may compare the values ​​of these system parameters in the under attack system performance data and the normal / baseline system performance data (see 804 and 808 for NIC activity and CPU activity, respectively). The changes (e.g., the increase in activity in this case) may then be quantified, documented, and recorded (see 806 and 810 for NIC activity and CPU activity, respectively). Thus, the system may record the NIC activity metric increase and the CPU activity metric increase for each attack (see 812). From the correlation of these system parameters with attacks, Nmap and Nessus attack signatures can then be developed, which can then be tested and updated over time to improve performance (see 814). The attack signatures can be based on characteristics and characteristic values ​​of NIC and CPU activity found to be indicative of Nmap and / or Nessus attacks.For example, the features and feature values ​​may include features and feature values ​​representing an increase in activity in terms of process switches per second (correlated to the time of Nmap and Nessus attacks, if applicable (shown at 816)) and features and feature values ​​representing an increase in activity in terms of process interrupts per second (correlated to the time of Nmap and Nessus attacks, if applicable (shown at 818)), and the signature may include thresholds for such increases, as well as time windows for such activity, including timing with respect to other measured activities observed in the system performance data. In another exemplary embodiment, the attack signature may be represented by an algorithm that defines how the system performance data may be tested to determine whether the features and feature values ​​match an attack pattern. Once the attack signature has been developed, it may be inserted into an intrusion detection dashboard of an attack monitoring application (see 820). The attack signature may then be tested against new attacks against the RHEL system (step 822) to evaluate how well the attack signature performs in detecting known attacks (step 824). If the attack signature performs adequately well according to an implementer-defined performance metric, the attack signature may be deployed for operational use (step 826). For example, the defined performance metric may require that steps 822 and 824 accurately detect attacks with a high degree of accuracy (e.g., 90%, 95%, 100%, etc.). If steps 822 and 824 do not produce adequate results (e.g., failing to detect known attacks (see 828)), the system may iteratively adjust the signature (see 830) until an adequately well-performing attack signature is developed.

[0072] FIG. 9 shows the overall process flow combining baseline and attack collection along with attack signature generation, where successfully tested attack signatures can be deployed for detection in production.

[0073] Figure 10 illustrates another process flow for collecting system performance data from a production RHEL host system that is to be tested for attacks, which generates the test system performance data discussed above in connection with Figures 2 and 3. Steps 1002, 1004, 1006, 1008, 1010, 1012, 1014, and 1016 of Figure 10 generally correspond to steps 402, 404, 406, 408, 410, 412, 414, and 416, respectively, of Figure 4, except that the RHEL system is a production system that is being monitored for cyber attacks based on the production attack signatures deployed as a result of step 826 of Figure 8B.

[0074] FIG. 11 shows plots from the IBM Nmon analyzer captured from a tested RHEL host system using the Nmon application while it was bombarded with Nmap and Nessus scans (and while the RHEL host system was running the Nmon, Collectl, and Monitorix applications to monitor performance). The plot in FIG. 11 shows captured performance data specific to the RHEL host system's NIC, with spikes 1100, 1102, and 1104 showing increased Ethernet card read activity. The X-axis shows time (equal to the 10 minute test period for each system) and the X-axis shows activity. The first spike 1100 correlates directly to the Nmap scan time. Then, about a minute later, there are two other notable spikes 1102 and 1104, which correlate directly to the Nessus scan start time.

[0075] FIG. 12 shows another plot from the IBM Nmon analyzer captured using the Nmon application from a RHEL host system under test as it was bombarded with Nmap and Nessus scans (and while the RHEL host system was running the Nmon, Collectl, and Monitorix applications to monitor performance). The plot in FIG. 12 covers the same time period and the same host system as the plot in FIG. 11, but the latter plot shows captured performance data focused specifically on the CPU activity of the system under test in terms of process switching activity (processor interrupts) per second. Box 1200 encloses a notable spike 1202 that correlates with the Nmap scan time. Box 1204 encloses two spikes 1206 and 1208, about one minute later, that correlate with the Nessus scan time.

[0076] FIG. 13 shows plots from the IBM Nmon analyzer captured from the retested RHEL host system using the Nmon application when the retested RHEL host system was bombarded with Nmap and Nessus scans (and while the RHEL host system was only running the Nmon application to monitor performance). The plot in FIG. 13 shows captured performance data specific to the RHEL host system's NIC, with spikes 1300, 1302, and 1304 showing increased Ethernet card read activity. The X-axis shows time (equal to the 10 minute test period for each system) and the X-axis shows activity. The first spike 1300 correlates to the Nmap scan time. Then, about a minute later, there are two other notable spikes 1302 and 1304, which directly correlate to the Nessus scan start time.

[0077] FIG. 14 shows another plot from the IBM Nmon analyzer captured using the Nmon application from the retested RHEL host system of FIG. 13 when it was bombarded with Nmap and Nessus scans (and while the RHEL host system was running only the Nmon application to monitor performance). The plot in FIG. 14 covers the same time period and the same host system as the plot in FIG. 13, but the plot in FIG. 14 shows the captured performance data specifically for the CPU activity of the system under test in terms of process switching activity (processor interrupts) per second. Box 1400 encloses a notable spike 1402 that correlates with the Nmap scan time. Box 1404 encloses two spikes 1406 and 1408, about one minute later, that correlate with the Nessus scan time.

[0078] Systems that exhibit such a significant and positive increase in Nmon data for Ethernet card read activity and processor switching activity per second, as reflected in the plots of Figures 11-14, can be marked as testing positive during analysis.

[0079] Figure 15 shows graphs from a Collectl-only dataset captured during retesting of an attacked RHEL host system, with plots focused on activity on the NIC and CPU (plot 1510 showing network packets inbound and plot 1512 showing processor interrupts per second for the CPU). Again, the X-axis shows time, which is equal to the 10 minute test period of the system, and the Y-axis shows activity. Box 1500 contains the first spike in NIC activity, which correlates to a spike in processor activity that directly matches the Nmap scan time. Box 1502 contains the next two spikes in NIC activity, which correlate to spikes in processor activity related to the Nessus scan time for the system. Thus, systems that show a significant increase in these activities can be marked as testing positive during analysis.

[0080] Appendix A: Examples of Collection, Testing, and Evaluation Procedures Create a full clone test system from the research baseline system configuration. Select the ORIGINAL RHEL BASELINE virtual machine and click Virtual Machine in the File menu. Select Create Full Clone from the drop-down menu. Save As: TEST_SYSTEM## and click Save. Once the cloning is complete, click the Play icon on the screen to boot the target machine.

[0081] Boot the Kali Linux attack system.

[0082] Log in to Kali Linux and start Nessus. Open a terminal window: Start Nessus. / etc / init.d / nessusd start

[0083] Once the RHEL target test system is running · Log in to the target test system.

[0084] Open a terminal window on your Kali Linux attack system. Run the ifconfig command to identify the IP address of the attacking system.

[0085] Open a terminal window on your RHEL target test system. Run the ifconfig command to identify the target system IP address.

[0086] In the terminal window of the Kali Linux attack system -Perform a ping command to the RHEL target IP address to verify network connectivity to the system.

[0087] In a terminal window on the RHEL target test system: ·Perform a ping command to the Kali Linux IP address to verify network connectivity to the system.

[0088] On the RHEL target system Start Monitorix and open a web page in Firefox. In a terminal window, type: Service monitorix start -> enter the administration password nimdanimda. Open the Firefox web browser and type: ·URL:Localhost:8080 / monitorix Hostname: localhost Select All graphs, Daily, OK. Start Nmon. Open a terminal window. Create a directory for the Nmon log. Type: sudo mkdir / home / nmon / testsys#_nmon (enter admin password if required: nimdanimda). Start Nmon. Type: sudo nmon ‐s 1 ‐c 600 ‐f ‐t ‐m / home / nmon / testsys#_nmon (enter admin password if required: nimdanimda). Start Collectl. Type: collect ‐scdimnt >> / home / admin / testsys#_collectl

[0089] On the Kali Linux attack system: Open a terminal window: ·Perform a ping command to the RHEL target IP address. ping XXX.XXX.XXX.XXX · Perform an Nmap scan of RHEL target IP addresses. ·nmap ‐A ‐sV ‐O ‐v XXX.XXX.XXX.XXX · Run a Nessus scan of the RHEL target IP addresses. Open Firefox and enter the URL: localhost:8834 Log in with your username and password and select Scans. · Run a Basic Scan on the RHEL target system by entering the IP address of the target system.

[0090] At the end of the 10 minute log collection period on the RHEL target system

[0091] On the RHEL target system -Exit Collectl. In the terminal window, press Ctrl-C. Copy Monitorix data from the Firefox web browser to a text application. Save the Monitorix text file in your Home\Documents folder. Collect Collectl, Nmon, and Monitorix text files and copy them to removable media for storage.

[0092] Thus, performance data may be collected via audit logs and system monitoring applications including Monitorix, Collectl, and Nmon. Once collected, the performance data is imported into a Microsoft Excel spreadsheet for analysis. This analysis involves comparing data collected from the uncompromised baseline system against data collected from the attacked or compromised system during the cyber attack. All baseline and test data is imported into a Microsoft Excel spreadsheet and plotted using IBM SPSS for graphical comparison and logistic regression analysis against positive and negative predictive value theory. The graphical data and plots within the IBM SPSS, Microsoft Excel, and IBM Nmon analysis applications allow for the examination and visual comparison of the uncompromised baseline system data with the attacked and compromised system data captured during the cyber attack testing. Visual and graphical comparisons within IBM SPSS, Microsoft Excel, and IBM Nmon analytics applications, as well as logistic regression analysis, allow practitioners to visually explore data for possible host-level system changes in resource usage, timing, behavior, and operating environment that may include performance fluctuations, spikes, or anomalies that may indicate a cyber attack.

[0093] Variable (host system parameter) specific data from each system is collected and compared against uncompromised baseline, test, and retest system performance data, as well as positive and negative predictive value theoretical benchmarks. As the variable data for each system is collected, it is separated into columns for variable, baseline benchmark, test, retest, positive predictive value, and negative predictive value.

[0094] The variable column data is divided into the following groups by row: Resource usage: central processing unit, power, random access memory, network cards Time: Time spent on each process Behavior: central processing unit, power, random access memory, network cards, and users · Environment: Network Object Model environment and command line interface

[0095] Examining the raw data helps determine if one variable is more indicative of the next attack or compromised system, or if the variable has no value in predicting the next cyber event.

[0096] Raw performance data is collected from each host-level system and includes each system's audit log, as well as performance data from the Monitorix, Collectl, and Nmon applications. Comma-separated values ​​(CSV) data is imported into a Microsoft Excel spreadsheet with columns associated with each variable, event, application, and log. Each row in the spreadsheet contains collected data for each host-level baseline system and each host-level operational system. If the performance data cannot be imported automatically, the practitioner can enter it manually by copying and pasting the recorded data into the correct cells in the Microsoft Excel spreadsheet. After the data is properly imported and verified, the Microsoft Excel spreadsheet is imported into IBM SPSS software for graphing, visualization, analysis, and review.

[0097] Graphical comparison, visualization, and logistic regression analysis allow practitioners to visually inspect performance data and recognize any changes (performance fluctuations, spikes, or anomalies) occurring in system resource usage, timing, behavior, and environment in host-level systems and peripherals. Additionally, timestamp correlations between Nmon, Collectl, host-level systems, and recorded test logs can be used to triangulate data and results to verify positive or false negative findings specific to cyber attacks. Timestamp triangulation is also used to validate and verify performance data in Microsoft Excel and IBM Nmon analysis software to ensure all times correspond to host-level system logs and cyber attack times.

[0098] Although the present invention has been described above with reference to exemplary embodiments thereof, various modifications may be made thereto that remain within the scope of the present invention. For example, rather than developing an attack signature for a given type of cyber attack, a system may instead be configured to collect normal-state system performance data using the techniques described herein from a host system known not to have been compromised by a cyber attack. The system may then collect test system performance data, which may then be statistically compared to the normal-state system performance data to determine whether any anomalies are present. Upon detection of anomalies in the test system performance data, these anomalies may be isolated and reported to a system administrator or other user for further review. While this approach is expected to have a higher false positive rate than the attack signature approach discussed above, the inventors believe that the anomaly detection approach can still provide value to users. These and other modifications to the present invention will be discernible upon review of the teachings herein.

Claims

1. A method for detecting cyberattacks based on host-level data analysis, comprising: accessing, by a processor, a cyberattack signature corresponding to a cyberattack, wherein the cyberattack signature includes data representing the impact of the cyberattack on the host computer system with respect to a plurality of internal system parameters of the host computer system; collecting, by the processor, test system performance data from a test host computer system, wherein the test system performance data includes a plurality of data values representing the internal system parameters of the test host computer system over time during operation of the test host computer system; comparing, by the processor, the collected test system performance data with the accessed cyberattack signature to determine whether there is a match with the accessed cyberattack signature in the collected test system performance data; generating, by the processor, a notification that the cyberattack has been detected on the test host computer system in response to a determination that the match exists; The method as described above.

2. The method according to claim 1, wherein the cyberattack includes an enumerated cyberattack.

3. The method according to claim 2, wherein the enumerated cyberattack includes an Nmap enumerated scan or a Nessus enumerated scan.

4. The step of collecting includes executing a performance monitoring application over time on the test host computer system, wherein the performance monitoring application includes Nmon, Collectl, and / or Monitorix. The method according to claim 1 or 2.

5. The method according to claim 1 or 2, wherein the internal system parameters include data representing processor usage activity.

6. The method according to claim 5, wherein the processor usage activity includes processor switching activity.

7. The method according to claim 5, wherein the processor usage activity includes processor interrupt activity.

8. The method according to claim 1 or 2, wherein the internal system parameters include network card activity, memory usage activity, RAM activity, power consumption, and / or elapsed time of the process.

9. The step of accessing includes accessing a plurality of cyber attack signatures, each cyber attack signature corresponding to a different type of cyber attack, and the data representing the impact of the corresponding cyber attack type on the host computer system with respect to a plurality of internal system parameters of the host computer system. The step of comparing includes comparing the collected test system performance data with each of the accessed cyber attack signatures to determine whether there is a match with any of the accessed cyber attack signatures in the collected system performance data. Generating the notification in response to a determination that there is a match with any of the accessed cyber attack signatures. The plurality of different cyber attack signatures represent their impacts on the host computer system with respect to different internal system parameters. The method according to claim 1 or 2.

10. Applying host-level data analysis to system performance data representing the operation of the host computer system when it is known that the host computer system is under a cyber attack and the operation of the host computer system when it is known that the host computer system is not under a cyber attack, with respect to a second plurality of internal system parameters of the host computer system, to find those of the second plurality of internal system parameters and their corresponding data values that are discriminably correlated with the cyber attack. Generating the cyber attack signature by the processor based on the internal system parameters and their corresponding data values that are found to be discriminably correlated with the cyber attack. The method according to claim 1 or 2, further comprising.

11. The step of applying includes Collecting normal state system performance data from the host computer system known not to have been subjected to the cyber attack, wherein the normal state system performance data includes a plurality of data values representing over time the second plurality of internal system parameters of the host computer system known not to have been subjected to the cyber attack during operation of the host computer system, and the collecting; Collecting attacked state system performance data from the host computer system known to have been subjected to the cyber attack, wherein the attacked state system performance data includes a plurality of data values representing over time the second plurality of internal system parameters of the host computer system known to have been subjected to the cyber attack during operation of the host computer system, and the collecting; Comparatively analyzing the normal state system performance data with the attacked state system performance data to identify the internal system parameters and their corresponding data values that are discriminably correlated with the attack; The method according to claim 10, comprising the above.

12. The method according to claim 11, wherein the step of comparatively analyzing includes performing a regression analysis on different combinations of the internal system parameters within the second plurality of internal system parameters.

13. The step of applying further includes: Creating a clone of the host computer system known not to have been subjected to the cyber attack; Executing the cyber attack on the cloned host computer system; Collecting attacked state signature performance data from the cloned host computer system on which the cyber attack has been executed; The method according to claim 10, comprising the above.

14. The cyber-attack includes an Nmap enumeration scan, and the cyber-attack signature of the Nmap enumeration scan includes at least two components out of the group consisting of (1) the activity of incoming network packets, (2) the process switches per second, (3) the processor usage, and (4) the processor interrupts increasing with respect to their normal state baselines over a predetermined period. The method according to claim 1 or 2.

15. The cyber-attack includes a Nessus enumeration scan, and the cyber-attack signature of the Nessus enumeration scan includes (1) over a first predetermined period, (i) the activity of incoming network packets, (ii) the process switches per second, (iii) the processor usage, and (iv) the processor interrupts increasing with respect to their normal state baselines, and then (2) these normal activities continuing over a second predetermined period, and then (3) over a third predetermined period, (i) the activity of incoming network packets, (ii) the process switches per second, (iii) the processor usage, and (iv) the processor interrupts increasing separately with respect to their said normal state baselines. The method according to claim 1 or 2.

16. The processor includes a plurality of processors. The processor includes a first processor and a second processor. The first processor generates the cyber-attack signature. The second processor executes the steps of accessing the cyber-attack signature, collecting the test system performance data, and comparing the collected test signature performance data with the accessed cyber-attack signature. The method according to claim 1 or 2.

17. A system for detecting a cyber-attack based on host-level data analysis, comprising a memory configured to store executable code, a processor cooperating with the memory, and wherein the processor, in response to the execution of the executable code, ​ Accessing a cyber-attack signature for responding to a cyber-attack, wherein the cyber-attack signature includes data representing the impact of the cyber-attack on a host computer system with respect to a plurality of internal system parameters of the host computer system, and said accessing; Collecting test system performance data from a test host computer system, wherein the test system performance data includes a plurality of data values representing the internal system parameters of the test host computer system over time during operation of the test host computer system, and said collecting; Comparing the collected test system performance data with the accessed cyber-attack signature to determine whether there is a match with the accessed cyber-attack signature in the collected test system performance data; Generating a notification that the cyber-attack has been detected on the test host computer system in response to the determination that the match exists; The system configured to perform the above.