Data enrichment method and system

By receiving and analyzing event information from targets of interest, and utilizing statistical models and large language models (LLM) to identify the impact of cyberattacks and mitigation measures, this approach addresses the problem of slow response speed in existing cybersecurity systems and improves the efficiency of identifying and mitigating cyber threats.

CN121986338APending Publication Date: 2026-05-05C2A SEC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
C2A SEC LTD
Filing Date
2024-10-15
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing cybersecurity systems and methods lack effective data enrichment tools when identifying cyberattacks and providing mitigation measures, making it difficult for SOC analysts to respond quickly and effectively to cyber threats.

Method used

By receiving event information associated with targets of interest, analyzing data using statistical models and large language models (LLM), identifying event impacts and mitigation measures, generating network models and threat intelligence related to targets of interest, and providing detailed event analysis and mitigation recommendations.

Benefits of technology

It improved the speed and effectiveness of SOC analysts' response to cyber threats, enhanced their ability to identify and mitigate cyberattacks, and improved the protection level of cybersecurity systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121986338A_ABST
    Figure CN121986338A_ABST
Patent Text Reader

Abstract

A data enrichment method includes: receiving event information associated with an event in a target of interest; based at least in part on the received event information, analyzing data associated with the target of interest to identify: an impact of the event within the target of interest, and existing one or more mitigation measures associated with the event; and outputting information about an event, the output information including the identified impact and the identified one or more mitigation measures, where the event information includes information about an abnormal behavior detected in the target of interest or information about a vulnerability detected in the target of interest, and wherein the analyzed data is based on a plurality of sources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure is primarily concerned with the field of cybersecurity, and more specifically with data enrichment. background

[0002] A Security Operations Center (SOC) is responsible for protecting an organization from cyber threats. SOC analysts typically monitor the organization's network around the clock and investigate any potential security incidents. If a cyberattack is detected, the SOC analyst may be responsible for taking any necessary steps to remediate it. In some examples, the SOC is operated by an external provider that offers remote monitoring and management services to the organization. One example of such a SOC is in the automotive cybersecurity field, where it is often referred to as a Vehicle SOC (SOC). Overview

[0003] Therefore, the primary objective of this invention is to overcome at least some of the shortcomings of existing cybersecurity systems and methods. This is provided in some examples through a data enrichment method, which includes receiving event information associated with an event in a target of interest. In some examples, based at least in part on the received event information, the method includes analyzing data associated with the target of interest to identify: the impact of the event within the target of interest; and / or one or more existing mitigation measures associated with the event. In some examples, the method also includes outputting information about the event, including the identified impact and the identified one or more mitigation measures. In some examples, the event information includes information about anomalous behavior detected in the target of interest or information about vulnerabilities detected in the target of interest. In some examples, the analyzed data is based on multiple sources.

[0004] Additional features and advantages of the invention will become apparent from the following figures and description.

[0005] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In case of conflict, the specification of this patent application (including the definitions) shall prevail. As used herein, unless the context clearly indicates otherwise, the articles “a” and “an” mean “at least one” or “one or more”. As used herein, “and / or” means any one or more items in a list connected by “and / or”. As an example, “x and / or y” means any element in the set of three elements {(x), (y), (x, y)}. In other words, “x and / or y” means “x, y or both x and y”. As in some examples, “x, y and / or z” means any element in the set of seven elements {(x), (y), (z), (x, y), (x, z), (y, z), (x, y, z)}.

[0006] Furthermore, unless explicitly stated otherwise, "or" refers to an inclusive "or," not an exclusive "or." For example, condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).

[0007] Additionally, the terms "a" or "an" are used to describe elements and components of embodiments of the inventive concept. This is done merely for convenience and to give the general meaning of the inventive concept, and "a" and "an" are intended to include one or at least one, and the singular includes the plural, unless it is obvious that it has a different meaning.

[0008] As used herein, the term “approximately” when referring to a measurable value (e.g., quantity, duration, etc.) means that it includes a deviation of + / -10%, more preferably + / -5%, even more preferably + / -1%, and still more preferably + / -0.1% from the specified value, because such deviation is suitable for performing the disclosed apparatus and / or methods.

[0009] As used herein, the term "fine-tuning" refers to a method of training the weights of a pre-trained model on new data, as is known to those skilled in the art. In some examples, as is known to those skilled in the art, fine-tuning involves feeding data into an LLM and prompting the LLM to fine-tune a predetermined subset of its weights based on a loss function. In some examples, as is known to those skilled in the art, the output of the LLM is fed into a second LLM that generates the input of the loss function. Note that this is merely one method for performing fine-tuning and is not intended to be limiting in any way.

[0010] The following embodiments and aspects of systems, tools, and methods are described and illustrated. These embodiments and aspects are intended to be exemplary and illustrative, and not to limit the scope. In various embodiments, one or more of the problems described above have been reduced or eliminated, while other embodiments address other advantages or improvements. Brief description of the attached diagram

[0011] To better understand the invention and to show how it can be practiced, reference will now be made to the accompanying drawings by way of example only, in which similar reference numerals throughout the drawings denote corresponding parts or elements.

[0012] Referring now specifically to the accompanying drawings, it is emphasized that the details shown are by way of example and are merely for the purpose of illustrative discussion of preferred embodiments of the invention, and are presented to provide the most useful and readily understood description of what is believed to be the principles and concepts of the invention. In this respect, no attempt is made to show the structural details of the invention in more detail than necessary for a basic understanding of the invention; the description taken in conjunction with the drawings will enable those skilled in the art to understand how the various forms of the invention can be embodied in practice.

[0013] In the attached diagram: Figure 1A A high-level block diagram of a data enrichment system according to some examples of this disclosure is shown; Figure 1B Some examples according to this disclosure are shown. Figure 1A A high-level diagram showing the connection between the data enrichment system and the data source; Figure 1C A high-level block diagram of a statistical model according to some examples of this disclosure is shown; Figure 1D It shows Figure 1A-Figure 1B A more detailed example of a data enrichment system is shown in the high-level diagram. Figure 1E Examples of some of the embodiments shown in this disclosure are illustrated. Figure 1A A high-level diagram showing the connection between the data enrichment system and the SOC; Figure 2A A high-level flowchart of a first data enrichment method according to some examples of this disclosure is shown; Figure 2B A high-level flowchart of a second data enrichment method according to some examples of this disclosure is shown; and Figure 3 A high-level flowchart of a method for performing security analysis according to some examples of this disclosure is shown. Detailed description of a specific example

[0014] In the following description, various aspects of this disclosure will be described. Specific configurations and details are set forth for purposes of explanation in order to provide a thorough understanding of the different aspects of this disclosure. However, it will be apparent to those skilled in the art that this disclosure can be practiced without the specific details presented herein. Furthermore, well-known features may have been omitted or simplified so as not to obscure this disclosure. In the accompanying drawings, similar reference numerals always refer to similar parts. To avoid excessive confusion due to too many reference numerals and leaders on a particular drawing, some components will be introduced through one or more drawings without being explicitly identified in each subsequent drawing containing that component.

[0015] Figure 1A A high-level block diagram of a data enrichment system 10 is shown. In some examples, the data enrichment system 10 includes: a data interface 20; a processor 30; and a memory 40. In some examples, each of the data interface 20, processor 30, and memory 40 is a component within a cloud server.

[0016] In some examples, data interface 20 includes software and / or hardware configured to communicate with an external network. In some examples, such as... Figure 1B As shown, data interface 20 communicates with data source 45. In some examples, data source 45 can be: a network; a predetermined sender within a cloud server containing data enrichment system 10, such as an original equipment manufacturer (OEM); or a predetermined sender within a network containing data enrichment system 10. Although data interface 20 is shown to communicate with only a single data source 45, this does not imply any limitation in any way, and data interface 20 can communicate with any number of data sources 45.

[0017] In some examples, as described below, data interface 20 receives log information from data source 45 and optionally receives log summaries.

[0018] In some examples, memory 40 stores multiple instructions that, when read by processor 30, cause processor 30 to perform one or more steps, as described below. While this document describes the data enrichment system 10 with respect to an example provided by a single processor 30, this is not intended to be limiting in any way, and the operation of processor 30 can be implemented by multiple processors 30 without exceeding the scope of this disclosure.

[0019] In some examples, segment 50 of memory 40 stores multiple statistical models. In some examples, each statistical model is associated with a corresponding target of interest within a target of interest. In some examples, each target of interest includes: one or more signals; and / or one or more components.

[0020] In some examples, each statistical model contained within segment 50 of memory 40 is generated based at least in part on network traffic received from the target of interest. In some examples, each statistical model contains statistical values ​​associated with the corresponding target of interest. In some examples, each statistical model includes a Bayesian network. In some examples, each statistical model includes a Markov random field model. In some examples, the statistical model represents the statistical behavior of a set of signals within the target of interest. In some examples, the accuracy of the statistical model is based on the risk value associated with the corresponding target of interest. For example, the corresponding statistical model is more accurate for higher-risk signals / components.

[0021] In some examples, segment 50 of memory 50 includes multiple statistical models 100. Figure 1C A high-level block diagram of an example of statistical model 53 is shown. In some examples, such as Figure 1C As shown, each statistical model 53 includes multiple statistical sub-models 54. As used herein, the term "statistical sub-model" means a statistical model that represents the statistical behavior of a corresponding subset of signals associated with a given statistical model 53. In some examples, for each statistical model 53, the model includes statistical relationships between signals from different sub-models 100.

[0022] In some examples, segment 51 of memory 40 includes configuration information associated with multiple targets of interest. In some examples, the configuration information includes one or a combination of the following: error reports of libraries or assets associated with the targets of interest; the software bill of materials (SBOM) of the targets of interest; the specification sheet of the targets of interest; a network database (e.g., an ARXML database, a DBC database, or other network database) including information about the communication attributes of the network; runtime configuration information of software components; or additional configuration information.

[0023] In some examples, segment 52 of memory 40 includes security threat data associated with the target of interest. In some examples, the security threat data includes attack trees or attack paths from Threat Analysis and Risk Assessment (TARA). As defined herein, the term "attack tree" or "attack path" is a hierarchical data structure of branches that represents a set of potential methods for implementing an event in which the security of a system is compromised or compromised in a specified manner. As defined herein, an "attack step" is any node in an attack tree or attack path. Note that although this disclosure references TARA, this is not intended to limit it in any way, and any other risk assessment method, such as the Risk Management Framework (RMF), Operationally Critical Threat, Asset and Vulnerability Assessment (OCTAVE), or other methods, may be used.

[0024] In some examples, security threat data includes data received from cyber threat intelligence. In some examples, processor 30 retrieves cyber threat intelligence via data interface device 20. In some examples, security threat data includes data received from multiple sources. In some examples, security threats... In some examples, memory 40 includes a corresponding network model for each target of interest. As used herein, the term "network model" refers to a model of the target of interest or assets within the target of interest. In some examples, the network model includes: a description of components within the target of interest; a description of the connections between the target of interest and other systems; and risk analysis data of the target of interest, such as attack paths associated with different threats. In some examples, the network model also includes: information about the configuration data of the asset / system; the specifications of the asset / system; and / or the operational status of the asset / system.

[0025] In some examples, segment 57 of memory 40 includes an attack path database. In some examples, the attack path database 57 includes multiple attack paths, each with a corresponding number of attack steps. As used herein, the term "attack path" refers to a sequence of attack steps. As used herein, the term "attack step" is any description of an attack on any resource or asset.

[0026] In some examples, each attack path has a corresponding risk level associated with it. In some examples, as used herein, the term "risk level" is defined as a predetermined function of the corresponding feasibility value and the corresponding impact value. In some examples, the impact value is a numerical indication of the impact a particular threat will have on the corresponding asset (or the project itself) if the corresponding threat is realized. In some examples, as is known to those skilled in the art, the impact value is assigned a numerical value.

[0027] In some examples, the feasibility value of an attack path is a predetermined function of the feasibility value of each attack step within it. Therefore, in some examples, the feasibility value of each attack step affects the overall feasibility value of the attack path. For instance, if one attack step has a low feasibility value, this can reduce the feasibility value of the corresponding attack path.

[0028] In some examples, the attack path database 57 also includes data on multiple mitigation measures, each associated with a corresponding attack step. In some examples, the mitigation measures include multiple security controls. As used herein, the term "security control" refers to a method used to reduce the feasibility of an attack step. For example, regarding an attack step defined as "Send malicious message," which may have a high feasibility rating, implementing a security control of "implementing a firewall to prevent unauthorized messages" would reduce the feasibility of the attack step to medium.

[0029] Figure 1D A more detailed example of the data enrichment system 10 is shown, which also includes a model generation device 60 for generating multiple statistical models. In some examples, the model generation device 60 is implemented as a plurality of corresponding instructions stored in memory 40, which, when executed by processor 30, cause processor 30 to perform the steps of model generation device 60.

[0030] In some examples, the data enrichment system 10 includes a statistical analysis device 65. In some examples, the statistical analysis device 65 is implemented as a plurality of corresponding instructions stored in memory 40, which, when executed by processor 30, cause processor 30 to perform the steps of the statistical analysis device 65.

[0031] In some examples, the data enrichment system 10 also includes a log summarization device 70. In some examples, the log summarization device 70 is implemented as a plurality of corresponding instructions stored in memory 40, which, when executed by processor 30, cause processor 30 to perform the steps of the log summarization device 70. In some examples, no separate log summarization device 70 is provided, and log summaries are received from an external source at data interface 20. In some examples, the log summarization device 70 summarizes the received log information according to predetermined summarization rules.

[0032] In some examples, the data enrichment system 10 also includes an anomaly analysis device 80. In some examples, the anomaly analysis device 80 is implemented as a plurality of corresponding instructions stored in memory 40, which, when executed by processor 30, cause processor 30 to perform the steps of the anomaly analysis device 80.

[0033] In some examples, the data enrichment system 10 also includes a data retrieval device 90. In some examples, the data retrieval device 90 includes a search engine. In some examples, the data retrieval device 90 is implemented as a plurality of corresponding instructions stored in memory 40, which, when executed by processor 30, cause processor 30 to perform the steps of the data retrieval device 90.

[0034] In some examples, the data retrieval device 90 includes connections to one or more data sources 95. In some examples, the data retrieval device 90 includes one or more of the following: a connection to the Internet; a connection to one or more databases; or a connection to one or more repositories. The databases and / or repositories may be stored within memory 40, in a separate memory, and / or on a different network, and the data retrieval device 90 includes appropriate connections for any of these options. In some examples, the data retrieval device 90 also communicates with external systems, processes, and / or networks.

[0035] In some examples, model generation device 60 receives network traffic data contained within network packets and learns a statistical model based on the received traffic data. In some examples, the network traffic data is received by data interface 20. In one example, where the statistical model includes a distribution with a corresponding mean and distribution, the model is learned by computing the corresponding distribution function (such as a normal distribution function). In another example, where the statistical model includes an autoregressive (AR) model, the coefficients of the model are learned by applying an optimization algorithm (such as maximum likelihood estimation) to the data.

[0036] In some examples, network traffic data is received from data source 45. In some examples, network traffic data is received from the OEM. In some examples, network traffic data is received directly from the system that generates the traffic.

[0037] In some examples, one or more statistical models are not generated by model generation device 60. Instead, in such examples, one or more statistical models are received from data source 45. For example, one or more statistical models may be received from an OEM.

[0038] In some examples, the statistical analysis device 65 similarly receives network traffic data contained within network packets and identifies statistical relationships between various signals within the network traffic data. As used herein, the term "statistical relationship between various signals" refers to the statistical relationship between the transmission of one signal and the transmission of another signal within the network. In some examples, the statistical relationship between two signals is identified by identifying the average time between the transmission of a first signal and the transmission of a second signal. In some examples, the statistical analysis device 65 determines a correlation function between the transmission times of the two signals to determine their statistical relationship.

[0039] In some examples, the data enrichment system 10 also includes a language analysis device 85. In some examples, the language analysis device 85 is implemented as a large language model (LLM) 85, and will be described as such unless otherwise stated. In some examples, as will be described below, the language analysis device 85 may be implemented as a natural language processor. An illustrative example of an LLM 85 (without any limitations) comes from the LLaMa-2 large language model family, which is commercially available from Meta AI.

[0040] In some examples, the LLM 85 is fine-tuned based on predefined data. Although the fine-tuning of the trained LLM 85 is described below, this is not intended to be limiting in any way. In some examples, as known to those skilled in the art, the steps for fine-tuning the LLM 85 described below are performed during the training process of the LLM 85, with appropriate modifications.

[0041] In some examples, a list of libraries associated with an asset is input into LLM 85, as described below, so that the data output by LLM 85 is associated with these libraries. As used herein, the term "associated with an asset" means a library accessed by the asset and / or a library accessed by which the asset uses data affected by the accessed library. For example, a first process accesses a library and uses it to obtain certain data, storing the data in a queue. A second process, as part of the asset, retrieves the data from the queue. Thus, in such an example, the library is associated with the asset because the asset retrieves the data associated with the library.

[0042] In some examples, when a list of libraries is used to prompt the LLM 85 output data, the list of libraries is input as a soft prompt to the LLM 85. As used herein, the term "soft prompt" refers to a learnable tensor cascaded with the input embedding.

[0043] In some examples, the asset's specification table is input into LLM 85, as described below, so that the data output by LLM 85 is related to the specification table. In some examples, when prompted to output data based on the specification table, the data in the specification table is input into LLM 85 as a soft prompt. In some examples, as described below, the asset's TARA data (such as TARA information) is input into LLM 85, so that the data output by LLM 85 is related to the TARA data. In some examples, when prompted to output data based on the TARA data, the TARA data or its predetermined parameters are input into LLM 85 as a soft prompt.

[0044] In some examples, the network model is input into the LLM 85 so that the data output by the LLM 85 is related to the network model. In other examples, when prompting the LLM 85 to output data based on the network model, the network model or a predetermined portion thereof is input into the LLM 85 as a soft prompt.

[0045] In some examples, such as Figure 1D As shown, the data enrichment system 10 also communicates with a response unit 98. In some examples, the response unit 98 is a Security Operations Center (SOC). In some examples, the response unit 98 is an incident response team. In some examples, communication with the response unit 98 is performed via data interface 20.

[0046] Figure 2A A high-level flowchart illustrating data enrichment methods based on some examples is shown. This article describes a data enrichment system 10. Figure 2A This method is described, but this is not intended to limit it in any way, and the method can be implemented using any suitable system. Please note that, without departing from the scope of this disclosure, Figure 2AAny step in the flowchart can be optional.

[0047] As used herein, the term “enrichment” is intended to have the same meaning as the term “improvement” and is not intended to be restrictive in any way.

[0048] In phase 900, information is received at data interface 20. In some examples, the received information includes event information. In some examples, the received event information includes information associated with anomalous behavior occurring in the target of interest or information about vulnerabilities detected in the target of interest. In some examples, the anomalous behavior is identified by an external source. As used herein, the term "anomalous behavior" refers to behavior outside the specification. For example, this could include unexpected signal values, optionally determined based on comparisons with statistical models, as described below. In some examples, detected vulnerability information is received from a vulnerability database, such as a CVE database.

[0049] In some examples, the received information includes log information associated with the target of interest. In some examples, the target of interest is at least a portion of any of a vehicle, a charging station, or various industrial equipment.

[0050] In some examples, the log information contains data about network traffic and / or events recorded within the target of interest. In some examples, the log information contains data about multiple targets of interest (such as multiple ECUs within a vehicle). In some examples, the data includes signal values, derived values ​​of the signal values, and events for each target of interest. For example, the data might contain the speed values ​​of the vehicle at predetermined time intervals. An example of an event might be the driver of the vehicle pressing a specific button.

[0051] In some examples, in phase 910, data associated with the target of interest is analyzed, at least in part, based on the received event information, to identify: the impact of the event within the target of interest, and one or more existing mitigation measures, such as security controls, associated with the event.

[0052] In some examples, the data associated with the target of interest being analyzed is based on multiple sources.

[0053] In some examples, multiple sources include two or more of the following: technical artifacts of the target of interest; product architecture associated with the target of interest; risk assessment analysis; software bill of materials (SBOM); specifications associated with the target of interest; bug reports; data received from electronic control units (ECUs) or other components within the target of interest; data received from sensors within the target of interest; and threat intelligence, such as information retrieved from online databases and other sources of information about threats.

[0054] In some examples, the impact of an event includes: the risk level of an attack path; and an indication of which part of the target of interest might be affected by the event.

[0055] In some examples, attack steps are identified within multiple attack paths associated with the target of interest, and the indication of which part of the target of interest might be affected is at least in part based on the attack path containing the identified attack steps. In some examples, each attack path is associated with a corresponding part of the target of interest.

[0056] In some examples, attack steps are identified in multiple attack paths associated with a target of interest, wherein the impact of an event includes an indication of which of the multiple parts of the target of interest may be affected by the event, and wherein the indication of which part of the target of interest may be affected is at least in part based on the attack path containing the identified attack steps.

[0057] In some examples, attack steps are identified within multiple attack paths associated with a target of interest, and the output information includes the attack paths containing the identified attack steps.

[0058] In some examples, identifying attack steps includes: identifying relevant components within the target of interest, at least in part based on the received event information and the network model of the target of interest; and identifying attack steps associated with the identified components.

[0059] In some examples, signals are received and anomalous behavior is identified within them. In other examples, information about the source and destination of the signal is used to identify the corresponding component (i.e., source and destination) and to identify the attack path or attack steps associated with that component.

[0060] In some examples, upon receiving vulnerability information (such as a CVE), the vulnerability information may include information about the corresponding software group. In such examples, the software group is identified within the SBOM, and thus the components associated with that software group are identified. In some examples, attack paths or attack steps associated with the identified components are identified.

[0061] In some examples, at stage 920, information about the event is output, including the identified impacts and one or more identified mitigation measures.

[0062] Figure 2B A high-level flowchart illustrating data enrichment methods based on some examples is shown. This article describes a data enrichment system 10. Figure 2B This method is described, but this is not intended to limit it in any way, and the method can be implemented using any suitable system. Please note that, without departing from the scope of this disclosure, Figure 2B Any step in the flowchart can be optional.

[0063] In phase 1000, information is received at data interface 20. In some examples, the received information includes event information. In some examples, the received event information includes information associated with anomalous behavior occurring in the target of interest or information about vulnerabilities detected in the target of interest. In some examples, the anomalous behavior is identified by an external source. As used herein, the term "anomalous behavior" refers to behavior that deviates from the normal range. For example, this may include unexpected signal values, optionally determined based on comparisons with statistical models, as described below. In some examples, detected vulnerability information is received from a vulnerability database, such as a CVE database.

[0064] In some examples, the received information includes log information associated with the target of interest. In some examples, the target of interest is at least a portion of any of a vehicle, a charging station, or various industrial equipment.

[0065] In some examples, the log information contains data about network traffic and / or events recorded within the target of interest. In some examples, the log information contains data about multiple targets of interest (such as multiple ECUs within a vehicle). In some examples, the data includes signal values, derived values ​​of the signal values, and events for each target of interest. For example, the data might contain the speed values ​​of the vehicle at predetermined time intervals. An example of an event might be the driver of the vehicle pressing a specific button.

[0066] In some examples, processor 30 extracts at least a predetermined portion of the data within the received information. In some examples, the extracted portion of data is data associated with a corresponding target of interest. In some examples, processor 30 extracts all the data within the received information.

[0067] In some examples, the received log information includes a summary of the log information. In some examples, the received log information includes both: at least a portion of the data contained within the log information; and a summary of the log information. In some examples, a summary of the log information is received from log summarization device 70. In some examples, a summary of the log information is received from data source 45. In some examples, a summary of the log information is provided by the same vendor as the log information itself.

[0068] In some examples, log information is received from the OEM, along with indications of suspicious anomalies, threats, attacks, or other security-related issues.

[0069] In some examples, log messages are received at predetermined intervals. In other examples, log messages are received at intervals of less than 10 seconds.

[0070] In some examples, this information may be provided along with configuration information about the target of interest. In some examples, the received configuration information may include, but is not limited to, one or more of the following: A. Operating system-related configuration settings, such as: i. The scheduling strategy used, such as preemptive round-robin, preemptive shortest remaining time first, non-preemptive; and / or ii. Stack size for several OS tasks to be scheduled.

[0071] B. Configuration settings related to application software components, such as: i. Communication ports that may be connected to other software components; and / or ii. Triggering API functions, such as triggering a loop function every 100 milliseconds.

[0072] C. Configuration settings related to network topology, such as: i. Communication nodes participating in the system; and / or ii. Baud rate of the CAN bus.

[0073] D. Configuration settings related to the data being transmitted, such as: i. The start bit and length of the signals contained in the Protocol Data Unit (PDU); and / or ii. Physical SI units associated with signal values.

[0074] E. Configuration settings related to the management program, such as: i. CPU allocated to the virtual environment; and / or ii. The amount of volatile memory (RAM) allocated for the virtual environment.

[0075] F. Configuration settings related to virtualization solutions, such as: i. Security policies to be applied to the virtualization environment, such as whether to allow client systems network access; and / or ii. Options for interaction between virtual environments, such as allowing copying and pasting between two virtualized desktop systems.

[0076] G. Configuration information that can be loaded at runtime, such as: i. A web server configuration file specifying the location of the webpage source or the server's listening port; and / or ii. Final view settings for graphical applications, such as Windows Calculator (Standard, Scientific, Programmer).

[0077] H. Startup configuration information, for example: i. Command-line options to pass to the application at startup, such as the option "-v" to launch the application in verbose mode, producing additional user-readable output about the activities it performs; and / or ii. Scheduler configurations for automatically starting applications, such as a corn job that performs a file system backup at midnight.

[0078] I. Runtime environment configuration information, such as: i. Set environment variables to pass information to the application, such as the location of the license file to be used; and / or ii. For the configuration settings of the sandbox solution, specify which parts of the file system the application can access.

[0079] J. Security-related configuration options, such as: i. Cryptographic primitives used by the target of interest; ii. Authentication mechanisms for network packets, such as AUTOSAR SecOC; iii. Use of certificates, for example, in diagnostics; iv. Protocols used for network security, such as using TLS to encrypt network traffic; v. Secure storage options, such as using an HSM as a storage for private keys.

[0080] In some examples, event information includes the ID of the signal exhibiting anomalous behavior, and the signal's value over a specific time period. In some examples, event information includes the path the signal takes through the system. In some examples, event information includes a request used to determine the likelihood of a malicious attack indicating anomalous behavior.

[0081] In some examples, in stage 1010, the anomaly analysis device 80 analyzes data associated with a corresponding target of interest to identify the presence or absence of non-malicious causes for the identified anomalous behavior. In some examples, the analysis is performed in relation to a network model of the target of interest stored in memory 40.

[0082] In some examples, data retrieval device 90 retrieves data associated with a corresponding target of interest from memory 40, data source 45, and / or data source 95. In some examples, data retrieval device 90 generates multiple keywords, and the search engine of data retrieval device 90 retrieves data associated with a corresponding target of interest based at least on the generated keywords. In some examples, keywords are extracted from received information such as signal IDs and / or systems / components associated with the corresponding signals.

[0083] In some examples, as described above, a list of libraries associated with the target of interest is input into LLM 85, such that the generated keywords are associated with the list of libraries. In a non-limiting example, LLM 85 can be prompted to generate keywords from the received information by requesting a list of keywords that can optionally be associated with the list of libraries. For example, if the received information contains information about vulnerabilities associated with multiple libraries, LLM 85 can generate only keywords related to the vulnerabilities associated with libraries in the list of libraries associated with the assets. In some examples, processor 30 inputs a list of libraries into LLM 85.

[0084] In some examples, as described above, the asset's specification table is entered into LLM 85, causing the generated keywords to be associated with the specification table. In some examples, as described above, TARA data associated with the asset (e.g., TARA) is entered into LLM 85, causing the generated keywords to be associated with the TARA data. In some examples, as described above, a network model of the asset or a system containing the asset is entered into LLM 85, causing the generated keywords to be associated with the network model.

[0085] In some examples, LLM 85 generates keywords by extracting words from the textual description of the received information and by providing terms associated with the received information. For example, LLM 85 can be fine-tuned to provide keywords that are context-dependent on the received information. This can include known terms or attack steps associated with terms contained within the received information. For instance, if the received information includes a description of a vulnerability in a first type of attack against a first function, LLM 85 can provide keywords related to other types of attacks against the first function. Additionally, LLM 85 can provide keywords related to the first type of attack but for different functions communicating with the first function.

[0086] Although examples of keywords generated by LLM 85 or LLM alone have been described above, this does not imply any limitation in any way, and processor 30 may initiate any suitable algorithm for generating keywords from the received information. As described above regarding LLM 85, such an algorithm may generate keywords associated with: a list of libraries, a specification table, and / or TARA data associated with assets.

[0087] In some examples, multiple keywords are generated by the corresponding algorithms of LLM 85 or processor 30, and processor 30 further filters the generated keywords based at least in part on the list of libraries, specification tables, TARA data, and / or network models.

[0088] As described above, in some examples, the retrieved data includes error reports associated with the target of interest. In some examples, the anomaly analysis device 80 analyzes the corresponding error reports to determine if there is an explanation for the anomalous behavior. For example, if an anomaly is indicated for a corresponding parameter, and there are error reports indicating that an error associated with the corresponding parameter has been identified, the anomaly analysis device 80 can determine that the identified anomalous behavior may have a non-malicious cause.

[0089] In some examples, the anomaly analysis device 80 can analyze the error report by analyzing the text description contained within it. In some examples, the anomaly analysis device 80 uses natural language processing, naive text parsing, large language models, etc., to analyze the text description. In some examples, the memory 40 already stores a list of text descriptions of parameters associated with various errors, where each description is associated with a corresponding parameter. The anomaly analysis device 80 then compares the text description of the received error report with the stored list of text descriptions to identify one or more parameters affected by the error.

[0090] In some examples, the anomaly analysis device 80 then compares the value associated with a parameter within the received log information or in supplementary information received with the log information with the associated value of the parameter in the error report. For example, if there is an error report indicating that a component may have failed at a temperature exceeding a predetermined threshold, the anomaly analysis device 80 compares the measured temperature value contained in the log information with the predetermined threshold. If the comparison indicates that the measured temperature does indeed exceed the predetermined threshold, the anomaly analysis device 80 concludes that the anomalous behavior identified in the corresponding component may have a non-malicious cause. It should be noted that the term "value associated with a parameter" is intended to include, but is not limited to, any of the following: the value of the parameter; and one or more other values ​​derived from the parameter value, such as Fourier transform windows, first-order difference estimators, or operators / functions for other applications.

[0091] In another example, the anomaly analysis device 80 extracts metadata, such as production year, production location, and nature of the problem, from error reports associated with a component / system; it then compares this metadata with the component / system where the anomaly was found. For example, if the error report indicates that a component with a production year prior to a predetermined year has certain problems that could explain the anomalous behavior, the anomaly analysis device 80 concludes that there may be a non-malicious cause for the anomalous behavior identified in the corresponding component, as described above.

[0092] In some examples, the score associated with the anomalous behavior is adjusted based on the reliability value of the error report, which indicates how reliable the error report is. In other examples, the score increases as the reliability value of the error report decreases, because the error report provides less explanation for the anomalous behavior due to its low reliability.

[0093] In some examples, the retrieved error report is retrieved along with its reliability value. In some examples, the anomaly analysis device 80 determines the reliability value for the error report. In some examples, the reliability value is determined based on the source of the error report. In some examples, the memory 40 stores multiple reliability values, each associated with a corresponding source of the error report. For example, if an error report is retrieved from an OEM's database, its reliability value may be higher than that of an error report retrieved from the internet.

[0094] In some examples, the data retrieval device 90 may optionally retrieve information about the age of a component that has transmitted a signal from a corresponding specification table or SBOM, as described above. In some examples, the information about the component's age includes the component's age and survival function values, which indicate the component's quality changes over its service life. In some examples, survival functions are provided in a lookup table. In some examples, the survival function values ​​are provided as results of a corresponding survival function indicating the component's survival probability. Note that survival function values ​​associated with a component can be extracted from the component's survival function and / or the survival function of the system / subsystem containing the component.

[0095] The above has described survival function values ​​associated with the age of a component; however, this is not intended to impose any limitations. In some examples, the data retrieval device 90 retrieves corresponding survival function values ​​for various physical properties associated with the component. For example, a survival function can be provided for the temperature of the component, indicating the probability of survival of the component or the system / subsystem containing the component for a given temperature value.

[0096] In some examples, survival function values ​​can also be correlated with other factors, such as general information about the component. For instance, a component (such as a sensor) may be known to have a high probability of failure, so its survival function value may be low, regardless of its age or other physical properties.

[0097] In some examples, the anomaly analysis device 80 analyzes survival function values ​​to determine whether an event can be explained by the age or other physical properties of a component. In some examples, if the survival function value t has exceeded a predetermined threshold, the anomaly analysis device 80 outputs an indication that the anomalous behavior may be due to the age or current state of the component. In some examples, the anomaly analysis device 80 sets a score for the event based at least in part on the component's survival function values. In some examples, the score is set as a predetermined function of one or more current survival function values ​​of the component. In some examples, one or more survival function values ​​are weighted using appropriate weights as a variable in the scoring function.

[0098] In some examples, the anomaly analysis device 80 determines whether the statistical model used to identify anomalous behavior (described below) might be unsuitable for the corresponding target or component of interest. For example, if the associated data includes indications generated based on traffic data from a system with 1-2 years of use, and the target of interest is outside this range, the anomaly analysis device 80 may determine that any anomalous behavior detected based on the statistical model is likely caused by non-malicious reasons because the anomaly was detected by an incorrect statistical model. In some examples, the score associated with the anomalous behavior is updated based on this determination.

[0099] In some examples, the anomaly analysis device 80 identifies components associated with the identified anomalous behavior within an associated SBOM or specification sheet, such as components sending traffic that includes anomalous behavior. In some examples, the anomaly analysis device 80 controls the data retrieval device 90 to retrieve data about the identified components from one or more data sources 95. In some examples, the retrieved data may include error reports, due dates, previous instability reports, etc. If a problem is identified in the corresponding component (such as an error or other structural problem), the anomaly analysis device 80 can determine that the identified anomalous behavior may have a non-malicious cause.

[0100] In some examples, the anomaly analysis device 80 analyzes the stored configuration information to determine if there is an explanation for the anomalous behavior.

[0101] For example, the anomaly analysis device 80 can retrieve the transmission attributes of a message from a network database. In an illustrative example, the message is configured with a 100 ms transmission period and an additional trigger mechanism for a change in one of the signals contained in the message. According to the corresponding statistical model, the message typically arrives with the expected 100 ms period. However, when the signal configured to trigger transmission upon change (e.g., a signal associated with a control warning lap) changes its value, the message is sent not only with the 100 ms period but also upon the change. Therefore, if the message is sent less than 100 ms later, the anomaly analysis device 80 can identify that the corresponding signal has been changed, and thus there is a non-malicious interpretation of premature message delivery. The anomaly analysis device 80 can output an indication of this interpretation.

[0102] On the other hand, if no additional trigger is configured to send the corresponding message, the anomaly analysis device 80 can output an indication that no non-malicious interpretation has been identified.

[0103] In some examples, the anomaly analysis device 80 can identify whether a specific value is valid within a specification table. For instance, in a scenario where the target of interest is an electric vehicle being charged at a charging station and an alert is received regarding an unencrypted credit card number, the anomaly analysis device 80 can analyze the relevant specification table to determine if there is a valid scenario where unencrypted credit card information would appear (e.g., during testing) and check if this is the case in the event information. If so, the anomaly analysis device determines that there is a non-malicious interpretation of the event.

[0104] In some examples, the anomaly analysis device 80 analyzes security threat data stored in memory 40 to identify the presence or absence of known security threats associated with the identified anomalous behavior. In some examples, the security threat data includes indications of corresponding outliers or indications of targets of interest associated with the anomalous behavior. In some examples, the security threat data includes attack trees or attack paths, which include such values. In some examples, when the anomaly analysis device 80 identifies security threat data associated with anomalous behavior, the anomaly analysis device 80 outputs an indication of the identified security threat.

[0105] In some examples, the anomaly analysis device 80 searches each of the attack tree or attack path to identify whether a signal associated with the anomaly is present in the attack tree / attack path. For example, the attack tree or attack path may be searched to identify whether it contains the ID of a signal exhibiting an anomalous value. In some examples, if a signal associated with the anomaly is present in the attack tree / attack path, the anomaly analysis device 80 outputs an indication of the presence of a potential security threat identified within the stored attack tree / attack path, as described above.

[0106] In some examples, the anomaly analysis device 80 identifies whether the system 10 has received an indication that an event existed in a level prior to the anomaly in the attack tree / attack path. For example, the anomaly analysis device 80 may identify whether the event's reception ID exists within these levels of the attack tree / attack path. In some examples, if the anomaly analysis device 80 identifies that an event did not occur within these levels of the attack tree / attack path, the anomaly analysis device 80 outputs an indication that no such event was identified. Alternatively or additionally, the anomaly analysis device 80 may not output an indication that a possible security threat has been identified because no attack path / attack tree has been identified as complete.

[0107] In some examples, the score associated with an event or anomalous behavior is increased based on the identification of events existing within such a hierarchy, and the score is decreased based on the absence of events existing within such a hierarchy.

[0108] In some examples, where the security threat data includes threat intelligence data, the anomaly analysis device 80 identifies whether a threat (e.g., a virus or other vulnerability) is associated with a component sending traffic containing anomalies. In some examples, vulnerabilities are identified from a Common Vulnerability Disclosure (CVE) system.

[0109] In some examples, when the LLM 85 is configured to summarize threat intelligence data, the LLM 85 generates a summary of the identified threat intelligence data. In some examples, the anomaly analysis device 80 inputs a request to the LLM 85 to generate this summary. In some examples, the LLM 85 sorts the identified threat intelligence data and / or the summary of the threat intelligence data. In some examples, the anomaly analysis device 80 inputs a request to the LLM 85 to perform sorting. In some examples, the LLM 85 is fine-tuned to sort the threat intelligence data. In some examples, sorting values ​​are assigned to each part of the threat intelligence data, indicating its sorting. In some instances, sorting values ​​are assigned by either the LLM 85 or the anomaly analysis device 80.

[0110] In some examples, the LLM 85 determines the relevance and / or summary of threat intelligence data. In some examples, the determined relevance is the relevance to data on anomalous behavior and / or events. For example, the LLM 85 may determine the degree of relevance of threat intelligence data to any of, but not limited to: a component; a system / subsystem containing the component; a signal exhibiting anomalous behavior; or the type of anomalous behavior or event. In some examples, the LLM 85 is fine-tuned to determine the relevance of threat intelligence data. In some examples, relevance values ​​are assigned to each part of the threat intelligence data, indicating its relevance. In some examples, relevance values ​​are assigned by the LLM 85 or the anomaly analysis device 80.

[0111] In some examples, where the language analysis device 85 is implemented as a natural language processor, ranking and / or relevance can be determined based on predetermined criteria, such as: the overall credibility and / or reliability of the threat intelligence data source; the credibility and / or reliability of the threat intelligence data source in relation to the type of the corresponding component, system, subsystem, anomalous behavior, and / or event; and / or the number of generated keywords appearing in the threat intelligence data.

[0112] In some examples, when an anomaly analysis device 80 identifies a threat, the anomaly analysis device 80 outputs an event indication of a potential security threat. In some examples, it further outputs information about the threats identified within the threat intelligence data. In some examples, the output information includes: a summary of the generated threat intelligence data; and further outputs the determined ranking.

[0113] Therefore, as mentioned above, in some examples, the output indicates whether there is a non-malicious interpretation of the event / abnormal behavior. In some examples, the output indicates whether there is a malicious interpretation of the event / abnormal behavior. In some examples, a score is generated for the event / abnormal behavior, which indicates the likelihood that the event / abnormal behavior indicates a malicious attack.

[0114] In some examples, as described above, the anomaly analysis device 80 determines a score for an event / abnormal behavior. In some examples, as described above, the score is determined at least in part based on a predetermined function comprising multiple variables. In some examples, the anomaly analysis device 80 utilizes the score as corresponding weights in a caching algorithm (such as a weighted least recently used (LRU) caching algorithm) to determine the order of events / identified abnormal behaviors. In some examples, as described above, information about events / abnormal behaviors is output based on their order.

[0115] In some examples, the anomaly analysis device 80 further filters data regarding events / abnormal behavior based at least in part on indications of non-malicious and / or malicious interpretations. In some examples, the anomaly analysis device 80 only outputs data regarding events / abnormal behavior to the response unit 98 if no non-malicious interpretation is identified. In some examples, in the case of a determined score, the anomaly analysis device 80 only outputs data regarding events / abnormal behavior to the response unit 98 if the score exceeds a corresponding threshold.

[0116] In some examples, where risk values ​​are assigned to corresponding signals and / or components, filtering is at least partially based on the assigned risk values. For example, if high-risk values ​​exist, less data will be filtered.

[0117] In some examples, data is filtered at least in part based on the received user input. In some examples, the user input allows selection of how the data is filtered. For example, the user can select one of several different options for filtering the data. In some examples, for each component, signal, and / or system, the user can select one of several different options for filtering the data, as will be further described below.

[0118] In some examples, at stage 1020, the anomaly analysis device 80 generates an adjustment request for adjusting log information and / or network traffic data sent to the data enrichment system 10. In some examples, a profile is generated for the data logger, intrusion detection system (IDS), and / or intrusion prevention system (IPS), indicating what adjustments are requested, and the adjustment request includes the generated profile.

[0119] In some examples, the adjustment request includes a request for additional logging information and / or network traffic data associated with the identified anomalous behavior. In some examples, the adjustment request identifies the component associated with the identified anomalous behavior and requests additional logging information and / or network traffic data associated with that component. In some examples, the adjustment request identifies the sender of one or more messages exhibiting the anomalous behavior.

[0120] In some examples, the adjustment request includes a request to add and / or remove one or more signals from received log information and / or network traffic data. For example, if anomalous behavior associated with a particular signal is identified, the adjustment request may include a request to provide additional log information and / or network traffic data including that signal. In another example, alternatively or additionally, if no anomalous behavior associated with a particular signal is identified for at least a predetermined period of time, the adjustment request may include a request to remove the signal from future log information and / or network traffic data.

[0121] In some examples, the adjustment request includes an indication of the risk value of the signal to be sent. Specifically, in some examples, each signal within the target of interest has a corresponding risk value associated with it. The risk value indicates the degree to which the corresponding signal is vulnerable to attack. In some examples, the adjustment request includes a request to send log information and / or network traffic data associated with signals having risk values ​​above a predetermined threshold.

[0122] For example, if abnormal behavior is detected regarding a specific signal and no non-malicious explanation is identified, the anomaly analysis device 80 may request log information and / or network traffic data associated with a signal having a lower risk value than the previously requested risk value. Conversely, if no abnormal behavior is detected within a predetermined time period, and / or a non-malicious explanation is identified for the detected abnormal behavior, the anomaly analysis device 80 may request log information and / or network traffic data associated with a signal having a higher risk value than the previously requested risk value.

[0123] In some examples, the adjustment request may optionally be sent to the data source 45 via the data interface 20.

[0124] In some examples, the adjustment request is generated based at least in part on a predetermined statistical relationship between two or more signals. In some examples, memory 40 already stores information about the statistical relationships between different signals.

[0125] For example, if abnormal behavior is detected regarding a specific signal and no non-malicious interpretation is identified, the anomaly analysis device 80 may request log information and / or network traffic data associated with signals that have a predetermined statistical relationship with the specific signal.

[0126] In some examples, the adjustment request is generated at least in part based on user input received from response unit 98.

[0127] In phase 1030, at least a portion of network traffic is received from the target of interest at data interface 20. It should be noted that phases 1000 and 1030 may be executed together or at different times. In some examples, phase 1030 is executed periodically, and phase 1000 is executed continuously.

[0128] In stage 1040, one or more statistical models stored in memory 30 are updated, at least in part, based on the network traffic received in stage 1010. In some examples, model generation device 60 inputs the received network traffic data into the corresponding statistical model and updates the model based on the newly received data. In some examples, the statistical model updates are performed periodically.

[0129] In some examples, where the received information in stage 1000 includes signal associated values, in stage 1050, the statistical analysis device 65 compares the received information about the log information in stage 1000 with the corresponding statistical model stored in memory 40. In some examples, the statistical analysis device 65 estimates the goodness of fit between the signal associated values ​​and the corresponding statistical model.

[0130] As stated above regarding the term "value associated with a parameter," the term "signal-associated value" is intended to include, but is not limited to, any of the following: the value of a signal; and one or more other values ​​derived from the signal value, such as Fourier transform windows, first differences estimators, or operators / functions for other applications. It should also be noted that the received signal-associated value does not imply limitation to a single signal at a specific time, but specifically implies the inclusion of multiple signals over a predetermined time period.

[0131] In some examples, at stage 1060, statistical analysis device 65 identifies anomalous behavior within the system, at least in part based on the results of the comparison in stage 1050. In some examples, anomalous behavior is identified by recognizing signal correlation values ​​within the received information that are not expected by the corresponding statistical model.

[0132] In some examples, at stage 1070, as described above, the statistical analysis device 65 and / or the anomaly analysis device 80 output information related to the identified anomalous behavior. In some examples, the information is output to the response unit 98. In some examples, the received information of stage 1000 associated with the identified anomalous behavior is output. In some examples, the information includes indications describing the anomalous behavior.

[0133] In some examples, as described above, the output information includes a score indicating the probability that the anomalous behavior is associated with a malicious attack and / or the probability that the anomalous behavior is not associated with a malicious attack. In some examples, as described above, the output information includes an indication of whether there is a non-malicious interpretation and / or a malicious interpretation of the anomalous behavior.

[0134] Figure 3 A high-level flowchart of the method for performing security analysis is shown. In phase 1100, data is sent to response unit 98. In some examples, as described above with respect to phases 1000-1070, the data may include log information indicating anomalous behavior and optionally an indication of whether the anomalous behavior has a non-malicious interpretation, a malicious interpretation, or no interpretation.

[0135] In some examples, the data also includes information received along with the information received in phase 1000. For example, if OEM supply log information is accompanied by indications of suspicious anomalies, threats, attacks, or other security-related issues, those indications are also output. In some examples, where the indication includes a risk level associated with the suspicious anomaly, the risk level is also output.

[0136] In phase 1110, security analysis of the data from phase 1000 is performed in response unit 98. In some examples, as is known to those skilled in the art, the security analysis in response unit 98 is performed by trained personnel and / or various machine learning models.

[0137] In phase 1120, input is received from response unit 98 at data interface 20. In some examples, the received input includes an indication of whether the identified anomalous behavior is the result of a security analysis caused by malicious activity. In some examples, the received input includes an indication of whether the problem causing the anomalous behavior (whether malicious or non-malicious) has been mitigated.

[0138] For example, if the anomalous behavior is caused by an attack, and the security vulnerability allowing the attack can be mitigated, the received input includes such an indication. In a similar example, if the anomalous behavior is caused by a bug, and the bug has been fixed, the received input includes such an indication.

[0139] In stage 1130, future analysis of the information received in stage 1000 is based at least in part on the input received in stage 1120. For example, a request for additional information as described above with respect to stage 1020 can be updated based on the received input. In an illustrative example, if the input received from response unit 98 includes an indication that an attack has been detected, additional information associated with the detected anomalous behavior is requested, as described above.

[0140] In some examples, if the input received from the response unit 98 includes an indication that a particular identified anomalous behavior is not due to malicious intent, the anomaly analysis device 80 determines that similar anomalous behavior identified in future information may have a non-malicious interpretation. In some examples, such information is then not sent to the response unit 98. Alternatively, this information is sent along with an indication that similar anomalous behavior was previously identified by the response unit 98 as non-malicious. In some examples, as described above, future filtering of data to be sent to the response unit 98 is at least in part based on the input received in stage 1120.

[0141] In some examples, one or more of the aforementioned statistical models may be adjusted, at least in part, based on input received from response unit 98. For example, if an analyst determines that the identification of anomalous behavior is due to a software or hardware update in a relevant component or system, model generation device 60 may update one or more corresponding models associated with the software or hardware update. In some examples, statistical model 53 may be updated by removing one or more statistical sub-models 54 and / or adding one or more statistical sub-models 54. In some examples, the input received from response unit 98 also includes instructions for updating the corresponding models. In some examples, based at least in part on such instructions regarding software or hardware updates, model generation device 60 relearns the model based on the network traffic data received in phase 1030.

[0142] In some examples, at least in part based on such an instruction of a software or hardware update, the model generation device 60 may replace one or more statistical sub-models 54 with a new statistical sub-model 54 associated with the software or hardware update. In some examples, when a statistical sub-model 54 is removed from a corresponding statistical model 53, the model generation device 60 removes the statistical relationships between the remaining statistical sub-models 54 and the signals in the removed statistical sub-models 54. In some examples, when a statistical sub-model 54 is added to a corresponding statistical model 53, the model generation device 60 learns the statistical relationships between the added statistical sub-model 54 and the signals in the other statistical sub-models 54.

[0143] In some examples, one or more statistical models can be updated similarly based on input from response unit 98, indicating that the corresponding model used is not suitable for a particular version of the target of interest. For example, an analyst may determine that a sub-model of a particular type of vehicle produces certain signal values ​​that differ from other sub-models of that vehicle. In this case, model generation device 60 can create a separate statistical model for this newly defined group. Similarly, anomaly analysis device 80 can create a new network model for this version of the target of interest. In some examples, a new network model is created by copying the original network model and adjusting it based on input received from response unit 98.

[0144] Some examples of the disclosed technology The following are some examples of the above-described embodiments. It should be noted that a feature of an isolated example, or a combination of one or more features of that example, or optionally a combination of one or more features of that example with one or more features of the examples below, also falls within the scope of this application.

[0145] Example 1. A data enrichment method comprising: receiving event information associated with an event in a target of interest; analyzing data associated with the target of interest, at least in part based on the received event information, to identify: the impact of the event within the target of interest, and one or more existing mitigation measures associated with the event; and outputting information about the event, the output information including the identified impact and the identified one or more mitigation measures, wherein the event information includes information about anomalous behavior detected in the target of interest or information about vulnerabilities detected in the target of interest, and wherein the analyzed data is based on multiple sources.

[0146] Example 2. The method described in any of the examples in this document, particularly Example 1, wherein the multiple sources include the technical artifacts of the target of interest and the product architecture associated with the target of interest.

[0147] Example 3. The method described in any of the examples in this document, particularly Example 1 or 2, wherein multiple sources include risk assessment analysis.

[0148] Example 4. The method described in any of the examples in this document, particularly any of Examples 1-3, wherein multiple sources include a Software Bill of Materials (SBOM).

[0149] Example 5. The method described according to any of the examples in this document, particularly Example 4, also includes identifying components within the SBOM associated with the event, wherein the data associated with the target of interest includes information about the identified components.

[0150] Example 6. The method described in any of the examples in this document, particularly any of Examples 1-5, wherein multiple sources include a specification sheet associated with the target of interest.

[0151] Example 7. The method described according to any of the examples in this document, particularly any of Examples 1-6, wherein multiple sources include error reports.

[0152] Example 8. The method described according to any of the examples in this document, particularly any of Examples 1-7, wherein the multiple sources include data received from an electronic control unit (ECU) within the target of interest.

[0153] Example 9. The method described in any of the examples in this document, particularly Example 8, wherein the multiple sources include data received from one or more sensors associated with the ECU.

[0154] Example 10. The method described in any of the examples in this document, particularly any of Examples 1-9, wherein multiple sources include threat intelligence.

[0155] Example 11. The method described according to any of the examples in this document, particularly any of Examples 1-10, wherein the impact of the event includes: risk impact; and an indication of which of the multiple parts of the target of interest may be affected by the event.

[0156] Example 12. The method described according to any of the examples herein, particularly Example 11, also includes identifying attack steps in a plurality of attack paths associated with a target of interest, wherein an indication of which part of the target of interest may be affected is based at least in part on the attack path containing the identified attack steps.

[0157] Example 13. The method according to any of the examples herein, particularly any of Examples 1-10, further includes identifying attack steps within a plurality of attack paths associated with the target of interest, wherein the impact of the event includes an indication of which of a plurality of parts of the target of interest may be affected by the event, and wherein the indication of which part of the target of interest may be affected is based at least in part on the attack path containing the identified attack steps.

[0158] Example 14. The method according to any of the examples in this document, particularly any of Examples 1-10, further includes identifying attack steps within a plurality of attack paths associated with the target of interest, wherein the output information includes the attack paths containing the identified attack steps.

[0159] Example 15. The method described according to any of the examples in this document, particularly any one of Examples 12-14, wherein the attack identification step includes: identifying a corresponding component within the target of interest based at least in part on the received event information and the network model of the target of interest; and identifying attack steps associated with the identified component.

[0160] Example 16. The method according to any of the examples in this document, particularly any one of Examples 1-15, further includes: analyzing event information associated with the target of interest, at least in part based on the received information, to identify the presence or absence of a non-malicious cause of the anomalous behavior, wherein the output information includes an indication of the presence or absence of the identified non-malicious cause.

[0161] Example 17. The method described according to any example in this document, particularly Example 16, further includes analyzing security threat data to identify the presence or absence of known security threats associated with the identified event, wherein the output information regarding the identified event includes an indication of the presence or absence of the identified known security threat.

[0162] Example 18. The method described in any of the examples in this document, particularly any of Examples 1-17, wherein the output information is output to an item or event manager.

[0163] Example 19. The method described in any of the examples in this document, particularly Example 18, wherein the output information is output to the Security Operations Center (SOC) or the incident response team.

[0164] Example 20. The method described according to any of the examples in this document, particularly Example 18 or 19, also includes receiving data from a matter or event manager, wherein the data analysis is based at least in part on the data received from the matter or event manager.

[0165] Example 21. The method according to any of the examples in this document, particularly any of Examples 1-20, further includes receiving log information associated with a target of interest; comparing the received log information with respect to one or more log messages with a corresponding one of a plurality of statistical models; and identifying anomalous behavior based at least in part on the result of the comparison with the statistical models.

[0166] Example 22. The method described according to any of the examples in this document, particularly Example 21, further includes: receiving a portion of network traffic from a target of interest; and updating one or more of a plurality of statistical models based at least in part on the received network traffic.

[0167] Example 23. The method described according to any of the examples in this document, particularly Example 21 or 23, wherein the received log information includes the values ​​of one or more signals.

[0168] Example 24. The method described according to any of the examples in this document, particularly any of Examples 21-23, wherein the received log information includes a summary.

[0169] Example 25. The method described according to any of the examples in this document, particularly any of Examples 21-24, further includes: outputting an adjustment request for adjusting the received log information, based at least in part on the identified anomalous behavior.

[0170] Example 26. The method described in any of the examples in this document, particularly Example 25, wherein the adjustment request is based at least in part on the identified anomalous behavior.

[0171] Example 27. The method described in any of the examples in this document, particularly Example 25, wherein the adjustment request is based at least in part on a statistical relationship between a first signal and a second signal, the first signal being associated with the identified anomalous behavior.

[0172] Example 28. The method described according to any of the examples in this document, particularly any of Examples 21-27, wherein each of the plurality of statistical models includes a plurality of corresponding statistical sub-models.

[0173] Example 29. The method described according to any example in this document, particularly Example 28, further includes: for a given one of a plurality of statistical models, removing one of the corresponding plurality of statistical sub-models.

[0174] Example 30. The method described according to any of the examples in this document, particularly Example 28 or 29, further includes: for a given one of a plurality of statistical models, adding the corresponding statistical submodel to the corresponding statistical model.

[0175] Example 31. A data enrichment system includes one or more processors and a memory, wherein the memory stores a plurality of instructions that, when executed by one or more processors, cause one or more processors to perform a plurality of steps, the plurality of steps including: receiving event information associated with an event in a target of interest; analyzing data associated with the target of interest, at least in part based on the received event information, to identify: the impact of the event within the target of interest, and one or more existing mitigation measures associated with the event; and outputting information about the event, the output information including the identified impact and the identified one or more mitigation measures, wherein the event information includes information about anomalous behavior detected in the target of interest or information about vulnerabilities detected in the target of interest, and wherein the analyzed data is based on multiple sources.

Claims

1. A data enrichment method, the method comprising: Receive event information associated with events in the target of interest; Based at least in part on the received event information, analyze the data associated with the target of interest to identify: the impact of the event within the target of interest, and one or more existing mitigation measures associated with the event; and Output information about the event, including the identified impacts and one or more identified mitigation measures. The event information includes information about anomalous behavior detected in the target of interest or information about vulnerabilities detected in the target of interest, and The data analyzed is based on multiple sources.

2. The method according to claim 1, wherein, The multiple sources include the technical artifacts of the target of interest and the product architecture associated with the target of interest.

3. The method according to claim 1 or 2, wherein, The multiple sources include risk assessment analysis.

4. The method according to any one of claims 1-3, wherein, The multiple sources include the Software Bill of Materials (SBOM).

5. The method of claim 4, further comprising identifying components within the SBOM associated with the event. in, The data associated with the target of interest includes information about the identified components.

6. The method according to any one of claims 1 to 5, wherein, The multiple sources include a specification sheet associated with the target of interest.

7. The method according to any one of claims 1-6, wherein, The multiple sources include error reports.

8. The method according to any one of claims 1 to 7, wherein, The multiple sources include data received from the electronic control unit (ECU) within the target of interest.

9. The method according to claim 8, wherein, The multiple sources include data received from one or more sensors associated with the ECU.

10. The method according to any one of claims 1 to 9, wherein, The multiple sources include threat intelligence.

11. The method according to any one of claims 1-10, wherein, The impact of the event includes: Risk impact; and An indication of which of the multiple parts of the target of interest might be affected by the event.

12. The method of claim 11, further comprising the step of identifying attacks within a plurality of attack paths associated with the target of interest. in, The indication of which part of the target of interest might be affected is at least in part based on an attack path that includes the identified attack steps.

13. The method according to any one of claims 1 to 10, further comprising the step of identifying attacks within a plurality of attack paths associated with the target of interest. in, The impact of the event includes an indication of which of the multiple parts of the target of interest might be affected by the event, and The indication of which part of the target of interest might be affected is at least in part based on the attack path that includes the identified attack steps.

14. The method according to any one of claims 1 to 10, further comprising the step of identifying attacks within a plurality of attack paths associated with the target of interest. in, The output information includes the attack path containing the identified attack steps.

15. The method according to any one of claims 12 to 14, wherein, The steps to identify the attack include: Based at least in part on the received event information and the network model of the target of interest, identify the corresponding components within the target of interest; and Identify the attack steps associated with the identified components.

16. The method according to any one of claims 1 to 15, further comprising: Based at least in part on the received information, the event information associated with the target of interest is analyzed to identify the presence or absence of non-malicious causes for the anomalous behavior. The output information includes an indication of the presence or absence of the identified non-malicious cause.

17. The method of claim 16, further comprising analyzing security threat data to identify the presence or absence of known security threats associated with the identified events. in, The output information regarding the identified events includes an indication of the presence or absence of the identified known security threats.

18. The method according to any one of claims 1-17, wherein, The output information is sent to the event or event manager.

19. The method according to claim 18, wherein, The output information is sent to the Security Operations Center (SOC) or the incident response team.

20. The method of claim 18 or 19, further comprising receiving data from the event or event manager. in, The analysis of the data is based, at least in part, on data received from the event or event manager.

21. The method according to any one of claims 1 to 20, further comprising: Receive log information associated with the target of interest; The received log information about one or more log messages is compared with a corresponding one of multiple statistical models; The anomalous behavior is identified, at least in part, based on the results of comparison with the statistical model.

22. The method of claim 21, further comprising: Receive a portion of network traffic from the target of interest; and One or more of the plurality of statistical models are updated, at least in part, based on the received network traffic.

23. The method according to claim 21 or 23, wherein, The received log information includes the values ​​of one or more signals.

24. The method according to any one of claims 21-23, wherein, The received log information includes a summary.

25. The method according to any one of claims 21-24, further comprising, at least in part, outputting an adjustment request for adjusting the received log information based on the identified anomalous behavior.

26. The method of claim 25, wherein, The adjustment request is based at least in part on the identified anomalous behavior.

27. The method according to claim 25, wherein, The adjustment request is based, at least in part, on a statistical relationship between a first signal and a second signal, the first signal being associated with the identified anomalous behavior.

28. The method according to any one of claims 21-27, wherein, Each of the plurality of statistical models includes a plurality of corresponding statistical sub-models.

29. The method of claim 28, further comprising: For a given statistical model, remove one of the corresponding statistical sub-models.

30. The method according to claim 28 or 29, further comprising: For each of the multiple statistical models, a corresponding statistical sub-model is added to that statistical model.

31. A data enrichment system, comprising one or more processors and memory, wherein, The memory stores a plurality of instructions that, when executed by the one or more processors, cause the one or more processors to perform a plurality of steps, the plurality of steps including: Receive event information associated with events in the target of interest; Based at least in part on the received event information, data associated with the target of interest is analyzed to identify: the impact of the event within the target of interest, and one or more existing mitigation measures associated with the event; and Output information about the event, including the identified impacts and one or more identified mitigation measures. The event information includes information about anomalous behavior detected in the target of interest or information about vulnerabilities detected in the target of interest, and The data analyzed is based on multiple sources.