System and method for providing threat intelligence

By combining large language models and network models to generate keywords and retrieve data sources, the problem of low efficiency in prioritizing network security traffic in existing technologies is solved, enabling more efficient threat intelligence generation and analysis, and improving network security data processing capabilities.

CN122003676APending Publication Date: 2026-05-08C2A SEC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
C2A SEC LTD
Filing Date
2024-09-05
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies suffer from inefficiency and insufficient information processing in prioritizing network security traffic, particularly lacking effective data integration and analysis methods when generating and utilizing threat intelligence.

Method used

By using a large language model (LLM) combined with network models and asset-related data, keywords are generated and retrieved from multiple data sources. The network model is then updated, and generative AI is used to aggregate and filter data to generate threat intelligence.

Benefits of technology

It improves the efficiency of threat intelligence generation and processing, provides more accurate and comprehensive cybersecurity analysis, and supports faster threat identification and response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122003676A_ABST
    Figure CN122003676A_ABST
Patent Text Reader

Abstract

A method for providing threat intelligence, the method comprising: generating a plurality of keywords based at least in part on a network model associated with an asset; searching one or more data sources to retrieve data associated with the network model based at least in part on the generated keywords, the retrieved data including information about threats and / or vulnerabilities associated with the asset; and updating the network model based at least in part on the retrieved data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally pertains to the field of threat intelligence. background

[0002] Threat intelligence involves collecting and analyzing multi-source cybersecurity data using various algorithms. By collecting and analyzing vast amounts of data on current cybersecurity threats and trends, threat intelligence providers gain usable data and insights to help their clients better detect and respond to threats.

[0003] Organizations have a wide range of threat intelligence needs, from low-level information about malware variants currently used in attack campaigns to high-level information designed to inform strategic investments and policymaking. Therefore, threat intelligence can be categorized into three different types: 1. Operational: Operational threat intelligence focuses on the tools (malware, infrastructure, etc.) and techniques that cyber attackers use to achieve their goals. This type of understanding helps analysts and threat hunters identify and understand attack campaigns.

[0004] 2. Strategic: Strategic threat intelligence is advanced and focuses on broad trends within the cyber threat landscape. This type of threat intelligence is geared towards managers (who typically do not have a cybersecurity background) who need to understand their organization's cyber risks as part of their strategic planning.

[0005] 3. Tactical: Tactical threat intelligence focuses on using compromise (IoC) indicators to identify specific types of malware or other cyberattacks. This type of threat intelligence is ingested by cybersecurity solutions and used to detect and block incoming or ongoing attacks. Overview

[0006] Therefore, the primary objective of this invention is to overcome at least some of the shortcomings of existing systems and methods for prioritizing network security traffic. In some examples, this is achieved through a method for providing threat intelligence that includes generating multiple keywords based at least in part on network models associated with assets.

[0007] In some examples, based at least in part on the generated keywords, the method involves searching one or more data sources to retrieve data associated with the network model. In some examples, the retrieved data includes information about threats and / or vulnerabilities associated with the assets.

[0008] In some examples, the method includes updating the network model, at least in part, based on the retrieved data.

[0009] In some examples, the method includes using a large language model (LLM) to summarize at least a portion of the retrieved data.

[0010] Additional features and advantages of the invention will become apparent from the following figures and description.

[0011] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In case of conflict, the specification of this patent application (including the definitions) shall prevail. As used herein, unless the context clearly indicates otherwise, the articles “a” and “an” mean “at least one” or “one or more”. As used herein, “and / or” means any one or more items in a list connected by “and / or”. As an example, “x and / or y” means any element in the set of three elements {(x), (y), (x, y)}. In other words, “x and / or y” means “x, y or both x and y”. As in some examples, “x, y and / or z” means any element in the set of seven elements {(x), (y), (z), (x, y), (x, z), (y, z), (x, y, z)}.

[0012] Furthermore, unless explicitly stated otherwise, "or" refers to inclusive or, not exclusive, or. For example, condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).

[0013] Additionally, the terms "a" or "an" are used to describe elements and components of embodiments of the inventive concept. This is done merely for convenience and to give the general meaning of the inventive concept, and "a" and "an" are intended to include one or at least one, and the singular includes the plural, unless it is obvious that it has a different meaning.

[0014] As used herein, the term “approximately” when referring to a measurable value (e.g., quantity, duration, etc.) means that it includes a deviation of + / -10%, more preferably + / -5%, even more preferably + / -1%, and still more preferably + / -0.1% from the specified value, because such deviation is suitable for performing the disclosed apparatus and / or methods.

[0015] The following embodiments and aspects of systems, tools, and methods are described and illustrated. These embodiments and aspects are intended to be exemplary and illustrative, and not to limit the scope. In various embodiments, one or more of the problems described above have been reduced or eliminated, while other embodiments address other advantages or improvements. Brief description of the attached diagram

[0016] To better understand the invention and to show how it can be practiced, reference will now be made to the accompanying drawings by way of example only, in which similar reference numerals throughout the drawings denote corresponding parts or elements.

[0017] Referring now specifically to the accompanying drawings, it is emphasized that the details shown are by way of example and are merely for the purpose of illustrative discussion of preferred embodiments of the invention, and are presented to provide the most useful and readily understood description of what is believed to be the principles and concepts of the invention. In this respect, no attempt is made to show the structural details of the invention in more detail than necessary for a basic understanding of the invention; the description taken in conjunction with the drawings will enable those skilled in the art to understand how the various forms of the invention can be embodied in practice.

[0018] In the attached diagram: Figure 1 A high-level block diagram of a threat intelligence system based on some examples of this disclosure is shown; Figure 2A A high-level flowchart of a first method for providing threat intelligence, based on some examples of this disclosure, is shown; Figure 2B A high-level block diagram of a module for fine-tuning an LLM, according to some examples of this disclosure, is shown; Figure 2C A high-level flowchart of a method for fine-tuning an LLM, according to some examples of this disclosure, is shown; Figure 2D A high-level flowchart of a method for updating TARA data based on one or more text attack paths, according to some examples of this disclosure, is shown. Figure 3 A high-level flowchart of a second approach for providing threat intelligence based on network models, according to some examples of this disclosure, is shown; Figure 4A A high-level block diagram of a system for identifying communication paths, according to some examples of this disclosure, is shown; and Figure 4B A high-level flowchart is shown for a method of identifying communication paths within a system using LLM. Detailed description of a specific example

[0019] In the following description, various aspects of this disclosure will be described. Specific configurations and details are set forth for purposes of explanation in order to provide a thorough understanding of the different aspects of this disclosure. However, it will be apparent to those skilled in the art that this disclosure can be practiced without the specific details presented herein. Furthermore, well-known features may have been omitted or simplified so as not to obscure this disclosure. In the accompanying drawings, similar reference numerals always refer to similar parts. To avoid excessive confusion due to too many reference numerals and leaders on a particular drawing, some components will be introduced through one or more drawings without being explicitly identified in each subsequent drawing containing that component.

[0020] As used herein, the term "asset" means a feature that performs a function or a resource used to perform a function. Such resources may include, for example, keys, configuration files, or any other suitable resources, including hardware or software components. In some examples, an asset may be one or more components, systems, and / or subsystems of a vehicle. An asset may also be a signaling asset, which is a message that can be sent from a source to a destination. In some examples, a signaling asset includes configuration information associated with a message that can be sent.

[0021] As used herein, the term "fine-tuning" refers to a method of training the weights of a pre-trained model on new data, as is known to those skilled in the art. In some examples, as is known to those skilled in the art, fine-tuning involves feeding data into an LLM and prompting the LLM to fine-tune a predetermined subset of its weights based on a loss function. In some examples, as is known to those skilled in the art, the output of the LLM is fed into a second LLM that generates the input of the loss function. Note that this is merely one method for performing fine-tuning and is not intended to be limiting in any way.

[0022] Figure 1 A high-level block diagram of a threat intelligence system 10 is shown. In some examples, the threat intelligence system 10 includes: a data interface 20; a data retrieval module 30; and a large language model (LLM) 40. In some examples, the threat intelligence system 10 includes one or more processors 11 and a memory 12. In some examples, the memory 12 stores instructions that, when read by the processor 11, cause the processor 11 to operate the data interface 20, the data retrieval module 30, and / or the LLM 40. Although only a single LLM 40 is described herein, this is not intended to be limiting in any way, and any number of LLMs can be used to perform the operations described below for the LLM 40. In some examples, the system 10 also includes a user input terminal 45.

[0023] In some examples, data interface 20 includes connections to one or more data sources 50. In some examples, data interface 20 includes one or more of the following: a connection to the Internet; a connection to one or more databases; or a connection to one or more repositories. Databases and / or repositories may be stored within memory 12, in separate memory, and / or on different networks, and data interface 20 includes appropriate connections for any of these options. In some examples, data interface 20 also communicates with external systems, processes, and / or networks.

[0024] In some examples, data source 50 includes any of the following: the Internet; the National Vulnerability Database (NVD); an Open Source Vulnerability (OSV) database; a database of exploit code; one or more code repositories; one or more specification tables associated with the asset, optionally stored on storage 12; and / or risk analysis, such as asset-associated Threat Analysis and Risk Assessment (TARA) data optionally stored on storage 12. Note that while this disclosure references TARA, this is not intended to limit in any way, and any other risk analysis method, such as the Risk Management Framework (RMF), Operational Critical Threat, Asset and Vulnerability Assessment (OCTAVE), or other methods, may be used.

[0025] In some examples, risk analysis data is stored on a dedicated database 60. In some examples, TARA data includes multiple attack trees or attack paths. As defined herein, the term "attack tree" or "attack path" is a branching hierarchical data structure that represents a set of potential methods for implementing an event in which the security of a system is compromised or compromised in a specified manner. As defined herein, an "attack step" is any node in an attack tree or attack path.

[0026] In some examples, system 10 also includes a video-to-text converter 70 that converts speech within a video into text. In some examples, converter 70 includes a speech recognition algorithm. In some examples, converter 70 also includes a corresponding algorithm for recognizing text within frames of the video.

[0027] In some examples, the data retrieval module 30 includes appropriate components and / or instruction sets for retrieving data from the data source 50. In some examples, the data retrieval module 30 includes a search engine. For example, the search engine may be based on a page ranking algorithm or a variant thereof. In some examples, the data retrieval module 30 includes a database management system.

[0028] An illustrative example of LLM 40 (without any restrictions) comes from the LLaMa-2 family of large language models, which is commercially available from Meta AI.

[0029] In some examples, the LLM 40 is fine-tuned based on predefined data. Although the fine-tuning of the trained LLM 40 is described below, this is not intended to be limiting in any way. In some examples, as known to those skilled in the art, the steps for fine-tuning the LLM 40 described below are performed during the training process of the LLM 40, with appropriate modifications.

[0030] In some examples, a list of configuration data (such as software libraries) associated with an asset is input into the LLM 40, as described below, so that the data output by the LLM 40 is associated with these libraries. As used herein with respect to configuration data, the term "associated with an asset" means configuration data accessed by the asset and / or accessed configuration data in which the asset uses information affected by that accessed configuration data. For example, a first process accesses a library and uses the accessed library to obtain certain data, storing the data in a queue. A second process, as part of the asset, retrieves the data from the queue. Thus, in such an example, the library is associated with the asset because the asset retrieves the data associated with the library.

[0031] In some examples, when the LLM 40 outputs data based on a list of configuration data, the list of configuration data is input to the LLM 40 as a soft hint. As used herein, the term "soft hint" refers to a learnable tensor cascaded with the input embedding.

[0032] In some examples, the asset's specification sheet is input into LLM 40, as described below, so that the data output by LLM 40 is related to the specification sheet. In some examples, when prompted to output data based on the specification sheet, the data in the specification sheet is input into LLM 40 as a soft prompt. In some examples, the asset's risk analysis data (such as TARA information) is input into LLM 40, so that the data output by LLM 40 is related to the risk analysis data, as described below. In some examples, when prompted to output data based on the risk analysis data, the risk analysis data or its predetermined parameters are input into LLM 40 as a soft prompt.

[0033] In some examples, a network model is input into the LLM 40, as described below, such that the data output by the LLM 40 is associated with the network model. As used herein, the term "network model associated with an asset" means a model of the asset, or a model of a larger system containing assets (e.g., vehicles). In some examples, the network model includes data about connectivity within the system. In some examples, connectivity is between components (such as ECUs) within the system. In some examples, connectivity is between assets within the system. In some examples, the connectivity data includes data about connectivity between software elements within the system, optionally without associated hardware data. In some examples, the connectivity data includes data about communication protocols between assets in the system.

[0034] In some examples, the network model includes any one or a combination of the following: risk analysis data for the asset / system, such as TARA data; information on the asset / system's configuration data; the asset / system's specifications; and / or the asset / system's operational status. In some examples, the network model is received from the asset's supplier (such as a vehicle manufacturer). In some examples, when prompted to output data based on the network model, the network model or a predetermined portion thereof is entered into the LLM 40 as a soft prompt.

[0035] Figure 2A A high-level flowchart is shown, based on some examples, of methods for providing threat intelligence using generative AI (i.e., leveraging one or more LLMs). The following methods are described for threat intelligence systems; however, this is not intended to limit them in any way, and any other suitable system can be used to perform these methods.

[0036] In some examples, in stage 1000, security information associated with the asset is received. In some examples, the security information includes: information about attacks on the asset; information about threats to the asset; information about vulnerabilities in the asset; and / or information about potential vulnerabilities in the asset. In some examples, the security information is input via user terminal 45. In some examples, the security information is received from an external server or from a dedicated program. In some examples, the security information is received at data interface 20. In some examples, the security information includes a textual description of an attack, vulnerability, or potential vulnerability.

[0037] In some examples, attacks on assets include detected attacks. Detected attacks can originate from a previous time period or from an attack currently being detected.

[0038] In some examples, the vulnerability is a known flaw in the corresponding asset against a specific attack. In other examples, information about the vulnerability is given in a predefined format, such as Public Vulnerability and Exposure (CVE) format.

[0039] In some examples, potential vulnerabilities are identified based on predefined vulnerability rules. In some examples, potential vulnerabilities are identified based on vulnerabilities identified in another asset similar to the current asset, according to predefined vulnerabilities. In some examples, potential vulnerabilities are identified based on data received from the asset that includes certain unexplained anomalies, according to predefined vulnerabilities.

[0040] In some examples, in stage 1010, multiple keywords are generated at least in part based on the security information received in stage 1000. In some examples, the keywords are generated by processor 12. In some examples, the keywords are generated by LLM 40. In some examples, where the security information includes a text description, the text description is input into LLM 40, and LLM 40 extracts the keywords from the text description. In some examples, a separate LLM is used to extract the keywords.

[0041] In some examples, multiple keywords are generated not based on the received security information, but at least in part based on information contained in the network model associated with the asset. In some examples, keywords are generated at least in part based on connectivity data from the network model. In some examples, keywords are generated at least in part based on risk analysis data from the network model. In some examples, security information is stored in the network model.

[0042] In some examples, as described above, a list of configuration data associated with an asset is input into LLM 40, causing generated keywords to be associated with that list of configuration data. In a non-limiting example, LLM 40 can be prompted to generate keywords from security information by requesting a list of keywords that can optionally be associated with the list of configuration data. For example, if the security information contains information about vulnerabilities associated with multiple libraries, LLM 40 can generate only keywords related to the vulnerabilities associated with libraries in the list of libraries associated with the asset. In some examples, processor 11 inputs a list of libraries into LLM 40.

[0043] In some examples, as described above, the asset's specification table is entered into LLM 40, causing the generated keywords to be associated with the specification table. In some examples, as described above, TARA data (e.g., TARA) associated with the asset is entered into LLM 40, causing the generated keywords to be associated with the TARA data. In some examples, as described above, a network model of the asset or a system containing the asset is entered into LLM 40, causing the generated keywords to be associated with the network model.

[0044] In some examples, LLM 40 generates keywords by extracting words from the textual description of the security information and providing terms associated with the security information. For example, LLM 40 can be fine-tuned to provide keywords that are context-dependent on the security information. This can include known terms or attack steps associated with terms contained within the security information. For instance, if the security information includes a description of a vulnerability in a first type of attack targeting a first function, LLM 40 can provide keywords related to other types of attacks targeting the first function. Additionally, LLM 40 can provide keywords related to the first type of attack but for different functions communicating with the first function.

[0045] Although examples of keywords generated by LLM 40 or LLM alone have been described above, this does not imply any limitation in any way, and processor 11 can initiate any suitable algorithm for generating keywords from received security information. As described above regarding LLM 40, such an algorithm can generate keywords associated with: lists of configuration data, specification tables, and / or TARA data associated with assets.

[0046] In some examples, multiple keywords are generated by the corresponding algorithms of LLM 40 or processor 11, and processor 11 further filters the generated keywords based at least in part on a list of configuration data, a specification table, TARA data, and / or a network model.

[0047] In some examples, the generated keywords include the attack steps of the attack tree. Specifically, in some examples, LLM 40 defines one or more attack trees or attack paths based on text descriptions.

[0048] In some examples, LLM 40 is fine-tuned to generate attack steps from security information. In one illustrative example, the input data used to fine-tune LLM 40 includes: a predetermined number of examples of textual descriptions of the security information; and a predetermined number of examples of textual attack steps. As used herein, the term "textual attack step" refers to a textual description of an attack step. Similarly, as used herein, the term "textual attack path" refers to a textual description of an attack path. Similarly, as used herein, the term "textual attack tree" refers to a textual description of an attack tree.

[0049] Figure 2B A high-level block diagram of module 100 for fine-tuning LLM 40 is shown. Module 100 includes: processor 110; memory 111; and additional LLM 120. In some examples, memory 111 stores a plurality of instructions that, when read by processor 110, cause processor 110 to perform the steps described below.

[0050] Figure 2C A high-level flowchart of the method for fine-tuning LLM 40 is shown. It should be noted that... Figures 2B-2C A non-limiting implementation for fine-tuning LLM 40 is described, and this is not intended to limit it in any way. In stage 1100, an example of a textual description of security information is input to LLM 40 by processor 110. In stage 1110, processor 110 fine-tunes a predetermined subset of weights of LLM 40 to generate textual attack steps from the textual description of the input security information. In some examples, stages 1100 and / or 1110 are performed at least in part based on one or more user inputs.

[0051] In stage 1120, in some examples, the output of LLM 40 is input to LLM 120, and the output of LLM 120 is fed into the loss function of LLM 40 to complete the fine-tuning process. However, it should be noted that the fine-tuning process can be performed in a variety of ways known to those skilled in the art.

[0052] In some examples, in phase 1020, the data retrieval module 30 searches one or more data sources 50 to retrieve data associated with the security information received in phase 1000. In some examples, the search is performed at least in part based on the keywords generated in phase 1010. It should be noted that the term "search" as used herein is not intended to be limited to querying any particular form of the corresponding data source 50, and any suitable method for querying and retrieving data from each data source 50 can be used. In some examples, the data retrieval module 30 performs a web crawl to retrieve data from the Internet. In some examples, the data retrieved by the data retrieval module 30 may include video data, and the video-to-text converter 70 converts the retrieved video into a text transcription. For example, the video data may include videos explaining / describing one or more vulnerabilities.

[0053] In some examples, processor 11 filters retrieved data to isolate predetermined portions of the retrieved data according to predefined rules. In some examples, predefined rules are used as part of the search parameters, causing the retrieved data to be restricted according to the predefined rules. In some examples, the isolated predetermined portions of the retrieved data are associated with configuration data related to the asset. For example, in the case of searching for various data associated with the described vulnerability, LLM 40 can filter the retrieved data to only data related to configuration data accessed by the asset.

[0054] In some examples, where a predefined rule is used as part of the search parameters, the predefined rule controls the search so that the retrieved data is associated with configuration data related to the asset.

[0055] In some examples, the predefined rules are at least partially based on the received network model. Therefore, in some examples, the retrieved data, or isolated portions thereof, is most relevant to the corresponding asset / system. For example, if the search retrieves information about a vulnerability in a library, but that vulnerability requires a specific dependency not present in the network model, the retrieved information can be discarded.

[0056] In some examples, the predetermined portion of the retrieved data is isolated from configuration data that is vulnerable to attacks, vulnerabilities, or potential vulnerabilities in the phase 1000 security information. In some examples, LLM 40 isolates data related to such libraries. In some examples, LLM 40 selects such configuration data only if it exists in the network model.

[0057] In some examples, process 12 sorts the retrieved data or isolated portions of the retrieved data according to a predetermined sorting rule. For example, in the case of retrieving data from the Internet, processor 12 may use a page sorting algorithm or a variant thereof to sort the web pages retrieved during the search. In some examples, processor 12 sorts the data based on its relevance to one or more parameters, such as relevance to attack steps, relevance to attack trees, relevance to configuration data, relevance to assets / systems, and / or relevance to any other appropriate parameters.

[0058] In some examples, in stage 1025, LLM 40 generates additional keywords, and data retrieval module 30 performs additional searches in data source 50 to obtain additional data. In some examples, the additional keywords include the configuration data identified in stage 1020. Therefore, in such examples, LLM 40 identifies configuration data (i.e., configuration data related to both security information and assets) in stage 1020 and then presents the identified configuration data as the basis for additional searches.

[0059] In some examples, additional searches are performed to identify additional vulnerabilities associated with the configuration data. Therefore, in some examples, LLM 40 filters the data retrieved from the additional searches to identify relevant vulnerabilities.

[0060] In some examples, additional searches are performed to identify which versions of the identified configuration data are most relevant. In some examples, version relevance is based on a predetermined score, and the most relevant versions are those with scores greater than a predetermined threshold. Therefore, for example, LLM 40 filters the data retrieved from the additional searches to identify one or more specific versions of the corresponding libraries most relevant to the attack. In some examples, library versions are identified based on searches performed in one or more code repositories, such as databases that exploit code.

[0061] Although phase 1025 is described prior to phase 1030 in this paper, this does not imply any limitation in any way. In some examples, additional searches of phase 1025 are performed after phase 1030, as described below.

[0062] In some examples, in phase 1030, LLM 40 summarizes at least a portion of the data retrieved in phases 1020 and / or 1025. In some examples, a portion of the retrieved data is associated with configuration data related to the asset, as described below. In some examples, as mentioned above, a portion of the retrieved data is also associated with additional configuration data that is related to an attack, vulnerability, or potential vulnerability. In some examples, the summarization of LLM 40 is done in Meta-Attack Language (MAL).

[0063] In some examples, LLM 40 is trained and / or fine-tuned to aggregate the retrieved data. In some examples, LLM 40 is prompted to aggregate the retrieved data at least in part based on additional information, including: configuration data associated with the asset; TARA data associated with the asset; specification tables associated with the asset; and / or network models associated with the asset. Therefore, in some examples, the aggregated data conforms to predefined information associated with the asset. In some examples, LLM 40 is further prompted to sort the aggregated data.

[0064] In some examples, the aggregation is based at least in part on information associated with previous attacks on the assets. In some examples, aggregation is performed to provide greater emphasis on information related to previous attacks, and / or the aggregation includes indications / aggregations of previous attacks.

[0065] In some examples, an additional search is performed, as described above regarding phase 1025. In some examples, the additional search is performed in response to a corresponding user input received at user input terminal 45. For example, after receiving summary data, the user can select a predetermined option to begin an additional search.

[0066] In some examples, in stage 1035, LLM 40 is prompted to analyze the summary data from stage 1030 to determine a confidence score for the data being provided. In some examples, a confidence score is determined for one or more of the data retrieval, data filtering, and data summarization processes. In some examples, LLM 40 is fine-tuned for this purpose. In some examples, the confidence score is determined by a second LLM (not shown). As used herein, the term "confidence score" refers to a value that indicates the likelihood that the data is accurate.

[0067] In some examples, summary data is output in stage 1040. In some examples, the summary data is output to the user terminal and / or the server.

[0068] In some examples, in phase 1050, the TARA of assets stored on memory 12 is updated by processor 11 at least in part based on the aggregated data from phase 1030. In some examples, the attack tree of the TARA is updated and / or new attack trees are added to the TARA based on threats identified in the retrieved data. For example, if a new threat is discovered, that threat can be added to the TARA along with the appropriate attack tree. In some examples, the attack tree or the risk level of a threat is updated based on the aggregated data.

[0069] In some examples, LLM 40 is fine-tuned to generate one or more text attack paths or attack trees, at least in part, based on summary data. In some examples, LLM 40 is fine-tuned to generate text attack paths / attack trees. In one illustrative example, the input data used to fine-tune LLM 40 includes a predetermined number of text attack paths / attack trees and a predetermined number of examples of summary data. In some examples, LLM 40 is first prompted to determine whether the summary data describes an attack.

[0070] In some examples, LLM 40 further prompts the assignment of a label indicating the category of each generated text attack step, selected from a predetermined list of labels optionally stored on memory 12. In some examples, each label provides a category that groups multiple attack steps together. An illustrative example of a label is "Unauthorized Access." This label would be assigned to any attack step that includes unauthorized access to a point. Another illustrative example of a label is "Vulnerability Discovery."

[0071] In some examples, processor 11 updates one or more attack trees or attack paths in the TARA data based on one or more text attack paths. As mentioned above, in some examples, the TARA data is stored in database 60. Figure 2D A high-level flowchart is shown for a method of updating TARA data based on one or more text attack paths.

[0072] In some examples, in stage 1150, the processor 11 or the corresponding user input prompts the LLM 40 to sort the attack steps of the text attack path against the attack tree or attack steps of the attack path in the TARA data to identify which attack step most closely matches the text attack step. In some examples, the LLM 40 is fine-tuned to match the text attack steps with the attack steps stored in the TARA data. In some examples, the stored attack steps are converted into text descriptions before the comparison.

[0073] While examples comparing text attack steps with TARA data have been described above, this is not intended to be limiting in any way. In some examples, other parts of the text attack path generated by LLM 40 are compared with portions of the attack path / attack tree in TARA data to identify corresponding parts.

[0074] Although LLM 40 has been described above, this does not imply any limitation in any way. In some examples, processor 11 uses a dedicated algorithm to perform a comparison between text attack steps and TARA data attack steps.

[0075] In some examples, at stage 1160, for each text attack step or text attack path segment, it is determined whether a similar attack step or attack path segment has already been found in the TARA data. If no similar attack step or attack path segment is found for one or more text attack steps or attack path segments, in some examples, at stage 1170, summary data associated with the text attack step (or text attack path) is output. In some examples, additional searches are performed on one or more data sources 50 to obtain additional information associated with the attack step. In some examples, LLM 40 also summarizes additional information.

[0076] In some examples, when identifying one or more attack steps or attack path segments in TARA data that are sufficiently similar to text attack steps, processor 11 performs one or more of the following steps.

[0077] In some examples, in stage 1180, for each identified attack step, processor 11 updates the feasibility and / or impact values ​​associated with the identified attack step. In some examples, prompting LLM 40 to generate updated feasibility / impact values ​​for the identified attack steps based on summary data. In some examples, LLM 40 is fine-tuned to generate feasibility and / or impact values ​​for attack steps based on summary data. In some examples, processor 11 runs a dedicated algorithm based on summary data to update the feasibility / impact values.

[0078] In some examples, in stage 1190, processor 11 updates one or more attack trees or attack paths in the TARA data using the identified attack steps. For example, the identified attack path portion or the identified attack step can be added to an existing attack tree / attack path. In some examples, the addition of attack steps / attack path portions can be performed based on one or more predetermined rules.

[0079] In some examples, at stage 1195, the identified attack steps or attack path portions are combined to generate one or more new attack trees / attack paths.

[0080] In some examples, the added attack tree and / or text attack path are first verified to see if they satisfy one or more of the original parameters and / or network model parameters of TARA. In some examples, processor 11 includes a logic engine configured to perform the verification.

[0081] In some examples, processor 11 updates the network model at least in part based on the retrieved data. For example, if the retrieved data includes information about vulnerabilities or attacks associated with a specific library, and that library is not present in the network model, processor 11 can update the network model based on including the missing library. In another example, if the risk analysis is updated, such as when the attack tree and / or risk level in the risk analysis are updated, processor 11 can utilize the updated risk analysis data to update the network model.

[0082] In some examples, as described above, the output of System 10 includes one or more of the following: summary data; configuration data associated with attacks, threats and / or vulnerabilities, such as a list of libraries; one or more text attack paths; updates to the TARA of assets; and updates to information in the network model (e.g., the risk level associated with an asset is updated).

[0083] In some examples, LLM 40 is prompted to summarize any updates made to the network model and outputs that summary.

[0084] Figure 3 A high-level flowchart for a method of providing threat intelligence based on a network model is shown. The method described herein relates to system 10; however, this is not intended to limit it in any way, and any suitable system can be used to perform the method. In phase 1200, security information associated with an asset is received, which includes a network model of the asset or a system including the asset.

[0085] In some examples, in stage 1210, processor 11 analyzes the network model to generate keywords for searching. In some examples, the keywords include configuration data of the network model. In some examples, the keywords are generated by LLM 40. In some examples, LLM 40 is fine-tuned to generate keywords from the network model. In some examples, the network model includes one or more text descriptions, and LLM 40 extracts keywords from the text descriptions, as described above regarding stage 1110.

[0086] In some examples, searches are performed periodically according to a predetermined schedule. In other examples, searches are performed in response to corresponding user input.

[0087] In some examples, in phase 1220, the data retrieval module 30 searches one or more data sources 50, at least in part, based on the generated keywords, to identify vulnerabilities associated with the network model of phase 1200.

[0088] In some examples, retrieved data is filtered according to predetermined rules to isolate predetermined portions of the retrieved data. In some examples, predetermined rules are used as part of the search parameters, causing the retrieved data to be restricted according to the predetermined rules. In some examples, the isolated predetermined portions of the retrieved data are associated with configuration data related to the asset. For example, in the case of searching for various data associated with a described vulnerability, LLM 40 can filter the retrieved data to only data related to configuration data accessed by the asset.

[0089] In some examples, where a predefined rule is used as part of the search parameters, the predefined rule controls the search so that the retrieved data is associated with configuration data related to the asset.

[0090] In some examples, the predefined rules are at least partially based on the received network model. Therefore, in some examples, the retrieved data, or isolated portions thereof, is most relevant to the corresponding asset / system. For example, if the search retrieves information about a vulnerability in a library, but that vulnerability requires a specific dependency not present in the network model, the retrieved information can be discarded.

[0091] In some examples, the predetermined portion of the retrieved data is isolated from libraries that are vulnerable to attacks, vulnerabilities, or potential vulnerabilities in the security information of stage 1200. In some examples, LLM 40 isolates data related to such libraries. In some examples, LLM 40 selects such libraries only if they exist in the network model.

[0092] In some examples, processor 12 sorts the retrieved data or isolated portions of the retrieved data according to a predetermined sorting rule. For example, in the case where the retrieved data is found from the Internet, processor 12 may use a page ranking algorithm or a variant thereof to sort the web pages retrieved during the search. In some examples, processor 12 sorts the data based on its relevance to one or more parameters, such as relevance to attack steps, relevance to attack trees, relevance to configuration data, relevance to assets / systems, and / or relevance to any other appropriate parameters.

[0093] In some examples, in stage 1230, LLM 40 summarizes at least a portion of the data retrieved in stage 1220. In some examples, processor 12 optionally prompts LLM 40 to summarize portions of the retrieved data associated with configuration data contained within the network model, as described above. In some examples, LLM 40 is fine-tuned to summarize the retrieved data. In some examples, as described above with respect to stage 1030, LLM 40 is prompted to summarize the retrieved data at least in part based on additional information including: configuration data associated with the asset, such as a library list; risk analysis data associated with the asset; a specification table associated with the asset; and / or a network model associated with the asset. In some examples, LLM 40 is further prompted to sort the summarized data.

[0094] In some examples, the aggregation is based at least in part on information associated with previous attacks on the asset or system. In some examples, aggregation is performed to provide greater emphasis on information related to previous attacks, and / or the aggregation includes indications / summaries of previous attacks.

[0095] Therefore, by using generative AI, especially Retrieval Enhanced Generation (RAG), information about vulnerabilities can be quickly retrieved and formulated into useful summary information, rather than requiring multiple analysts to work for a long time.

[0096] In some examples, in stage 1235, LLM 40 is prompted to analyze the summary data from stage 1230 to determine a confidence score for the data being provided, as described above. In some examples, a confidence score is determined for one or more of the data retrieval, data filtering, and data summarization processes. In some examples, LLM 40 is fine-tuned for this purpose. In some examples, the confidence score is determined by a second LLM (not shown).

[0097] In some examples, additional searches are performed, as described above regarding phase 1025.

[0098] In some examples, in phase 1240, the TARAs (such as TARAs) of assets / systems stored on memory 12 are updated by processor 11 at least in part based on the aggregated data of phase 1230, as described above with respect to phases 1150-1195.

[0099] In some examples, the network model is updated at least in part based on the retrieved data. For example, if the retrieved data includes information about vulnerabilities or attacks associated with a particular library, and that library does not exist in the network model, the processor 11 can update the network model based on including the missing library.

[0100] In some examples, in stage 1250, the output is a summary of the data from stage 1230.

[0101] Figure 4A A high-level block diagram of a system 200 for identifying communication paths is shown. In some examples, system 200 includes: a processor 11; a memory 12; a data interface 20; an LLM 210; and an inference engine 220. Although system 200 is described herein as including an LLM different from the LLM 40 described above, this is not intended to limit it in any way, and LLM 40 can perform the operations of LLM 210 without departing from the scope of this disclosure.

[0102] Figure 4B A high-level flowchart is shown for a method of using LLM to identify paths within a system, such as a transportation system. In some examples, system documentation is received in stage 1400. In some examples, the documentation includes, but is not limited to, one or more of the following: a specification sheet; a list of requirements; a TARA; source code; build scripts; and / or makefiles.

[0103] In some examples, in stage 1410, LLM 210 translates the received system documentation from stage 1400 into a formal specification language, such as Action Timing Logic (TLA+). In some examples, LLM 40 is fine-tuned and / or trained to perform this translation.

[0104] In some examples, in stage 1420, processor 11 feeds the transformed document from stage 1410 into inference engine 220. In some examples, processor 11 is configured to use one or more protocols from the transformed document to query 220 how to get from each point in the system to other points in the system.

[0105] In some examples, in stage 1430, the information from stage 1420 can be used to verify the generated attack tree or attack path. As used herein with respect to the attack tree or attack path, "verification" means that the attack tree / attack path can be executed in the relevant system. For example, if an attack tree or attack path is generated (as described above), but one or more attack steps of the attack tree cannot be executed due to a lack of capability to reach the two points in the system using the appropriate protocol, then the attack tree or attack path is not verified and can be discarded. In some examples, information about the attack path or attack tree is output, indicating that one or more attack steps cannot be executed in the system.

[0106] In some examples, information from stage 1420 can be used in stage 1440 to simulate traffic within the system. In some examples, information about the simulated traffic is output.

[0107] Some examples of the disclosed technology The following are some examples of the above-described embodiments. It should be noted that a feature of an isolated example, or a combination of one or more features of that example, or optionally a combination of one or more features of that example with one or more features of the examples below, also falls within the scope of this application.

[0108] Example A1. A method for providing threat intelligence, the method comprising: generating a plurality of keywords based at least in part on a network model associated with an asset; searching one or more data sources at least in part on the generated keywords to retrieve data associated with the network model, the retrieved data including information about threats and / or vulnerabilities associated with the asset; and updating the network model at least in part on the retrieved data.

[0109] Example A2. The method according to Example A1 further includes using a first large language model (LLM) to aggregate at least a portion of the retrieved data, wherein updates to the network model are based at least in part on the aggregated data.

[0110] Example A3. The method according to Example A1 or A2, wherein the network model includes risk analysis of assets, and wherein updating the network model includes updating the risk analysis.

[0111] Example A4. The method according to any one of Examples A1-A3, wherein updating the network model includes updating one or more attack paths.

[0112] Example A5. The method according to Example A4 further includes: identifying one or more attack steps of one or more attack paths based at least in part on the retrieved data, wherein the update of one or more attack paths is based at least in part on the identified one or more attack steps.

[0113] Example A6. The method according to any one of Examples A1-A5, wherein updating the network model includes adding one or more attack paths.

[0114] Example A7. The method according to any one of Examples A3-A6 further includes: utilizing a second LLM to generate one or more text attack paths based at least in part on the retrieved data, wherein the network model is updated based at least in part on the generated one or more text attack paths.

[0115] Example A8. The method according to Example A7 further includes using an LLM to aggregate at least a portion of the retrieved data, wherein updates to the network model are based at least in part on the aggregated data, and one or more text attack paths are based at least in part on the aggregated data.

[0116] Example A9. The method according to Example A7 or A8 further includes: receiving a system document; converting the received system document into a canonical language using an LLM; feeding the converted document into an inference engine; querying the inference engine for paths within the system; and verifying one or more output text attack paths based at least in part on the query.

[0117] Example A10. According to any one of Examples A7-A9 of Example 2, wherein the second LLM is the first LLM.

[0118] Example A11. The method according to any one of Examples A1-A10 further includes receiving security information, which includes information about attacks on assets, information about threats to assets, information about vulnerabilities in assets, or information about potential vulnerabilities in assets, wherein the generation of keywords is based at least in part on the received security information.

[0119] Example A12. The method described in Example A11, wherein a portion of the retrieved data is associated with configuration data related to the asset.

[0120] Example A13. The method described in Example A11, wherein the retrieved data is associated with any configuration data that is vulnerable to attack, vulnerabilities, or potential vulnerabilities.

[0121] Example A14. The method described in Example A13 further includes outputting information about configuration data that is vulnerable to attacks, vulnerabilities, or potential vulnerabilities.

[0122] Example A15. The method according to any one of Examples A13-A14 further includes performing additional searches in one or more data sources to: identify one or more most relevant versions of configuration data that are vulnerable to attack, vulnerabilities, or potential vulnerabilities.

[0123] Example A16. The method according to any one of Examples A11-A15 further includes performing additional searches in one or more data sources to identify data regarding additional vulnerabilities associated with configuration data, which is associated with a portion of the retrieved data.

[0124] Example A17. The method according to Example A16 further includes receiving user input regarding retrieved data associated with the received security information, wherein an additional search responds to the received user input.

[0125] Example A18. The method according to Example A1 further includes receiving security information, including information about attacks on assets, information about threats to assets, information about vulnerabilities in assets, or information about potential vulnerabilities in assets, wherein the generation of keywords is based at least in part on the received security information, wherein the received security information includes a text description, and the method further includes inputting the text description into an LLM, wherein multiple keywords are extracted from the text description by the LLM.

[0126] Example A19. The method according to any one of Examples A1-A18, wherein one or more data sources include one or more of the following: the Internet; the dark web; the National Vulnerability Database (NVD); the Open Source Vulnerability (OSV) database; a database of exploit code; a code repository; a specification sheet associated with the asset; and a risk analysis associated with the asset.

[0127] Example A20. The method according to Example A1 further includes sorting the retrieved data using an LLM, wherein updates to the network model are based at least in part on the sorting of the retrieved data.

[0128] Example A21. The method described in Example A20, wherein the sorting includes page sorting.

[0129] Example A22. The method described in Example A2, wherein the aggregation is based at least in part on information associated with previous attacks on the assets.

[0130] Example A23. The method described in Example A2 further includes determining and outputting confidence scores for the summarized data.

[0131] Example A24. The method described in Example A23, wherein the confidence score is determined by a first LLM.

[0132] Example A25. The method according to any one of Examples A1-A24, wherein the retrieved data includes video and / or audio data, and the method further includes converting the video and / or audio data into text.

[0133] Example A26. The method described in Example A1 further includes using an LLM to aggregate updates to the network model.

[0134] Example A27. The method according to any one of Examples A1-A26, wherein the network model includes data on connectivity within the system including assets.

[0135] Example A28. The method described in Example A27, wherein the network model further includes data about connections to and from the system.

[0136] Example A29. The method described in Example A27 or A28, wherein the network model further includes data about the software implementation of the assets.

[0137] Example A30. The method according to any one of Examples A27-A29, wherein the network model further includes data on threats associated with the asset.

[0138] Example A31. A threat intelligence system for providing threat intelligence, the threat intelligence system comprising one or more processors and a memory, wherein the memory has a plurality of instructions stored in the memory, the instructions, when executed by one or more processors, causing one or more processors to perform a method according to any one of Examples A1 to A30.

[0139] Example B1. A method for providing threat intelligence, the method comprising: receiving security information associated with an asset; generating a plurality of keywords based at least in part on the received security information; searching one or more data sources to retrieve data associated with the received security information based at least in part on the generated keywords, the retrieved data including information about threats and / or vulnerabilities associated with the asset; and using a large language model (LLM) to aggregate at least a portion of the retrieved data.

[0140] Example B2. The method described in any of the examples in this document, particularly Example B1, wherein the security information includes the network model.

[0141] Example B3. The method described in any of the examples in this document, particularly Example B1, wherein the security information includes information about attacks on the asset, information about threats to the asset, information about vulnerabilities in the asset, or information about potential vulnerabilities in the asset.

[0142] Example B4. As described in any of the examples in this document, particularly Example B3, a portion of the retrieved data is associated with configuration data related to the asset.

[0143] Example B5. As described in any of the examples in this document, particularly Example B3, the retrieved data is associated with any configuration data that is vulnerable to attack, vulnerabilities, or potential vulnerabilities.

[0144] Example B6. The method described in any of the examples in this document, particularly Example B5, also includes outputting information about configuration data that is vulnerable to attack, vulnerabilities, or potential vulnerabilities.

[0145] Example B7. The method as described in any of the examples in this document, particularly any of Examples B5-B6, further includes performing additional searches in one or more data sources to identify one or more of the most relevant versions of configuration data that are vulnerable to attack, vulnerabilities, or potential vulnerabilities.

[0146] Example B8. The method described in any of the examples herein, particularly any of Examples B3-B7, further includes performing additional searches in one or more data sources to: identify data regarding additional vulnerabilities associated with configuration data related to a portion of the retrieved data; and summarize the retrieved data regarding the additional vulnerabilities.

[0147] Example B9. The method described in any of the examples herein, particularly Example B8, further includes receiving user input regarding summary data associated with the received security information, wherein an additional search is performed in response to the received user input.

[0148] Example B10. The method, as in any of the examples in this document, particularly any one of Examples B1-B9, also includes outputting summarized data.

[0149] Example B11. The method as described in any of the examples herein, particularly any one of Examples B1-B10, also includes updating the risk analysis of the assets based at least in part on the aggregated data.

[0150] Example B12. The method described in any of the examples in this document, particularly Example B11, wherein updating the risk analysis includes updating one or more attack trees and / or adding one or more attack trees.

[0151] Example B13. The method described in any of the examples in this document, particularly Examples B11 or B12, wherein the LLM is fine-tuned to generate one or more text attack paths, based at least in part on the aggregated data, and the risk analysis is updated based at least in part on the generated one or more text attack paths.

[0152] Example B14. The method described in any of the examples in this document, particularly any of Examples B1-B12, wherein the LLM is fine-tuned, at least in part, based on the summarized data, to generate and output one or more text attack paths.

[0153] Example B15. The method described in any of the examples herein, particularly Example B14, further includes: receiving a system document; converting the received system document into a canonical language using an LLM; feeding the converted document into an inference engine; querying the inference engine for paths within the system; and verifying one or more text attack paths output based at least in part on the query.

[0154] Example B16. The method as described in any of the examples herein, particularly examples B1-B15, wherein the received security information includes a text description, and the method further includes inputting the text description into an LLM, wherein multiple keywords are extracted from the text description by the LLM.

[0155] Example B17. The method as described in any of the examples herein, particularly any one of Examples B1-B16, wherein one or more data sources include one or more of the following: the Internet; the dark web; the National Vulnerability Database (NVD); the Open Source Vulnerability (OSV) database; a database of exploit code; a code repository; a specification sheet associated with the asset; and a risk analysis associated with the asset.

[0156] Example B18. The method as described in any of the examples herein, particularly examples B1-B17, wherein the LLM sorts the retrieved data, and a portion of the retrieved data is selected based at least in part on its sorting.

[0157] Example B19. The method described in any of the examples in this document, particularly Example B18, wherein the sorting includes page sorting.

[0158] Example B20. The method as described in any of the examples herein, particularly examples B1-B19, wherein the aggregation is based at least in part on information associated with previous attacks on the assets.

[0159] Example B21. The method as described in any of the examples herein, particularly any of Examples B1-B20, further includes updating the network model associated with the asset based at least in part on the retrieved data.

[0160] Example B22. The method as described in any of the examples herein, particularly any of Examples B1-B21, further includes determining and outputting confidence scores for the summarized data.

[0161] Example B23. The method described in any of the examples in this document, particularly Example B22, wherein the confidence score is determined by LLM.

[0162] Example B24. The method as described in any of the examples herein, particularly examples B1-B23, wherein the retrieved data includes video data, and the method further includes converting the video data into text.

[0163] Example B25. A threat intelligence system comprising one or more processors and a memory, wherein the memory stores a plurality of instructions which, when executed by one or more processors, cause one or more processors to perform a method comprising: receiving security information associated with an asset; generating a plurality of keywords based at least in part on the received security information; searching one or more data sources at least in part on the generated keywords to retrieve data associated with the received security information, the retrieved data including information about threats and / or vulnerabilities associated with the asset; and summarizing at least a portion of the retrieved data using a large language model (LLM).

[0164] It should be understood that certain features of the invention described in the context of independent embodiments for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for brevity may also be provided individually or in any suitable sub-combination.

[0165] Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although similar or equivalent methods to those described herein may be used in the practice or testing of this invention, suitable methods are described herein.

[0166] All publications, patent applications, patents, and other references mentioned herein are incorporated herein by reference in their entirety. In case of any conflict, the specification of this patent application (including definitions) shall prevail. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0167] Those skilled in the art will recognize that the present invention is not limited to what has been specifically shown and described above. Rather, the scope of the invention is defined by the claims and includes combinations and sub-combinations of the various features described above, as well as variations and modifications thereto that would occur to those skilled in the art upon reading the foregoing description.

Claims

1. A method for providing threat intelligence, the method comprising: Multiple keywords are generated, at least in part, based on a network model associated with the assets; Based at least in part on the generated keywords, one or more data sources are searched to retrieve data associated with the network model, including information about threats and / or vulnerabilities associated with the asset; and The network model is updated based at least in part on the retrieved data.

2. The method of claim 1, further comprising using a first large language model (LLM) to aggregate at least a portion of the retrieved data, wherein updates to the network model are based at least in part on the aggregated data.

3. The method according to claim 1 or 2, wherein, The network model includes risk analysis of the assets, and The update of the network model includes the update of the risk analysis.

4. The method according to any one of claims 1-3, wherein, The network model update includes updating one or more attack paths.

5. The method of claim 4, further comprising identifying one or more attack steps of the one or more attack paths based at least in part on the retrieved data. in, The update of the one or more attack paths is based at least in part on the one or more attack steps identified.

6. The method according to any one of claims 1-5, wherein, The network model update includes adding one or more attack paths.

7. The method according to any one of claims 3-6, further comprising: The second LLM is used to generate one or more text attack paths, at least in part, based on the retrieved data. The network model is updated based at least in part on one or more generated text attack paths.

8. The method of claim 7, further comprising using the LLM to aggregate at least a portion of the retrieved data. in, The network model is updated at least in part based on the aggregated data, and The one or more text attack paths are at least partially based on the aggregated data.

9. The method according to claim 7 or 8, further comprising: Receive system documents; The received system documents are converted into a standard language using the LLM; The transformed documents are fed into the inference engine; Query the inference engine for paths within the system; and Based at least in part on the query, verify one or more text attack paths in the output.

10. The method according to any one of claims 7-9 when dependent on claim 2, wherein, The second LLM is the first LLM.

11. The method according to any one of claims 1-10, further comprising receiving security information, said security information including information about attacks on the asset, information about threats to the asset, information about vulnerabilities in the asset, or information about potential vulnerabilities in the asset. in, The generation of the keywords is based at least in part on the received security information.

12. The method according to claim 11, wherein, The portion of the retrieved data is associated with configuration data related to the asset.

13. The method according to claim 11, wherein, The retrieved data is associated with any configuration data that is vulnerable to the aforementioned attacks, vulnerabilities, or potential vulnerabilities.

14. The method of claim 13, further comprising outputting information about the configuration data that is susceptible to the attack, vulnerability, or potential vulnerability.

15. The method of any one of claims 13-14, further comprising performing additional searches in the one or more data sources to: Identify one or more of the most relevant versions of the configuration data that are vulnerable to attacks, vulnerabilities, or potential vulnerabilities.

16. The method of any one of claims 11-15, further comprising performing additional searches in the one or more data sources to identify data regarding additional vulnerabilities associated with the configuration data, the configuration data being associated with said portion of the retrieved data.

17. The method of claim 16, further comprising receiving user input regarding the retrieved data, the retrieved data being associated with the received security information. in, The additional search response is based on the received user input.

18. The method of claim 1, further comprising receiving security information, the security information including information about attacks on the asset, information about threats to the asset, information about vulnerabilities in the asset, or information about potential vulnerabilities in the asset. in, The generation of the keywords is based at least in part on the received security information. The received security information includes a text description, and the method further includes inputting the text description into an LLM. Specifically, the LLM extracts the multiple keywords from the text description.

19. The method according to any one of claims 1-18, wherein, The one or more data sources include one or more of the following: the Internet; the dark web; the National Vulnerability Database (NVD); the Open Source Vulnerability (OSV) database; a database of exploit code; a code repository; a specification sheet associated with the asset; and a risk analysis associated with the asset.

20. The method of claim 1, further comprising sorting the retrieved data using an LLM, wherein updates to the network model are based at least in part on the sorting of the retrieved data.

21. The method according to claim 20, wherein, The ranking includes page sorting.

22. The method according to claim 2, wherein, The summary is based, at least in part, on information associated with previous attacks on the assets.

23. The method of claim 2, further comprising determining and outputting a confidence score for the summarized data.

24. The method according to claim 23, wherein, The confidence score is determined by the first LLM.

25. The method according to any one of claims 1-24, wherein, The retrieved data includes video and / or audio data, and the method further includes converting the video and / or audio data into text.

26. The method of claim 1, further comprising using LLM to aggregate updates to the network model.

27. The method according to any one of claims 1-26, wherein, The network model includes data about connectivity within the system that includes the asset.

28. The method according to claim 27, wherein, The network model also includes data about connections to and from the system.

29. The method according to claim 27 or 28, wherein, The network model also includes data about the software implementation of the assets.

30. The method according to any one of claims 27-29, wherein, The network model also includes data on threats associated with the asset.

31. A threat intelligence system for providing threat intelligence, the threat intelligence system comprising one or more processors and a memory, wherein the memory has a plurality of instructions stored therein, the instructions, when executed by the one or more processors, causing the one or more processors to perform the method according to any one of claims 1 to 30.