Live tracking of cyber components and ai guided automated vulnerability detection & fixing

WO2025097073A8PCT designated stage expired Publication Date: 2025-07-03UNIVERSITY OF MAINE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/054291
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-11-01
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Modern complex automated network systems face challenges in identifying and mitigating vulnerabilities due to the increasing diversity of software, hardware, and devices, as well as the complexity of their connections and the sheer number of devices involved.

Method used

A live mapping framework is generated for large enterprise-wide networks, tracking key attributes of each device, such as type, software packages, network connections, and usage information. This framework uses generative and explainable AI for penetration testing on virtual replicas of the networks, identifying vulnerabilities and generating solution sets for mitigation.

Benefits of technology

The solution enables efficient identification and mitigation of network vulnerabilities, reducing the need for extensive human resources and minimizing the risk of cyber threats, while also allowing for routine penetration testing without the need for third-party examination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024054291_03072025_PF_FP_ABST
    Figure US2024054291_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A live mapping framework may be generated for large enterprise- wide networks and systems. The live mapping framework includes identification of key attributes of every device connected to the network. The key attributes may include device type, installed software packages, network connection information, device usage information, and other information. Frequent updates to the live mapping framework and trending over time enable the identification of system vulnerabilities. Generative and explainable Al may additional be used for penetration testing of virtual replicas of the enterprise- wide networks, leading to the generation of solution sets addressing the identified vulnerabilities.
Need to check novelty before this filing date? Find Prior Art

Description

LIVE TRACKING OF CYBER COMPONENTS AND Al GUIDED AUTOMATED VULNERABILITY DETECTION & FIXINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 547,265 filed November 3, 2023 and U.S. Provisional Application No. 63 / 596,127 filed November 3, 2023, the disclosures of each of which are hereby incorporated by reference in their entireties.FIELD

[0002] The present disclosure relates generally to tracking of components in a cyber environment, identifying network vulnerabilities, and taking mitigating action.BACKGROUND

[0003] Modern complex automated network systems and facilities include many digital and electronic components which are interconnected with each other in a multitude of ways. These electronic components can be of different types and may include: sensors including cameras, compute units such as servers, terminals and laptops, Intemet-of-Things (loT) devices, Raspberry Pi’s, digitally controlled mechanical actuators, 3D printers, and / or networking devices. These devices typically run on a number of different electronic hardware platforms and use a wide range of software components. Therefore, identifying and implementing a single cybersecurity strategy presents challenges.SUMMARY

[0004] As the types of software, hardware, means of connecting, and sheer number of devices all increase over time, solutions and strategies for protecting such devices and the networks they are connected to are becoming more and more complex. For example, vulnerabilities often originate from outdated hardware and software elements. Patching, isolating, and upgrading such components is possible, but only if the administrators are aware of the situation. Hence it is important to keep track of each and every component that getsconnected to the automated manufacturing ecosystem. This tracking process is, however, not an easy task and may require extensive human resources to fulfill the requirements. This effort becomes almost intractable as the overall system grows larger, which is likely to be the case for future Industry 4.0 systems.

[0005] In addition, for an Industry 4.0 system, it is important to perform routine penetration testing towards ensuring its safety against real threats. However, performing regular penetration testing is a time consuming and human-resource intensive activity. Additionally, allowing third party entities to examine a secure infrastructure can be undesirable for many organizations.

[0006] The present embodiments provide for, inter alia, tracking of components in a cyber environment, identifying network vulnerabilities, and taking mitigating action. The present disclosure describes a live mapping framework that may be generated for large enterprise-wide networks and systems. The live mapping framework includes identification of key attributes of every device connected to the network. The key attributes may include device type, installed software packages, network connection information, device usage information, and other information. Frequent updates to the live mapping framework and trending over time enable the identification of system vulnerabilities. Generative and explainable Al may additionally be used for penetration testing of virtual replicas of the enterprise-wide networks, leading to the generation of solution sets addressing the identified vulnerabilities.

[0007] In one aspect, the disclosure encompasses methods of penetration testing on a digital replica of a system (e.g., an Industry 4.0 system, an loT system), the methods comprising: accessing a first system state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; updating the map of the digital replica using the second system state; and identifying one or more vulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

[0008] In some embodiments, the methods comprise identifying a solution set for mitigating the one or more identified vulnerabilities of the digital replica.[00091 In some embodiments, the methods comprise applying the solution set to the system.

[0010] In some embodiments, the methods comprise mitigating the one or more vulnerabilities, wherein mitigating the one or more vulnerabilities comprises isolating (e.g., temporarily isolating) one or more devices of the system and / or applying an update (e.g., a firmware update, a software update) to a device of the system.

[0011] In some embodiments, the methods comprise determining an action pathway fitness (e.g., an action pathway fitness score).

[0012] In some embodiments, the action pathway fitness is determined based on one or more of: discovering the one or more vulnerabilities; severity of the one or more vulnerabilities; discovering one or more new devices on a network of the digital replica; exploiting, by the penetration test model, the one or more vulnerabilities; disclosing of credentials (e.g., user credentials, administrator credentials); and inserting an exploit (e.g., a back door).

[0013] In some embodiments, the methods comprise obtaining feedback of a user based on one or more of: the first system state, the second system state, and the determined action pathway.

[0014] In some embodiments, the methods comprise updating the action pathway based on feedback of a user (e.g., a human expert) (e.g., by merging the feedback of the user and the action pathway).

[0015] In some embodiments, the methods comprise creating (or updating) a feedback dictionary based on the first system state, the second system state, the action pathway, and the action pathway fitness (e.g., for later use in training / re-training the penetration test model).

[0016] In some embodiments, the methods comprise retraining the penetration test model based on the feedback dictionary.

[0017] In some embodiments, the methods comprise determining a defense pathway based on applying a defender model to the first system state.

[0018] In some embodiments, the methods comprise receiving one or more maps of the system and creating the digital replica based on the one or more maps of the system.

[0019] In some embodiments, the digital replica comprises a plurality of networked components mimicking devices of the system, wherein the plurality of network componentscomprise one or more of hardware emulated components, software emulated components, and software simulated components.

[0020] In some embodiments, the first system state is a vectorized version of the digital replica.

[0021] In some embodiments, the methods comprise vectorizing the digital replica using one-hot encoding, linear encoding, and / or neural network encoding.

[0022] In some embodiments, determining the action pathway comprises: identifying one or more action options based on the penetration test model and the first system state of the digital replica; creating one or more action pathway options based on the action options; scoring each of the one or more action pathway options; and selecting, based on the score, one of the one or more action pathway options to become the action pathway.

[0023] In another aspect, the disclosure encompasses methods of penetration testing of on a digital replica of a system (e.g., an Industry 4.0 system, an loT system), the methods comprising: receiving a first system state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; determining a defense pathway based on applying a defense model to the second system state the digital replica; executing the defense pathway on the second system state to produce a third system state; updating the map of the digital replica using the third system state; and identifying one or more vulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

[0024] In some embodiments, the methods comprise identifying a solution set (e.g., based on the updated system map) based on, at least, the one or more identified vulnerabilities of the digital replica.

[0025] In some embodiments, the methods comprise applying the solution set to the digital replica.

[0026] In some embodiments, the methods comprise mitigating one or more vulnerabilities of the system based on, at least, the identified one or more vulnerabilities of the digital replica.[00271 In some embodiments, the methods comprise determining an action pathway fitness (e.g., an action pathway fitness score).

[0028] In some embodiments, the methods comprise obtaining feedback of a user based on one or more of: the first system state, the second system state, the third system state, the action pathway, and the defense pathway.

[0029] In some embodiments, the methods comprise updating the action pathway fitness and feedback of a user (e.g., by merging).

[0030] In some embodiments, the methods comprise creating (or updating) a feedback dictionary based on the first system state, the second system state, the third system state, the action pathway, and the action pathway fitness (e.g., for later use in training I retraining the penetration test model).

[0031] In some embodiments, the methods comprise retraining the penetration test model and / or the defender model based on the feedback dictionary.

[0032] In some embodiments, the defense pathway comprises one or more of: scanning a network of the digital replica for an anomaly; modifying a topology (e.g., connections of) of the map of the digital replica; and disconnecting one or more hosts (e.g., devices) from the network of the digital replica.

[0033] In another aspect, the disclosure encompasses tracking systems on a networked system (e.g., an loT system, an Industry 4.0 system), the tracking system comprising: a plurality of tracking agents, wherein each of the plurality of tracking agents are communicatively connected to at least one or more components on the networked system and configured to transmit (e.g., periodically transmit) data corresponding to the at least one or more components; and a central database, wherein the central database is communicatively connected to the plurality of tracking agents and configured to receive (e.g., via an encrypted means) (e.g., directly, indirectly) (e.g., periodically receive) data from the plurality of tracking agents.

[0034] In some embodiments, at least one of the plurality of tracking agents is a software -based tracking agent installed on a component of the networked system.

[0035] In some embodiments, at least one of the plurality of tracking agents is a physical device communicatively connected to the networked system.[00361 In some embodiments, the tracking agent is configured to extract data from a processor of a component.

[0037] In some embodiments, the processor of the component comprises a hardware performance counter.

[0038] In some embodiments, at least one of the plurality of tracking agents is configured to extract data from one or more components neighboring (e.g., connected to) the tracking agent.

[0039] In some embodiments, the data from the one or more components neighboring the tracking agent comprises one or more open ports of the neighboring component(s), network traffic, operating system(s) of the neighboring component(s), and / or one or more services running on the networked system.

[0040] In some embodiments, the central database is or comprises an aggregation agent for generating a map (e.g., in real time, a spatio-temporal map) of the networked system.

[0041] In some embodiments, the central database is configured to create and / or update a digital twin of the networked system based on the generated map.

[0042] In some embodiments, the aggregation agent is configured to perform an assessment based on a digital twin created from the data from the tracking agents to identify one or more vulnerabilities, anomalies, and / or policy violations.

[0043] In another aspect, the disclosure encompasses methods of using tracking systems described herein, the methods comprising: pinging, by the tracking agent, at least one of a networked device, server, and computer to establish a communication connection therewith, executing, by the tracking agent, at least one information retrieval routine to retrieve at least one piece of information from the networked device, server, or computer; and storing, in the central database, the at least one piece of information.

[0044] In some embodiments, the at least one piece of information comprises at least one of an IP address, a device name, device current network interface configuration information, operating system information, USB device identity, USB bus information, USB usage history, one or more kernel / buffer messages, names and / or types of software running on a system, a list of upgradable software running on a system, usage history for a program or application, internetwork activity information, and hardware information.[00451 In some embodiments, the methods further comprise taking at least one mitigating action as a result of the at least one piece of information retrieved from the tracking device, the at least one mitigating action comprising at least one of (1) reassigning an IP address of the networked device, server, or computer to a quarantined IP address, and (2) renaming (i.e., changing the host name of) the networked device, server, or computer.

[0046] In another aspect, the disclosure encompasses methods of identifying attacks (e.g., a cyber attack) on a system, the methods comprising: creating a virtual replica (digital twin) of the system (e.g., an loT system) and generating a normal state (e.g., not attacked) profile of the virtual replica; simulating an attack on the virtual replica using an artificial intelligence framework (e.g., a generative artificial intelligence framework); and creating a model for detecting a deviation of (one or more aspects of) a test profile of the virtual replica from the normal state profile using the simulated attack.

[0047] In some embodiments, the methods further comprises: receiving one or more test profiles of the virtual replica; and detecting an attack on the system by applying the deviation to the one or more test profiles of the virtual replica.

[0048] In another aspect, the disclosure encompasses methods for generating solutions to system vulnerabilities, the methods comprising: generating a profile (e.g., a state) of the system at a time point; (e.g., automatically) creating a virtual replica of the system at the time point; carrying out a penetration test on the virtual replica (e.g., using an Al framework); identifying one or more vulnerabilities of the virtual replica based on the penetration test; and generating a first set of solutions corresponding to (e.g., based on) the one or more vulnerabilities of the virtual replica.

[0049] In some embodiments, the methods further comprise implementing at least one mitigating action on the system based on the first set of solutions (e.g., to enhance the robustness of the system) (e.g., to address one or more vulnerabilities of the system).

[0050] In some embodiments, the methods comprise applying the first set of solutions to the virtual replica and updating the virtual replica.

[0051] In some embodiments, the methods comprise: carrying out a second vulnerability test on the virtual replica (e.g., using a generative Al framework); identifying one or more vulnerabilities based on the second vulnerability test; generating a second set of solutionscorresponding to (e.g., based on) the second vulnerability test; and updating the virtual replica based on the one or more vulnerabilities based on the second vulnerability test.

[0052] In some embodiments, the methods comprise: applying a penetration test (e.g., an Al guided penetration test) to the virtual replica; and determining the virtual replica meets a threshold for robustness (e.g., by preventing and / or resisting attacks).

[0053] In some embodiments, the methods comprise carrying out the vulnerability test on the virtual replica (e.g., the first attack, the second attack) using a generative Al method (e.g., a generative Al (GAI) framework).

[0054] In some embodiments, the generative Al method comprises: applying a penetration test model to a system state corresponding to the profile of the system; identifying one or more action options (e.g., one or more attacks) based on the system state and the penetration test model; creating one or more action pathways based on the one or more action options; scoring the one or more action pathways (e.g., based on a numerical scale); selecting at least one of the one or more action pathways for the vulnerability test based on, at least, the score(s) of the one or more action pathways.

[0055] In some embodiments, the methods comprise training (or re-training) the penetration test model comprising training the penetration test model based on a feedback dictionary.

[0056] In some embodiments, the feedback dictionary comprises one or more of: the selected action pathway, the profile of the system at the time point, and a fitness of the penetration test model.

[0057] In another aspect, the disclosure encompasses methods of creating a digital twin of a system (e.g., an loT system, an Industry 4.0 system), the method comprising: receiving a representation of the system (e.g., based on tracking data from a tracking system (e.g., from a plurality of tracking agents)); identifying one or more portions of the representation for hardware emulation; (e g., automatically) mapping one or more hardware elements (e.g., an loT device) to correspond to each of the one or more portions of the representation for hardware emulation; identifying one or more portions of the representation for software simulation and / or emulation; (e.g., automatically) mapping one or more virtual systems to correspond to each of the one or more portions of the representation for software simulation and / or emulation;networking the one or more virtual systems and the one or more hardware elements based on the representation of the system; and initiating network traffic on the network.

[0058] In some embodiments, the methods comprise receiving tracking data from a tracking system (e.g., as described herein) and creating the representation of the system based on the tracking data.

[0059] In some embodiments, the methods comprise updating the representation of the system using tracking data from a tracking system.

[0060] In some embodiments, the methods comprise receiving and / or updating the representation of the system using a controller.

[0061] In some embodiments, the methods comprise visualizing one or more aspects of the digital twin (e.g., network traffic, a map of the digital twin) using a controller.

[0062] In some embodiments, the methods comprise carrying out a penetration test on the digital twin (e.g., as described in embodiments 1-26).

[0063] In another aspect, the disclosure encompasses Al-guided cyber vulnerability mitigation methods comprising: creating a live map network framework; creating a snapshot of the network framework at a first time point; creating a virtual replica of the network framework; using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on the virtual replica; identifying vulnerabilities of the virtual replica; and generating one or more solution sets to address the identified vulnerabilities.

[0064] In some embodiments, the methods further comprise applying the one or more solution sets to the virtual replica.

[0065] In some embodiments, the methods further comprise determining if the virtual replica passes one or more robustness tests.

[0066] In some embodiments, the methods further comprise: if the virtual replica does not pass the one or more robustness tests, using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on an updated version of the virtual replica; and if the virtual replica does pass the one or more robustness tests, addressing real system vulnerabilities based on knowledge and / or insights obtained from identifying vulnerabilities of the virtual replica.[00671 In another aspect, the disclosure encompasses methods of generating a live device tracking map, the method comprising: installing a tracking system on one or more loT devices, the tracking system configured to capture one or more attributes of the one or more loT devices and transmit the one or more attributes to a central tracking system; communicatively coupling the one or more loT devices to a network housing the central tracking system, and periodically transmitting, by the tracking system, the one or more attributes to the central tracking system based on a predetermined user-defined update frequency.

[0068] In another aspect, the disclosure encompasses multi-network generative adversarial networks (GANs) comprising: a first neural network trained for threat identification; and a second neural network trained for vulnerability exploitation, wherein each of the first neural network and the second neural network are trained via unsupervised procedural generation; and wherein the unsupervised procedural generation used for training of each of the first neural network and the second neural network comprises at least one mitigation activity to distinguish between normal and anomalous behavior.

[0069] In some embodiments, the first neural network and the second neural network are mutually trainable.BRIEF DESCRIPTION OF THE DRAWING

[0070] The foregoing and other objects, aspects, features, and advantages of the present disclosure will become more apparent and better understood by referring to the following descriptions taken in conjunction with the accompanying drawings, in which:

[0071] FIG. 1 is a schematic diagram of an enterprise network, according to aspects of the present embodiments.

[0072] FIG. 2 is a schematic diagram of a live mapping framework, according to aspects of the present embodiments.

[0073] FIG. 3 illustrates a method for tracking unit information extraction, according to aspects of the present embodiments.[00741 FIG. 4 illustrates a method for loT system map generation, according to aspects of the present embodiments.

[0075] FIG. 5 illustrates a method for anomaly detection, according to aspects of the present embodiments.

[0076] FIG. 6 is a schematic diagram of a device and network tracking hierarchy, according to aspects of the present embodiments.

[0077] FIG. 7 is a schematic diagram of a device information tracking over time, according to aspects of the present embodiments.

[0078] FIG. 8 is a schematic diagram of an automated Al-guided penetration testing framework, according to aspects of the present embodiments.

[0079] FIG. 9 illustrates an Al-guided cyber vulnerability mitigation method, according to aspects of the present embodiments.

[0080] FIG. 10 illustrates an Al model re-training method, according to aspects of the present embodiments.

[0081] FIG. 11 illustrates an Al-guided penetration testing method, according to aspects of the present embodiments.

[0082] FIG. 12 illustrates an action pathway generation method, according to aspects of the present embodiments.

[0083] FIG. 13 illustrates a method of creating a digital twin I virtual replica of a system, according to aspects of the present embodiments.

[0084] FIG. 14 is an illustrative embodiment of an exemplary hardware / software cmulation / simulation setup, according to aspects of the present embodiments.

[0085] FIG. 15 illustrates a component of a method of an Al guided penetration testing, according to aspects of the present embodiments.

[0086] FIG. 16 illustrates an Al-guided penetration testing method, according to aspects of the present embodiments.[00871 FIG. 17 is an illustrative example of an interaction between an attack model (AM) and a defense model (DM), according to aspects of the present embodiments.

[0088] FIG. 18 shows an exemplary embodiment of a method of extracting information from a real system according to aspects of the present embodiments.

[0089] FIG. 19 shows exemplary systems and methods used for detecting attacks on a system, according to aspects of the present embodiments.

[0090] FIG. 20 is a visual representation of complexities of an automated manufacturing system / facility, according to aspects of the present embodiments.

[0091] FIG. 21 shows an illustrative architecture of a CAST-Map framework according to aspects of the present embodiments.

[0092] FIG. 22 provides a K-Means visualization when the number of clusters is set to five (top panel) and three (bottom panel). The data is for 2000 randomly generated sub-systems and the computed cluster weights are specified inside the box.

[0093] FIG. 23 is a block diagram of an exemplary cloud computing environment, used in certain embodiments.

[0094] FIG. 24 is a block diagram of an example computing device and an example mobile computing device used in certain embodiments.

[0095] The features and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings.DEFINITIONS

[0096] About or Approximately: The term “about” or “approximately”, when used herein in reference to a value, refers to a value that is similar, in context to a stated reference value. In general, those skilled in the art, familiar with the context, will appreciate the relevant degree of variance encompassed by “about” or “approximately” in that context. For example, in some embodiments, the term “about” or “approximately” may encompass a range of values thatare within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referred value.

[0097] Action Pathway: As used herein, the term “action pathway” refers to a set of steps taken to move a state of a system (e.g., a sub-system) from one state to another. In certain embodiments, a system state can be described by, but not limited to, states of hardware (e.g., devices), software, and networking makeup (e.g., connections within a network, impact of an event on a network (e.g., malicious activity)). In certain embodiments, an action pathway and steps of an action pathway can be represented using the following representation or variations thereof: A =[S1, S2....].

[0098] Artificial Intelligence (Al): As used herein, the term “artificial intelligence (Al)” refers to the intelligence of machines or software, as opposed to the intelligence of humans. Artificial intelligence leverages computers and machines to mimic the problem-solving and decision-making capabilities of the human mind.

[0099] Digital Twin As used herein, the term “digital twin” and “virtual replica” are used interchangeably to refer to a simulated and / or emulated system. In certain embodiments, a digital twin is a one-to-one replica of a real system. In certain embodiments, a digital twin includes simulated software and / or emulated physical devices (e.g., hardware).

[0100] Edge Device’. As used herein, the term “edge device” refers to any piece of hardware that controls data flow at the boundary between two networks. Edge devices fulfill a variety of roles, depending on what type of device they are, but they essentially serve as network entry or exit points. Some common functions of edge devices include, but are not limited to, the transmission, routing, processing, monitoring, filtering, translation and storage of data passing between networks.

[0101] Explainable Al As used herein, the term “explainable Al” refers to a set of processes and methods that allows human users to comprehend and trust the results created by machine learning algorithms. Explainable Al is used to describe an Al model, its expected impact and potential biases. It helps characterize model accuracy, fairness, transparency and outcomes in Al-powered decision making. Explainable Al also helps an organization adopt a responsible approach to Al development.[01021 Generative Al: As used herein, the term “generative Al” refers to models or algorithms that create brand-new output, such as text, photos, videos, code, data, and / or 3D renderings, from the vast amounts of data they are trained on. The models generate new content by referring back to the data they have been trained on, making new predictions.

[0103] Improved, increased or reduced: As used herein, these terms, or grammatically comparable comparative terms, indicate values that are relative to a comparable reference measurement. For example, in some embodiments, an assessed value achieved with an agent of interest may be “improved” relative to that obtained with a comparable reference agent.Alternatively or additionally, in some embodiments, an assessed value achieved in a subject or system of interest may be “improved” relative to that obtained in the same subject or system under different conditions (e.g., prior to or after an event such as administration of an agent of interest), or in a different, comparable subject (e.g., in a comparable subject or system that differs from the subject or system of interest in presence of one or more indicators of a particular disease, disorder or condition of interest, or in prior exposure to a condition or agent, etc.). In some embodiments, comparative terms refer to statistically relevant differences (e.g., that are of a prevalence and / or magnitude sufficient to achieve statistical relevance). Those skilled in the art will be aware, or will readily be able to determine, in a given context, a degree and / or prevalence of difference that is required or sufficient to achieve such statistical significance.

[0104] Industry 4.0 : As used herein, the term “industry 4.0”, also called the fourth industrial revolution, is the next phase in the digitization of the manufacturing sector. Industry 4.0 describes manufacturers who are integrating new technologies, including Internet of Things (loT), cloud computing and analytics, Al and machine learning into production facilities and throughout operations.

[0105] Internet of Things (loT) As used herein, the term “internet of things (loT)” describes the network of physical objects, i.e., “things”, that are embedded with sensors, software, and other technologies for the purpose of connecting and exchanging data with other devices and systems over the internet without human intervention. These devices range from ordinary household objects to sophisticated industrial tools.[01061 loT Device’. As used herein, the term “loT device” refers to anything that has a sensor attached to it and can transmit data from one object to another or to people with the help of the internet and / or other means of connectivity. loT devices include wireless sensors, software, actuators, computer devices and more. loT devices may be attached to a particular object that operates through the internet, enabling the transfer of data among objects or people automatically without human intervention.

[0107] Reference: As used herein, the term “reference” describes a standard or control relative to which a comparison is performed. For example, in some embodiments, an agent, animal, individual, population, sample, sequence or value of interest is compared with a reference or control agent, animal, individual, population, sample, sequence or value. In some embodiments, a reference or control is tested and / or determined substantially simultaneously with the testing or determination of interest. In some embodiments, a reference or control is a historical reference or control, optionally embodied in a tangible medium. Typically, as would be understood by those skilled in the art, a reference or control is determined or characterized under comparable conditions or circumstances to those under assessment. Those skilled in the art will appreciate when sufficient similarities are present to justify reliance on and / or comparison to a particular possible reference or control.

[0108] Risk: As used herein, the term “risk” describes a probability of an event occurring. In some embodiments, a risk is a probability of a negative event occurring. In some embodiments, a risk describes scale or magnitude of harm. In some embodiments, a risk is used to describe how likely an attempt would be made to exploit a vulnerability.

[0109] Threat: As used herein, the term “threat” describes a vector which could be used to exploit a vulnerability. In some embodiments, a threat is a user (e.g., an unauthorized user), software (e.g., malware), hardware (e.g., a malicious device), or a mode through which a vulnerability is exploited.

[0110] Vulnerability: As used herein, the term “vulnerability” describes a weakness, flaw, or another shortcoming, e.g., in a system, hardware (e.g., a device or a component thereof), or software. In some embodiments, a vulnerability is a technical vulnerability based in softwareand / or hardware. In some embodiments, a vulnerability exists or is present due to potential actions (or inactions) of a user.DETAILED DESCRIPTIONA. LIVE TRACKING OF CYBER COMPONENTS

[0111] As the types of software, hardware, means of connecting, and sheer number of devices all increase over time, solutions and strategies for protecting such devices and the networks they are connected to are become more and more complex. FIG. 1 is a schematic diagram of an enterprise network (100), according to aspects of the present embodiments. As shown in FIG. 1, in some embodiments, the network (100) may include a plurality of wide-area networks (WAN) (102). Each wide-area network (102) of the plurality of wide-area networks (102) may include a plurality of local area networks (LANs) (104). Each LAN (104) may include one or more servers (106). Each server may include a plurality of nodes and / or node clusters (108). Each node or node cluster (108) may include a plurality of connections including (but not limited to) additional nodes I node clusters / sub-nodes (108), computers (114), robotic arms (116), lab equipment (118), drones (120), mobile devices (122), local controllers (for example, programable logical controllers) (128), and / or other network connectable devices. Accordingly, enterprise networks may expand very quickly and may be structurally and operationally complex.

[0112] FIG. 2 is a schematic diagram of a live mapping framework or system (200), according to aspects of the present embodiments. In some embodiments, the system (200) may include automated framework for tracking all electronic components (both hardware and software) that are connected to a target Industry 4.0 system. In some embodiments, a tracking system (212) may be incorporated inside every electronic device. In some embodiments, the tracking system (212) periodically extracts necessary information about the device such as hardware specifications, software packages installed, USB connections being used, and other key attributes. In some embodiments, the extraction frequency may be selectable by the user (forexample, the extraction frequency may include a predetermined user-defined update frequency). The information extracted from the devices may then be forwarded to a control system which may perform one or more of the following tasks: create a live map of how every electronic device is connected within the network; and populate a database with information about specific devices based on the reported data using the tracking systems. Based on this collected data, the system (200) enables: (1) an understanding of the need for software / hardware upgrades depending on latest vulnerability findings; (2) detection of policy violations; (3) detection of suspicious connected devices; and (4) insight generation for mitigation of vulnerabilities and prevention of tasks. In some embodiments, mitigation of vulnerabilities includes, but is not limited to, isolating (e.g., temporarily isolating) one or more components (e.g., devices) of a system and applying an update (e.g., a firmware update, a software update) to a device of a system.

[0113] Referring still to FIG. 2, devices connected within system (200) (and including the embedding tracking system (212)) may include drones, routers, databases, production / fabrication I manufacturing equipment, robots, laboratory equipment, robotic aims, desktop computers, programmable logical controllers (PLCs), laptop computers, small phones, tables, other wi-fi enabled devices, as well as other connectable devices. In addition, a system (200) may include connections to utilities such as power, water, gas infrastructure, cloud-based systems, remote factories, human-in-the-loop devices, and other external systems. In some embodiments, a system may track one or more of the following pieces of information about a given device (e.g., an IoT / edge device) via a tracking unit (212) or (T): information about any devices connected to Input / Output Ports including (but not limited to): USB; Bluetooth; ethemet, general purpose input / output (GPIO), universal asynchronous receiver / transmitter (UART), Pmod connectors, microSD card connectors, and / or USB-JTAG. In addition, in some embodiments, a system may track one or more of the following pieces of information about a given device (e.g., an loT / edge device) via a tracking unit (212) or (T): operating system version, build, and other details; network interfaces; software packages installed on the device including versions, builds, and installation / last-updatc timcstamp(s); hardware (processor, memory, graphical processing unit, etc.) including version and / or build information; information aboutany virtual hosts on a specific device; metadata including uptime, owner information, GPS tracking data, etc.; and / or installed packages. Once this information is collected by the system (200), the following capabilities become enabled: network discovery, active node enumeration, and network layout generation (i.e., graphical network mapping). Insights can then be generated by correlating identified risks to a number of possible attributes (i.e., network connection information, device type, software version, device usage time period information, etc.), thereby enabling appropriate mitigating actions to be taken.

[0114] FIG. 3 illustrates a method (314) for tracking unit information extraction deployed on every edge / IoT device, according to aspects of the present embodiments. In some embodiments, tracking units (e.g., tracking agents) are deployed on every edge / IoT device connected to an loT system (100). If an operating system of a specific device (D) does not support the installation of a tracker, then the loT device(s) immediately adjacent (in a networked system (100)) to D may become responsible for tracking D. FIG. 3 highlights this process. A tracking unit (T) (e.g., a tracking agent) can be realized as: (1) a software which can be installed onto the system or (2) as a physical hardware plugin. Both embodiments are capable of performing the same functions. As illustrated in FIG. 3, the method (314) begins at step 316. At step 318, the method (314) may include waiting for a period of time Tl, which may include an administrator-define time period. At step 320, the method (314) may include collecting edge / loT device information, as described herein. At step 322, the method (314) may include transmitting information to one or more databases within the system (100). In some embodiments, the transmitted information includes one or more relevant timestamps, as well as other potential metadata (for example, nodal location, device ID, installed software, etc.). At step 324, the method (314) may include determining if the system (100) should stop tracking the device in question. If the system (100) determines that it should stop tracking the device in question, the method (314) proceeds to step 326, at which point the method includes discontinuing tracking of the device in question. If the system (110) determines that it should continue to track the device in question, the method (314) (at step 328) proceeds back to step 318 to repeat steps 318, 320, 322, and 324.[01151 FIG. 4 illustrates a method 400 for loT system map generation, according to aspects of the present embodiments. An loT system map generation process (400) may include collecting and / or displaying network-level information and internal information of every loT / edge device in an loT system. As illustrated in FIG. 4, the method (400) begins at step 432. At step 434, the method (400) may include waiting for a period of time T2, which may include an administrator-define time period. At step 436, the method (400) may include building and / or refining a network map. At step 438, the method (430) may include mapping edge I ToT device information on the network map. At step 440, the method (400) may include storing edge / loT device information in one or more databases (e.g., a central database). At step 440, the method (400) may include determining if a system (200) should stop processing (for example, updating, for example, querying the network for new data) the network map. If the system (200) determines that it should stop processing the network map, the method (400) proceeds to step (444), at which point the method includes discontinuing building I processing of the network map (and may include displaying the network map, for example in a graphical user interface (GUI)). If the system (200) determines that it should continue to process the network map, the method (400) (at step 446) proceeds back to step 434 to repeat steps 434, 436, 438, 440, and 442.

[0116] Referring still to FIG. 4, device-level tracking data may be collected in a database node for additional processing. As shown in FIG. 4, tracker information may be combined with the network layout information (for example, at step 438) in order to create an loT system living map. According to aspects of the present embodiments, this map may be generated at a periodicity of T2 (specified by the system admin). This information, for each time stamp, may be stored in the database for future use during the analysis phase.

[0117] FIG. 5 illustrates a method (500) for anomaly detection, according to aspects of the present embodiments. In some embodiments, the method (500) of anomaly / policy violation detection may include using temporal and differential loT system map data. As illustrated in FIG. 5, the method (500) begins at step 548. At step 552, the method (500) may include: inputting an loT system map at a first time point, XI; inputting the loT system map at a second time point, X2; assessing device usage information in view of device usage policies; and compiling a vulnerability list. At step 554, the method (500) may include detecting anomaliesand / or policy violations within the network (i.e., based on the first time point, the second time point, policies, etc.) At step 556, the method (500) may include detecting anomalies and / or policy violations in or associated with newly added and / or updated software packages. At step 558, the method (500) may include detecting anomalies and / or policy violations in or associated with newly connected device(s) on inputs and / or outputs of other devices (i.e., other devices connected to a system (200)). At step 560, the method (500) may include detecting vulnerable hardware and / or software packages. At step 562, the method (500) may include generating one or more status reports. At step 564, the method (500) terminates. As illustrated in FIG. 5, the present disclosure, in some embodiments, describes a specific embodiment where two loT system maps at two different timestamps are analyzed. However, the overall process may be applicable for analyzing multiple (for example, more than 2) system maps over several time stamps. Additional timestamp data allows for more accurate analysis, at the cost of computation efficiency. In some embodiments, additional timestamp data is collected only when one or more variation thresholds is met (for example, when one or more data attribute differs a certain amount in comparison to previous timestamp data). In some embodiments, additional timestamp data is used to identify anomalies which can occur over longer periods of time. In some embodiments, once an anomaly and / or policy violation is detected, the topology of the network can be changed. For example, a host (or more than one host) can be isolated (e.g., removed) from the network until an appropriate fix (e.g., a firmware update, a software update, or the like) is applied, either remotely or directly to the host device.

[0118] FIG. 6 is a schematic diagram of a device and network tracking hierarchy (600), according to aspects of the present embodiments. As illustrated in FIG. 6, loT device and network data tracking can be done in a hierarchical manner based on network sub-networks. This enables sub-net level privacy and can expedite both the map building process and the analysis process using distributed computing. Stated otherwise, the data collection and storing processes can be carried out in a distributed hierarchical manner. The hierarchical approach disclosed herein provides one or more potential advantages. For example, distributed processing can significantly speed up the data collection and analysis process, as discussed. In addition, certain sub networks might include restricted access. That is, internal information cannot be sentoutside the network. Hence, for those sub networks, the analysis can be executed in a self- contained manner.

[0119] FIG. 7 is a schematic diagram of a device information tracking over time, according to aspects of the present embodiments. Tracking edge / IoT device information over time enables detection of maliciously connected devices, unauthorized software package installation, suspicious I / O connections, vulnerable electronic hardware, and other potential threats to the system. By instantiating an loT system map that considers and reconciles different timestamps via cross-referencing, detection of malicious input / output (I / O) connections, package installations, and device connections may be achieved. The stored hardware / software information may also be cross checked with the current vulnerability lists to ensure security. It is believed that, as of the time of this disclosure, no such frameworks exist currently that can accurately and with high scalability track cyber / digital aspects (network, hardware, software) of an Industry 4.0 and digital manufacturing system. The system-wide live map of every device including attribute information such as connection details, device type, and software packages may be temporally tracked to generate insights. In addition, in accordance with the present disclosure, system vulnerabilities may be identified and mitigated by enabling algorithms and methodologies used to detect anomalous behavior and policy violations based on multiple loT system maps generated at different timestamps (temporal and differential strategies) in the context of Industry 4.0 and digital manufacturing.B. Al Guided Automated Vulnerability Detection & Fixing

[0120] FIG. 8 is a schematic diagram of an automated Al-guided penetration testing framework (800), according to aspects of the present embodiments. In some embodiments, a framework (800) may include a Generative Artificial Intelligence (Al) guided framework (800) for automated penetration testing (“pen testing”) and fixing of Industry 4.0 systems using a digital twin approach. In some embodiments, a framework may include: a live mapper framework (832) used to create a snapshot (e.g., a profile) (834) of a given Industry 4.0 system at a given timestamp. This information may then be used to automatically to create a virtual systemreplica (836) (digital twin, digital replica) of the entire system (810). A Generative Al framework (842) may be used for automatically carrying out attacks on the entire system (810) with the goal of discovering different vulnerabilities. Based on the vulnerabilities (844) discovered, a set of solutions (846) may be generated, which may then be applied back to the virtual replica (838). In some embodiments, the framework (832) may include continuing this loop until the virtual system replica (838) (digital twin, digital replica) is hardened (e.g., against cyber threats) beyond a certain threshold determined based on different metrics. Finally, in some embodiments, the framework (832) may include updating / fixing (at step 840) a real base system (810) (installed network) based on the knowledge obtained through this analysis, in order to make the system (810) more robust and attack resilient.

[0121] A state of the loT system (inside the digital twin 838 and in the real world) must be captured in a formal manner (e.g., by vectorization of the system state) to enable automated penetration testing and repair. All aspects of an loT system may be captured inside this state representation including (but not limited to): the loT network, loT devices, and / or auxiliary knowledge. In some embodiments, the loT network (810) covers aspects such as (but not limited to): connectivity, interface types used, networking protocols used, network settings, and / or QoS (quality of service) settings. In some embodiments, loT devices include aspects such as (but not limited to): hardware information, software packages installed, communication patterns, attached devices / components, information relating to internally hosted virtual systems. In some embodiments, auxiliary knowledge includes certain information and insights of the overall system as perceived by the penetration testing framework at a given point in time. These data points may be stored for comparison to future points in time, and insight generation. In some embodiments, auxiliary knowledge includes (but is not limited to): IP and MAC addresses of discovered loT devices / hosts, information regarding which hosts / devices have been compromised and are under control of the penetration testing framework, information about which hosts have known vulnerabilities, and / or information about leaked credentials.Additionally, the state must be represented as a vector for machine learning and artificial intelligence (Al) use. In some embodiments, this vectorized state may be performed following strategies such as (but not limited to): (1) one-hot encoding; (2) linear encoding; and / or (3)neural network encoding. In certain embodiments, neural network encoding inputs raw data (e.g., a graph, data points) into a neural network model and extracts the output from different layers of the neural network to create a set of features. In some embodiments, a neural network model is a pretrained neural network model. In some embodiments, a set of features are extracted as a vector (e.g., of floating-point numbers).

[0122] FIG. 9 illustrates an Al-guided cyber vulnerability mitigation method (900), according to aspects of the present embodiments. At step 906, the method (900) may include creating a live map framework. At step 908, the method (900) may include creating a network snapshot at time Tl. At step 912, the method (910) may include automatically creating a virtual replica (i.e., digital twin) (838) or the loT system (810). At step 914, the method (900) may include using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests. At step 916, the method (900) may include identifying vulnerabilities. At step 918, the method (910) may include generating one or more solution sets. At step 920, the method (910) may include applying the one or more solution sets to the virtual replica (i.e., digital twin) (838). At step 922, the method (900) may include determining if the virtual replica passes the robustness test. If the virtual replica does not pass the robustness test, at step 926, the method (900) may include returning to step 914. If the virtual replica does pass the robustness test, at step 924, the method (900) may include addressing real system vulnerabilities based on the knowledge and insights obtained from the vulnerability testing of the virtual replica (i.e., digital twin) (e.g., based on a solution set generated) (838).

[0123] FIG. 10 illustrates an Al model re-training method (1000), according to aspects of the present embodiments. As illustrated in FIG. 10, the method (1000) starts at step 1002. At step 1004, the method (1000) may include inputting a feedback dictionary (FD) and an Al pen test model (M). At step 1006, the method (1000) may include retraining the model M based on the feedback dictionary (FD) and the Al pen test (i.e., penetration testing) model (M). At step 1008, the method (1000) terminates. As shown in FIG. 10, the model M may be further readjusted periodically (as determined by the administrator) using a feedback dictionary (FD), which is generated and / or updated during the automated penetration testing process.[01241 FIG. 11 illustrates an Al-guided penetration testing method (1100), according to aspects of the present embodiments. As illustrated in FIG. 11, the method (1100) starts at step 1158. At step 1162, the method 1160 may include inputting an loT system map I (i.e., a digital replica) and an Al pen test (i.e., penetration testing) model M. At step 1164, the method (1100) may include vectorizing the system state I. At step 1166, the method (1100) may include choosing (i.e., determining) action pathways based on the model and state. At step 1168, the method (1100) may include updating the state based on the action pathways and the existing state (e.g., by execution of the action pathway). In some embodiments, an action pathway is defined as a set of steps taken to move the state of the system from one state to another state. At step 1170, the method (1100) may include computing one or more action pathway finesses (i.e., fitness score) based on the original state, the new state and the action pathway. At step 1172, the method (1100) may include getting human feedback based on the original state, the new state, and the action pathway. At step 1174, the method (1100) may include merging feedback (i.e., the human feedback with the computed action pathway). At step 1176, the method (1100) may include storing the state, action pathway and merged feedback in the feedback dictionary, FD. At step 1178, the method (1100) may include updating a state (i.e., the previous or original state) to a new state. At step 1180, the method (1100) may include updating the loT system map based on the new (updated) state. At step 1182, the method (1100) may include assessing if one or more goals have been reached. If the one or more goals have not been reached, the method (1100) may include returning to step 1164. If the one or more goals have been reached, at step 1184, the method (1100) may include extracting penetration testing knowledge from the loT system map. At step 1186, the method (1100) may include outputting the extracted knowledge (i.e., from step 1184). In certain embodiments, knowledge from the test can be used to identify vulnerabilities of a real system based on the loT (i.e., acting as a digital replica of the real system). At step 1188, the method (1100) terminates.

[0125] Referring still to FIG. 11, in some embodiments, explainable Al techniques may be used for extracting the reasonings behind various action pathways constructed by the Al model M. These human-understandable explanations behind each Al decision enable the human expert to provide more accurate feedback faster. As illustrated in FIG. 11, the automatedpenetration testing process is initiated by vectorizing the current loT system state (inside the digital twin). Then the Al model (M) analyzes the state (S) to determine the best next action pathway (i.e. , a set of actions to be canned out by the penetration testing framework), for example at step 1166. Next this action pathway (A) is executed and the new state (S2) is attained at step 1168. The goodness / fitness of this action pathway is captured at step 1170 based on metrics such as (but not limited to): whether or not the actions lead to the discovery of vulnerabilities; the potency of the discovered vulnerabilities; whether any new loT devices on the network got discovered; whether any vulnerability was exploited successfully; whether any user / admin credentials were exposed; and / or whether any backdoor(s) were successfully inserted. This metric-based fitness value is then augmented with additional feedback from human experts (optionally) towards creating FM, at steps 1172 and 1174. Parameters [S, A, S2, FM] are then stored in a dictionary (e.g., a feedback dictionary) for later use during the Al model retraining process, at step 1176. S and the loT system (I) are next updated at step 1180. At step 1182, the method (1100) deteimines whether to terminate the penetration testing process or not. This decision is guided by different goals such as: whether all sensitive nodes got compromised; whether a certain time has elapsed since the start of this penetration testing process; whether all sensitive data were compromised; and / or whether all possible backdoors were successfully inserted. If all set goals (by the admin) are reached, then the knowledge (K) obtained through the penetration testing process (on the digital twin) is captured (at step 1184) for later use. The capture knowledge K is later used to update the actual real system (on which the digital twin is based) with the goal of mitigating all discovered vulnerabilities. If all goals are not reached, the process may be repeated in a new round.

[0126] FIG. 12 illustrates an action pathway generation method (1200), according to aspects of the present embodiments. The action pathways may be computed as shown in FIG. 12. Based on the current state of the loT system (inside the digital twin), the best possible actions may be computed using the Al model M. This model (M) may use generative Al techniques that create effective actions. These actions may then be stitched to form different action pathways (AP). Each action pathway may in turn be scored. Finally, the optimal action pathway is returned from this process.

[0127] Referring still to FIG. 12, the method (1200) starts at step 1210. At step 1220, the method (1200) may include inputting an loT system state S and an Al pen test model M. At step 1230, the method (1200) may include identifying action options (AO) based on the loT system state and the Al pen test model. At step 1240, the method (1200) may include creating actin pathways based on the action options identified in step 1230. At step 1250, the method (1200) may include scoring the action pathways (for example, on a numerical scale). At step 1260, the method (1200) may include selecting an optimal action pathway based on the score and the action pathways. At step 1270, the method (1200) terminates.

[0128] According to aspects of the present embodiments, an Al model (M) used during a penetration testing process may be initially trained on a custom dataset prepared by an administrator. Generative and explainable Al guided vulnerability detection and exploitation provide effective and transparent approaches to vulnerability identification. As it pertains to Industry 4.0 systems, which may include diverse sensors and digital components, generative and explainable Al guided vulnerability detection quickly identify risks to the system and generate solution sets for mitigating identified risks and / or vulnerabilities. The present disclosure provides automated knowledge translation from the virtual digital twin to the real system. In addition, the disclosed Al techniques, action pathway generation flow, Al model training process, proposed state capturing technique offer a holistic approach to threat identification and risk management.

[0129] FIG. 13 illustrates a method (1300) of creating a digital twin I virtual replica of a system (e.g., an loT system, an Industry 4.0 system). As shown in FIG. 13, a digital twin can be created based on detailed information obtained from tracking a real loT or Industry 4.0 system. For example, an input to this process can be an loT system representation (IR) (1304). Based on an IR, a framework determines which portions of the loT system to map to real hardware (1308) (HE) and which portion is to be represented through software simulation (1310) / emulation (1308) (SE). Next, the software simulation environment is initialized (1312) (I), the emulated / simulated hosts are spawned (1314), and the hardware elements (HE) are programmed based on the components of the IR being mapped to them (1316). Next, the networking between the simulated and emulated components (software -based or hardware -based) are defined (1318)and instantiated (1322). After, services and processes are created on the hosts (both softwarebased hosts and hardware-based hosts) and network traffic is generated based on the IR.

[0130] Referring still to FIG. 13, the method (1300) of creating a digital twin starts at step 1302. At step 1304, a representation of a system (e.g., an loT system representation (IR), an industry 4.0 system) is input. A representation of a system can be based on, for example, tracking data from a tracking system (e.g., using tracking agents within a system). At step 1306, the method (1300) may include identifying portions of a representation of a system for hardware emulation (HE). At step 1308, the method (1300) may include identifying portions of a representation of a system are identified for software emulation. In some embodiments, a system may not include portions needing hardware or software emulation. At step 1310, the method (1300) may include identifying portions of a representation (e.g., an IR, an industry 4.0 system) that will be used for simulation. At step 1312, the method (1300) may include initializing a simulation I emulation environment (I) (e.g., using software). At step 1314, the method (1300) may include adding software emulated and / or simulated hosts to an initialized simulation / emulation environment (I). At step 1316, the method (1300) may include programming hardware elements to mimic hardware emulated (HE) components and adding the HE components to an initialized emulation / simulation environment (I). At step 1318, the method (1300) may include creating networking and connectivity among emulated and / or simulated systems (e.g., all of the emulated and simulated systems). At step 1320, the method (1300) may include starling processes and / or services across emulated and / or simulated systems. At step 1322, the method (1300) may include initiating network traffic between simulated and / or emulated hosts. At step 1324, the method (1300) outputs a simulation / emulation of the system (e.g., an loT system, an Industry 4.0 system).

[0131] According to aspects of the present embodiments, a live map may be translated or transformed into a virtual replica / digital twin of a system. Devices identified in the live map can be emulated using either real hardware (R) or simulated (S). For each device (r) that is to be mapped using real hardware (R), a real target loT device (t) which closely matches its specification can be identified and software elements of the device (r) can be mapped onto the loT device (t). For each device that is to be simulated (s), a virtual system (v) with the samehardware resources as the device(s) can be created inside a simulation environment. Exemplary simulation environments include Microsoft Azure. Software elements of each simulated device can also be mapped onto the virtual system. Digital twin mapped devices, both virtual and real, can then be connected together based on the network topology of the real system. Network traffic based on behavior observed in the real system can also be simulated. Processes of mapping a real device onto a target device and creating simulated devices in a virtual system can be carried out automatically.

[0132] FIG. 14 is an illustrative embodiment of an exemplary hardware / software emulation / simulation setup (1400) (e.g., as used in FIG. 13). A set of physical devices (e.g., in a physical setup layer) (1410) can be connected via a router that in turn can be connected to a control system / controller (1420). A simulation platform can also be connected to the controller. The physical layer (1410) and the simulation elements (1430) can together mimic the behavior of a real loT system / application (1440). The controller can allow automated configuration of this setup to realize the target goal. The data, about the state of the emulated / simulated system, from the controller can be visualized via AR / VR techniques (1450) for providing expert feedback necessary during the Al training process.

[0133] FIG. 15 illustrates a component of a method (1500) of an Al guided penetration testing. As illustrated in FIG. 15, an attack model (AM) is initialized (1502) and a defender model (1504) is initialized. In certain embodiments, a model (e.g., an attack model, a defender model) can be initialized using training data, random weights. An attacker can be identified as belonging to or being on a “red team.” A defender model (DM) can be a random model, an AI- based model, or a human-inspired model. A defender can be identified as belonging to or being on a “blue team.” An attacker model then performs an action (1504) (e.g., an attack action) on a virtual system replica (digital twin) (1506) of a system described herein (e.g., an loT system, an Industry 4.0 system). A defender model (DM) ( 1508) can perform (at the same time or a later time) a defense move (e.g., a counter action (CA)) (1510) as part of a defense pathway on the virtual system replica (digital twin) of the system. A defense move (e.g., a counter action (CA)) (1510) of a defense pathway can include performing one or more actions including, but not limited to, scanning the network of a virtual system replica (digital twin) for an anomaly,modifying the topology of the network of a virtual system replica (digital twin) (e.g., by isolating a device), bringing down a host (or more than one host), and disconnecting one or more hosts from the system (e.g., edge devices, and attempting recovery (e.g., of a system, of a sub-system, of devices in a system). In the present example, a host can refer to a device on a network or a sub-system of a network. The effectiveness of the action(s) performed (1504) by the attack model AM (1502) and / or the counter-action(s) performed (1510) by the defender model (1508) can be observed (1512) based on a change in state of the virtual system replica (digital twin). A goodness / fitness score can be computed based on actions performed by attack and defender models. In some embodiments, human feedback is provided to update and / or generate a goodness / fitness score. Feedback (1516) can be used to update an attack model (AM) (1518) and / or a defender model (DM) (1520). After updating one or both of the attack model (AM) or defender model (DM), the process can be repeated to increase the robustness of the attack model and / or defender model.

[0134] Referring still to FIG. 15, a component (1500) of Al guided penetration testing is described. An attack model (AM) (1502) and the defender model (DM) (1502) are both initialized based on preliminary training data or using random weights. Next, AM makes an attack move (A) on the digital twin and DM makes a defense move (D). Based on these moves the digital twin changes state and a score is computed to determine the goodness / fitness (FM) of both A and D. This feedback is then used to update AM and DM towards making them more robust in the future.

[0135] FIG. 16 illustrates an Al-guided penetration testing method (1600), according to aspects of the present embodiments. As illustrated in FIG. 16, the method (1600) starts at step 1601. At step 1602, the method (1600) may include inputting an loT system map I and an Al pen test (i.e., penetration testing) attack model AM and a defender model DM. At step 1604, the method (1600) may include vectorizing the state. The loT system map can be a map of a digital replica at a first time point or state. At step 1606, the method (1600) may include choosing action pathways based on the Al pen test attack model AM and state. At step 1608, the method (1600) may include updating the state (to SI) based on the action pathways and the existing state. In some embodiments, an action pathway is defined as a set of steps taken to move thestate of the system from one state to another state. At step 1610, the method (1600) may include choosing (e.g., determining) a defense pathway (D) based on a defender model DM and state of the system. At step 1612, the method (1600) may include updating the state (to S2) based on the action pathways, defense pathway, and an updated state. At step 1614, the method (1600) may include computing one or more action pathway finesses (i.e. , fitness score) based on the original state, new state(s), action pathway, and defense pathway. At step 1616, the method (1600) may include getting human feedback based on the original state, new state(s), action pathway, and defense pathway. At step 1618, the method (1600) may include merging feedback (i.e., the human feedback with the computed action pathway). At step 1620, the method (1000) may include storing the state, updated state(s), action pathway, and merged feedback in a feedback dictionary, FD. In some embodiments, the information is used to create a feedback dictionary. In some embodiments, the information is used to update an already existing feedback dictionary. At step 1622, the method (1600) may include updating a state (i.e., the previous or original state) to a new state. At step 1624, the method (1600) may include updating the loT system map based on the new (updated) state. At step 1626, the method (1600) may include updating the Al pen test (i.e., penetration testing) attack model AM and defender model DM based on the merged feedback. At step 1628, the method (1600) may include assessing if one or more goals have been reached. If the one or more goals have not been reached, the method (1600) may include returning to step 1604. If the one or more goals have been reached, at step 1628, the method (1600) may include extracting penetration testing knowledge from the loT system map. At step 1632, the method (1600) may include outputting the extracted knowledge (i.e., from step 1630). At step 1634, the method (1600) terminates.

[0136] Referring still to FIG. 16, in some embodiments, this framework can start by vectorizing the current loT system state (inside the digital twin) (1604). Then the Al pen test attack model (AM) will analyze the state (S) to determine the best next action pathway (set of actions to be carried out by the penetration testing framework) (1606). Next, the determined action pathway (A) is executed, and the new state (SI) is attained (1608). The Defender Model (DM) next takes an action (D) (1610). Based on D, the state of the system changes to S2 (1612).[01371 The goodness / fitness of D and A (1614) are captured based on metrics including, but not limited to:1. Did the actions lead to the discovery of vulnerabilities?2. The potency of the discovered vulnerabilities.3. Did any new loT devices on the network get discovered?4. Did any vulnerability get exploited successfully?5. Did any user / admin credentials got exposed?6. Did any backdoor(s) get inserted?7. Did the defense action hinder the attack?

[0138] This metric-based fitness value is then augmented with additional feedback from human experts (optionally) (1614) towards creating merged feedback FM. [S, A, SI, D, S2, FM] are stored in a dictionary (1620) (e.g., a feedback dictionary) for later use during the Al model retraining process. S and the loT system (I) are next updated (1622, 1624). Next this framework decides whether to terminate the penetration testing process or not (1628). This decision is guided by assessing different goals including, but not limited to:1. Did all sensitive nodes got compromised?2. Did a certain time has elapsed since the start of this penetration testing process?3. Did all sensitive data get compromised?4. Did all possible backdoors got inserted?5. Did the attacker lose all foothold?

[0139] If all set goals (by the admin) are reached (“YES” pathway), then this framework captures (i.e., extracts) knowledge (K) obtained through the penetration testing process (on the digital twin) for later use (1630). The output, K, (1632) is later used to update the actual real system (on which the digital twin is based off) towards mitigating all discovered vulnerabilities.If all goals are not reached ("NO" pathway), then this framework will move to the next round of the same process. An example interaction between AM and DM is shown in FIG. 17.

[0140] FIG. 17 is an illustrative example of an interaction (1700) between an attack model (AM) and a defense model (DM). In certain embodiments, interactions between AM and DMs are used in methods of Al guided penetration testing (e.g., on a digital replica) as shown in, for example, FIGs. 15 and 16. In FIG. 17, “red” is an attacker from the attack model (AM), while “blue” is a defender from the defense model (DM). At step 1710, an attacker (red) can scan a network of a system (e.g., a network of an loT system, a network of an Industry 4.0 system). In step 1715, a defender can track network traffic of the system (e.g., to identify suspicious network traffic). At this point the attacker has not gained access to a host (represented by circles) in the network. In step 1720, the attacker gains access to a host. The accessed host is shaded in the network shown in FIG. 17. In step 1725, blue (the defender) can see (e.g., identifies) the connection made by red (the attacker) to the host in real time. At step 1730, red gains access to a second host of the network. In FIG. 17, two hosts are shaded to show red has gained access to two, separate hosts. At step 1735, blue starts to take action(s) (e.g., a defense action, a counter-action) to remediate the access of the attacker on the network. In some embodiments, an action can include, but is not limited to, modification of a network (e.g., the topology of the network) and disconnection of a host from a network. At step 1740, red loses access to one of the two hosts to which it has gained access. At step 1745, blue removes red from the host.

[0141] According to aspects of the present embodiments, interplay between an attack model (AM) and a defense model (DM) (e.g., as shown in FIG. 17) can allow for a framework to come up with attack vectors (for a target system) that can be used to deal damage as well as remain hidden from standard defense strategies. "Stealth attacks” are a new class of attack vector that can be automatically generated using this framework.C. Exemplary Tracking System

[0142] The present example describes exemplary tracking systems used in embodiments described herein. A tracking agent of a tracking system is installed (e.g., remotely installed,physically installed) on sub-systems and / or devices of a system. A tracking agent can use different methods and techniques to extract relevant information from a device and transmit the information (e.g., using encryption, e.g., end-to-end encryption) to a centralized system (e.g., a central database) for aggregation and / or analysis. In some embodiments, there may be one or more intermediary agents (e.g., an intermediary aggregation agent) to which data can be transmitted prior to being passed to a centralized system for tracking.

[0143] In some embodiments, a tracking agent extracts information from a processor (or more than one processor) of a device (e.g., an loT device) of a system (or sub-system). For example, a tracking agent can extract information from a hardware performance counter of a processor. In certain embodiments, a processor of a device has at least one, two, three, four, six, eight, eighteen or more hardware counters. Tracking agents can be a software-based component (e.g., software only) or a physical device (e.g., a stand-alone device) that is a part of a system (or sub-system), inclusive of a software component.

[0144] A tracking agent (e.g., installed on a device) can be used to extract relevant information from a neighboring device. For example, a tracking agent can extract relevant information as described above from a neighboring device. Additionally or alternatively, a tracking agent can be used to observe open ports of a neighboring device, look at network traffic, and look at services running on a system or sub-system network.

[0145] FIG. 18 shows an exemplary embodiment of a method (1800) of extracting information from a real system. A real system of the present example has multiple sub systems (e.g., Sub System 1, Sub System 2, etc.) (1810). For each of the sub systems (1815), information (1825) can be extracted which can be used to, for example, recreate an environment (1820) by mapping the system. In some embodiments, information extracted from a sub-system can attributes including, but is not limited to, internet protocol (IP) address, operating system information, sub-system serial number, connected USB (Universal Serial Bus) interfaces, network interfaces, connected device interfaces, USB history, installed packages. Exemplary Linux commands which can be used to extract information are provided as shown in FIG. 18. Based on the present specification and descriptions herein, a person of skill in the art would be able to implement these methods on a system. Extracted information can be transferred to a centralized system where it can be used to create a digital twin of the real system (1830).Extracted information can be transmitted encrypted using end-to-end encryption or other suitable means of encryption A digital twin (1830) can be used as described herein for, for example, penetration testing (“Pen testing”) and detection of system and / or sub-system vulnerabilities, among other uses.

[0146] Referring to Table I, in some embodiments, the tracking agent may first ping a given device (i.e., using the device’s IP address or the device name (or identifier) if the IP address is not known). The tracking agent may then use an ifconfig command (or equivalent such as the Unix “IP” command or ipconfig) to display the current network interface configuration information (i.e., IPv4 address, IPv6address, subnet Mask, default gateway, etc.) for the device in question (i.e., as it applies to both physical and virtual network interfaces). In some embodiments, the tracking agent may instead or additionally employ the ipaddr command (or equivalent) to display network configuration information associated with a connected device such as MAC address, IPv4 address, and / or IPv6 address. In some embodiments, the tracking agent pings a network or server and may employ the lsb_release command (or equivalent such as os-release and / or hostnamectl) in order to verify operating system information including the type, version, release name, software distributor, etc. of an operating system being used. In some embodiments, the tracking agent may use a Isusb command (or equivalent) to scan the system in order to identify USB devices and / or connections currently in operation or employed within the system, as well as the associated USB bus information. In some embodiments, for an identified USB device or bus, the tracking agent may query the USB history (for example, to identify suspicious behavior on a system) via a file activity page, a registry, using Windows PowerShell, via an associated device manager, and / or using a program such as Binalyze. In some embodiments, where (for example) other types of devices (i.e., non-USB devices) are employed, the tracking agent may use a dmesg command (or equivalent) to print a kernel message buffer via the device driver associated with the device.

[0147] Still referring to Table I, in some embodiments, the tracking agent may use an “apt list” command (or equivalent) to identify the names and / or types of software running on a system, including software that is upgradeable, thereby identifying potential vulnerabilities associate with software that has yet to be upgraded and may be susceptible to one or more nonehacks. In some embodiments, the tracking agent may further use the “ / var / log / apt / history.log” command (or equivalent) to retrieve the usage history for a program or application. In some embodiments, the tracking agent may use the netstat or net-tools commands (or equivalent) to make periodic cross-system (i.e., inter-network) queries to assess network operational information such as network interface information, as well as items such as the number of packets received, transmitted, and / or dropped within a given timeframe. By periodically monitoring intra-network activity over time, normal usage models can be developed to understand how much information is typically being exchanged. If larger than normal amounts of information (i.e., anomalous behavior) is detected, the system can identify the times and locations of such behavior to help identify potential suspicious activities. In some embodiments, the tracking agent may use commands such as neofetch or Ishw (or equivalent) to identify hardware information such as processor type, RAM, GPU, peripherals, and / or chipsets, and may compare these to “expected” hardware types, again to identify any anomalies. The tracking agent may also add time stamps to any collected information such that temporal trends may be captured and identified, as disclosed herein. In some embodiments, mitigating actions may be implemented by piggy-backing off of tracking agents. For example, in some embodiments, if a suspicious device is identified, a device API may be pinged by the tracking agent, and then reassigned to, for example, a quarantined IP address (thereby temporarily suspending the device access to the network) via the ifconfig or IP commands (or equivalent). In another example, where isolation of an entire network or server is needed, the hostnamectl command (or equivalent) can be used to temporarily rename a server, thereby creating a mismatch and mitigating potential threats.D. Exemplary Embodiment of Al Guided Vulnerability Detection

[0148] FIG. 19 shows exemplary systems and methods (1900) used for detecting attacks on a system as described herein.

[0149] In FIG. 19, a real system (1905) can be rc-crcatcd as a digital twin (1910) using simulation and / or emulation. As described herein, a digital twin (1910) can be a one-to-onerecreation of a real system (1905). A digital twin can include information extracted from a real system which encompasses, but is not limited to, cyber physical systems, software components, intemet-of-thing (loT) components, embedded hardware, operation technology (OT) components, information technology (IT) components, networking, networking information, and data traffic of a network mimicking normal operation. Information from a digital twin of the real system can be used to create or compile a normal profile data (1930) of the digital twin (1930). Normal profile data (1930) of a digital twin can be based on a profile of network traffic (e.g., prior to a simulated attack) (1915), a profile of inter-device communication patterns (1920), and a profile of intra-device inter-software / hardware patterns (1925). Profiles used to create a normal profile can be created using machine learning techniques (e.g., unsupervised machine learning, one-class classification techniques / class-modelling, and the like) and statistics-based techniques.

[0150] A digital twin can then be used for simulation of attacks on the real system (1935). For example, Al (artificial intelligence) frameworks can be used to simulate attacks on a digital twin. An attack on a digital twin can be used to generate one or more profiles from the attacked digital twin, thus creating a profile of a system under attack. Comparing the normal profile data of a system in normal operation (e.g., not under attack) with a profile of an attacked system can reveal discrepancies between the two sets of data. Analyses (1935) of the simulated attacks can be used to create a model for detecting deviations of a system from its normal behavior (1940). Models trained for detecting deviations from normal (1940) can be used to detect a real attack (1950) on a real system (1905) under a real attack (1945).E. Experimental Example of CAST-MAP Generation / . INTRODUCTION

[0151] Major industries such as additive manufacturing and semiconductors fabrication are rapidly moving towards Industry 4.0 setup and beyond. However, this progress is being impeded by mounting cybersecurity concerns that can lead to design / data tampering, data theft,and denial-of-service. Adding to this problem is the complexity and diversity of the digital infrastructure that is necessary to realize the Industry 4.0 dream. A cybersecurity risk analysis and mitigation of an Industry 4.0 system is not possible without first knowing the exact topology of all the electronic hardware, software, computing, sensors, and edge components that are being used in the said Industry 4.0 system. This is a particularly challenging task due to the diversity of hardware / software and the dynamic nature of these systems. An automated framework that can tackle this challenge by creating a live / dynamic spatio-temporal map of all the cyber / digital components in an Industry 4.0 system. An unsupervised learning guided scheme for computing the system’s cybersecurity risk / criticality value by leveraging these spatio-temporal maps. In the present example, the methods and systems described are implemented and validated to demonstrate the effectiveness of this framework using a real Industry 4.0 testbed and a randomized large scale system generation method.

[0152] Industry 4.0 has had large scale adoption already underway across diverse sectors including automobile, additive manufacturing, semiconductors fabrication, renewable energy, and aerospace. At its core, Industry 4.0 is about automatic collection of data from diverse sensors on the manufacturing bed, movement of data over the network, data storage, data analytics, and actuation based on inferencing. To achieve these steps a large number of sensors (e.g., optical, airflow, profilometer, proximity sensor), electronic operational devices (e.g., robotic arms, 3D printers), edge devices (e.g., Raspberry Pi, Jetson, Orin), and information technology infrastructure (e.g., networking devices, data hubs, compute units) are utilized.These units are made up of electronic hardware and software. FIG. 20 is a visual representation of complexities of an automated manufacturing system / facility.

[0153] Keeping a track of these cyber units in an Industry 4.0 system without automation can be a challenging and almost impossible task for a large scale setup. At the same time, automating this process is also not trivial due to the complexity of the system, presence of diverse electronic components, non-uniform communication methodologies, and overall scale. However, tracking these electronic units over time and understanding their communication patterns is a crucial step for cybersecurity threat detection and risk management.

[0154] Cybersecurity of a manufacturing facility (e.g., for semiconductors, additive manufacturing, aerospace) is critical for ensuring robust operation in this era of increasing cyberthreats. Successful cyber attacks on these facilities can lead to, among other things, manufacturing halt for undetermined amount of time, creation of a physically unsafe operational environment, intellectual property theft, data leakage, and loss of credibility. Depending on the industry, such an event can potentially be also life threatening.

[0155] Hence, to aid in the process of cybersecurity threat and risk analysis / management, presented herein is a Cybersecurity Assessment using Spatio-Temporal Mapping (CAST-Map) of a target ToT system. This graph / map contains information about individual cyber components in the system, their connection / communication patterns, and the evolution of these components over time. Also presented herein is an unsupervised learning scheme that utilizes this spatiotemporal graph information for calculating a system- wide cybersecurity criticality value that is more accurate than other state-of-the-art criticality score estimation techniques. Use of both the tracking and criticality assessment techniques provides for a highly robust framework. In the present example, a live Industry 4.0 testbed was successfully tracked using the tracking framework and taking inspiration from this tracked system. A randomized Industry 4.0 system generation scheme is also described that can generate systems with a very large number of subsystems (up to 2000 sub systems tested) of varying parameters. The randomized Industry 4.0 system generation scheme is then used to evaluate the criticality assessment framework.

[0156] In summary, the following objectives were achieved in the present experiment:• Analyze in detail the concerns associated with ensuring the security of large scale Industry 4.0 facilities.• Develop a system architecture and associated algorithms for creating a spatio-temporal evolution map of a target Industry 4.0 system through live tracking of all cyber components and events.• Develop an unsupervised learning algorithm to more accurately estimate the criticality score of a connected Industry 4.0 electronic system utilizing the spatio-temporal evolution map.• Implement this tracking and criticality estimation techniques as highly robust and parameterized frameworks.• Evaluate these frameworks using a real Industry 4.0 additive manufacturing testbed and a randomized largescale system generation procedure.2. MOTIVATIONS i. Current State of Cybersecurity: Industry 4.0

[0157] The Industry 4.0 ecosystem stands on six pillars: (1) Predictive Engineering; (2) Data; (3) Manufacturing Technology and Process; (4) Resource Sharing and Networking; (5) Materials and (6) Sustainability. Out of these six, pillars 1-4 involves the use of Operation Technologies (e.g., 3D printers, robots, drones), Information Technologies (e.g., high performance computing, data servers), networking devices, edge devices, and sensors. Each of these devices / systems are also internally comprise of electronic hardware and software making an Industry 4.0 setup extremely complex from the digital / cyber perspective. Securing this complex system is not an easy task and cyber attacks on an Industry 4.0 setup can lead to massive financial, brand-value, and safety concerns. Unique Industry 4.0 Cybersecurity concerns arises due to the following reasons:

[0158] 1) Connectivity and Automation: A smart manufacturing setup deeply relies on large scale connectivity among digital components and automation of most manufacturing steps. These communication links among different machineries, sensors, and electronic devices can be subject to a wide range of cyber attacks. Furthermore, automation and integration to cloud increases the attack surface further.

[0159] 2) Complex Interactions Between IT / OT: Merging of IT and OT systems creates additional complexity and cybersecurity attack surfaces leading to attacks such as data breach, side channels, and unauthorized access.

[0160] 3) Lack of Standards: OT security guidelines such as NIST SP 800-82r3 were recently proposed. However, there is still a severe lack of standard protocols for integrating IT and OT.

[0161] 4) Incompatible Communication Protocols: Industry 4.0 systems arc highly heterogeneous in nature. Hence the communication protocols used by different devices rarely exactly match. This mismatch opens up the system against a wide range of cybersecurity attacks.

[0162] 5) Physical Access: Industry 4.0 systems are easier to physically access compared to systems residing in data centers (e.g., High Performance Computers). This is because, daily human interactions with the OT / IT systems are required to achieve the manufacturing goals. This leads to a higher risk of physical / hardware tampering attacks on these Industry 4.0 digital components. These attacks can lead to cyber threats such as data leakage through side channel analysis, denial-of-service through hardware Trojans, and intellectual property theft through reverse engineering. ii. Improvements Over Related Works

[0163] The present CAST-Map framework was designed to be compatible with a large set of loT and Industry 4.0 systems. Different techniques have been proposed to estimate cybersecurity risk and system criticality for loT systems, however these techniques rely on manually annotating most of the system threat parameters before any computation is possible. Such manual weight / scoring can be time / resource consuming and inaccurate. The present example addresses both of these limitations through the CAST-Map framework by leveraging unsupervised learning techniques. iii. Motivation: Automated Cyber Components Tracking

[0164] To effectively secure an Industry 4.0 facility / system from the cyber perspective, it should be determined what cyber / digital components are present in the system, how they are connected, what software they are running, what underlying electronic hardware are being used in the devices, and how are the devices communicating are important to know. Without accurately knowing this information, a cybersecurity analysis and risk management is not possible. Moreover, an active Industry 4.0 system will dynamically evolve over time with new hardware being connected, software being updated, and old hardware being disconnected.Hence, tracking the cyberspace of an Industry 4.0 facility / system over time as well as performing robust risk analysis / management is important. iv. Motivation: Learning Guided Criticality Weighting

[0165] When electronic sub- systems in a large connected network experiences a negative event (e.g., a malicious software install, denial-of-service), those negative events are oftenconsidered for estimating the future risk / criticality of that sub-system. However, electronic subsystems that did not get affected by these negative events do not suffer any penalty in terms of their risk / criticality scores. This strategy may not be optimal because if: sub-system-A experienced a malicious event; and sub-system-B is similar to sub-system-A in terms of hardware / software and networking; then sub-system-B may also be vulnerable to the same malicious event. An unsupervised learning technique might be able to group sub-systems together that are similar and have their risk / criticality scores weighted depending on the joint malicious event experiences of the whole cluster.3. METHODOLOGY

[0166] In the present example, a framework is described that can help an Industry 4.0 facility / system cybersecurity expert obtain an evolving spatio-temporal map of all the digital components in the systems along with their hardware / software makeup and communication patterns. Also described is an un- supervised learning methodology for more accurately (compared to the present, state-of-the-art systems) calculating cybersecurity risk / criticality scores of an loT system leveraging the spatio-temporal map.

[0167] FIG. 21 shows an illustrative architecture of a CAST-Map framework as described in the present example. Tracking agents collect local data and push them to a central server for aggregation, graph generation, and subsequent threat analysis. Steps within the architecture are described and exemplified herein. i. Live Spatio-Temporal Map Generation Framework

[0168] To initialize a mapping framework, a set of tracking agents are deployed across all Operational Technology (OT), Information Technology (IT), and edge devices. In the present example, a tracking agent is a software package that can periodically collect information about the host system as well as nearby systems in the network. This tracking software is programmed to collect different information about the host system as shown in Table I.Table I: Information gathering procedure used by the CAST-Map framework.

[0169] Moreover, the tracking agent also collects information about nearby systems in the network via tools such as Nmap and gathers information about open ports, services, and operating systems. Collection of individual data types can be turned off to reduce computation and data transfer overheads. All collected information, from each tracking agent, is periodically (using cron, a command-line utility) sent out, end-to-end encrypted, to an aggregation agent located in a central system (e.g., a central tracking system). The aggregation agent in the central system is responsible for generating spatio-temporal map(s) of an entire Industry 4.0 setup and performing cybersecurity analysis. ii. Asset Criticality Estimation

[0170] The cybersecurity risk of a system is formulated as a function of asset criticality scores. The Key Performance Indicator (KPI) weights of each sub-asset inside each sub-system is computed. KPIs include: confidentiality (C), availability (A), integrity (I), reliability (R), authorization (ATH), authenticity (AUT), privacy (p), maintainability (M), conformance (CON), accountability (ACC). For each sub-asset, we assign a score 0 / 1 (i.e., using one hot-encoding) for each of these Key Performance Indicators depending on the respective needs of the said subasset. KPI weight (KPIW) of a sub-asset is the summation of all these indicator values (see Eqn. 1). Here n is the total number of Key Performance Indicators used.

[0171] Using KPIW of each sub-asset we can compute the asset criticality (ACj) of the jth asset as shown in Eqn. 2. Here Nj is the total number sub-assets inside the jth asset and KPIWk is the KPI weight of the kth sub-asset.

[0172] Next the combined asset criticality score (ACC) of an entire sub-system is computed by averaging all asset criticality scores as shown in Eqn. 3. Here P is the total number of assets.

[0173] Next we compute the system level asset criticality score (ACCS) as shown in Eqn.4. Here Q is the total number of sub-systems inside a given system, Sk is the jth sub-system, Class(Sk) indicates the classification of the sub-system Sk inside the spatio-temporal map based on hardware / software similarity, and WeightQ is a pre-computed value indicating an heightened / lowered criticality of a given class of sub-systems.Hi. Determining the Class / Weight of a Sub-System

[0174] The steps necessary to compute the weights of each sub-system (Sk) for Eqn. 4 is provided in Algorithm 1 below:Algorithm 1 : Clustering loT Sub-SystemsInput: [Map]Output: Weights1 Features Empty Dictionary2 Events «— Empty Dictionary3 Class Empty Dictionary4 Weights Empty Dictionaryfor i *— 0 to size(Map) do e Fs ^ encode(Map[i]. Software)7 Fh *— encode(Map[i] .Hardware) Fc ^ encode(Map[i], Connectivity) Fa encode(Map[i]. Physical Access) F = Fs + Fh+ Fc + Fa Features [i] = F Events[i] encode(Map[i]. Events) P PCA(n components) Features Mod P.f it transform(Features) Centroids *— Cluster(Features Mod) for i 0 to size(Map) do Class[findClass(Map[i], Centroids)], append(i) for j 0 to size(Class) do io Weights[j] getWeight(Class[j], Events) Weights. min max normalize() return Weights

[0175] This procedure takes Map as the input which contains all the tracked information about a target loT system. For each sub-system (Sk or Map[i]), software, hardware, connectivity, physical access, and event level details are extracted, followed by the encoding of each of these components into a vector form for subsequent machine learning. The encoding process is as follows:

[0176] Software: The presence and absence of a software can be encoded as a binary vector (Fsi). The vector dimension can be vs for vs number of specific packages being tracked. Moreover we will construct another vector (FS2) of dimension vs such that Fsi[z] = 1 if the package i is outdated (determined based on open database look-up) and Fsi[ / ] = 0 if the package is not outdated. Fsiand FS2 are concatenated together to form Fs.

[0177] Hardware: vh number of enumerated and tracked unique hardware types can be present in a sub-system. Presence and absence of a given hardware component can be indicated by 1 / 0 respectively (e.g., one-hot encoded) inside a vector, Fhi of dimension vh. Moreover, another vector (Fh2) of dimension vh can be constructed such that Fh2[z] = 1 if the hardware i isknown to have a vulnerability (based on open database search) and Fh2[z] = 0 if the hardware is not known to be vulnerable.

[0178] Connectivity: This can be a k dimensional vector Fcsuch that Fc[z] = 1, if this sub-system is directly connected to thesub-system. Otherwise, Fc[z] = 0. k is the total number of sub-systems in the system.

[0179] Physical Access: This is a one-dimensional vector indicating the level of physical access a sub-system has. Fa= 0 if the sub-system is located in a high security facility, Fa= 1 if the sub-system is located in a publicly accessible area. More levels can be defined and added for specific use-cases without affecting the overall framework.

[0180] Events: This is a ve dimensional vector such that Events[i] - X, where X is the number of times the event Events[i\ has occurred over a time lapse T (user defined hyperparameter), ve indicates the total number unique events that are being tracked.

[0181] Once the features are extracted, a principal component analysis (PCA) condenses the information in lesser dimensions for more accurate and fast analysis (lines 13-14). The subsystem are clustered based on their Features (line 15). Inside the dictionary Class, the subsystem IDs are stored against the class they belong to using the calculated cluster centroides (lines 17). A weight value is calculated (using Algorithm 2) for each class based on the events occurred for all the sub-systems belonging to that specific class (line 19).Algorithm 2: Computing Cluster Weight - getWeightQInput: [Class X, Events]Output: Weight XI W ^ O2 for sub system 6 Class X do3 for event G Events[sub system] do4 W <— W + (event.count x event.impact)5 return W

[0182] The weights are scaled to fall between [0, 1] using the standard min max normalization method (line 20). Inside Algorithm 2, for each sub system belonging to a givenclass / cluster (Class X) and for each event associated with that sub system, a product of the number of such events (event. count) and an impact value of that event (event. impact) are calculated. This product is then aggregated inside W to generate the raw weight value for the cluster / class. iv. Other Applications of the loT Tracking Framework[01831 The loT Tracking framework can be used for other purposes besides estimating asset criticality and risk.

[0184] 1) Vulnerable Hardware / Software: An automated program is formulated that can detect potentially vulnerable hardware components (e.g., processors and sensors) and software packages by referring to a vulnerability database.

[0185] 2) Inventory Management: The spatio-temporal map of all the Industry 4.0 cyber systems can also allow us to keep a very accurate inventory. This step is crucial from functional perspective of an Industry 4.0 facility / system. This tracking can allow efficient determination of unused electronic devices, reallocation of cyber resources, and efficient maintenance.4. RESULTSAND DISCUSSIONS

[0186] The effect of different system level variations and clustering parameters on the criticality score (ACCS) calculation will be discussed. i. Experimental Setup

[0187] The ACC value (Eqn. 3) can vary between [0,10] inclusive depending on a given sub-system’s assets and key performance indicators. To keep from being restricted to a very specific local setup, a large number of sub-systems (up to 2000) with different random properties within acceptable ranges are randomly generated. To effectively generate these random systems, a real Industry 4.0 testbed (using CAST-Map) is tracked to determine the nature of hardware / software being used and networking patterns. The random system generator uses this information to populate the software packages, hardware components, and network topology for crafted systems. The ACC value is kept within the predefined threshold mentioned above. Thephysical access is varied between [0, 1] and the frequency of events (New Connection, New USB Device, Privilege Escalation, Malicious Install) is randomly varied between [0, 10], These randomly initialized, large set of sub-systems are then analyzed by the automation framework for computing the ACCS metric. ii. Automated Computation of ACCS Metric

[0188] The automated criticality estimation framework used Principal Component Analysis (n components = 2) to create an array of features which were then clustered using both K-Mean and Affinity Propagation techniques. Both techniques were used with a different criticality index set and number of sub-systems. The criticality index set indicates the range the ACC value with vary between. The ACCS metric values in Table II are an average over 10 runs of randomly generated systems. K-Mean was used with different number of clusters which can be seen in FIG. 22.

[0189] FIG. 22 provides a K-Means visualization when the number of clusters is set to five (top panel) and three (bottom panel). The data is for 2000 randomly generated sub-systems and the computed cluster weights are specified inside the box.

[0190] Here, N indicates the number of sub-systems. The overall trend shows a decreasing ACCS value when using K-Mean for the clustering as the ratio of sub-systems to number of clusters increases. This trend is present because more clusters will lead to smaller localities preventing long-range information flow from negative events associated with the subsystems. So a higher than optimal cluster / sub-system can lead to potential underestimation of the cybersecurity criticality value. Conversely a lower than optimal cluster / sub-system value can lead to an overestimation of the criticality value. The ACCS value increases as the criticality index set changes from [0,10] to [5,10] to [8,10], This is expected because if a system has more critical assets then the overall criticality score will also be high.Table II: ACCS Values for different system settings, clustering techniques, and other parameters.

[0191] Affinity Propagation was used with a different damping value (0.5, 0.7, 0.9). Compared to K-Means, the results obtained using Affinity Propagation are generally more stable across different system sizes (for a given damping value). This is a good indicator that, this algorithm is able to automatically adapt to the nature of the data much better than what K-Mean can do. The ACCS value increases as the criticality index set changes as before. Also, as the damping factor increases, the criticality value shrinks. This is because the amount of messages passing decreases (with higher damping) leading to a higher number of clusters and a potential underestimation of the system-level criticality value.Hi. Additional Research Directions

[0192] The generated spatio-temporal map of an Industry 4.0 setup is a rich source of different types of information. Although, the framework supports many Industry 4.0 devices, there is room to add support for more niche and complex components. Furthermore, this framework can be enhanced to support other Internet-of-Things (loT) setups beyond Industry 4.0. The efficacy of this framework in larger Industry 4.0 and loT testbeds can also be tested. This map can be used to perform post-attack forensics and more precise cybersecurity risk assessments. To support forensics, a tracking agent can be used to collect forensics data such as Event Log, LNK Shell Item, User Account Profiling, NTUSER.DAT, Jump List, Shell Item, USB device history, Windows Error Reporting, Windows Search Database, Remote Desktop Protocol Cache, and Windows registry.5. CONCLUSION

[0193] In this Example, an automated framework (CAST-Map) was presented that can generate and dynamically update a map of all digital / cyber components in an Industry 4.0 setup. Such a spatio-temporal map allows for more robust cybersecurity risk analysis and management enabling wider adoption of Industry 4.0 setups. An unsupervised learning approach is also presented, which clusters and assigns weights to the criticality scores of each sub-systems based on their similarities and associated negative cybersecurity events. The example provides detailed descriptions of the algorithms used to realize this framework and evaluates it using a randomlarge scale loT system generation procedure, which takes inspiration from a real Industry 4.0 testbed. More sophisticated machine learning approaches can also be investigated, allowing.F. Software, Computer System, and Network Environment

[0194] Certain embodiments described herein make use of computer algorithms in the form of software instructions executed by a computer processor. In certain embodiments, the software instructions include a machine learning module, also referred to herein as artificial intelligence or artificial intelligence software. As used herein, a machine learning module refers to a computer implemented process (e.g., a software function) that implements one or more specific machine learning algorithms, such as an artificial neural network (ANN), random forest, decision trees, support vector machines, and the like, in order to determine, for a given input, one or more output values. In certain embodiments, the input comprises alphanumeric data which can include numbers, words, phrases, or lengthier strings, for example. In certain embodiments, the one or more output values comprise values representing numeric values, words, phrases, or other alphanumeric strings. In certain embodiments, the one or more output values comprise an identification of one or more response strings (e.g., selected from a database).

[0195] In certain embodiments, machine learning modules implementing machine learning techniques are trained, for example using datasets that include categories of data described herein. Such training may be used to determine various parameters of machine learning algorithms implemented by a machine learning module, such as weights associated with layers in neural networks. In certain embodiments, once a machine learning module is trained, e.g., to accomplish a specific task such as identifying certain response strings, values of determined parameters are fixed and the (e.g., unchanging, static) machine learning module is used to process new data (e.g., different from the training data; e.g., infer a result) and accomplish its trained task without further updates to its parameters (e.g., the machine learning module does not receive feedback and / or updates). In certain embodiments, machine learning modules may receive feedback, e.g., based on automated review of accuracy or human user review of accuracy, and such feedback may be used as additional training data, to dynamicallyupdate the machine learning module. In certain embodiments, two or more machine learning modules may be combined and implemented as a single module and / or a single software application. In certain embodiments, two or more machine learning modules may also be implemented separately, e.g., as separate software applications. A machine learning module may be software and / or hardware. For example, a machine learning module may be implemented entirely as software, or certain functions of an ANN module may be carried out via specialized hardware (e.g., via an application specific integrated circuit (ASIC), field programmable gate arrays (FPGAs)).

[0196] In certain embodiments, machine learning modules implementing machine learning techniques may be composed of individual nodes (e.g., units, neurons). A node may receive a set of inputs that may include at least a portion of a given input data for the machine learning module and / or at least one output of another node. A node may have at least one parameter to apply and / or a set of instructions to perform (e.g., mathematical functions to execute) over the set of inputs. In certain embodiments, node instructions may include a step to provide various relative importance to the set of inputs using various parameters, such as weights. The weights may be applied by performing scalar multiplication (e.g., or other mathematical function) between a set of inputs values and the parameters, resulting in a set of weighted inputs. In certain embodiments, a node may have a transfer function to combine the set of weighted inputs into one output value. A transfer function may be implemented by a summation of all the weighted inputs and the addition of an offset (e.g., bias) value. In certain embodiments, a node may have an activation function to introduce non-linearity into the output value. Non-limiting examples of the activation function include Rectified Linear Activation (ReLu), logistic (e.g., sigmoid), hyperbolic tangent (tanh), and softmax. In certain embodiments, a node may have a capability of remembering previous states (e.g., recurrent nodes). Previous states may be applied to the input and output values using a set of learning parameters.

[0197] A layer is a building block in a deep learning architecture composed of nodes. In particular, a layer is a set of nodes that receives data input (e.g., weighted or non-weighted input), transforms it (e.g., by carrying out instructions, e.g., applying a set of functions e.g., lineal' and / or non-linear functions), and passes transformed values as output (e.g., to the nextlayer). In certain embodiments, the set of nodes in a particular layer may share the same parameters and instructions without interacting with each other. A machine learning module may be composed of at least one layer (e.g., ordered). Examples of types of layers include convolutional layers (e.g., layers with a kernel, a matrix of parameters that is slid across an input to be multiplied with multiple input values to reduce them to a single output value); fully connected (FC) layers (e.g., all nodes are connected to all outputs of the previous layer); recurrent layers, long / short term memory (LSTM) layers, gated recurrent unit (GRU) layers (e.g., nodes with the various abilities to memorize and apply their previous inputs and / or outputs); batch normalization (BN) layers (e.g., layers that normalize a set of outputs from another layer, allowing for more independent learning of individual layers); activation layer (e.g., layers with nodes that only contain an activation function); (un)pooling layers [e g., layers that reduce (increase) dimensions of an input by summarizing (splitting) input values in defined patches).

[0198] In certain embodiments, the performance of a machine learning module may be characterized by its ability to produce an output data that reproduces an input data with specific accuracy. To achieve specific accuracy, a training process is performed to find optimal parameters, such as weights, for every node in every layer of the machine learning module. In certain embodiments, the training process of a machine learning module may involve using output data to calculate an objective function (e.g., cost function, loss function, error function) that needs to be optimized (e.g., minimized, maximized). For example, a machine learning objective function may be a combination of a loss function and regularization parameter. The loss function is related to how well the output is able to predict the input. The loss function may take various forms, like mean squared error, mean absolute error, binary cross -entropy, categorical cross-entropy, for example. The regularization term may be needed to prevent overfitting and improve generalization of the training process. Typical regularization techniques include LI Regularization or Lasso Regression, L2 Regularization or Ridge Regression, and Dropout (e.g., dropping layer outputs at random during training process).

[0199] In certain embodiments, objective function optimization of a machine learning module may involve finding at least one (e.g., all) of the present global optima (e.g., as opposed to local optima). A typical algorithm for objective function optimization follows principles ofmathematical optimization for a multi-variable function and relies on achieving specific accuracy of the process. Examples of objective function optimization algorithms include gradient descent, nonlinear’ conjugate gradient, random search, Levenberg-Marquardt algorithm, limited-memory Broyden-Fietcher-Goldfarb-Shanno algorithm, pattern search, basin hopping method, Krylov method, Adam method, genetic algorithm, particle swarm optimization, surrogate optimization, and simulated annealing.

[0200] In certain embodiments, available input data includes training data and validation data, e.g., where the validation data is separate and non-overlapping with the training data. Training data is used during the training process to optimize a model, whereas validation data is used to check the accuracy of the model while operating on previously unseen data. In certain embodiments, training data is divided into batches (e.g., portions) that is sequentially used (e.g., in random order) as sets of inputs to train a model. In certain embodiments, a model is trained multiple times (e.g., epochs) on the entire set of training data.

[0201] As shown in FIG. 23, an implementation of a network environment 2300 for use in providing systems, methods, and architectures as described herein is shown and described. In brief overview, referring now to FIG. 23, a block diagram of an exemplary cloud computing environment 2300 is shown and described. The cloud computing environment 2300 may include one or more resource providers 2302a, 2302b, 2302c (collectively, 2302). Each resource provider 2302 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, exemplary computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 2302 may be connected to any other resource provider 2302 in the cloud computing environment 2300. In some implementations, the resource providers 2302 may be connected over a computer network 2308. Each resource provider 2302 may be connected to one or more computing device 2304a, 2304b, 2304c (collectively, 2304), over the computer network 2308.[02021 The cloud computing environment 2300 may include a resource manager 2306. The resource manager 2306 may be connected to the resource providers 2302 and the computing devices 2304 over the computer network 2308. In some implementations, the resource manager 2306 may facilitate the provision of computing resources by one or more resource providers 2302 to one or more computing devices 2304. The resource manager 2306 may receive a request for a computing resource from a particular computing device 2304. The resource manager 2306 may identify one or more resource providers 2302 capable of providing the computing resource requested by the computing device 2304. The resource manager 2306 may select a resource provider 2302 to provide the computing resource. The resource manager 2306 may facilitate a connection between the resource provider 2302 and a particular computing device 2304. In some implementations, the resource manager 2306 may establish a connection between a particular resource provider 2302 and a particular computing device 2304. In some implementations, the resource manager 2306 may redirect a particular computing device 2304 to a particular resource provider 2302 with the requested computing resource.

[0203] FIG. 24 shows an example of a computing device 2400 and a mobile computing device 2450 that can be used to implement the techniques described in this disclosure. The computing device 2400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The mobile computing device 2450 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to be limiting.

[0204] The computing device 2400 includes a processor 2402, a memory 2404, a storage device 2406, a high-speed interface 2408 connecting to the memory 2404 and multiple highspeed expansion ports 2410, and a low-speed interface 2412 connecting to a low-speed expansion port 2414 and the storage device 2406. Each of the processor 2402, the memory 2404, the storage device 2406, the high-speed interface 2408, the high-speed expansion ports 2410, and the low-speed interface 2412, are interconnected using various busses, and may bemounted on a common motherboard or in other manners as appropriate. The processor 2402 can process instructions for execution within the computing device 2400, including instructions stored in the memory 2404 or on the storage device 2406 to display graphical information for a GUI on an external input / output device, such as a display 2416 coupled to the high-speed interface 2408. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system). Thus, as the term is used herein, where a plurality of functions are described as being performed by “a processor”, this encompasses embodiments wherein the plurality of functions are performed by any number of processors (one or more) of any number of computing devices (one or more). Furthermore, where a function is described as being performed by “a processor”, this encompasses embodiments wherein the function is performed by any number of processors (one or more) of any number of computing devices (one or more) (e.g., in a distributed computing system).

[0205] The memory 2404 stores information within the computing device 2400. In some implementations, the memory 2404 is a volatile memory unit or units. In some implementations, the memory 2404 is a non-volatile memory unit or units. The memory 2404 may also be another form of computer-readable medium, such as a magnetic or optical disk.

[0206] The storage device 2406 is capable of providing mass storage for the computing device 2400. In some implementations, the storage device 2406 may be or contain a computer- readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. Instructions can be stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 2402), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices such as computer- or machine- readable mediums (for example, the memory 2404, the storage device 2406, or memory on the processor 2402).[02071 The high-speed interface 2408 manages bandwidth-intensive operations for the computing device 2400, while the low- speed interface 2412 manages lower bandwidth-intensive operations. Such allocation of functions is an example only. In some implementations, the highspeed interface 2408 is coupled to the memory 2404, the display 2416 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 2410, which may accept various expansion cards (not shown). In the implementation, the low- speed interface 2412 is coupled to the storage device 2406 and the low-speed expansion port 2414. The low-speed expansion port 2414, which may include various communication ports (e.g., USB, Bluetooth®, Ethernet, wireless Ethernet) may be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0208] The computing device 2400 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 2420, or multiple times in a group of such servers. In addition, it may be implemented in a personal computer such as a laptop computer 2422. It may also be implemented as part of a rack server system 2424. Alternatively, components from the computing device 2400 may be combined with other components in a mobile device (not shown), such as a mobile computing device 2450. Each of such devices may contain one or more of the computing device 2400 and the mobile computing device 2450, and an entire system may be made up of multiple computing devices communicating with each other.

[0209] The mobile computing device 2450 includes a processor 2452, a memory 2464, an input / output device such as a display 2454, a communication interface 2466, and a transceiver 2468, among other components. The mobile computing device 2450 may also be provided with a storage device, such as a micro-drive or other device, to provide additional storage. Each of the processor 2452, the memory 2464, the display 2454, the communication interface 2466, and the transceiver 2468, are interconnected using various buses, and several of the components may be mounted on a common motherboard or in other manners as appropriate.

[0210] The processor 2452 can execute instructions within the mobile computing device 2450, including instructions stored in the memory 2464. The processor 2452 may beimplemented as a chipset of chips that include separate and multiple analog and digital processors. The processor 2452 may provide, for example, for coordination of the other components of the mobile computing device 2450, such as control of user interfaces, applications run by the mobile computing device 2450, and wireless communication by the mobile computing device 2450.

[0211] The processor 2452 may communicate with a user through a control interface 2458 and a display interface 2456 coupled to the display 2454. The display 2454 may be, for example, a TFT (Thin-Film-Transistor Liquid Crystal Display) display or an OLED (Organic Light Emitting Diode) display, or other appropriate display technology. The display interface 2456 may comprise appropriate circuitry for driving the display 2454 to present graphical and other information to a user. The control interface 2458 may receive commands from a user and convert them for submission to the processor 2452. In addition, an external interface 2462 may provide communication with the processor 2452, so as to enable near area communication of the mobile computing device 2450 with other devices. The external interface 2462 may provide, for example, for wired communication in some implementations, or for wireless communication in other implementations, and multiple interfaces may also be used.

[0212] The memory 2464 stores information within the mobile computing device 2450. The memory 2464 can be implemented as one or more of a computer-readable medium or media, a volatile memory unit or units, or a non-volatile memory unit or units. An expansion memory 2474 may also be provided and connected to the mobile computing device 2450 through an expansion interface 2472, which may include, for example, a SIMM (Single In Line Memory Module) card interface. The expansion memory 2474 may provide extra storage space for the mobile computing device 2450, or may also store applications or other information for the mobile computing device 2450. Specifically, the expansion memory 2474 may include instructions to carry out or supplement the processes described above, and may include secure information also. Thus, for example, the expansion memory 2474 may be provide as a security module for the mobile computing device 2450, and may be programmed with instructions that permit secure use of the mobile computing device 2450. In addition, secure applications may beprovided via the SIMM cards, along with additional information, such as placing identifying information on the SIMM card in a non-hackable manner.

[0213] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random-access memory), as discussed below. In some implementations, instructions are stored in an information carrier. The instructions, when executed by one or more processing devices (for example, processor 2452), perform one or more methods, such as those described above. The instructions can also be stored by one or more storage devices, such as one or more computer- or machine-readable mediums (for example, the memory 2464, the expansion memory 2474, or memory on the processor 2452). In some implementations, the instructions can be received in a propagated signal, for example, over the transceiver 2468 or the external interface 2462.

[0214] The mobile computing device 2450 may communicate wirelessly through the communication interface 2466, which may include digital signal processing circuitry where necessary. The communication interface 2466 may provide for communications under various modes or protocols, such as GSM voice calls (Global System for Mobile communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS messaging (Multimedia Messaging Service), CDMA (code division multiple access), TDMA (time division multiple access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may occur, for example, through the transceiver 2468 using a radio-frequency. In addition, short-range communication may occur, such as using a Bluetooth®, Wi-Fi™, or other such transceiver (not shown). In addition, a GPS (Global Positioning System) receiver module 2470 may provide additional navigation- and location-related wireless data to the mobile computing device 2450, which may be used as appropriate by applications running on the mobile computing device 2450.

[0215] The mobile computing device 2450 may also communicate audibly using an audio codec 2460, which may receive spoken information from a user and convert it to usable digital information. The audio codec 2460 may likewise generate audible sound for a user, such as through a speaker, e.g., in a handset of the mobile computing device 2450. Such sound mayinclude sound from voice telephone calls, may include recorded sound (e.g., voice messages, music files, etc.) and may also include sound generated by applications operating on the mobile computing device 2450.

[0216] The mobile computing device 2450 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a cellular telephone 2480. It may also be implemented as part of a smart-phone 2482, personal digital assistant, or other similar mobile device.

[0217] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0218] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer- readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0219] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to thecomputer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0220] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0221] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0222] In some implementations, certain modules described herein can be separated, combined or incorporated into single or combined modules. Any modules depicted in the figures are not intended to limit the systems described herein to the software architectures shown therein.ALTERNATIVE EMBODIMENTS

[0223] Embodiment 1. A method of penetration testing on a digital replica of a system (e.g., an Industry 4.0 system, an loT system), the method comprising: accessing a first system state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; updating the map of the digital replica using the second system state; and identifying one or morevulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

[0224] Embodiment 2. The method of embodiment 1, wherein the method comprises identifying a solution set for mitigating the one or more identified vulnerabilities of the digital replica.

[0225] Embodiment 3. The method of embodiment 2, wherein the method comprises applying the solution set to the system.

[0226] Embodiment 4. The method of any one of embodiments 1-3, wherein the method comprises mitigating the one or more vulnerabilities, wherein mitigating the one or more vulnerabilities comprises isolating (e.g., temporarily isolating) one or more devices of the system and / or applying an update (e.g., a firmware update, a software update) to a device of the system.

[0227] Embodiment 5. The method of any one of embodiments 1-4, wherein the method comprises determining an action pathway fitness (e.g., an action pathway fitness score).

[0228] Embodiment 6. The method of embodiment 5, wherein the action pathway fitness is determined based on one or more of: discovering the one or more vulnerabilities; severity of the one or more vulnerabilities; discovering one or more new devices on a network of the digital replica; exploiting, by the penetration test model, the one or more vulnerabilities; disclosing of credentials (e.g., user credentials, administrator credentials); and inserting an exploit (e.g., a back door).

[0229] Embodiment 7. The method of any one of embodiments 1-6, wherein the method comprises obtaining feedback of a user based on one or more of: the first system state, the second system state, and the determined action pathway.

[0230] Embodiment 8. The method of embodiment 1-7, wherein the method comprises updating the action pathway based on feedback of a user (e.g., a human expert) (e.g., by merging the feedback of the user and the action pathway).

[0231] Embodiment 9. The method of any one of embodiments 5-8, comprising creating (or updating) a feedback dictionary based on the first system state, the second system state, theaction pathway, and the action pathway fitness (e.g., for later use in training / re-training the penetration test model).

[0232] Embodiment 10. The method of embodiment 9, wherein the method comprises retraining the penetration test model based on the feedback dictionary.

[0233] Embodiment 11. The method of any one of embodiments 1-10, wherein the method comprises determining a defense pathway based on applying a defender model to the first system state.

[0234] Embodiment 12. The method of any one of embodiments 1-11, wherein the method comprises receiving one or more maps of the system and creating the digital replica based on the one or more maps of the system.

[0235] Embodiment 13. The method of any one of embodiments 1-12, wherein the digital replica comprises a plurality of networked components mimicking devices of the system, wherein the plurality of network components comprise one or more of hardware emulated components, software emulated components, and software simulated components.

[0236] Embodiment 14. The method of any one of embodiments 1-13, wherein the first system state is a vectorized version of the digital replica.

[0237] Embodiment 15. The method of embodiment 14, wherein the method comprises vectorizing the digital replica using one-hot encoding, linear encoding, and / or neural network encoding.

[0238] Embodiment 16. The method of any one of embodiments 1-15, wherein determining the action pathway comprises: identifying one or more action options based on the penetration test model and the first system state of the digital replica; creating one or more action pathway options based on the action options; scoring each of the one or more action pathway options; and selecting, based on the score, one of the one or more action pathway options to become the action pathway.

[0239] Embodiment 17. A method of penetration testing of on a digital replica of a system (e.g., an Industry 4.0 system, an loT system), the method comprising: receiving a firstsystem state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; determining a defense pathway based on applying a defense model to the second system state the digital replica; executing the defense pathway on the second system state to produce a third system state; updating the map of the digital replica using the third system state; and identifying one or more vulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

[0240] Embodiment 18. The method of embodiment 17, wherein the method comprises identifying a solution set (e.g., based on the updated system map) based on, at least, the one or more identified vulnerabilities of the digital replica.

[0241] Embodiment 19. The method of embodiment 18, wherein the method comprises applying the solution set to the digital replica.

[0242] Embodiment 20. The method of embodiments 17-19, wherein the method comprises mitigating one or more vulnerabilities of the system based on, at least, the identified one or more vulnerabilities of the digital replica.

[0243] Embodiment 21. The method of any one of embodiments 17-20, wherein the method comprises determining an action pathway fitness (e.g., an action pathway fitness score).

[0244] Embodiment 22. The method of any one of embodiments 17-21, wherein the method comprises obtaining feedback of a user based on one or more of: the first system state, the second system state, the third system state, the action pathway, and the defense pathway.

[0245] Embodiment 23. The method of embodiment 22, wherein the method comprises updating the action pathway fitness and feedback of a user (e.g., by merging).

[0246] Embodiment 24. The method of any one of embodiments 17-23, comprising creating (or updating) a feedback dictionary based on the first system state, the second system state, the third system state, the action pathway, and the action pathway fitness (e.g., for later use in training / retraining the penetration test model).[02471 Embodiment 25. The method of embodiment 24, comprising retraining the penetration test model and / or the defender model based on the feedback dictionary.

[0248] Embodiment 26. The method of any one of embodiments 17-25, wherein the defense pathway comprises one or more of: scanning a network of the digital replica for an anomaly; modifying a topology (e.g., connections of) of the map of the digital replica; and disconnecting one or more hosts (e.g., devices) from the network of the digital replica.

[0249] Embodiment 27. A tracking system on a networked system (e.g., an loT system, an Industry 4.0 system), the tracking system comprising: a plurality of tracking agents, wherein each of the plurality of tracking agents are communicatively connected to at least one or more components on the networked system and configured to transmit (e.g., periodically transmit) data corresponding to the at least one or more components; and a central database, wherein the central database is communicatively connected to the plurality of tracking agents and configured to receive (e.g., via an encrypted means) (e.g., directly, indirectly) (e.g., periodically receive) data from the plurality of tracking agents.

[0250] Embodiment 28. The tracking system of embodiment 27, wherein at least one of the plurality of tracking agents is a software-based tracking agent installed on a component of the networked system.

[0251] Embodiment 29. The tracking system of embodiments 27 or 28, wherein at least one of the plurality of tracking agents is a physical device communicatively connected to the networked system.

[0252] Embodiment 30. The tracking system of any one of embodiments 27-29, wherein the tracking agent is configured to extract data from a processor of a component.

[0253] Embodiment 31 . The tracking system of embodiment 30, wherein the processor of the component comprises a hardware performance counter.

[0254] Embodiment 32. The tracking system of embodiments 27-31, wherein at least one of the plurality of tracking agents is configured to extract data from one or more components neighboring (e.g., connected to) the tracking agent.[02551 Embodiment 33. The tracking system of embodiment 32, wherein the data from the one or more components neighboring the tracking agent comprises one or more open ports of the neighboring component(s), network traffic, operating system(s) of the neighboring component(s), and / or one or more services running on the networked system.

[0256] Embodiment 34. The tracking system of embodiments 27-33, wherein the central database is or comprises an aggregation agent for generating a map (e.g., in real time, a spatiotemporal map) of the networked system.

[0257] Embodiment 35. The tracking system of embodiment 34, wherein the central database is configured to create and / or update a digital twin of the networked system based on the generated map.

[0258] Embodiment 36. The tracking system of embodiments 34 or 35, wherein the aggregation agent is configured to perform an assessment based on a digital twin created from the data from the tracking agents to identify one or more vulnerabilities, anomalies, and / or policy violations.

[0259] Embodiment 37. A method of using the tracking system of any one of embodiments 27-36 comprising: pinging, by the tracking agent, at least one of a networked device, server, and computer to establish a communication connection therewith, executing, by the tracking agent, at least one information retrieval routine to retrieve at least one piece of information from the networked device, server, or computer; and storing, in the central database, the at least one piece of information.

[0260] Embodiment 38. The method of embodiment 37, wherein the at least one piece of information comprises at least one of an IP address, a device name, device current network interface configuration information, operating system information, USB device identity, USB bus information, USB usage history, one or more kernel I buffer messages, names and / or types of software running on a system, a list of upgradable software running on a system, usage history for a program or application, inter-network activity information, and hardware information.

[0261] Embodiment 39. The method of embodiment 37, further comprising taking at least one mitigating action as a result of the at least one piece of information retrieved from thetracking device, the at least one mitigating action comprising at least one of (1) reassigning an IP address of the networked device, server, or computer to a quarantined IP address, and (2) renaming (i.e., changing the host name of) the networked device, server, or computer.

[0262] Embodiment 40. A method of identifying attacks (e.g., a cyber attack) on a system, the method comprising: creating a virtual replica (digital twin) of the system (e.g., an loT system) and generating a normal state (e.g., not attacked) profde of the virtual replica; simulating an attack on the virtual replica using an artificial intelligence framework (e.g., a generative artificial intelligence framework); and creating a model for detecting a deviation of (one or more aspects of) a test profile of the virtual replica from the normal state profile using the simulated attack.

[0263] Embodiment 41. The method of embodiment 40, wherein the method further comprises: receiving one or more test profiles of the virtual replica; and detecting an attack on the system by applying the deviation to the one or more test profiles of the virtual replica.

[0264] Embodiment 42. A method for generating solutions to system vulnerabilities, the method comprising: generating a profile (e.g., a state) of the system at a time point; (e.g., automatically) creating a virtual replica of the system at the time point; carrying out a penetration test on the virtual replica (e.g., using an Al framework); identifying one or more vulnerabilities of the virtual replica based on the penetration test; and generating a first set of solutions corresponding to (e.g., based on) the one or more vulnerabilities of the virtual replica.

[0265] Embodiment 43. The method of embodiment 42, wherein the method comprises implementing at least one mitigating action on the system based on the first set of solutions (e.g., to enhance the robustness of the system) (e.g., to address one or more vulnerabilities of the system).

[0266] Embodiment 44. The method of embodiment 42 or 43, wherein the method comprises applying the first set of solutions to the virtual replica and updating the virtual replica.

[0267] Embodiment 45. The method of embodiments 42-44, wherein the method comprises: carrying out a second vulnerability test on the virtual replica (e.g., using a generative Al framework); identifying one or more vulnerabilities based on the second vulnerability test;generating a second set of solutions corresponding to (e.g., based on) the second vulnerability test; and updating the virtual replica based on the one or more vulnerabilities based on the second vulnerability test.

[0268] Embodiment 46. The method of embodiments 42-45, wherein the method comprises: applying a penetration test (e.g., an Al guided penetration test) to the virtual replica; and determining the virtual replica meets a threshold for robustness (e.g., by preventing and / or resisting attacks).

[0269] Embodiment 47. The method of any one of embodiments 42-46, wherein the method comprises carrying out the vulnerability test on the virtual replica (e.g., the first attack, the second attack) using a generative Al method (e.g., a generative Al (GAI) framework).

[0270] Embodiment 48. The method of embodiment 47, wherein the generative Al method comprises: applying a penetration test model to a system state corresponding to the profde of the system; identifying one or more action options (e.g., one or more attacks) based on the system state and the penetration test model; creating one or more action pathways based on the one or more action options; scoring the one or more action pathways (e g., based on a numerical scale); selecting at least one of the one or more action pathways for the vulnerability test based on, at least, the score(s) of the one or more action pathways.

[0271] Embodiment 49. The method of embodiment 48, wherein the method comprises training (or re-training) the penetration test model comprising training the penetration test model based on a feedback dictionary.

[0272] Embodiment 50. The method of embodiment 49, wherein the feedback dictionary comprises one or more of: the selected action pathway, the profile of the system at the time point, and a fitness of the penetration test model.

[0273] Embodiment 51. A method of creating a digital twin of a system (e.g., an loT system, an Industry 4.0 system), the method comprising: receiving a representation of the system (e.g., based on tracking data from a tracking system (e.g., from a plurality of tracking agents)); identifying one or more portions of the representation for hardware emulation; (e.g., automatically) mapping one or more hardware elements (e.g., an loT device) to correspond toeach of the one or more portions of the representation for hardware emulation; identifying one or more portions of the representation for software simulation and / or emulation; (e.g., automatically) mapping one or more virtual systems to correspond to each of the one or more portions of the representation for software simulation and / or emulation; networking the one or more virtual systems and the one or more hardware elements based on the representation of the system; and initiating network traffic on the network.

[0274] Embodiment 52. The method of embodiment 51, wherein the method comprises receiving tracking data from a tracking system (e.g., as described in embodiments 27-36) and creating the representation of the system based on the tracking data.

[0275] Embodiment 53. The method of embodiment 51 or 52, wherein the method comprises updating the representation of the system using tracking data from a tracking system.

[0276] Embodiment 54. The method of embodiments 51-53, wherein the method comprises for receiving and / or updating the representation of the system using a controller.

[0277] Embodiment 55. The method of embodiments 51-54, wherein the method comprises visualizing one or more aspects of the digital twin (e.g., network traffic, a map of the digital twin) using a controller.

[0278] Embodiment 56. The method of embodiments 51-55, wherein the method comprises carrying out a penetration test on the digital twin (e.g., as described in embodiments 1- 26).

[0279] Embodiment 57. An Al-guided cyber vulnerability mitigation method comprising: creating a live map network framework; creating a snapshot of the network framework at a first time point; creating a virtual replica of the network framework; using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on the virtual replica; identifying vulnerabilities of the virtual replica; and generating one or more solution sets to address the identified vulnerabilities.

[0280] Embodiment 58. The method of embodiment 57, further comprising applying the one or more solution sets to the virtual replica.[02811 Embodiment 59. The method of embodiment 58, further comprising determining if the virtual replica passes one or more robustness tests.

[0282] Embodiment 60. The method of embodiment 59, further comprising: if the virtual replica does not pass the one or more robustness tests, using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on an updated version of the virtual replica; and if the virtual replica does pass the one or more robustness tests, addressing real system vulnerabilities based on knowledge and / or insights obtained from identifying vulnerabilities of the virtual replica.

[0283] Embodiment 61. A method of generating a live device tracking map, the method comprising: installing a tracking system on one or more loT devices, the tracking system configured to capture one or more attributes of the one or more loT devices and transmit the one or more attributes to a central tracking system; communicatively coupling the one or more loT devices to a network housing the central tracking system, and periodically transmitting, by the tracking system, the one or more attributes to the central tracking system based on a predetermined user-defined update frequency.

[0284] Embodiment 62. A multi-network generative adversarial network (GAN) comprising: a first neural network trained for threat identification; and a second neural network trained for vulnerability exploitation, wherein each of the first neural network and the second neural network are trained via unsupervised procedural generation; and wherein the unsupervised procedural generation used for training of each of the first neural network and the second neural network comprises at least one mitigation activity to distinguish between normal and anomalous behavior.

[0285] Embodiment 63. The multi-network generative adversarial network (GAN) of embodiment 62, wherein the first neural network and the second neural network are mutually trainable.EQUIVALENTS

[0286] Those skilled in the ail will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the disclosure described herein. Therefore, the scope of the present disclosure is not intended to be limited to the above Description.

[0287] Elements of different implementations described herein may be combined to form other implementations not specifically set forth above. Elements may be left out of the processes, computer programs, databases, etc. described herein without adversely affecting their operation. In addition, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. Various separate elements may be combined into one or more individual elements to perform the functions described herein.

[0288] Throughout the description, where apparatus and systems are described as having, including, or comprising specific components, or where processes and methods are described as having, including, or comprising specific steps, it is contemplated that, additionally, there are apparatus, and systems of the present invention that consist essentially of, or consist of, the recited components, and that there are processes and methods according to the present invention that consist essentially of, or consist of, the recited processing steps.

[0289] It should be understood that the order of steps or order for performing certain action is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.

[0290] While the invention has been particularly shown and described with reference to specific preferred embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims.

Claims

CLAIMSWe claim:

1. A method of penetration testing on a digital replica of a system, the method comprising: accessing a first system state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; updating the map of the digital replica using the second system state; and identifying one or more vulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

2. The method of claim 1, wherein the method comprises identifying a solution set for mitigating the one or more identified vulnerabilities of the digital replica.

3. The method of claim 2, wherein the method comprises applying the solution set to the system.

4. The method of claim 1, wherein the method comprises mitigating the one or more vulnerabilities, wherein mitigating the one or more vulnerabilities comprises isolating one or more devices of the system and / or applying an update to a device of the system.

5. The method of claim 3, wherein the method comprises determining an action pathway fitness.

6. The method of claim 5, wherein the action pathway fitness is determined based on one or more of: discovering the one or more vulnerabilities; severity of the one or more vulnerabilities; discovering one or more new devices on a network of the digital replica; exploiting, by thepenetration test model, the one or more vulnerabilities; disclosing of credentials; and inserting an exploit.

7. The method of claim 5, wherein the method comprises obtaining feedback of a user based on one or more of: the first system state, the second system state, and the determined action pathway.

8. The method of claim 7, wherein the method comprises updating the action pathway based on feedback of a user.

9. The method of claim 7, comprising creating (or updating) a feedback dictionary based on the first system state, the second system state, the action pathway, and the action pathway fitness.

10. The method of claim 9, wherein the method comprises retraining the penetration test model based on the feedback dictionary.

11. The method of claim 1, wherein the method comprises determining a defense pathway based on applying a defender model to the first system state.

12. The method of claim 1, wherein the method comprises receiving one or more maps of the system and creating the digital replica based on the one or more maps of the system.

13. The method of claim 12, wherein the digital replica comprises a plurality of networked components mimicking devices of the system, wherein the plurality of network components comprise one or more of hardware emulated components, software emulated components, and software simulated components.

14. The method of claim 1, wherein the first system state is a vectorized version of the digital replica.

15. The method of claim 14, wherein the method comprises vectorizing the digital replica using one-hot encoding, linear encoding, and / or neural network encoding.

16. The method of claim 5, wherein determining the action pathway comprises: identifying one or more action options based on the penetration test model and the first system state of the digital replica; creating one or more action pathway options based on the action options; scoring each of the one or more action pathway options; and selecting, based on the score, one of the one or more action pathway options to become the action pathway.

17. A method of penetration testing of on a digital replica of a system, the method comprising: receiving a first system state of the digital replica corresponding to a map of the digital replica; determining an action pathway based on applying a penetration test model to the first system state of the digital replica; executing the action pathway on the first system state to produce a second system state; determining a defense pathway based on applying a defense model to the second system state the digital replica; executing the defense pathway on the second system state to produce a third system state; updating the map of the digital replica using the third system state; and identifying one or more vulnerabilities of the digital replica of the system based on, at least, the updated map of the digital replica.

18. The method of claim 17, wherein the method comprises identifying a solution set based on, at least, the one or more identified vulnerabilities of the digital replica.

19. The method of claim 18, wherein the method comprises applying the solution set to the digital replica.

20. The method of claim 17, wherein the method comprises mitigating one or more vulnerabilities of the system based on, at least, the identified one or more vulnerabilities of the digital replica.

21. The method of claim 20, wherein the method comprises determining an action pathway fitness.

22. The method of claim 21, wherein the method comprises obtaining feedback of a user based on one or more of: the first system state, the second system state, the third system state, the action pathway, and the defense pathway.

23. The method of claim 17, wherein the method comprises updating the action pathway fitness and feedback of a user.

24. The method of claim 17, comprising creating (or updating) a feedback dictionary based on the first system state, the second system state, the third system state, the action pathway, and the action pathway fitness.

25. The method of claim 24, comprising retraining the penetration test model and / or the defender model based on the feedback dictionary.

26. The method of claim 17, wherein the defense pathway comprises one or more of: scanning a network of the digital replica for an anomaly; modifying a topology of the map of the digital replica; and disconnecting one or more hosts from the network of the digital replica.

27. A tracking system on a networked system, the tracking system comprising: a plurality of tracking agents, wherein each of the plurality of tracking agents are communicatively connected to at least one or more components on the networked system and configured to transmit data corresponding to the at least one or more components; anda central database, wherein the central database is communicatively connected to the plurality of tracking agents and configured to receive data from the plurality of tracking agents.

28. The tracking system of claim 27, wherein at least one of the plurality of tracking agents is a software-based tracking agent installed on a component of the networked system.

29. The tracking system of claim 27, wherein at least one of the plurality of tracking agents is a physical device communicatively connected to the networked system.

30. The tracking system of claim 27, wherein the tracking agent is configured to extract data from a processor of a component.

31. The tracking system of claim 30, wherein the processor of the component comprises a hardware performance counter.

32. The tracking system of claim 27, wherein at least one of the plurality of tracking agents is configured to extract data from one or more components neighboring the tracking agent.

33. The tracking system of claim 32, wherein the data from the one or more components neighboring the tracking agent comprises one or more open ports of the neighboring component(s), network traffic, operating system(s) of the neighboring component(s), and / or one or more services running on the networked system.

34. The tracking system of claim 27, wherein the central database is or comprises an aggregation agent for generating a map of the networked system.

35. The tracking system of claim 34, wherein the central database is configured to create and / or update a digital twin of the networked system based on the generated map.

36. The tracking system of claim 34, wherein the aggregation agent is configured to perform an assessment based on a digital twin created from the data from the tracking agents to identify one or more vulnerabilities, anomalies, and / or policy violations.

37. A method of using the tracking system of claim 27 comprising: pinging, by the tracking agent, at least one of a networked device, server, and computer to establish a communication connection therewith, executing, by the tracking agent, at least one information retrieval routine to retrieve at least one piece of information from the networked device, server, or computer; and storing, in the central database, the at least one piece of information.

38. The method of claim 37, wherein the at least one piece of information comprises at least one of an IP address, a device name, device current network interface configuration information, operating system information, USB device identity, USB bus information, USB usage history, one or more kernel I buffer messages, names and / or types of software running on a system, a list of upgradable software running on a system, usage history for a program or application, internetwork activity information, and hardware information.

39. The method of claim 37, further comprising taking at least one mitigating action as a result of the at least one piece of information retrieved from the tracking device, the at least one mitigating action comprising at least one of (1) reassigning an IP address of the networked device, server, or computer to a quarantined IP address, and (2) renaming the networked device, server, or computer.

40. A method of identifying attacks on a system, the method comprising: creating a virtual replica (digital twin) of the system and generating a normal state profile of the virtual replica; simulating an attack on the virtual replica using an artificial intelligence framework; andcreating a model for detecting a deviation of (one or more aspects of) a test profile of the virtual replica from the normal state profile using the simulated attack.

41. The method of claim 40, wherein the method further comprises: receiving one or more test profiles of the virtual replica; and detecting an attack on the system by applying the deviation to the one or more test profiles of the virtual replica.

42. A method for generating solutions to system vulnerabilities, the method comprising: generating a profile of the system at a time point; creating a virtual replica of the system at the time point; carrying out a penetration test on the virtual replica; identifying one or more vulnerabilities of the virtual replica based on the penetration test; and generating a first set of solutions corresponding to the one or more vulnerabilities of the virtual replica.

43. The method of claim 42, wherein the method comprises implementing at least one mitigating action on the system based on the first set of solutions.

44. The method of claim 42, wherein the method comprises applying the first set of solutions to the virtual replica and updating the virtual replica.

45. The method of claim 42, wherein the method comprises: carrying out a second vulnerability test on the virtual replica; identifying one or more vulnerabilities based on the second vulnerability test; generating a second set of solutions corresponding to the second vulnerability test; and updating the virtual replica based on the one or more vulnerabilities based on the second vulnerability test.

46. The method of claim 42, wherein the method comprises: applying a penetration test to the virtual replica; and determining the virtual replica meets a threshold for robustness.

47. The method of claim 42, wherein the method comprises carrying out the vulnerability test on the virtual replica using a generative Al method.

48. The method of claim 47, wherein the generative Al method comprises: applying a penetration test model to a system state corresponding to the profile of the system; identifying one or more action options based on the system state and the penetration test model; creating one or more action pathways based on the one or more action options; scoring the one or more action pathways; selecting at least one of the one or more action pathways for the vulnerability test based on, at least, the score(s) of the one or more action pathways.

49. The method of claim 48, wherein the method comprises training (or re-training) the penetration test model comprising training the penetration test model based on a feedback dictionary.

50. The method of claim 49, wherein the feedback dictionary comprises one or more of: the selected action pathway, the profile of the system at the time point, and a fitness of the penetration test model.

51. A method of creating a digital twin of a system, the method comprising: receiving a representation of the system; identifying one or more portions of the representation for hardware emulation;mapping one or more hardware elements to correspond to each of the one or more portions of the representation for hardware emulation; identifying one or more portions of the representation for software simulation and / or emulation; mapping one or more virtual systems to correspond to each of the one or more portions of the representation for software simulation and / or emulation; networking the one or more virtual systems and the one or more hardware elements based on the representation of the system; and initiating network traffic on the network.

52. The method of claim 51, wherein the method comprises receiving tracking data from a tracking system and creating the representation of the system based on the tracking data.

53. The method of claim 52, wherein the method comprises updating the representation of the system using tracking data from a tracking system.

54. The method of claims 53, wherein the method comprises for receiving and / or updating the representation of the system using a controller.

55. The method of claim 54, wherein the method comprises visualizing one or more aspects of the digital twin using a controller.

56. The method of claim 55, wherein the method comprises carrying out a penetration test on the digital twin.

57. An Al-guided cyber vulnerability mitigation method comprising: creating a live map network framework; creating a snapshot of the network framework at a first time point; creating a virtual replica of the network framework;using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on the virtual replica; identifying vulnerabilities of the virtual replica; and generating one or more solution sets to address the identified vulnerabilities.

58. The method of claim 57, further comprising applying the one or more solution sets to the virtual replica.

59. The method of claim 58, further comprising determining if the virtual replica passes one or more robustness tests.

60. The method of claim 59, further comprising: if the virtual replica does not pass the one or more robustness tests, using a generative artificial intelligence (GAI) framework to automatically carry out virtual vulnerability tests on an updated version of the virtual replica; and if the virtual replica does pass the one or more robustness tests, addressing real system vulnerabilities based on knowledge and / or insights obtained from identifying vulnerabilities of the virtual replica.

61. A method of generating a live device tracking map, the method comprising: installing a tracking system on one or more loT devices, the tracking system configured to capture one or more attributes of the one or more loT devices and transmit the one or more attributes to a central tracking system; communicatively coupling the one or more loT devices to a network housing the central tracking system, and periodically transmitting, by the tracking system, the one or more attributes to the central tracking system based on a predetermined user-defined update frequency.

62. A multi-network generative adversarial network (GAN) comprising: a first neural network trained for threat identification; anda second neural network trained for vulnerability exploitation, wherein each of the first neural network and the second neural network are trained via unsupervised procedural generation; and wherein the unsupervised procedural generation used for training of each of the first neural network and the second neural network comprises at least one mitigation activity to distinguish between normal and anomalous behavior.

63. The multi-network generative adversarial network (GAN) of claim 62, wherein the first neural network and the second neural network are mutually trainable.