Method for predicting the health state of a distributed network by means of an artificial neural network
By identifying and evaluating sites and assets in a distributed network using artificial neural networks, and combining a feedforward network trained by backpropagation with an aging frequency, the problem of false conclusions caused by static analysis is solved, enabling dynamic prediction and risk assessment of the health status of a distributed network.
Patent Information
- Application Number
- CN202110727283.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-29
- Filing Date
- 2021-06-29
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2041-06-29
AI Technical Summary
Existing technologies for assessing the health of distributed networks suffer from problems such as false realism and incorrect conclusions due to static analysis, and cannot effectively predict the risk impact and abnormal health status of the system over time.
By employing artificial neural network methods, the health status and infection risk of assets and sites are assessed by identifying sites, assets, and links in a distributed network. A feedforward artificial neural network trained by backpropagation is used for prediction, and aging frequency is combined with event frequency tracking to achieve dynamic prediction of network health status.
It enables dynamic assessment and prediction of the health status of distributed networks, providing accurate predictions of future health status. It can assess the evolution of networks based on risk and health status, reducing errors and improving prediction accuracy.
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the field of security methods and security systems in distributed network management, in particular to distributed networks. In particular, the present invention relates to a method for predicting the health state of a distributed network by using artificial neural networks. BACKGROUND
[0002] A site represents a physical location where a certain number of network reachable assets are located.
[0003] An asset is a physical (or virtual, for example a virtual machine) network supported device that is physically connected within the network of a site. An asset can be a computer, a tablet, a printer, or any other type of device capable of communicating in a network such as TCP / IP.
[0004] Moreover, assets can communicate or have the possibility to communicate with other assets. In this case, they have a common link that simulates the fact that one asset can communicate with another asset through the network in a certain protocol. A computer network can have several components between assets and there are different device types (routers, firewalls, application firewalls, etc.) that can prohibit all or some protocols between two assets. For this reason, a link needs to have "from" and "to" assets and a protocol.
[0005] Due to the nature of networking software, one or more vulnerabilities can affect one or more assets and, likewise, are usually subject to attacks that compromise their security.
[0006] In the world of cybersecurity, it is common to assess the security posture of a given asset or system in a static way, by looking at its current health, vulnerabilities and security measures to prevent various types of interruptions.
[0007] A complex method to assess system vulnerabilities is to assess the entire system and each asset, a scoring system of the CVSS type includes three groups of metrics: base, temporal and environmental. The base group represents the inherent quality of the vulnerability constant over time and across user environments, the temporal group reflects the characteristics of the vulnerability that change over time, while the environmental group represents the characteristics of the vulnerability that are unique to the user environment. In summary, the base metric produces a score that can then be modified by scoring the temporal and environmental metrics.
[0008] In any case, analyzing these aspects in isolation and statically can give a false sense of reality and can lead to incorrect conclusions.
[0009] Therefore, it would be desirable to have a method capable of predicting the health state of the sites in a distributed network. Moreover, it would be desirable to have a method capable of better predicting how risks can affect the health state of the system by analyzing the evolution of the system over time as a whole. Finally, it would be desirable to have a method capable of preventing anomalous health states related to changes in the vulnerability of assets.
[0010] Likewise, it would be desirable to have a device capable of better predicting how risks can affect the health state of the system by analyzing the evolution of the system over time as a whole. SUMMARY
[0011] It is an object of the present invention to provide a method for predicting the health state of a distributed network through an artificial neural network capable of minimizing the aforementioned drawbacks.
[0012] Therefore, a method for predicting the health state of a distributed network through an artificial neural network is described according to the present invention, the method comprising a phase of identifying objects in the distributed network, comprising the following steps:
[0013] - identifying, by a computerized data processing unit operatively connected to the distributed network, one or more sites in the distributed network;
[0014] - identifying, by the computerized data processing unit, one or more assets of each identified site;
[0015] - identifying, by the computerized data processing unit, links between the identified assets, wherein a link is defined by a data packet exchanged in the distributed network, the data packet having a protocol field related to a sender asset, a protocol field related to a receiver asset and a protocol field allowing communication between the sender asset and the receiver asset, and wherein for each of the links, the sender asset and the receiver asset define nodes and the connection between the sender asset and the receiver asset defines the link between the nodes, the link having a direction from the sender asset to the receiver asset;
[0016] - storing, in a storage unit of the permanent type operatively connected to the data processing unit, the identified sites of the distributed network, the identified assets and the identified links;
[0017] wherein the method for predicting the health state further comprises a phase of evaluating the actual health state of each of the identified assets in an actual iteration, comprising the following steps:
[0018] - evaluating, by said computerized data processing unit, an actual asset health status level of each of said identified assets according to a set of predefined asset health status values ranging from a worst asset health status to a best asset health status;
[0019] - evaluating, by said computerized data processing unit, an actual asset infection risk of each of said identified assets according to a set of predefined asset infection risk values ranging from a maximum asset infection risk to no asset infection risk;
[0020] - said artificial neural network operated by said computerized data processing unit computes an actual asset infection factor of each of said identified assets as a probability that an infection of said asset can spread to other assets according to said identified links;
[0021] wherein said method for predicting health status further comprises a phase of evaluating said actual health status of each of said identified sites in said actual iteration, comprising the steps of:
[0022] - evaluating, by said computerized data processing unit, said actual site health status level of each of said identified sites as equal to a minimum actual asset health status value of said assets in said site;
[0023] - evaluating, by said computerized data processing unit, said actual site infection risk of each of said identified sites as equal to a maximum asset infection risk value of said assets in said site; and
[0024] wherein said method for predicting health status further comprises a phase of predicting, by said artificial neural network operated by said computerized data processing unit, a subsequent health status of each of said identified sites in a subsequent iteration according to a prediction function based on a set of prediction values comprising said actual asset health status level, said actual asset infection risk, said actual asset infection factor, said actual site health status level and said actual site infection risk.
[0025] Therefore, the method according to the present application allows to evaluate a network of actual sites according to risk and health status and provides a prediction about how it will perform in the near future. By using an artificial neural network, it is possible to define a machine learning method where the prediction is based on learning events in actual status.
[0026] said phase of evaluating said actual health status of each of said identified assets and said phase of evaluating said actual health status of each of said identified sites are performed within a predetermined learning time interval,
[0027] wherein said actual asset health status level of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status level of each of said identified sites and said actual site infection risk of each of said identified sites are stored in said storage unit.
[0028] The predetermined learning time interval defines a planned time for the computation of the actual iteration, so that the artificial neural network can be trained within said learning time interval.
[0029] The phase of evaluating the actual health status of each identified asset and the phase of evaluating the actual health status of each identified site are performed within a predetermined learning time interval,
[0030] wherein said actual asset health status level of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status level of each of said identified sites and said actual site infection risk of each of said identified sites comprise a plurality of values defined at predetermined learning instants comprised in said predetermined learning time interval.
[0031] In this way, the changes to the assets or sites are evaluated at predetermined instants.
[0032] The phase of evaluating the actual health status of each of said identified assets and the phase of evaluating the actual health status of each of said identified sites are performed within a predetermined learning time interval, and
[0033] wherein said actual asset health status level of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status level of each of said identified sites and said actual site infection risk of each of said identified sites comprise a plurality of values defined at changes during said predetermined learning time interval.
[0034] In this way, the changes to the assets or sites are evaluated at changes to the assets or sites.
[0035] The phase of predicting the subsequent health status of each identified site is performed within a predetermined prediction time interval.
[0036] The predetermined prediction time interval defines a planned time for the computation of the next iteration, so that the artificial neural network can predict the health status within said prediction time interval.
[0037] Within predetermined learning time intervals, phases are performed to evaluate the actual health status of each of the identified assets and phases to evaluate the actual health status of each of the identified sites.
[0038] The phase of predicting the subsequent health status of each of the identified sites is performed within a predetermined prediction time interval.
[0039] The prediction time interval is equal to the learning time interval.
[0040] Therefore, the prediction range corresponds to the training range.
[0041] The artificial neural network is a feedforward type trained using backpropagation.
[0042] In this way, information moves from the input node to the output node only in one direction. There are no loops or cycles in the network. The output value is compared with the actual value to calculate the value of a predetermined error function. The error is then fed back through the network. Using this information, the algorithm adjusts the weights of each connection to reduce the value of the error function by a small amount.
[0043] Artificial neural networks are 3-hidden-layer networks, which have at least as many neurons in the hidden layers as there are in a set of predictions.
[0044] Different artificial neural networks are used for each of the identified sites.
[0045] By defining this number of layers and neurons, one can approximate the location of each type of artificial neural network.
[0046] The set of predicted values also includes an aging frequency value for each of the identified assets in the identified sites.
[0047] For each of the assets, the computerized data processing unit calculates the aging frequency value for the next iteration by applying a predetermined decay factor to the infection factor of the actual asset in the actual iteration.
[0048] Therefore, aging frequency allows tracking the frequency of entities and events over time and can be viewed as a synapse of an artificial neural network.
[0049] The artificial neural network operated by the computerized data processing unit calculates the infection factor for each of the identified assets as the maximum value between the actual asset vulnerability factor and the actual asset diffusion factor of the asset, where the actual asset vulnerability factor is the probability that a vulnerability affects the asset, and the actual asset diffusion factor is the probability that another asset attacks the identified asset according to the identified link. Detailed Implementation
[0050] The present invention relates to a method for predicting the health state of a distributed network by means of artificial neural networks.
[0051] The method according to the present invention can be used in physical or virtual infrastructures or automation systems, in particular industrial automation systems, for example industrial processes for manufacturing production, industrial processes for power generation, infrastructures for distributing fluids (water, oil and gas), infrastructures for power generation and / or transmission, infrastructures for transport management.
[0052] In the present invention, the term "site" refers to a physical location where a certain number of network-reachable assets are located.
[0053] In the present invention, the term "asset" refers to a physical or virtual network- supported device that is physically connected to a site. An asset can be a computer, a tablet, a printer, or any other type of device capable of communicating in a network such as TCP / IP.
[0054] In the present invention, the term "link" refers to a model that represents the communication between two resources on a network by means of a certain protocol. Assets can communicate or have the possibility to communicate with other assets. If an asset can communicate with another asset, they have a common link as described above. A computer network can have several components between assets and there are different device types (routers, firewalls, application firewalls, etc.) that can prohibit all or some protocols between two assets. For this reason, a link needs to have "from" and "to" assets, as well as a protocol, because it cannot be guaranteed that if asset a can be connected to asset b with a protocol b the same can happen to asset a . It is also useful to represent links because it is possible to create a reachability graph of assets that, in turn, can be used to understand how an infection propagates on a network.
[0055] Therefore, a distributed network can connect several sites that, in turn, can be provided with one or more assets. The latter can create an interconnected network through links, as described above.
[0056] The method according to the present invention allows to identify the above-mentioned elements by means of a plurality of phases and by means of a prediction function implemented by artificial neural networks, to predict the health state of a distributed network. In particular, the scope of the present invention is to predict the health state of a distributed network on two subsequent iterations, namely the actual iteration and the subsequent iteration.
[0057] In the present invention, the term "actual iteration" means an iteration that is still running and used in the learning phase of the artificial neural network. In this regard, the term "learning time interval" in the present invention means the time interval according to the learning phase for the artificial neural network.
[0058] In the present invention, the term "subsequent iteration" means an iteration that is still not running and will be used in the prediction phase of the artificial neural network. In this regard, the term "prediction time interval" in the present invention means the time interval according to the prediction phase for the artificial neural network.
[0059] Due to the nature of network software, one or more vulnerabilities can affect an asset.
[0060] In the present invention, the term "vulnerability" means a potential security issue that a given hardware or software product (or a combination thereof) can have in a given version. A given vulnerability can be exploited in several different ways, and one of these vulnerabilities is via a network with one or more protocols, where these protocols are first used to infect an asset or spread the infection to more assets (the protocols of the first and second can be different). It is important to note that this representation takes into account the existence of vulnerabilities that can be exploited in other ways (for example, delivery of malware via USB key), but in this case the set of protocols will be empty.
[0061] In the present invention, the term "infection" means the appearance of certain malware within a network, affecting one (or more) assets, usually due to some form of vulnerability. Another characteristic of an infection is the infection factor (I-factor), represented by the probability P that the infection can spread to another asset, assuming it is also affected by the same vulnerability.
[0062] The method according to the present invention allows to evaluate the network of actual sites according to the risk and health status, and provides a prediction about how it will behave in the near future. By using an artificial neural network, a machine learning method can be defined, where the prediction is based on learning events in the actual state, as described herein.
[0063] The method for predicting the health status of a distributed network by means of an artificial neural network according to the present invention comprises three main phases, in particular a phase of identifying the objects in the distributed network, a subsequent phase of evaluating the actual health status of each identified asset in an actual iteration, a subsequent phase of evaluating the actual health status of each identified site in an actual iteration, and, finally, a phase of predicting the subsequent health status of each identified site in a subsequent iteration and by means of an artificial neural network.
[0064] The method is preferably performed by using one or more computerized data processing units, and in particular, the artificial neural network is operated by one or more of said computerized data processing units.
[0065] The phase of identifying objects in the distributed network comprises a first step of identifying, by a computerized data processing unit operatively connected to the distributed network, one or more sites in the distributed network, and a second step of identifying, by the computerized data processing unit, one or more assets of each of the identified sites.
[0066] Therefore, the distributed network can comprise one or more sites, which in turn can comprise one or more assets.
[0067] The phase of identifying objects in the distributed network further comprises a step of identifying, by the computerized data processing unit, links between the identified assets, wherein a link is defined by a data packet exchanged in the distributed network having a protocol field related to a sender asset, a protocol field related to a receiver asset and a protocol field allowing communication between the sender asset and the receiver asset, and wherein for each link, the sender asset and the receiver asset define a node and the connection between the sender asset and the receiver asset defines a link between the nodes, with a direction from the sender asset to the receiver asset.
[0068] Finally, a further step is performed, i.e. storing the identified sites, the identified assets and the identified links of the distributed network in a storage unit of the permanent type operatively connected to the data processing unit.
[0069] Therefore, the above phases of identifying objects in the distributed network allow defining the entire structure of the distributed network to be predicted, taking into account all the connections between the objects. In particular, these are the main entities and data structures. In a computer network, these entities evolve over time according to several events, thus changing the status of one or more of the involved entities.
[0070] The cyber-security bulletin of a site comprises two different values, which are its health status level and its infection risk. As mentioned for the sites, the assets themselves also have a cyber-security bulletin, which comprises a health status level and an infection risk representing the same concepts but focused on the specific asset. In the following, according to the above values, the main events affecting the evolution of the cyber-security posture are described.
[0071] In the present application, the term "health status level" means an encoded value regarding the health status of the object, i.e. the site or the asset. Preferably, the health status level is a number selected in a predetermined range which allows expressing an encoded value of the health from a worst value to a best value. In particular, in the present application, the health status level is a decimal number between the number 0 and the number 10, wherein the number 0 represents a very poor (worst) health status, while the number 10 represents a good (best) health status. A poor health status means that some infection, typically a malware, is active in the site on one or more assets, or some other form of functional degradation occurs due to cyber-security issues.
[0072] Therefore, the health status level assessed for an asset is denoted as asset health status level, while the health status level for a site is denoted as site health status level. Moreover, as mentioned above, considering the type of iteration, the health status level of an asset can be assessed in the actual iteration as an actual asset health status level or actual health asset of the asset, and in the subsequent iteration as a subsequent asset health status level or subsequent health asset of the asset. For a site considering the actual iteration, with the necessary modifications, as an actual site health status level or actual health asset of the site, and in the subsequent iteration, as a subsequent site health status level or subsequent health asset of the site.
[0073] In the present application, the term "infection risk" means an encoded value regarding the risk of infection of the object, i.e. the site or the asset. Preferably, the infection risk is a number selected in a predetermined range which allows expressing an encoded value of the infection risk from a maximum value to a minimum value. In particular, in the present application, the infection risk is a decimal number between the number 0 and the number 10, wherein the number 0 represents the risk of not being infected (best), while the number 10 represents the certainty of being infected (worst).
[0074] In assessing the above values, the method according to the present application comprises a phase of assessing the actual health status of each identified asset in the actual iteration. In particular, this phase comprises a step of evaluating, by the computerized data processing unit, the actual asset health status level of each identified asset according to a set of predefined asset health status values ranging from a worst asset health status to a best asset health status. A further step is performed of evaluating, by the computerized data processing unit, the actual asset infection risk of each identified asset according to a set of predefined asset infection risk values ranging from a maximum asset infection risk to no asset infection risk. Finally, a step is performed of calculating, by the computerized data processing unit operating a neural network, the actual asset infection factor of each identified asset as the probability that the asset infection can propagate to other assets according to the identified links.
[0075] Considering the values evaluated or calculated for each asset of a site, the method according to the present application comprises a phase of evaluating the actual health status of each identified site in the real iteration. In particular, such phase comprises a first step of evaluating, by the computerized data processing unit, the actual site health status level of each identified site as equal to the minimum actual asset health status value of the assets in the site, and a second step of evaluating, by the computerized data processing unit, the actual site infection risk of each identified site as equal to the maximum asset infection risk value of the assets in the site.
[0076] A high infection risk causes an increase in the health status level in a short time, while a site with a low infection risk can have a good health status level.
[0077] Based on the above ranges, when an infection affects an asset, the corresponding health status level decreases by the health impact value of the infection (a decimal number between the number 0 and the number 10), where the number 10 represents the maximum damage to the asset health status level. The health impact derives from the vulnerability used for the infection by considering their maximum health impact.
[0078] In the present application, the term "vulnerability" means the inability of an object to withstand the effects of an adverse environment. A vulnerability is characterized by a set of conditions (e.g. software version) that need to be present on an asset in order to have a risk factor available.
[0079] The above phases, i.e. the phase of evaluating the actual health status of each identified asset and the phase of evaluating the actual health status of each identified site, allow the training (or learning phase) of the artificial neural network design to perform the method, as described in more detail below. An artificial neural network (ANN) is a computational system inspired by biological neural networks. Such systems learn to perform tasks by considering examples, usually without being programmed with specific rules for the task. ANNs are based on a collection of connected units or nodes called artificial neurons (or simply neurons) which loosely model the neurons in a biological brain. Each connection can transmit a signal to other neurons. The reception of a signal triggers an activation of the receiving neuron, which in turn may relay signals to further neurons. Layers of neurons sometimes model different levels of processing. Signals travel from the first (input) layer to the last (output) layer, possibly after traversing the layers multiple times.
[0080] In an embodiment, the artificial neural network of the present application is of the feedforward type trained with backpropagation.
[0081] A feedforward neural network is an artificial neural network in which connections between nodes do not form cycles, in which information moves in only one direction, from the input nodes, forward through the hidden nodes (if any) and to the output nodes. There are no cycles or loops in the network. The output values are compared to the actual values to compute the value of some predetermined error function. The error is then fed back through the network. Using this information, the algorithm adjusts the weights of each connection in such a way as to reduce the value of the error function by some small amount.
[0082] In an embodiment, the artificial neural network is a 3 hidden layer network having at least as many neurons in the hidden layers as the number of the set of predicted values. In particular, the number of sites to be evaluated defines the number of artificial neural networks to be used, with a different artificial neural network being used for each identified site.
[0083] By defining such a number of layers and neurons, each site having its own artificial neural network can be approximated.
[0084] In an ANN multi-layer utilizing backpropagation, the output values are compared to the correct answers to compute the value of some predetermined error function. The error is then fed back through the network by various techniques. Using this information, the algorithm adjusts the weights of each connection in such a way as to reduce the value of the error function by some small amount. After repeating this process for a sufficient number of training cycles, the network will typically converge to some state in which the computed error is small, and the ANN has learned something about the target function.
[0085] In an embodiment, the phase of evaluating the actual health status of each identified asset and the phase of evaluating the actual health status of each identified site are performed within a predetermined learning time interval, wherein the actual asset health status rating of each identified asset, the actual asset infection risk of each identified asset, the actual asset infection factors of each identified asset, the actual site health status rating of each identified site and the actual site infection risk of each identified site are stored in a storage unit.
[0086] The predetermined learning time interval defines a scheduled time for the computation of actual iterations, so that the artificial neural network can be trained within said learning time interval.
[0087] In particular, the phase of evaluating the actual health status of each identified asset and the phase of evaluating the actual health status of each identified site are performed within a predetermined learning time interval, wherein the actual asset health status rating of each identified asset, the actual asset infection risk of each identified asset, the actual asset infection factors of each identified asset, the actual site health status rating of each identified site and the actual site infection risk of each identified site comprise a plurality of values defined at predetermined learning instants within the predetermined learning time interval.
[0088] In this way, the changes to the assets or sites are evaluated at predetermined time instants.
[0089] Alternatively or in combination with the above features, the phase of evaluating the actual health status of each identified asset and the phase of evaluating the actual health status of each identified site are performed within a predetermined learning time interval, wherein the actual asset health status rating of each identified asset, the actual asset infection risk of each identified asset, the actual asset infection factor of each identified asset, the actual site health status rating of each identified site and the actual site infection risk of each identified site comprise a plurality of values defined at the time of the change during the predetermined learning time interval.
[0090] In this way, the changes to the assets or sites are evaluated at the time of the change of the assets or sites.
[0091] In one embodiment, the set of prediction values further comprises an aging frequency value for each identified asset in the identified site, wherein the aging frequency value for each asset for the next iteration is computed by the computerized data processing unit by applying a predetermined decay factor to the actual asset infection factor in the actual iteration.
[0092] Hence, the aging frequency allows to track the frequency of entities and events over time and can be seen as the synapse of an artificial neural network
[0093] The aging frequency allows to track the frequency of entities and events over time and can be seen as the synapse of an artificial neural network. Indeed, this data structure is the basis of the learning and prediction algorithms that allow to understand the current and future behavior of the system. The aging frequency can be used to track the frequency of single objects or to compute correlation matrices. In both cases, the main idea is that this data structure represents the knowledge of a given event whose importance decreases over time. For example, when tracking the probability of an asset to infect another asset, we can represent it as a matrix AgingFrequencyProbabilityofContagion(Asset i ,Asset j ) whose values can be initialized with a certain quantity (say, 0.5). When iterating to the next period of the aging frequency, each value of the matrix is updated with a decay factor that reduces all probabilities by the value of the decay factor. In the case of a decay factor of 0.01, at each iteration, AgingFrequencyProbabilityofContagion(Asset i ,Asset j ) is adjusted, so that in the case of a previous iteration of 0.5, the new value will be 0.49. Different aging frequency structures (tracking different objects) can use different decay factors.
[0094] After the learning phase, the method for predicting the health status further comprises the following phase: in a subsequent iteration, predicting, by the artificial neural network operated by the computerized data processing unit, a subsequent health status of each of the identified sites according to a prediction function based on a set of predicted values comprising the actual asset health status level, the actual asset infection risk, the actual asset infection factor, the actual site health status level and the actual site infection risk.
[0095] Preferably, the infection factor is calculated for each identified asset by the artificial neural network operated by the computerized data processing unit as the maximum between the actual asset vulnerability factor of the asset, which is the probability that a vulnerability affects the asset, and the actual asset diffusion factor of the asset, which is the probability that another asset attacks the identified asset according to the identified link.
[0096] Therefore, the prediction function uses a machine learning method, preferably with a feedforward artificial neural network (ANN) trained with backpropagation, which is used to build a model to understand the relationships between all the considered factors of each asset (among them and over time).
[0097] Some events to be evaluated can be defined by a "connection", which occurs whenever an asset communicates with another asset using a given protocol and application. When this event occurs, the link is created or updated accordingly.
[0098] Another event can be defined by an "attack", which can occur at a given time by an attacker asset or an external attacker on a target asset, thus creating a new infection. The attack uses one or more vulnerabilities. When an infection is created, these updates are triggered in the method:
[0099] - updating the health status level of the infected asset;
[0100] - updating the aging frequency
[0101] o the aging frequency ProbabilityOfBeingExploited(漏洞) = 1
[0102] The probability that an asset with a given vulnerability is affected by it is increasing;
[0103] o the aging frequency Asset(ProbabilityOfBeingAttacked) = 1
[0104] The probability that an asset to be attacked is high.
[0105] Moreover, when a new software is installed in the system, an event "software change" can occur, which is either a completely new software or an upgrade of already installed software. Sometimes, the update of software is called a "patch". The patch or a series of patches can be due to the will to remove an infection from the asset. When the software is installed or upgraded, the asset can have solved some vulnerabilities or can have appeared new ones. The risk factor of the asset is updated: its risk is computed by finding the maximum of the risks in the vulnerabilities that affect it. If the software change event is removing an infection, the following updates are also performed:
[0106] - update the health status level of the infected asset as described;
[0107] - update the aging frequency:
[0108] aging frequency ProbabilityOfBeingExploited(漏洞) = 0
[0109] If the event does not leave a single asset vulnerable to the given vulnerability.
[0110] aging frequency ProbabilityOfBeingExploited(资产) = 0
[0111] If the event fixes all the vulnerabilities present in the asset.
[0112] Finally, when an infected asset spreads its infection to another asset, an event "contamination" can occur. This event is similar to an attack, but it is tracked differently in order to be able to better predict the future evolution of the system. Several updates are triggered in this method, similar but different from the attack:
[0113] - update the health status level of the infected asset as described;
[0114] - update the aging frequency:
[0115] aging frequency ProbabilityOfBeingExploited(漏洞) = 1
[0116] The probability that an asset with the given vulnerability can be affected by it is increasing.
[0117] aging frequency ProbabilityOfContagion(AssetX,AssetY) = 1
[0118] The probability that asset X may infect asset Y is increased
[0119] aging frequency Asset(ProbabilityOfBeingAttacked) = 1
[0120] The probability that the asset to attack is high.
[0121] The method of the invention allows to compute the health state grade and the infection risk of a site (based on the same computation for the corresponding assets), and the computation of these two values over time allows to track and predict the cyber-security posture of a complex, geographically distributed and interconnected network of networks.
[0122] The idea of the method is ideal, in the case where everything starts on site at time 0 (first iteration) with a new and secure software installation and it has an ideal case where all assets have an infection risk with a value equal to 0 and a health state grade with a value equal to 10.
[0123] From the second iteration, this initial and ideal case is rapidly deteriorated by some assets infected by external actors: these events are driven by the existence and evolution of vulnerabilities and how big the attack surface of those vulnerabilities is, for example, if there are any defensive measures to prevent them.
[0124] From the second iteration, assets can be contaminated by other infected assets. This flow is mainly driven by the I-factors of the ongoing infection and the measures that can prevent the spread of the infection in situ. Of course, the external actor infection flow remains active from the second iteration and beyond.
[0125] In one embodiment, the phase of predicting a subsequent health state of each identified site is performed within a predetermined prediction time interval. In particular, the predetermined prediction time interval defines a planned time for the computation of the next iteration, so that the artificial neural network can predict the health state for the prediction time interval.
[0126] Preferably, the phase of evaluating an actual health state of each identified asset and the phase of evaluating an actual health state of each identified site are performed within a predetermined learning time interval, wherein the phase of predicting a subsequent health state of each identified site is performed within a predetermined prediction time interval, and wherein the prediction time interval is equal to the learning time interval. Thus, the prediction range corresponds to the training range.
[0127] The prediction function tries to understand what happens in the next iteration, which happens after the predetermined prediction time interval. The prediction time interval of the function can be set to e.g. 24 hours - the method will try to predict the state of the system in the next 24 hours, assuming that all entities and data structures are updated to the current state. It is important to note that if the prediction time interval needs to be changed, the whole learning needs to start from scratch.
[0128] As already described, the prediction function uses a machine learning method, which has a feed-forward artificial neural network trained with backpropagation, which is used to build a model for each asset to understand the relationships between all considered factors (between them and over time), i.e.:
[0129] f Asset_a (x) = y
[0130] where "x" is called the pattern of a given set of features of the asset and "y" is the estimated new health status rank.
[0131] Preferably, the "x" vector is a pattern of features as described herein:
[0132] - the current health status rank of asset a ;
[0133] - the aging frequency a ProbabilityOfBeingExploited(Asset_a) of asset a ;
[0134] - the highest "n" value of the aging frequency ProbabilityOfBeingExploited(漏洞) ProbabilityOfContagion(Asset_b,Asset_a) b a
[0135] - the highest "n" value of the aging frequency ProbabilityOfContagion(Asset_b,Asset_a) b a
[0136] - the highest "n" value of the I-factor active infection on asset b b a
[0137] The artificial neural network used to estimate f Asset_a is a 3 hidden layer network with at least the same number of neurons in the hidden layers as the number of features, i.e. 3*n+2.
[0138] The method is trained in this way. At any given time, for asset a , we make Xa have the last "m" entries to allow the method to evolve over time and not to be biased towards past behavior.
[0139] In the first iteration (actual), the pattern for observed behavior is recorded. The method adds to the available pattern Xa the pair (x,y), computes the features of "x" taking into account the previous health status rank and "y" as the current health status rank.
[0140] When at least "z" iterations have been made (learning phase), where "z" is a parameter set during the learning phase, the method starts to predict behavior. For each asset a , it trains itself to estimate f Asset_a Split, take random 2 / 3 of Xa and use the remaining 1 / 3 to verify its performance, using some form of metric, like overall accuracy, not described in detail. If the overall prediction accuracy is higher than a predetermined amount, i.e. 0.9, which means that the prediction error on the test set is less than 10%, the predicted value of the health status class of the asset is f Asset_a (x) = y. In any case, at each iteration, the actual observation of (x, y) is added to Xa to improve the future prediction in further iterations.
[0141] For each aging frequency table, for each entry, an attenuation factor is applied to the next iteration.
[0142] The above steps allow to predict the subsequent health status class of each asset. The subsequent health status class of a site is equal to the minimum predicted health status class of the assets that compose it.
[0143] The above method allows a complete, unsupervised operation of the algorithm. In case more complex asset to site combination functions are needed, for example to give less weight to mostly isolated assets, some more steps are needed and a human expert is needed to provide knowledge to the system to understand the required aggregation strategy.
[0144] The method according to the present application thus allows to compute an automatic announcement about the state of the network of sites according to the risk and the current health status and provides a prediction about how it will perform in the near future.
Claims
1. A method for predicting a health state of a distributed network by an artificial neural network, characterized in that, The method comprises a phase of identifying objects in the distributed network, comprising the steps of: - identifying, by a computerized data processing unit operatively connected to the distributed network, one or more sites in the distributed network; - identifying, by the computerized data processing unit, one or more assets of each identified site; - identifying, by the computerized data processing unit, links between the identified assets, wherein a link is defined by data packets exchanged in the distributed network, having a protocol field related to a sender asset, a protocol field related to a receiver asset and a protocol field allowing communication between the sender asset and the receiver asset, and wherein for each said link, the sender asset and the receiver asset define nodes and the connection between the sender asset and the receiver asset defines the link between the nodes, the link having a direction from the sender asset to the receiver asset; - storing, in a storage unit of the permanent type operatively connected to the data processing unit, the identified sites of the distributed network, the identified assets and the identified links; wherein the method for predicting health status further comprises a phase of evaluating, in an actual iteration, an actual health status of each said identified asset, comprising the steps of: - evaluating, by the computerized data processing unit, an actual asset health status level of each said identified asset according to a set of predefined asset health status values ranging from a worst asset health status to a best asset health status; - evaluating, by the computerized data processing unit, an actual asset infection risk of each said identified asset according to a set of predefined asset infection risk values ranging from a maximum asset infection risk to a null asset infection risk; - calculating, by the artificial neural network operated by the computerized data processing unit, an actual asset infection factor of each said identified asset as a probability that an infection of the asset can spread to other assets according to the identified links; wherein the method for predicting health status further comprises a phase of evaluating, in the actual iteration, the actual health status of each said identified site, comprising the steps of: - evaluating, by the computerized data processing unit, an actual site health status level of each said identified site as equal to the minimum actual asset health status value of the assets in the site; - evaluating, by the computerized data processing unit, an actual site infection risk of each said identified site as equal to the maximum asset infection risk value of the assets in the site; and wherein the method for predicting health status further comprises a phase of predicting, by the artificial neural network operated by the computerized data processing unit, a subsequent health status of each said identified site in a subsequent iteration according to a prediction function based on a set of prediction values comprising the actual asset health status level, the actual asset infection risk, the actual asset infection factor, the actual site health status level and the actual site infection risk.
2. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, characterized in that, said phase of evaluating the actual health status of each of said identified assets and said phase of evaluating the actual health status of each of said identified sites are performed within a predetermined learning time interval, and wherein said actual asset health status rating of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status rating of each of said identified sites and said actual site infection risk of each of said identified sites are stored in said storage unit.
3. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, characterized in that, said phase of evaluating the actual health status of each of said identified assets and said phase of evaluating the actual health status of each of said identified sites are performed within a predetermined learning time interval, and wherein said actual asset health status rating of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status rating of each of said identified sites and said actual site infection risk of each of said identified sites comprise a plurality of values defined at predetermined learning instants comprised in said predetermined learning time interval.
4. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, characterized in that, said phase of evaluating the actual health status of each of said identified assets and said phase of evaluating the actual health status of each of said identified sites are performed within a predetermined learning time interval, and wherein said actual asset health status rating of each of said identified assets, said actual asset infection risk of each of said identified assets, said actual asset infection factors of each of said identified assets, said actual site health status rating of each of said identified sites and said actual site infection risk of each of said identified sites comprise a plurality of values defined at varying times during said predetermined learning time interval.
5. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, wherein, said phase of predicting the subsequent health status of each of said identified sites is performed within a predetermined prediction time interval.
6. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, wherein, said phase of evaluating the actual health status of each of said identified assets and said phase of evaluating the actual health status of each of said identified sites are performed within a predetermined learning time interval, said phase of predicting the subsequent health status of each of said identified sites is performed within a predetermined prediction time interval, and wherein said prediction time interval is equal to the learning time interval.
7. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, wherein, said artificial neural network is of the feed-forward type trained with backpropagation.
8. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, wherein, said artificial neural network is a 3-hidden layer network having at least as many neurons in the hidden layers as the number of prediction values of said set of prediction values, and wherein a different artificial neural network is used for each of said identified sites.
9. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, wherein, said set of prediction values further comprises an aging frequency value for each of said identified assets of said identified sites, wherein, for each of said assets, said aging frequency value for the next iteration is computed by said computerized data processing unit by applying a predetermined decay factor to said actual asset infection factors in said actual iteration.
10. The method for predicting a health state of a distributed network by an artificial neural network according to claim 1, characterized in that, The artificial neural network operated by the computerized data processing unit computes, for each of the identified assets, the infection factor as a maximum between an actual asset vulnerability factor of the asset, which is a probability that a vulnerability affects the asset, and an actual asset diffusion factor of the asset, which is a probability that another asset attacks the identified asset according to the identified link.
Citation Information
Patent Citations
Plant Process Management System with Normalized Asset Health
CN106873548A
Anomaly Detection to Implement Security Protection of a Control System
US20120210158A1