Transfer device and transfer method
Patent Information
- Application Number
- PCT/JP2023/039736
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-08
AI Technical Summary
When a new network (NW), the knowledge of the NW cannot be constructed due to insufficient data, and even similar NWs cannot be directly transferred due to topological differences.
By designing a transfer device, the device can transfer a decision tree of a source NW with a hierarchy to another target NW with the same hierarchy, and adjust the variables and thresholds in the decision tree through the device and variable correspondence calculation to adapt to the conditions of the target NW.
The transfer of knowledge from one NW to another NW is implemented, which solves the problem of lack of data in new NWs and improves the feasibility of knowledge transfer between NWs.
Smart Images

Figure JP2023039736_08052025_PF_FP_ABST
Abstract
Description
Transfer device and transfer method
[0001] The present disclosure relates to a transfer device and a transfer method.
[0002] Network operators and the like who provide multiple network solutions manage and operate multiple networks, and therefore utilize technologies that enable the management and operation of multiple networks (for example, see Patent Document 1).
[0003] Furthermore, an anomaly detection model for detecting anomalies (e.g., failures, equipment breakdowns, increased load, human error, etc.) occurring in a network (hereinafter referred to as NW) is created using a decision tree, and NW anomalies are detected using this anomaly detection model. In addition to anomaly detection, decision trees are also used in the fields of NW maintenance and construction, such as fault isolation, response procedure acquisition, and NW configuration design. These represent knowledge related to NW maintenance and construction, such as anomaly detection, isolation, response procedures, NW configuration design, etc., in the form of a decision tree.
[0004] Japanese Patent Application Laid-Open No. 2002-290405
[0005] However, when a new network is constructed, there is little data on that network, so knowledge of that network cannot be constructed. On the other hand, although it is possible to transfer knowledge of a certain network to knowledge of a newly constructed network, even if those networks are similar, there are differences such as differences in the topology of those networks, so the transfer cannot be performed as is.
[0006] The present disclosure has been made in consideration of the above points, and provides a technology that enables knowledge of one network to be transferred as knowledge of another network.
[0007] A transfer device according to one aspect of the present disclosure is a transfer device for transferring a first decision tree related to a first network having a hierarchical structure to a second network having a hierarchical structure consisting of the same number of levels as the first network, and includes: a first correspondence calculation unit that calculates, for each level of the hierarchical structure, a first correspondence representing a correspondence between a first device belonging to the level of the first network and a second device belonging to the level of the second network; a second correspondence calculation unit that calculates, for each level of the hierarchical structure, based on the first correspondence relationship, a second correspondence representing a correspondence between a first variable representing an attribute of an observation value observed in a first device belonging to the level and a second variable representing an attribute of an observation value observed in a second device belonging to the level; and a transfer unit that creates a second decision tree related to the second network by replacing a first variable of a conditional expression set for a node included in the first decision tree with a second variable based on the second correspondence relationship.
[0008] A technology is provided that allows knowledge of one network to be transferred as knowledge of another network.
[0009] 1 is a diagram illustrating an example of a network; FIG. 2 is a diagram illustrating an example of a transfer source network; FIG. 3 is a diagram illustrating an example of a transfer destination network; FIG. 4 is a diagram illustrating an example of a decision tree for detecting an abnormality in the transfer source network; FIG. 5 is a diagram illustrating an example of the hardware configuration of a transfer device according to the present embodiment; FIG. 6 is a diagram illustrating an example of the functional configuration of a transfer device according to the present embodiment; FIG. 7 is a flowchart illustrating an example of decision tree creation and transfer processing; FIG. 8 is a diagram illustrating an example of device correspondence relationship; FIG. 9 is a diagram illustrating an example of an adjacency matrix; FIG. 10 is a diagram illustrating an example of variable correspondence relationship; FIG. 11 is a flowchart illustrating an example of a decision tree creation process for the transfer destination network; FIG. 12 is a diagram illustrating an example of a decision tree for detecting an abnormality in the transfer destination network.
[0010] An embodiment of the present invention will be described in detail below with reference to the drawings. In the following embodiment, an anomaly detection is taken as an example, and a transfer device 10 will be described that can transfer an anomaly detection model of a certain network as an anomaly detection model of another network, assuming an anomaly detection model as network knowledge. However, assuming an anomaly detection model as network knowledge is just one example, and network knowledge is not limited to an anomaly detection model, and may be, for example, fault isolation, response procedures, network configuration design, etc.
[0011] Note that transferring an anomaly detection model of a certain network as an anomaly detection model of another network means making the anomaly detection model of a certain network available as an anomaly detection model of the other network. Transfer may be called, for example, "porting," "transplantation," or "migration."
[0012] In the following, the anomaly detection model is represented by a decision tree. The network that is the source of a decision tree is called the "source network," and the network that is the destination of a decision tree is called the "destination network."
[0013] The source network is assumed to be a network from which sufficient learning data has been collected to create a decision tree, which is an anomaly detection model. On the other hand, the destination network is assumed to be a network from which sufficient learning data has not been collected to create a decision tree, which is an anomaly detection model, due to its use, for example, as a newly constructed network or a network that has just been put into operation.
[0014] <Problem Setting> First, an example of a network is shown in Figure 1. As shown in Figure 1, a network is generally composed of various servers (web servers, application servers, database servers, etc.), network devices (routers, gateways, core network devices, etc.), various terminals, etc., and has a hierarchical structure. In the network shown in Figure 1, each terminal belongs to the first layer, each network device belongs to the second layer, and each server belongs to the third layer.
[0015] Hereinafter, elements constituting a network, such as terminals, network devices, and servers, will be referred to simply as "devices," and each device will be assigned identification information that allows it to be uniquely identified within the network. Information identifying observed values collected from each device (in other words, the attributes and metric names of the observed values) will be referred to as "variables." Specific examples of variables include CPU (Central Processing Unit) utilization, memory utilization, traffic volume, temperature, and logs.
[0016] In the following embodiment, the source network is mainly assumed to be NW-A shown in Fig. 2, and the destination network is mainly assumed to be NW-B shown in Fig. 3. The source network may be a network that is actually in operation, or may be a network that is constructed as a template for the destination network.
[0017] 2 is composed of devices a1 to a5, devices a11 to a13, and devices a21 to a22, which form layers 1 to 3. Furthermore, observed values of variables X1 to X5 are collected from devices a1 to a5, respectively, and similarly observed values of variables X11 to X13 are collected from devices a11 to a13, respectively, and observed values of variables X21 to X22 are collected from devices a21 to a22.
[0018] 3 is composed of devices b1 to b6, devices b11 to b14, and devices b21 to b23, which form layers 1 to 3. Observed values of variables Y1 to Y6 are collected from devices b1 to b6, respectively, and similarly, observed values of variables Y11 and Z11 to Y14 and Z14 are collected from devices b11 to b14, respectively, and observed values of variables Y21 to Y23 are collected from devices b21 to b23, respectively.
[0019] In this way, it is assumed that the number of devices is not necessarily the same between the transfer source NW (NW-A) and the transfer destination NW (NW-B).
[0020] Furthermore, the decision tree TA shown in FIG. 4 is assumed as a decision tree for detecting anomalies in the NW-A shown in FIG. 2 . The decision tree TA shown in FIG. 4 is created using a known algorithm such as CART (Classification and Regression Trees) using observation data (X1, X2, X3, X4, X5, X11, X12, X13, X21, X22) collected from the NW-A and a label indicating whether the entire NW-A was normal or abnormal at that time. In the decision tree TA shown in FIG. 4 , at a node where a conditional expression expressed as "variable < threshold" is set, if the condition is met, the tree branches to a left child node, and if not, to a right child node. Furthermore, a value of 0 set at a leaf node indicates that the entire system (NW-A) is normal, and a value of 1 indicates that the entire system is abnormal. As a result, when certain observation data of the NW-A is obtained, the decision tree TA can determine whether the entire NW-A is normal or abnormal, thereby realizing anomaly detection in the NW-A. Here, θ1, θ2, and θ3 are thresholds. Note that the term "node" can refer to a node of a decision tree (root node, intermediate node, leaf node) or to a device in a network.
[0021] At this time, the objective is to transfer the decision tree TA of the transfer source NW (NW-A) as an anomaly detection model for the transfer destination NW without using the observation data of the transfer destination NW (NW-B).
[0022] <Assumptions> This embodiment is based on the following assumptions: Note that when simply referred to as "NW", this refers to both NW-A and NW-B.
[0023] Premise 1: Topology information representing the topology of the network (i.e., the network configuration) is available (or may be known).
[0024] Assumption 2: The number of layers in NW-A and the number of layers in NW-B can be acquired as topology information (or may be known), and the number of layers in NW-A and the number of layers in NW-B are the same.
[0025] Assumption 3: The performance and capacity of each link in the network are assumed to be equal.
[0026] Assumption 4: The performance of each device belonging to the same layer in each layer of the network is assumed to be equal.
[0027] Assumption 5: The number of links connected to each device belonging to the same layer in each layer of the network is the same.
[0028] Assumption 6: The topology (graph structure) between adjacent layers in the network is symmetric. In other words, even if devices in the i-th layer are arbitrarily swapped, the topology between the i-th layer and the i+1-th layer remains unchanged. Similarly, even if devices in the i+1-th layer are arbitrarily swapped, the topology between the i-th layer and the i+1-th layer remains unchanged. A typical example of a graph that satisfies this symmetry is a complete bipartite graph.
[0029] Assumption 7: The performance and capacity of each link in NW-B are assumed to be the same as the performance and capacity of the corresponding link in NW-A. Note that if a link in NW-B and a link in NW-A are links between the same layers, these links correspond to each other.
[0030] Assumption 8: The performance of the devices belonging to each layer of NW-B is assumed to be the same as the performance of the devices belonging to the corresponding layer of NW-A. Note that if a layer of NW-B and a layer of NW-A are on the same hierarchical level, these layers correspond to each other.
[0031] Assumption 9: For the conditional expressions set for the root node and intermediate nodes of the decision tree TA of NW-A, there are one or more types of thresholds in the conditional expressions related to variables at the same level, and the variables for the same type of threshold are symmetrical. Here, the same type of threshold refers to thresholds that are equal to each other at the same level (or thresholds that are equal with a certain margin of error allowed), while different types of thresholds refer to thresholds that are not equal to each other at the same level (or thresholds that are not equal even with a certain margin of error allowed). Furthermore, the symmetry of variables for the same type of threshold means that in conditional expressions containing the same type of threshold, the conclusion of the decision tree does not change even if the variables included in those conditional expressions (note that these variables are variables at the same level) are swapped.
[0032] However, with regard to the above premise 9, it is not necessary for all variables to be symmetrical with respect to the same type of threshold. That is, for example, for some types of thresholds, the variables with respect to those thresholds may not be symmetrical with respect to those thresholds. In this case, the conditional expressions for the asymmetric variables will appear as surplus terms, which will be described later.
[0033] <Example of Hardware Configuration of Transfer Device 10> An example of the hardware configuration of the transfer device 10 according to this embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the hardware configuration of the transfer device 10 according to this embodiment.
[0034] 5, the transfer device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a RAM (Random Access Memory) 105, a ROM (Read Only Memory) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.
[0035] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the transfer device 10 does not necessarily have to have at least one of the input device 101 and the display device 102, for example.
[0036] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.
[0037] The communication I / F 104 is an interface through which the transfer device 10 communicates with other devices. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. The processor 108 is, for example, a CPU, a GPU (Graphics Processing Unit), or any of various other arithmetic devices.
[0038] 5 is an example, and the hardware configuration of the transfer device 10 is not limited to this. For example, the transfer device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various hardware other than the hardware shown in the figure.
[0039] <Example of Functional Configuration of Transfer Device 10> An example of the functional configuration of the transfer device 10 according to this embodiment will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the functional configuration of the transfer device 10 according to this embodiment.
[0040] As shown in FIG. 6 , the transfer device 10 according to this embodiment includes a topology information acquisition unit 201, a decision tree creation unit 202, a device correspondence calculation unit 203, a variable correspondence calculation unit 204, a decision tree transfer unit 205, and an output unit 206. Each of these units is realized, for example, by a process in which one or more programs installed in the transfer device 10 are executed by the processor 108 or the like. The transfer device 10 according to this embodiment also includes a learning data storage unit 207. The learning data storage unit 207 is realized, for example, by the auxiliary storage device 107 or the like. Note that the learning data storage unit 207 may also be realized, for example, by a database or the like communicatively connected to the transfer device 10.
[0041] The topology information acquisition unit 201 acquires topology information representing the topology (graph structure) of the transfer source network and topology information representing the topology (graph structure) of the transfer destination network. Note that the topology information acquisition unit 201 can acquire the topology information of the transfer source network and the transfer destination network by using, for example, SNMP (Simple Network Management Protocol) or a protocol unique to the network device.
[0042] The decision tree creation unit 202 creates a decision tree as an anomaly detection model for the source network using the learning data stored in the learning data storage unit 207. The decision tree creation unit 202 may create a decision tree for the source network using a known algorithm such as CART. In CART, a decision tree is created by dividing the tree so that statistical bias is increased using an index value called the Gini coefficient (impurity). Here, the learning data refers to observation data to which a label indicating whether the entire source network is normal or abnormal is assigned.
[0043] The device correspondence calculation unit 203 calculates a device correspondence that represents a correspondence between devices that configure the transfer source network and devices that configure the transfer destination network, using topology information of the transfer source network and topology information of the transfer destination network. Here, the device correspondence is represented by a set of pairs of a set of one or more devices included in the transfer source network and a set of one or more devices included in the transfer destination network.
[0044] The variable correspondence calculation unit 204 calculates a variable correspondence relationship that represents a correspondence relationship between variables of the transfer source NW and variables of the transfer destination NW, using the device correspondence relationship. Here, the variable correspondence relationship is represented by a set of pairs of a set of one or more variables of the transfer source NW and a set of one or more variables of the transfer destination NW.
[0045] The decision tree transfer unit 205 uses the decision tree of the transfer source network and the variable correspondence relationship to create a decision tree that serves as an anomaly detection model for the transfer destination network. Here, the decision tree transfer unit 205 includes a pattern extraction unit 211, a surplus term deletion unit 212, a variable substitution unit 213, and a threshold substitution unit 214. However, the decision tree transfer unit 205 does not necessarily have to include the surplus term deletion unit 212.
[0046] The pattern extraction unit 211 extracts, for each layer of the decision tree of the transfer source network, patterns of conditional expressions in which variables of devices belonging to that layer are exchangeable (hereinafter also referred to as exchangeable patterns). An exchangeable pattern is a logical expression that expresses a conditional expression in which, when the decision tree of a certain network is expressed as a logical expression, the conclusion of the decision tree does not change even if variables of devices belonging to the same layer of the network are exchanged. The surplus term deletion unit 212 deletes conditional expressions that cannot be expressed as exchangeable patterns (i.e., conditional expressions related to variables that are not exchangeable with variables belonging to the same layer; hereinafter referred to as surplus terms) when the logical expression of the decision tree of the transfer source network is expressed as an exchangeable pattern. The variable substitution unit 213 uses the variable correspondence relationship and the logical expression after the surplus terms have been deleted from the logical expression of the decision tree of the transfer source network to replace the exchangeable patterns included in the logical expression with exchangeable patterns of the transfer destination network. The threshold substitution unit 214 replaces the thresholds of each conditional expression included in the logical expression after the substitution with the exchangeable patterns of the transfer destination network. As a result, a decision tree expressed by the logical formula after the threshold value has been replaced is obtained as the decision tree of the transfer destination NW.
[0047] The output unit 206 outputs the decision tree of the transfer destination NW to a predetermined output destination, such as a storage area of the auxiliary storage device 107, or another device or equipment communicably connected to the transfer device 10.
[0048] The learning data storage unit 207 stores a set of learning data (learning data set) used to create a decision tree of the source network.
[0049] <Creating a decision tree and transferring process> Hereinafter, a process of creating a decision tree TA as an anomaly detection model for NW-A, which is a transfer source NW, and then transferring the decision tree TA as an anomaly detection model for NW-B, which is a transfer destination NW, will be described with reference to Fig. 7. Fig. 7 is a flowchart showing an example of the creating a decision tree and transferring process.
[0050] First, the topology information acquisition unit 201 acquires the topology information of NW-A, which is the source NW (step S101).
[0051] Next, the topology information acquisition unit 201 acquires the topology information of NW-B, which is the transfer destination NW (step S102).
[0052] Next, the decision tree creation unit 202 creates a decision tree TA as an anomaly detection model for NW-A by using the learning data stored in the learning data storage unit 207 and a known algorithm such as CART (step S103). Note that if the decision tree TA has already been created, step S103 does not need to be executed.
[0053] Next, the device correspondence calculation unit 203 calculates the device correspondence between the devices constituting the NW-A and the devices constituting the NW-B using the topology information of the NW-A and the topology information of the NW-B (step S104). The device correspondence calculation unit 203 may, for example, identify the devices belonging to each layer of the NW-A and the devices belonging to each layer of the NW-B, and then calculate the device correspondence, which is a correspondence between sets of devices belonging to the same layer between the NW-A and the NW-B. That is, for example, when i = 1, 2, ..., N (where N is the number of layers of the NW-A and NW-B), the device correspondence calculation unit 203, for each i, identifies the devices belonging to the i-th layer of the NW-A and the devices belonging to the i-th layer of the NW-B, and then calculates the device correspondence, which is a correspondence between the set of devices belonging to the i-th layer of the NW-A and the set of devices belonging to the i-th layer of the NW-B. Hereinafter, the device correspondence will be represented by g.
[0054] For example, between NW-A shown in FIG. 2 and NW-B shown in FIG. 3, the device correspondence shown in FIG. 8 is calculated. The device correspondence shown in FIG. 8 is composed of correspondence g1, correspondence g2, and correspondence g3. That is, g = {g1, g2, g3}. Correspondence g1 is a correspondence representing that {device a1, device a2, device a3, device a4, device a5} corresponds to {device b1, device b2, device b3, device b4, device b5, device b6}. Correspondence g2 is a correspondence representing that {device a11, device a12, device a13} corresponds to {device b11, device b12, device b13, device b14}. Correspondence g3 is a correspondence representing that {device a21, device a22} corresponds to {device b21, device b22, device b23}.
[0055] Here, the devices belonging to each layer of the NW can be identified by, for example, the following identification method 1 or 2.
[0056] Method 1 for identifying devices belonging to each layer of a network: When information indicating which devices belong to each layer of a network is given in advance, the device correspondence calculation unit 203 can identify devices belonging to each layer of the network from that information. This is because, for example, in a network with a hierarchical structure (GC floors, ZC floors, etc.), the network is designed with the layers in mind, so it is obvious which device belongs to which layer, and it is possible to obtain information indicating which device belongs to which layer.
[0057] Method 2 for identifying devices belonging to each layer of the network: It is assumed that each device belonging to the first layer of the network is connected to only one device in the second layer. An example of such a case is when each device belonging to the first layer of the network is not multihomed. It is also assumed that each device belonging to the highest layer is connected to multiple devices belonging to the layer immediately below it. In this case, the device correspondence calculation unit 203 can automatically identify devices belonging to each layer of the network by following steps 1 to 3 below. It is assumed that the number of layers of the network is N.
[0058] Step 1: The device correspondence calculation unit 203 calculates the adjacency matrix of the network. Hereinafter, if a certain device i and a certain device j are connected in the network, the (i, j) element of the adjacency matrix is assumed to be 1, and if not, the (i, j) element of the adjacency matrix is assumed to be 0. Note that the network is an undirected graph, and therefore can be expressed by an adjacency matrix.
[0059] Step 2: The device correspondence calculation unit 203 determines that a device corresponding to a row (or column) having only one 1 in the adjacency matrix belongs to the first layer.
[0060] Step 3: Using the adjacency matrix, the device correspondence calculation unit 203 determines that, for i = 1, ..., N, among the devices that do not belong to any of the layers from the first layer to the i-th layer, the device that is connected to the device in the i-th layer belongs to the i+1-th layer.
[0061] As an example, the adjacency matrix U of NW-A shown in Figure 2 is shown in Figure 9. In the adjacency matrix U shown in Figure 9, the first to fifth rows and the first to fifth columns correspond to devices a1 to a5, respectively, the sixth to eighth rows and the sixth to eighth columns correspond to devices a11 to a13, respectively, and the ninth to tenth rows and the ninth to tenth columns correspond to devices a21 to a22, respectively. In this case, by the above steps 1 to 3, it is determined that devices a1 to a5 belong to the first layer, devices a11 to a13 belong to the second layer, and devices a21 to a22 belong to the third layer.
[0062] Returning to the description of FIG. 7 , following step S104, the variable correspondence relationship calculation unit 204 uses the device correspondence relationship g to calculate a variable correspondence relationship that represents the correspondence relationship between the variables of NW-A and the variables of NW-B (step S105). The variable correspondence relationship calculation unit 204 calculates, as the variable correspondence relationship, a correspondence relationship that associates sets of variables of devices belonging to the same layer between NW-A and NW-B, using the sets of devices that are associated with each other by the device correspondence relationship g. That is, for example, when i=1, 2, ..., N (where N is the number of layers of NW-A and NW-B), the variable correspondence relationship calculation unit 204 calculates, for each i, a correspondence relationship that associates a set of variables of devices belonging to the i-th layer of NW-A with a set of variables of devices belonging to the i-th layer of NW-B. However, at this time, if there is a device that has multiple types of variables among the devices belonging to the i-th layer of NW-B, the variable correspondence relationship calculation unit 204 calculates the variable correspondence relationship using only one type of variable among those multiple types of variables.
[0063] For example, device b11 of NW-B shown in Fig. 3 has two types of variables: variable Y11 and variable Z11. Similarly, devices b12 to b14 also have two types of variables: variables Y12 to Y14 and variables Z12 to Z14, respectively. In this case, the variable correspondence calculation unit 204 calculates a correspondence that associates the variable set {X11, X12, X13} of the second layer of NW-A shown in Fig. 3 with either the variable set {Y11, Y12, Y13, Y14} or {Z11, Z12, Z13, Z14}.
[0064] Here, whether the variable set {X11, X12, X13} is associated with either the variable set {Y11, Y12, Y13, Y14} or the variable set {Z11, Z12, Z13, Z14} can be determined based on the semantic similarity of the names of the variable types (variable names), etc.
[0065] For example, the variable names of variables X11, X12, and X13 are assumed to be "CPU usage" of devices a11, a12, and b13, respectively. On the other hand, the variable names of variables Y11, Y12, Y13, and Y14 are assumed to be "CPU usage" of devices b11, b12, b13, and b14, respectively, and the variable names of variables Z11, Z12, Z13, and Z14 are assumed to be "traffic volume" of devices b11, b12, b13, and b14, respectively. In this case, the semantic similarity between "CPU usage" and "CPU usage" is closer than the semantic similarity between "CPU usage" and "traffic volume," so it is determined that the variable set {X11, X12, X13} is associated with the variable set {Y11, Y12, Y13, Y14}.
[0066] Note that the semantic similarity of variable names can be calculated by any method that calculates the similarity between words or sentences using natural language processing, etc. For example, a dictionary, word2vec (cosine similarity between word vectors), BERT, ontology, etc. can be used.
[0067] Hereinafter, the variable correspondence relationship will be represented by h. For example, between NW-A shown in FIG. 2 and NW-B shown in FIG. 3, the variable correspondence relationship shown in FIG. 10 is calculated. However, in the example shown in FIG. 10, it is assumed that it is determined that the variable set {X11, X12, X13} corresponds to the variable set {Y11, Y12, Y13, Y14}. That is, h = {h1, h2, h3}. The correspondence relationship h1 is a correspondence relationship representing that {X1, X2, X3, X4, X5} corresponds to {Y1, Y2, Y3, Y4, Y5, Y6}. The correspondence relationship h2 is a correspondence relationship representing that {X11, X12, X13} corresponds to {Y11, Y12, Y13, Y14}. The correspondence relationship h3 is a correspondence relationship indicating that {X21, X22} corresponds to {Y21, Y22, Y23}.
[0068] Returning to the description of FIG. 7 , following step S105, the decision tree transfer unit 205 uses the decision tree TA of NW-A and the variable correspondence relationship h to create a decision tree (hereinafter also referred to as decision tree TB) that serves as an anomaly detection model for NW-B (step S106). Here, the variables of devices belonging to the i-th layer of NW-A correspond to the variables of the i-th layer of NW-B, but the correspondence is not one-to-one. For example, in the example shown in FIG. 10 , variables X1 to X5 correspond to variables Y1 to Y6, but the correspondence is not one-to-one, and it is not possible to determine which of variables Y1 to Y6 each variable X1 to X5 corresponds to. On the other hand, due to the assumptions satisfied by NW-A, in the decision tree TA, among the variables of devices belonging to the same layer, variables corresponding to the same type of threshold are interchangeable. Therefore, in the decision tree TB, among the variables of devices belonging to the same layer, variables corresponding to the same type of threshold must also be interchangeable. Therefore, the decision tree transfer unit 205 expresses the decision tree TA as a logical formula, extracts exchangeable patterns from the logical formula, and creates a decision tree TB by replacing the variables and thresholds of the conditional expressions included in the exchangeable patterns. Note that the binary classification decision tree TA can be equivalently expressed as a logical formula. Details of the process of creating the decision tree TB for the transfer destination NW in this step will be described later.
[0069] The output unit 206 outputs the decision tree TB to a predetermined output destination (step S107).
[0070] <<Decision Tree Creation Process for Transfer Destination NW>> Hereinafter, a process for creating a decision tree TB from the variable correspondence relationship h and the decision tree TA will be described with reference to Fig. 11. Fig. 11 is a flowchart showing an example of a decision tree creation process for a transfer destination NW.
[0071] First, the pattern extraction unit 211 extracts exchangeable patterns from the logical formula of the decision tree TA (step S201).
[0072] Here, the conditional expression expressed as X<θ with respect to the variable X is f θLet (X) be the number of variables of a device belonging to a certain i-th layer. For simplicity, it is assumed that there are two types of thresholds for the variables of the i-th layer, and the thresholds for the variables X1, ..., Xn are θi and φi. In this case, the following S i1 ~S i8 corresponds to an interchangeable pattern in which the conclusion does not change even if the variables X1, . . . , Xn are interchanged.
[0073] S i1 = f θi (X1) AND f θi (X2) AND...ANDf θi (Xn) S i2 = NOT (f θi (X1) AND f θi (X2) AND...ANDf θi (Xn)) S i3 = f θi (X1) ORf θi (X2)OR...ORf θi (Xn) S i4 = NOT (f θi (X1) ORf θi (X2)OR...ORf θi (Xn)) S i5 = f φi (X1) AND f φi (X2) AND...ANDf φi (Xn) S i6 = NOT (f φi (X1) AND f φi (X2) AND...ANDf φi (Xn)) S i7 = f φi (X1) ORf φi (X2)OR...ORf φi (Xn) S i8 = NOT (f φi (X1) ORf φi (X2)OR...ORf φi (Xn)) where AND represents logical product, OR represents logical sum, and NOT represents negation. For simplicity, the above assumes that there are two types of thresholds for the variables of the i-th layer, but even if there are three or more types, interchangeable patterns can be defined in the same way.
[0074] For example, if the decision tree TA of the NW-A shown in FIG. 2 is as shown in FIG. 4, the exchangeable patterns of the first layer of the NW-A are as follows:
[0075] S 11 = f θ1 (X1) AND f θ1 (X2) AND...ANDf θ1 (X5) S 12 = NOT (f θ1 (X1) AND f θ1 (X2) AND...ANDf θ1 (X5)) S 13 = f θ1 (X1) ORf θ1 (X2)OR...ORf θ1 (X5) S 14 = NOT (f θ1 (X1) ORf θ1 (X2)OR...ORf θ1 (X5)) Similarly, the exchangeable patterns for the second layer of NW-A shown in FIG.
[0076] S 21 = f θ2 (X11) AND f θ2 (X12) ANDf θ2 (X13) S 22 = NOT (f θ2 (X11) AND f θ2 (X12) ANDf θ2 (X13)) S 23 = f θ2 (X11)ORf θ2 (X12) ORf θ2 (X13) S 24 = NOT (f θ2 (X11)ORf θ2 (X12) ORf θ2 (X13)) Similarly, the exchangeable patterns for the third layer of NW-A shown in FIG.
[0077] S 31 = f θ3 (X21) ANDf θ3 (X22) S 32 = NOT (f θ3 (X21) ANDf θ3(X22)) S 33 = f θ3 (X21) ORf θ3 (X22) S 34 = NOT (f θ3 (X21) ORf θ3 (X22)) Note that an exchangeable pattern can be said to be a combination of conditional expressions (more precisely, conditional expressions including the same type of threshold) relating to each variable of the i-th layer of NW-A using one of AND, NAND, OR, and NOR.
[0078] Next, the surplus term deletion unit 212 deletes conditional expressions (surplus terms) that cannot be expressed in exchangeable patterns when the logical expressions of the decision tree TA are expressed in exchangeable patterns (step S202).
[0079] For example, the decision tree TA shown in FIG. 4 has p(1)=f θ3 (X21) ANDf θ3 (X22)AND(f θ1 (X1) ORf θ1 (X2) ORf θ1 (X3) ORf θ1 (X4) ORf θ1 (X5) ORf θ2 (X11)). p(1) is a logical expression that indicates abnormality (1) when true and normality (0) when false. In this case, the exchangeable pattern S i1 ~S i4 Using this, p(1) = S 31 (S 13 +f θ2 (X11)). Therefore, f θ2 (X11) is a surplus term, and the surplus term deletion unit 212 deletes p(1)=S 31 (S 13 +f θ2 (X11) to obtain the surplus term f θ2 (X11) is deleted. This allows us to use the absorption law to obtain p(1) = S 31 S 13 +S 31 = S 31 is obtained.
[0080] Here, since NW-A is composed of "layers composed of a set of exchangeable devices," ideally, "conditional expressions related to non-exchangeable variables" do not appear in the decision tree TA, and no excess terms appear. However, since the decision tree TA is created using an algorithm such as CART, there is a possibility that conditional expressions related to non-exchangeable variables may appear as excess terms. In particular, if the maximum depth of the tree can be set as a parameter of an algorithm such as CART, a node having a conditional expression related to a non-exchangeable variable may appear depending on the value. Since the decision tree TB also needs to be exchangeable between variables of devices belonging to the same layer, excess terms may be deleted from the logical formula of the decision tree TA so that the logical formula of the decision tree TB does not include excess terms.
[0081] If there are no surplus terms when the logical formula of the decision tree TA is expressed as an exchangeable pattern, the above step S202 is not executed. Also, even if there are surplus terms, the above step S202 does not have to be executed.
[0082] Next, the variable replacement unit 213 uses the variable correspondence relationship h to replace each exchangeable pattern contained in the logical formula of the decision tree TA after deleting excess terms as necessary with an exchangeable pattern of NW-B (step S203).
[0083] Here, it is assumed that the correspondence between the set of variables {X1, ..., Xn} of the devices belonging to the i-th layer of NW-A and the set of variables {Y1, ..., Ym} of the devices belonging to the i-th layer of NW-B is included in the variable correspondence relationship h. In this case, the exchangeable pattern S i1 ~S i8 are the exchangeable patterns T of the i-th layer of the following destination network. i1 ~T i8 is replaced by
[0084] T i1 = f θi (Y1) ANDf θi (Y2) AND...ANDf θi (Ym) T i2 = NOT (f θi (Y1) ANDf θi (Y2) AND...ANDfθi (Ym)) T i3 = f θi (Y1) ORf θi (Y2) OR...ORf θi (Ym) T i4 = NOT (f θi (Y1) ORf θi (Y2) OR...ORf θi (Ym)) T i5 = f φi (Y1) ANDf φi (Y2) AND...ANDf φi (Ym) T i6 = NOT (f φi (Y1) ANDf φi (Y2) AND...ANDf φi (Ym)) T i7 = f φi (Y1) ORf φi (Y2) OR...ORf φi (Ym) T i8 = NOT (f φi (Y1) ORf φi (Y2) OR...ORf φi (Ym)) Note that n=m or n≠m may also be true.
[0085] For example, when NW-A, NW-B, and the variable correspondence relationship h are as shown in FIG. 10, the exchangeable pattern S of the first layer of NW-A is 11 ~S 14 , the second layer of interchangeable patterns S 21 ~S 24 , the third layer of interchangeable patterns S 31 ~S 34 are as follows, respectively.
[0086] S 11 = f θ1 (X1) AND f θ1 (X2) AND...ANDf θ1 (X5) S 12 = NOT (f θ1 (X1) AND f θ1 (X2) AND...ANDf θ1 (X5)) S 13 = f θ1 (X1) ORf θ1(X2)OR...ORf θ1 (X5) S 14 = NOT (f θ1 (X1) ORf θ1 (X2)OR...ORf θ1 (X5)) S 21 = f θ2 (X11) AND f θ2 (X12) ANDf θ2 (X13) S 22 = NOT (f θ2 (X11) AND f θ2 (X12) ANDf θ2 (X13)) S 23 = f θ2 (X11)ORf θ2 (X12) ORf θ2 (X13) S 24 = NOT (f θ2 (X11)ORf θ2 (X12) ORf θ2 (X13)) S 31 = f θ3 (X21) ANDf θ3 (X22) S 32 = NOT (f θ3 (X21) ANDf θ3 (X22)) S 33 = f θ3 (X21) ORf θ3 (X22) S 34 = NOT (f θ3 (X21) ORf θ3 (X22)) At this time, the exchangeable pattern S of the first layer of NW-A 11 ~S 14 are the exchangeable patterns T of the first layer of NW-B as follows: 11 ~T 14 is replaced by
[0087] T 11 = f θ1 (Y1) ANDf θ1 (Y2) AND...ANDf θ1 (Y6) T 12 = NOT (f θ1 (Y1) ANDf θ1 (Y2) AND...ANDf θ1 (Y6) T13 = f θ1 (Y1) ORf θ1 (Y2) OR...ORf θ1 (Y6) T 14 = NOT (f θ1 (Y1) ORf θ1 (Y2) OR...ORf θ1 (Y6)) Similarly, the exchangeable pattern S of the second layer of NW-A 21 ~S 24 are the exchangeable patterns T of the second layer of NW-B as follows: 21 ~T 24 is replaced by
[0088] T 21 = f θ2 (Y11) ANDf θ2 (Y12) ANDf θ2 (Y13) ANDf θ2 (Y14) T 22 = NOT (f θ2 (Y11) ANDf θ2 (Y12) ANDf θ2 (Y13) ANDf θ2 (Y14) T 23 = f θ2 (Y11) ORf θ2 (Y12) ORf θ2 (Y13) ORf θ2 (Y14) T 24 = NOT (f θ2 (Y11) ORf θ2 (Y12) ORf θ2 (Y13) ORf θ2 (Y14)) Similarly, the exchangeable pattern S of the third layer of NW-A 31 ~S 34 are the exchangeable patterns T of the third layer of NW-B as follows: 31 ~T 34 is replaced by
[0089] T 31 = f θ3 (Y21) ANDf θ3 (Y22) ANDf θ3 (Y23) T 32 = NOT (f θ3 (Y21) ANDf θ3 (Y22) ANDfθ3 (Y23) T 33 = f θ3 (Y21) ORf θ3 (Y22) ORf θ3 (Y23) T 34 = NOT (f θ3 (Y21) ORf θ3 (Y22) ORf θ3 (Y23)) As a result, for example, in step S202, p(1)=S 31 If p(1)=T 31 = f θ3 (Y21) ANDf θ3 (Y22) ANDf θ3 (Y23) is obtained.
[0090] If step S202 is not executed despite the existence of a surplus term, the variable replacement unit 213 may replace the variable included in the surplus term present in the decision tree of the transfer source NW with any one of the variables of the transfer destination NW corresponding to the variable. For example, the correspondence relationship between the set of variables {X1, ..., Xn} of devices belonging to the i-th layer of NW-A and the set of variables {Y1, ..., Ym} of devices belonging to the i-th layer of NW-B is included in the variable correspondence relationship h, and θ If (X1) is a surplus term, the variable substitution unit 213 substitutes the surplus term f θ The variable X1 included in (X1) may be replaced with any one of the variables included in {Y1, ..., Ym}. In this case, the variable replacement unit 213 may replace the surplus term f θ The variable X1 contained in (X1) may be replaced with a variable Y1 having the same variable number.
[0091] Finally, the threshold replacement unit 214 replaces the threshold of each conditional expression included in the logical expression of the decision tree TA after replacing the exchangeable patterns (step S204).
[0092] Here, the threshold replacement unit 214 may replace the threshold using, for example, one of the following threshold replacement methods 1 to 3. Hereinafter, in the decision tree TA, all thresholds θi in the conditional expressions related to the variables of the devices in the i-th layer of NW-A are assumed to be equal. Furthermore, the thresholds after replacement are represented by adding "'" such as "θi'".
[0093] Threshold substitution method 1: The same value as the threshold value θi before substitution is set as the threshold value θi′ after substitution. In other words, θi′=θi. In this case, p(1)=f in step S203 above. θ3 (Y21) ANDf θ3 (Y22) ANDf θ3 When (Y23) is obtained, the logical formula of the decision tree TB is p(1) = f θ3' (Y21) ANDf θ3' (Y22) ANDf θ3' (Y23), where θ3′=θ3. Replacement method 1 is an effective method when, for example, the source network and the destination network are very similar, and makes it possible to quickly obtain a decision tree TB.
[0094] Threshold Replacement Method 2 The replaced threshold θi' is obtained by multiplying the threshold θi before replacement by the ratio between the number of devices in the i-th layer of the transfer source network and the number of devices in the i-th layer of the transfer destination network. Alternatively, the replaced threshold θi' is obtained by multiplying the threshold θi before replacement by the ratio between the number of links connected to the devices in the i-th layer of the transfer source network and the number of links connected to the devices in the i-th layer of the transfer destination network. In other words, the number of devices in the i-th layer of the transfer source network (or the number of links connected to the devices in the i-th layer of the transfer source network) is set to ai, and the number of devices in the i-th layer of the transfer destination network (or the number of links connected to the devices in the i-th layer of the transfer destination network) is set to bi. In this case, θi' = (ai / bi)θi (or θi' = (bi / ai)θi may also be used).
[0095] When the value obtained by multiplying the threshold value θi before replacement by the ratio of the number of devices in the i-th layer of the transfer source network to the number of devices in the i-th layer of the transfer destination network is set as the threshold value θi′ after replacement, p(1)=f θ3 (Y21) ANDf θ3 (Y22) ANDf θ3 When (Y23) is obtained, the logical formula of the decision tree TB is p(1) = f θ3' (Y21) ANDf θ3' (Y22) ANDf θ3'(Y23), where θ3' = (2 / 3)θ3 (or θ3' = (3 / 2)θ3 is also acceptable). Replacement method 2 is an effective method for replacing thresholds in conditional expressions related to variables that share resources among devices in the same tier, such as traffic volume.
[0096] Threshold Replacement Method 3: A value calculated by an external device, program, or the like is set as the threshold θi'. That is, when a value ci calculated by an external device, program, or the like is given, θi' = ci is set. Replacement Method 3 has the advantage that, for example, the method (algorithm) for calculating the converted threshold θi' by an external device, program, or the like can be arbitrarily changed. Note that, in order to calculate the threshold θi' by an external device, program, or the like, the threshold replacement unit 214 may provide arbitrary information (for example, the logical formula of the decision tree TA after replacing the exchangeable patterns) to the external device, program, or the like.
[0097] As a result of the above, a decision tree expressed by a logical formula after replacing the threshold value is obtained as the decision tree TB. For example, in step S203 above, p(1)=f θ3 (Y21) ANDf θ3 (Y22) ANDf θ3 When (Y23) is obtained, the decision tree TB shown in FIG. 12 is obtained.
[0098] <Another Example of Threshold Replacing Method> In the above embodiment, in the decision tree TA, the thresholds θi in the conditional expressions related to the variables of the devices in the i-th layer of the NW-A are all set to be equal, but each threshold θi may be set to be equal while allowing for a certain predetermined error. This is because the decision tree TA is created by an algorithm such as CART, and therefore each threshold θi is not necessarily strictly equal.
[0099] At this time, the threshold substitution unit 214 may substitute the threshold using, for example, the above threshold substitution method 3 or any of the following threshold substitution methods 4 to 5. Hereinafter, it is assumed that the correspondence between the set of variables {X1, ..., Xn} of the devices in the i-th layer of NW-A and the set of variables {Y1, ..., Ym} of the devices belonging to the i-th layer of NW-B is included in the variable correspondence relationship h. Also, the threshold of the conditional expression related to the variable Xk (kε{1, ..., n}) is defined as θi (k)Let's say.
[0100] Threshold value replacement method 4: When m≦n, the threshold value θi before replacement (k) (kε{1, . . . , n}) (k) Let's say.
[0101] If m>n, then for k∈{1,...,n}, the threshold θi before replacement (k) (kε{1, . . . , n}) (k) On the other hand, for k∈{n+1,...,m}, θi (1) , ..., θi (n) The average value of (θi (1) +...+θi (n) ) / n is replaced with the threshold value θi′ (k) Let's say.
[0102] Threshold Replacement Method 5: When m≦n, the ratio of the number of devices in the i-th layer of the transfer source network to the number of devices in the i-th layer of the transfer destination network is used as the threshold θi before replacement. (k) The threshold value θi′ after replacing the value multiplied by (k) Alternatively, the ratio of the number of links connected to the device in the i-th layer of the transfer source network to the number of links connected to the device in the i-th layer of the transfer destination network is set as the threshold value θi before replacement. (k) The threshold value θi′ after replacing the value multiplied by (k) That is, the number of devices in the i-th layer of the transfer source network (or the number of links connected to the devices in the i-th layer of the transfer source network) is ai, and the number of devices in the i-th layer of the transfer destination network (or the number of links connected to the devices in the i-th layer of the transfer destination network) is bi. In this case, θi' (k) = (ai / bi) θi (k) (or θi' (k) = (bi / ai) θi (k) is also acceptable.)
[0103] If m>n, then θi′ for k∈{1, . . . , n} (k) = (ai / bi) θi (k) (or θi' (k) = (bi / ai) θi (k)) where ai is the number of devices in the i-th layer of the source network (or the number of links connected to devices in the i-th layer of the source network), and bi is the number of devices in the i-th layer of the destination network (or the number of links connected to devices in the i-th layer of the destination network). On the other hand, for k∈{n+1,...,m}, θi (1) , ..., θi (n) The average value of the above is multiplied by ai / bi (or bi / ai) to obtain the threshold value θi′ after substitution. (k) Let's say.
[0104] <Summary> As described above, the transfer device 10 according to the present embodiment does not create a decision tree TB from observation data of the transfer destination NW (NW-B), but creates a decision tree TB for the transfer destination NW (NW-B) by replacing the variable names and thresholds of the conditional expressions included in the decision tree TA of the transfer source NW (NW-A). Therefore, by using the transfer device 10 according to the present embodiment, knowledge (decision tree TB) of the NW-B can be obtained without using learning data of the NW-B, that is, before the operation of the NW-B begins. In particular, when sufficient learning data of the transfer source NW has been collected and the transfer source NW is a NW that serves as a template for the transfer destination NW, a highly accurate decision tree TB can be obtained before the operation of the NW-B begins.
[0105] After the transfer device 10 according to the present embodiment creates a decision tree TB for NW-B, if observation data is collected after the start of operation of NW-B, the decision tree TB may be additionally learned as the initial state of the decision tree TB. This is expected to improve the accuracy of the decision tree TB.
[0106] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0107] 10 Transfer device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Topology information acquisition unit 202 Decision tree creation unit 203 Device correspondence calculation unit 204 Variable correspondence calculation unit 205 Decision tree transfer unit 206 Output unit 207 Learning data storage unit 211 Pattern extraction unit 212 Surplus term deletion unit 213 Variable substitution unit 214 Threshold substitution unit
Claims
1. A transfer device for transferring a first decision tree related to a first network having a hierarchical structure to a second network having a hierarchical structure consisting of the same number of levels as the first network, comprising: a first correspondence calculation unit that calculates, for each level of the hierarchical structure, a first correspondence representing a correspondence between a first device belonging to the level of the first network and a second device belonging to the level of the second network; a second correspondence calculation unit that calculates, for each level of the hierarchical structure, a second correspondence representing a correspondence between a first variable representing an attribute of an observation value observed at a first device belonging to the level and a second variable representing an attribute of an observation value observed at a second device belonging to the level, based on the first correspondence; and a transfer unit that creates a second decision tree related to the second network by replacing a first variable of a conditional expression set for a node included in the first decision tree with a second variable based on the second correspondence.
2. The transfer device of claim 1, wherein the transfer unit creates the second decision tree by replacing, for each level of the hierarchical structure, a first variable included in the conditional expression relating to a first variable of a first device belonging to the level with a second variable of a second device belonging to the level.
3. The transfer device according to claim 1 or 2, wherein the transfer unit creates the second decision tree by further replacing a threshold value for a first variable included in the conditional expression with a predetermined threshold value.
4. A transfer method for transferring a first decision tree related to a first network having a hierarchical structure to a second network having a hierarchical structure consisting of the same number of levels as the first network, the transfer method being executed by a computer, comprising: a first correspondence calculation step of calculating, for each level of the hierarchical structure, a first correspondence representing a correspondence between a first device belonging to the level of the first network and a second device belonging to the level of the second network; a second correspondence calculation step of calculating, for each level of the hierarchical structure, based on the first correspondence, a second correspondence calculation step of calculating, for each level of the hierarchical structure, a second correspondence representing a correspondence between a first variable representing an attribute of an observation value observed in a first device belonging to the level and a second variable representing an attribute of an observation value observed in a second device belonging to the level; and a transfer step of creating a second decision tree related to the second network by replacing a first variable of a conditional expression set in a node included in the first decision tree with a second variable based on the second correspondence.
Citation Information
Patent Citations
Creation method of decision tree in industrial Ethernet fault diagnosis method
CN104506340A
Power communication network fault detection method based on transfer learning
CN110995475A
Method and device for diagnosing grounding fault of power distribution network
CN114355240A
Network anomaly detection method and device
CN115242600A