Event identification device, event identification method, and program
Patent Information
- Application Number
- PCT/JP2025/005427
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2026-08-27
Smart Images

Figure JP2025005427_27082026_PF_FP_ABST
Abstract
Description
Event identification device, event identification method, and program
[0001] This disclosure relates to an event identification device, an event identification method, and a program.
[0002] Under the current trend of digital transformation (DX), pioneering ICT (Information and Communication Technology) technologies are being demonstrated. In the future, it is expected that there will be an increase in large-scale solution projects that are limited to specific regions or time periods, such as smart cities and special zones like the World Expo. Such large-scale solution projects require a network as infrastructure, and it is anticipated that there will be a need for its maintenance and operation.
[0003] Research is progressing on the operation and management of managed networks (MNs) in such large-scale solutions. Specifically, Network Operation Centers (NOCs), which are organizations within telecommunications carriers, remotely monitor and manage MNs across the country.
[0004] Furthermore, as an example of operational management, a second-time automated intervention method using the Case Based Reasoning (CBR) methodology has been proposed (Non-Patent Literature 1). CBR is an inference method that, when there is a new problem to be solved, searches a database (DB) that organizes and stores known (past) problems (obstacles) and their solutions (measures) in a common format, extracts similar cases, and modifies their solutions to obtain a solution to the new problem. In this case, the obtained solution is stored in the DB as a new case along with the problem.
[0005] The second-time automatic remediation system is based on the CBR (Continuous Breakdown Record) and, for the first time a fault occurs, the NOC (Network Operation Control) operator takes remediation measures and records the details. If a similar fault occurs again, the MN (Network Management) remediation execution device automatically takes remediation measures. In this case, the operator determines whether the newly occurring type of fault (hereinafter, "type of fault" may be referred to as "fault type") is similar to the type of fault that has occurred in the past.
[0006] Furthermore, in CBR, a single failure can be represented by a single vector of metrics (traffic volume, CPU usage, etc.) generated at that time. Then, when a new failure occurs in a given environment T, the distance (similarity) between the vectors of the new failure and past failures is calculated to determine whether the new failure is close to (similar to) past failures. In this case, since the metric value at a given point in time rarely perfectly matches the past value, it is necessary to determine the distance (similarity) to past metrics with a certain "range".
[0007] Takaaki Moriya, Takashi Mukai, Manabu Nishio, Ai Tsunoda, and Ken Kanishima, Proposal of a measure generation method using past performance in future managed networks, IEICE Technical Report ICM2024-2, 2024.5
[0008] However, the distance mentioned above is a relative comparison between a new failure and a past failure. Moreover, if a past failure has only occurred once or a few times in the current environment T, then the comparison is only with that number of failures. Therefore, it is unreliable to determine that the same type of failure as a past failure is occurring based solely on the distance between vectors based on metric values. In other words, the reliability of the "width" of the distance is low. Consequently, it is difficult to say that the distance between vectors based solely on metric values reflects the certainty (plausibility, likelihood) that the same type of failure as a past failure is truly occurring now, and there is a risk of low identification accuracy.
[0009] This disclosure is made in view of the circumstances described above and aims to improve the accuracy of determining the type of current failure.
[0010] To achieve the above objective, this disclosure provides an action execution device for executing action in response to an event occurring in a communication device in a predetermined network, comprising: a determination unit that determines whether there is one or more past predetermined metric information where the distance to metric information indicating the usage status of the predetermined communication device as a vector is less than or equal to a threshold at the time the event that is the target of determination for the predetermined communication device occurs; a conversion unit that, if there is past predetermined metric information that is less than or equal to the threshold, converts the metric information related to the event that is the target of determination and the predetermined metric information using a metric information conversion function from a first environment to a second environment; and a machine learning model that has been machine-learned in the second environment to execute the converted determination The device for executing countermeasures includes: an estimation unit that estimates the probability of occurrence for each type of event with respect to metric information relating to an elephant event, and estimates the probability of occurrence for each type of event with respect to the predetermined metric information after conversion; and an identification unit that identifies a predetermined type of event for which the probability of occurrence of the predetermined event type is the highest after conversion, identifies a specific changed metric information for which the probability of occurrence of the predetermined event type is the highest among the predetermined metric information after conversion, and identifies the predetermined event type for the event that is the target of determination if the probability of occurrence of the predetermined event type for which the conversion relates to the metric information after conversion is equal to or greater than a predetermined value relative to the probability of occurrence of the predetermined metric information after conversion.
[0011] As explained above, this disclosure has the effect of improving the accuracy of determining the type of current failure.
[0012] This is an overall configuration diagram of the communication system according to the embodiment. This is an electrical hardware configuration diagram of the action execution device according to the embodiment. This is a diagram showing the functional configuration of the action execution device according to the embodiment. This is a conceptual diagram showing labeled data (training data). This is a sequence diagram showing a series of processes for determining the type of fault. This is a flowchart showing the fault type determination process. This is a flowchart showing the fault type determination process.
[0013] Hereinafter, embodiments of the present invention will be described based on the drawings. Note that the present invention is not limited to the embodiments shown below, and various modifications are possible without departing from the technical idea of the present invention.
[0014] [Overall Configuration of Embodiment] First, the overall configuration of the communication system according to the embodiment will be described using FIG. 1. FIG. 1 is an overall configuration diagram of the communication system according to the embodiment.
[0015] As shown in FIG. 1, the communication system 10 of the present embodiment is constructed by a communication device 20, a monitoring and collection server 30, a measure execution device 50, and an input device 80.
[0016] The communication device 20, the monitoring and collection server 30, and the measure execution device 50 are installed in the special zone X such as a smart city or an expo. The communication device 20, the monitoring and collection server 30, and the measure execution device 50 can communicate with each other via a LAN (Local Area Network) 90 included in the MN constructed within the special zone W. Also, the communication device 20, the monitoring and collection server 30, and the measure execution device 50 can communicate with the input device 80 via the LAN 90 and a WAN (Wide Area Network) 100 outside the special zone W.
[0017] The input device 80 is installed in a NOC (Network Operation Center) which is an organization of a communication carrier or the like. The input device 80 is operated by the operator Y of the NOC.
[0018] In FIG. 1, one input device 80 is connected to each LAN 90 of a plurality of special zones W via the WAN 100, and the measure execution device 50 within each special zone W can be remotely operated. Note that the LAN 90 and the WAN 100 are an example of a communication network. A partial connection form of the communication network may be either wireless or wired.
[0019] <Communication Device>The communication device 20 is a target for measures in case of a failure occurring within the LAN 90, and is, for example, a server, a modem, a router, a switch, a hub, a communication terminal (such as a PC). In FIG. 1, for convenience of explanation, one communication device 20 is shown, but this communication device 20 means a single or multiple communication devices.
[0020] <Monitoring and Collection Server>The monitoring and collection server 30 monitors the communication devices 20 within the LAN 90 and collects metric information (also referred to as "metrics"). The monitoring and collection server 30 includes an SNMP (Simple Network Management Protocol) manager and a log management server. In the case where some failure occurs in the communication device 20, the communication device 20 may use SNMP Trap or the like to transmit alarm information to the monitoring and collection server 30. Further, the monitoring and collection server 30 may have a function of transmitting alarm information to the measure execution device 50.
[0021] "Metric information" is an example of information (status information) indicating the usage (operation) status of the communication device 20 when a failure occurs, and is information that can be used to characterize the status. For example, it includes traffic volume, memory usage, OS (Operating System) version, the time when the failure occurred (or was detected), etc.
[0022] Also, the metric information can be regarded as a multi-dimensional vector for each time (snapshot). Note that for metrics other than numbers, the monitoring and collection server 30 quantifies them by, for example, making them categorical variables.
[0023] Further, the monitoring and collection server 30 has a metric DB (Data Base) 31. In this metric DB 31, at least the type of failure that occurred when this metric information was collected by the monitoring and collection server 30 (including the type of failure that occurred for the first time in environment T), the metric information at this time, and the occurrence time of this failure (including the date and time of occurrence) are managed in association with each other.
[0024] Furthermore, "alarm information" is information that notifies that a failure has occurred, and includes information such as the time the failure occurred or was detected, the source of the alarm, and the content of the alarm. Alarm information may also include detection logs and metrics.
[0025] <Measure Execution Device> The measure execution device 50 plays a central role in this embodiment and is a device that executes recovery measures in response to a failure (event) that occurs in the communication device 20 in a predetermined network such as LAN 90.
[0026] In this embodiment, for the first time a type of failure occurs, operator Y manually takes recovery measures to the communication device 20 via the input device 80. However, if the same type of failure as the first one occurs again, the recovery action execution device 50 within the LAN 90 will autonomously (spontaneously) take recovery measures to the communication device 20. Furthermore, if the recovery action execution device 50 is unable to take measures, or if the communication device 20 does not recover even after taking measures, the NOC's input device 80 can transmit information regarding the measures to be taken for the communication device 20 to the recovery action execution device 50, thereby allowing the recovery action execution device 50 to take measures remotely.
[0027] In Figure 1, an RCA (Root Cause Analysis) server may be installed either inside or outside the LAN 90. The RCA server analyzes alarm information and other data acquired from the monitoring and data collection server 30 through root cause analysis, deriving information including fault location information and fault type information, and transmits this information to the corrective action execution device 50 as fault information. The type and format of information included in the fault information will vary depending on the specifications of the RCA server 70. "Fault location information" indicates the communication device 20 where the fault occurred or the specific faulty part of the communication device 20. "Fault type information" is derived from the alarm content in the alarm information and indicates the type of fault, such as a buffer overflow error or link down.
[0028] <Input Device> The input device 80 periodically (for example, every 30 seconds) acquires metric information of the communication device 20 from the monitoring and collection server 30. If operator Y determines that a failure has occurred in the communication device 20 based on the metric information, the input device 80 remotely controls the communication device 20 to perform recovery measures for the first predetermined failure (event) that occurred in the communication device 20.
[0029] [Hardware Configuration] Next, the electrical hardware configuration of the remediation execution device 50 will be explained using Figure 2. Figure 2 is an electrical hardware configuration diagram of the remediation execution device according to the embodiment.
[0030] As shown in Figure 2, the action execution device 50 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a processor 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are all interconnected by a bus 1010 such as a data bus.
[0031] The program that enables processing on the computer is provided on a recording medium 1001, such as a CD-ROM or memory card. When the recording medium 1001 containing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001; it may also be downloaded from another computer via the communication network 100. The auxiliary storage device 1002 stores the installed program as well as necessary files and data.
[0032] When a program startup command is received, the memory device 1003 reads the program from the auxiliary storage device 1002 and stores it. The processor 1004 implements the functions related to the memory device 1003 according to the program stored in the memory device 1003. The processor 1004 may include not only a CPU (Central Processing Unit) but also a GPU (Graphics Processing Unit).
[0033] The interface device 1005 is used as an interface for connecting to a communication network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like, programmed by the user. The input device 1007 consists of a keyboard and mouse, buttons, or a touch panel, and is used to input various operation instructions. The output device 1008 outputs the calculation results to the outside.
[0034] If the communication device 20 is a server or the like, the monitoring and collection server 30, input device 80, and RCA server have the same hardware configuration as the action execution device 50, so their explanation will be omitted.
[0035] [Functional Configuration of the Action Execution Device] Next, the functional configuration of the action execution device 50 will be explained using Figure 3. Figure 3 is a diagram relating to an embodiment and mainly shows the functional configuration of the action execution device.
[0036] As shown in Figure 3, the remediation execution device 50 includes a transmitting / receiving unit 51, a distance determination unit 52, a metrics conversion unit 53, a fault occurrence probability estimation unit 54, a fault type identification unit 55, and a remediation execution unit 56. Each of these units is a function implemented by instructions from the processor 1004 in Figure 2 based on a program.
[0037] Furthermore, the action execution device 50 has a storage unit 60 constructed by an auxiliary storage device 1002 or a memory device 1003. The storage unit 60 stores an action information DB 62. The storage unit 60 also stores a conversion function g acquired from the input device 80 and a trained machine learning model M.
[0038] <Measures Information DB> The measures information DB62 manages, at a minimum, the type of failure and the details of the recovery measures taken for that type of failure, in association with each other.
[0039] <Machine Learning Model> The machine learning model M is a neural network (NN) that takes metric information (vectors) as input data and outputs the type of failure (one-hot vector) in a pre-trained environment (hereinafter referred to as "environment S") which is a pre-built or arbitrary MN.
[0040] Environment S can also be described as a simulated environment. Ideally, Environment S should closely simulate the real environment (hereinafter referred to as "Environment T").
[0041] Furthermore, the input device 80 collects metric information from the environment S using SNMP (Simple Network Management Protocol), etc., and prepares labeled data (training data) as shown in Figure 4. A vector listing the m types of metric values that can be obtained in the environment S at the same time is denoted as metric z. For example, z = (z1: traffic amount of each port, z2: CPU usage rate, ... z m (A vector listing memory usage, etc.) T Each sample of z is assigned a single label. The labels are: normal (failure type is "normal"), failure (failure type) A, failure (failure type) B, etc. This training data consists of multiple samples of z (see the vertical direction of the table in Figure 4). As mentioned above, "normal" is treated the same as a failure type.
[0042] The input device 80 then uses the data shown in Figure 4 to learn the classification of the types of failures using a neural network (NN). The learned machine learning model M, which is a neural network f, takes metrics (vectors) z in the environment S as input and interprets this to output a vector representing the probability of each type of failure occurring, as shown below. The output is a vector with the number of elements = "number of types of failures + 1" (where "1" represents the normal class). For convenience, the learned neural network will be represented as a function f from now on, i.e., y = f(z).
[0043] f(z) = (Probability of normal, probability of A, probability of B, ...) TFurthermore, for convenience, from now on, we will refer to the scalar representing the probability that the output vector y of f is of type A as f. A This is denoted as (z). The output layer uses the softmax function. As a result, the sum of the outputs is 1, and each output is interpreted as representing the probability that the input z belongs to each class.
[0044] <Conversion Function> The conversion function g is generated by the input device 80 and then transmitted from the input device 80 to the action execution device 50. The input device 80 collects samples of metrics under normal conditions in environment T of the MN to be verified and generates a conversion function g for metrics from environment T to environment S.
[0045] Furthermore, the dimensionality of the metrics in environment S and the metrics in environment T are the same, and the type of metric for the i-th element of the former and the i-th element of the latter are the same (where i is a positive number). For example, it is assumed that the same model is used in both environment S and environment T.
[0046] Here, we will explain one example of how to create a transformation function. Here, we assume that both z and x follow a multivariate normal distribution, and that the transformation function g from x to z is an affine transformation (scaling and translation). Then, we will determine the coefficients of the affine transformation from the obtained sample data of z and x. Further details will be explained below.
[0047] Here, we denote the metrics (D-dimensional vector) in environment S as z, and the metrics (D-dimensional vector) in environment T as x.
[0048] (Assumption) Assume that x and z each follow a multivariate normal distribution. Although this is a strong assumption, it has the advantage of being able to take into account the correlation between elements of x (or z) (for example, between CPU usage and traffic volume).
[0049] (Key point) Consider x to be related to z by translating and scaling it.
[0050] (Method of creation) S1: Assume that z follows a multivariate normal distribution with a D-dimensional mean vector μ and a D×D covariance matrix Σ.
[0051] From the data collected in environment S, μ and Σ are estimated using the maximum likelihood method.
[0052] Similarly, S2: x is assumed to follow a multivariate normal distribution with mean vector μ' and covariance matrix Σ', and the mean vector μ' and covariance matrix Σ' are estimated using the maximum likelihood method from the normal data in environment T.
[0053] S3: Assume that the following affine transformation relationship holds between x and z.
[0054] S4: From the formula for linear transformation of multivariate normal distributions (quoted from https: / / healy.econ.ohio-state.edu / kcb / Ma103 / Notes / Lecture11.pdf), the following system of equations holds.
[0055] Since μ, μ', Σ, and Σ' in S5 are obtained from S1 and S2, we solve the system of equations in S4 for the following coefficients.
[0056] S6: From the above, the following transformation function g is obtained.
[0057] Note that the method for creating the transformation function g is not limited to the method of finding the coefficients of this affine transformation. For example, it could be a simple method of transforming the scale based on the maximum and minimum values of corresponding elements of z and x. Alternatively, it could be a method of finding the distribution (histogram) of each element of z and x from the sample data. In this case, it is assumed that the corresponding elements of z and x are related by z = ax + b, and a and b are found from the sample data.
[0058] Furthermore, if the transformation method is the same as that of the transformation function g, a transformation table can be used instead of a function.
[0059] <Functional Configuration> Next, the functional configuration of the action execution device 50 will be explained using Figure 3.
[0060] (Transmitting / receiving unit) The transmitting / receiving unit 51 communicates data with the input device 80 via the LAN 90 and WAN 100.
[0061] (Distance Determination Unit) The distance determination unit 52 determines whether there is one or more past predetermined metric information where the distance to the metric information showing the usage status of the predetermined communication device 20 as a vector at the time of a failure that is the target of determination for the predetermined communication device 20 is less than or equal to a threshold th. The distance determination unit 52 may also act as a similarity determination unit and determine whether there is one or more past predetermined metric information where the similarity to the information showing the usage information of the predetermined communication device 20 as a vector at the time of a failure that is the target of determination for the predetermined communication device 20 is greater than or equal to a threshold.
[0062] (Metrics conversion unit) When there is predetermined past metric information that is below the threshold th, the metrics conversion unit 53 uses a metric information conversion function g from environment T (an example of a first environment) to environment S (an example of a second environment) to convert the metric information related to the fault to be judged and the predetermined metric information.
[0063] (Fault Occurrence Probability Estimation Unit) The fault occurrence probability estimation unit 54 uses a machine learning model M trained in the environment S to estimate the occurrence probability for each type of fault with respect to the converted metric information related to the faults to be judged, and also estimates the occurrence probability for each type of fault with respect to predetermined converted metric information.
[0064] (Fault Type Identification Unit) The fault type identification unit 55 identifies a predetermined fault type that has the highest probability of occurrence related to the converted metric information subject to judgment, and also identifies a specific converted metric information that has the highest probability of occurrence of a predetermined event type among the converted predetermined metric information. Furthermore, the fault type identification unit 55 identifies a predetermined fault type for the fault subject to judgment if the probability of occurrence related to the converted metric information for a predetermined event type is equal to or greater than a predetermined value v relative to the probability of occurrence related to the converted predetermined metric information.
[0065] (Action Execution Unit) The action execution unit 56 searches the action information DB 62 using a predetermined fault type identified by the fault type identification unit 55 as a search key, and reads the corresponding action information. The action execution unit 56 then executes the action content according to the command sequence contained in the action information, thereby performing recovery measures on the communication device 20.
[0066] [Processing according to the embodiment] Next, the processing according to this embodiment will be described with reference to Figures 5 to 7.
[0067] Figure 5 is a sequence diagram showing a series of processes for determining the type of failure.
[0068] S11: The input device 80 generates a machine learning model M by performing machine learning.
[0069] S12: The input device 80 creates a conversion function g.
[0070] S13: The input device 80 transmits the trained machine learning model M generated in process S11 and the conversion function g created in process S12 to the action execution device 50. As a result, the transmitting / receiving unit 51 of the action execution device 50 receives the trained machine learning model M and the conversion function g and stores them in the storage unit 60.
[0071] Note that the input device 80 does not generate the machine learning model M; instead, another computer, such as a computer for generating machine learning models, may generate it. In this case, this other computer transmits the machine learning model M to the action execution device 50 either directly to the action execution device or via the input device 80. Similarly, the input device 80 does not generate the conversion function g; instead, another computer, such as a computer for creating conversion functions, may generate it. In this case, this other computer transmits the conversion function g to the action execution device 50 either directly to the action execution device or via the input device 80.
[0072] S14-1: The monitoring and collection server 30 monitors each communication device 20 in the LAN 90 and transmits the metrics obtained to the input device 80 at predetermined intervals (for example, 30 seconds). Each of these metrics includes environmental identification information (special zone identification information) for identifying the environment T (special zone W), and communication device identification information (monitoring target identification information) for identifying the communication device 20.
[0073] S14-2: The monitoring and collection server 30 monitors each communication device 20 in the LAN 90 and transmits the metrics obtained to the action execution device 50 at predetermined intervals (for example, 30 seconds). Each of these metrics includes environmental identification information (special zone identification information) for identifying the environment T (special zone W), and communication device identification information (monitoring target identification information) for identifying the communication device 20.
[0074] S15: The input device 80 detects whether any malfunction has occurred in a predetermined communication device 20 based on the metrics of each communication device 20, when the operator Y is an Ops (Operation system), etc. Alternatively, as shown in Non-Patent Document 1, the action execution device 50 may receive the above-mentioned alarm information from the monitoring and collection server 30, identify that some malfunction has occurred in a predetermined communication device 20 among the communication devices 20, and the input device 80 may receive notification from the action execution device 50 that some malfunction has occurred in the predetermined communication device 20. The type of malfunction is not specified.
[0075] Then, operator Y determines the type of failure that occurred in environment T, which is LAN 90, and whether it is the first time the failure has occurred in environment T (human judgment). In this case, operator Y may make the judgment using a trained machine learning model M in the input device 80, or may make the judgment based on their own experience. If it is a first-time failure A, operator Y denotes the current time t as A1 and enters metrics x into the memory device 1003 in the input device 80. A1 This information is stored in memory (human processing). The same processing is performed for failures B, C, etc.
[0076] S16: The input device 80 transmits information to the monitoring and collection server 30 about the type of fault that occurred for the first time in environment T. This information includes the time of occurrence (including the date and time of occurrence) and metric x. tIt is included. As a result, the monitoring and collection server 30 receives information on the type of the newly occurred failure. Also, the monitoring and collection server 30 stores and manages in association with the metrics DB 31 the type of the newly occurred failure, the metrics information at this time, and the occurrence time (including the date and time of occurrence) of this newly occurred failure.
[0077] S17: On the other hand, the measure execution device 50 uses the current metric x of the failure to be determined of each communication device 20 collected by the transmission / reception unit 51 every predetermined time (for example, 30 seconds) by the above-described process S14-2 to perform a failure type determination process. Here, the failure type determination process will be described using FIGS. 6 and 7. FIGS. 6 and 7 are flowcharts showing the failure type determination process. t to perform a failure type determination process. Here, the failure type determination process will be described using FIGS. 6 and 7. FIGS. 6 and 7 are flowcharts showing the failure type determination process.
[0078] S31: The distance determination unit 52 searches the metrics DB 31 of the monitoring and collection server 30 via the transmission / reception unit 51 using the current metric x to be determined as a search key to determine whether there is a past metric x t whose distance from the metric x to be determined is equal to or less than a threshold th. Here, x t is a vector of metrics when failure X (failure X = failure A, failure B,...) occurred for the i-th time in the past. If there is no metric x Xi whose distance from the metric x to be determined is equal to or less than the threshold th (NO), the process proceeds to process S38 described later. Xi is a vector of metrics when failure X (failure X = failure A, failure B,...) occurred for the i-th time in the past. And if there is no metric x t whose distance from the metric x to be determined is equal to or less than the threshold th (NO), the process proceeds to process S38 described later. Xi If there is no metric x whose distance from the metric x to be determined is equal to or less than the threshold th (NO), the process proceeds to process S38 described later.
[0079] S32: In process S31, if there is a metric x t whose distance from the metric x to be determined is equal to or less than the threshold th (YES), the distance determination unit 52 selects a predetermined metric x Xi whose distance from the metric x to be determined is equal to or less than the threshold th. t and the metric x to be determined, Xi selects a predetermined metric x whose distance from the metric x to be determined is equal to or less than the threshold th.
[0080] S33: The metrics conversion unit 53 uses the conversion function g to convert the metric x to be determined tThis involves determining the metric X in the pre-trained environment S (an example of the second environment) from environment T (an example of the first environment), which is special zone W. t Similarly, the metrics conversion unit 53 uses the conversion function g to convert a predetermined metric x in the environment T of the special zone W. Xi A predetermined metric X in environment S Xi Convert to this.
[0081] S34: The failure probability estimation unit 54 uses the trained machine learning model M to calculate each of the converted metrics (X t , X Xi Based on this, each metric after conversion (X t , X Xi For this, the probability of occurrence for each type of failure (each type of failure) is estimated. For example, Metric X t Regarding this, the probability of normal occurrence is estimated to be 5%, the probability of failure type A is 80%, the probability of failure type B is 10%, the probability of failure type C is 5%, and so on. Also, for example, metric X A1 Regarding this, the probability of normal occurrence is estimated to be 5%, the probability of failure type A is 70%, the probability of failure type B is 20%, the probability of failure type C is 5%, and so on. Also, for example, metric X B1 Regarding this, the estimated probabilities are as follows: 5% for normal occurrence, 40% for type A, 50% for type B, and 5% for type C.
[0082] S35: The fault type identification unit 55 determines the metric X to be judged after conversion. t Identify a specific type of failure (for example, failure type A (link down, etc.)) that has the highest probability of occurring.
[0083] S36: Proceeding to Figure 7, the fault type identification unit 55 converts the predetermined metric X Xi Among these, the specific metric after conversion (hereinafter referred to as metric X) has the highest probability of occurring a predetermined type of failure (for example, failure type A (link down, etc.)). A1 Identify (the one to be used as)
[0084] For example, for type A, if i = 1 to 3, then metric X A1 The probability of type A in this case is 70%, metric XA2 The probability of type A in this case is 40%, metric X A3 If the probability of failure type A occurring is 50%, the failure type identification unit 55 determines the specific converted metric X that has the highest probability of failure type A occurring. A1 Identify.
[0085] S37: The fault type identification unit 55 determines the metric X to be determined after conversion for a predetermined fault type. t The probability of occurrence related to the converted specific metric X A1 It is determined whether the probability of occurrence related to this is greater than or equal to a predetermined value v.
[0086] For example, MetricsX t The probability of failure type A occurring is 80%, and specific metric X A1 The probability of failure type A occurring is 70%, and specific metric X A1 If the predetermined value v for the probability of occurrence related to this is ±0%, then 80(%)≧70(%)±0(%), and therefore metric X t The probability of occurrence related to the converted specific metric X A1 It is determined that the probability of occurrence related to this is greater than or equal to a predetermined value v (YES).
[0087] Also, MetricsX t The probability of occurrence related to this is 80%, and for a specific metric X A1 The probability of occurrence related to this is 85%, and specific metric X A1 If the predetermined value v for the probability of occurrence related to this is -10%, then 80(%) ≥ 85(%) - 10(%), so metric X t The probability of occurrence related to the converted specific metric X A1 It is determined that the probability of occurrence related to this is equal to or greater than a predetermined value (YES).
[0088] Furthermore, the fault type identification unit 55 does not use the predetermined value v described above, but simply determines the metric X to be judged after conversion for a predetermined fault type. t The probability of occurrence related to the converted specific metric X A1 It is also possible to determine whether the above is true, but this determination is included in the processing for the case where the predetermined value v = ±0 as described above.
[0089] S38: In processing S37, the metric X to be judged after conversion. t The probability of occurrence related to the converted specific metric X A1 If the probability of occurrence related to the above is greater than or equal to a predetermined value v (YES), the fault type identification unit 55 identifies a predetermined fault type (for example, fault type A) for the fault that is the subject of the determination.
[0090] S39: The fault type identification unit 55 transmits the metrics x of the target to be determined before conversion by the conversion function g to the metrics DB 31 of the monitoring and collection server 30 via the transmitting and receiving unit 51. t The system stores the judgment result (a predetermined type of failure (e.g., failure type A), the probability of occurrence (e.g., 80%)) in association with the data.
[0091] S40: On the other hand, in processing S37, the metric X to be judged after conversion t The probability of occurrence related to the converted specific metric X A1 If the probability of occurrence related to is not greater than or equal to a predetermined value v (NO), or if, in process S31, the metric x to be judged t If there are no metrics whose distance to the threshold th is less than or equal to (NO), the fault type identification unit 55 does not identify (or cannot identify) the fault type for the fault being judged.
[0092] S18: Next, returning to Figure 5, the transmitting / receiving unit 51 transmits the fault type determination result as a response to process S16. If a predetermined fault type is identified in process S38, the determination result indicates, for example, that "the fault type is link down." If the fault type is not identified (or cannot be identified) in process S40, the determination result indicates, for example, that "the fault type is unknown (or cannot be identified)."
[0093] S19: If a predetermined type of failure is identified in processing S38, the action execution unit 56 refers to the action information DB 62, reads information on the recovery measures corresponding to the identified type of failure, and executes the recovery measures on the communication device 20, for example, as shown in Non-Patent Document 1.
[0094] [Main Effects of the Embodiment] As described above, according to this embodiment, when a type of failure occurs for the first time in environment T, operator Y manually takes recovery measures to the communication device 20 from the input device 80. Subsequently, if the same type of failure as the first failure occurs again, the action execution device 50 autonomously (spontaneously) takes recovery measures to the communication device 20 within MN (LAN 90). In this case, first, as a narrowing down of the comparison target, the distance determination unit 52 selects a predetermined metric whose distance to the metric to be determined is less than or equal to a threshold. Then, the failure type identification unit 55 identifies a specific metric after conversion among the predetermined metrics after conversion by the conversion function g that has the highest probability of occurrence of a predetermined failure type. Furthermore, the failure type identification unit 55 identifies a predetermined failure type for the failure of the metric to be determined if the probability of occurrence related to the converted metric for a predetermined failure type is greater than or equal to a predetermined value relative to the probability of occurrence related to the converted predetermined metric. This has the effect of improving the accuracy of determining the type of current failure, even if the number of failures that have occurred in the current environment T in the past is small, such as only one to a few times.
[0095] [Supplement] The present invention is not limited to the embodiments described above, and may also have configurations or processes (operations) as shown below.
[0096] (1) The action execution device 50 can also be implemented by a computer and a program, but this program can be recorded on a (non-temporary) recording medium or provided via a network such as the WAN 100.
[0097] (2) The hardware processor 1004 may be single or multiple.
[0098] (3) In the above embodiment, "fault" is an example of "event". Therefore, fault information is an example of event information, fault occurrence probability estimation unit 54 is an example of an event occurrence probability estimation unit, and fault type identification unit 55 is an example of an event type identification unit.
[0099] 10 Communication system 30 Monitoring and collection server 31 Metrics DB (Example of Metrics Management Unit) 50 Action Execution Device 51 Transmitting / receiving unit (Example of Transmitting Unit, Example of Receiving Unit) 52 Distance Determination Unit (Example of Determination Unit) 53 Metrics Conversion Unit (Example of Conversion Unit) 54 Failure Occurrence Probability Estimation Unit (Example of Estimation Unit) 55 Failure Type Identification Unit (Example of Identification Unit) 56 Action Execution Unit (Example of Action Unit, Example of Execution Unit) 60 Storage Unit 62 Action Information DB (Example of Action Information Management Unit) 80 Input Device M Trained machine learning model g Conversion function
Claims
1. An event identification device for identifying the type of event that occurs in a communication device in a predetermined network, comprising: a determination unit that determines whether there is one or more past predetermined metric information at the time of the occurrence of an event that is the target of determination for the predetermined communication device, the distance to the metric information that shows the usage status of the predetermined communication device as a vector is below a threshold; a conversion unit that, if there is past predetermined metric information that is below the threshold, converts the metric information related to the event to be determined and the predetermined metric information using a metric information conversion method function from a first environment to a second environment; and an estimation unit that estimates the occurrence probability for each type of event with respect to the converted metric information related to the event to be determined, and estimates the occurrence probability for each type of event with respect to the converted predetermined metric information, An event identification device having: an identification unit that identifies a predetermined event type for which the probability of occurrence of the converted metric information to be judged is the highest, identifies a specific changed metric information for which the probability of occurrence of the converted predetermined event type is the highest among the converted predetermined metric information, and identifies the predetermined event type for the event to be judged if the probability of occurrence of the converted metric information to be judged for the predetermined event type is equal to or greater than a predetermined value relative to the probability of occurrence of the converted predetermined metric information.
2. The event identification device according to claim 1, wherein the machine learning model is a neural network that takes the metric information as input data and the event type as output data.
3. An event identification method performed by a computer that takes action in response to an event occurring in a communication device in a predetermined network, wherein the computer includes: a determination process that determines whether there is one or more past predetermined metric information where the distance to the metric information indicating the usage status of the predetermined communication device as a vector is below a threshold at the time the event that is the target of determination for the predetermined communication device occurs; a conversion process that, if there is past predetermined metric information below a threshold, converts the metric information related to the event to be determined and the predetermined metric information using a metric information conversion function from a first environment to a second environment; and an estimation process that estimates the occurrence probability for each type of event with respect to the converted metric information related to the event to be determined and the occurrence probability for each type of event with respect to the converted predetermined metric information, An event identification method that includes: identifying a predetermined event type for which the probability of occurrence of the metric information to be judged after conversion is the highest; identifying a specific changed metric information for which the probability of occurrence of the predetermined event type is the highest among the predetermined metric information after conversion; and, if the probability of occurrence of the metric information to be judged after conversion for the predetermined event type is less than or equal to a predetermined value relative to the probability of occurrence of the predetermined metric information after conversion, performing an identification process to identify the predetermined event type for the event to be judged.
4. A program that causes a computer to perform the method described in claim 3.