Failure management apparatus and failure management method

The fault management device employs a two-stage analysis process with trained models to efficiently analyze fault causes in complex communication networks, reducing the operational burden by minimizing the required information and resource usage.

WO2026013939A1PCT designated stage Publication Date: 2026-01-15MITSUBISHI ELECTRIC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/036090
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2024-10-09
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

The increasing complexity of communication networks, including 5G systems and local 5G systems, has led to a rise in the diversity of failure factors and index values, necessitating a large amount of information for accurate fault analysis, which increases operational burden.

Method used

A fault management device and method that employs a two-stage analysis process using a network information acquisition unit, a network information control unit, and a fault cause analysis unit to reduce the amount of information needed by narrowing down candidate fault causes and determining their occurrence through trained models, specifically using unsupervised and supervised learning techniques.

Benefits of technology

Reduces the amount of information required for fault analysis in communication networks, minimizing resource usage such as communication bandwidth, computational resources, and memory capacity while accurately identifying fault causes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024036090_15012026_PF_FP_ABST
    Figure JP2024036090_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A failure management apparatus (10) is characterized by comprising: a network information acquisition unit (11) that acquires network state information from a communication apparatus (3); a network information control unit (14) that changes the network state information transmitted by the communication apparatus (3); and a failure factor analysis unit (13) that analyzes a failure factor on the basis of the acquired network state information. The failure management apparatus is characterized in that the failure factor analysis unit (13): performs a process for analyzing the failure factor in two stages composed of first stage analysis processing for selecting failure factor candidates from among failure factors capable of occurring in the communication network (2), and second stage analysis processing for determining, for each of the failure factor candidates selected in the first stage analysis processing, whether or not the failure factor has occurred; identifies the network state information acquired in each of the stages; and instructs the network information control unit (14) to acquire the identified network state information.
Need to check novelty before this filing date? Find Prior Art

Description

Fault management device and fault management method

[0001] The present disclosure relates to a fault management device and a fault management method for analyzing the cause of a fault in a communication network.

[0002] In recent years, communication networks have become larger and more complex, and the applications that use communication networks have become more diverse. Furthermore, as multiple applications run on a single communication network system, the burden of operation and management of communication networks has increased. Therefore, there is a demand for technologies that support the operation and management of communication networks to reduce the operational burden.

[0003] Fault management, one of the main functions of operation and management of communication networks, consists of a cycle of detecting faults, analyzing the causes of the detected faults, proposing countermeasures to the faults, deciding on them, and then implementing them.In the field of fault management, too, there is a demand for reducing the operational burden.

[0004] For example, Patent Document 1 discloses a technology in which an estimation model is constructed by machine learning using input data generated from time-series data of received power in wireless communication between wireless stations, and the constructed estimation model is used to estimate factors causing degradation in communication quality, and the factors causing degradation in communication quality are then estimated from the constructed estimation model and the input data.

[0005] International Publication No. 2022 / 195755

[0006] In the above-described conventional technology, the received power of wireless communication is used as information about the communication network, and the cause of communication quality degradation is assumed to be related to the radio wave propagation of wireless signals. However, in recent years, the configuration of communication networks has become more complex. For example, in communication network systems such as 5G systems and local 5G systems, wireless communication is used between user terminals and base stations, but wired communication is typically used between base stations and core networks. In addition, a technique called C / U separation is being implemented to separate control data communication from application data communication. Furthermore, advances in virtualization technology have led to the realization of some functions being implemented by software, and a technology called slicing has made it possible to virtually use one physical communication network as multiple communication networks.

[0007] As the configuration of such communication networks becomes more complex, the types of failure factors in communication networks and the index values ​​that represent the state of communication networks become more diverse. Therefore, in order to accurately estimate the cause of a failure in a communication network, the amount of information that needs to be obtained from the communication network increases.

[0008] The present disclosure has been made in consideration of the above, and aims to provide a fault management device that can reduce the amount of information obtained from a communication network when analyzing the cause of a fault in the communication network.

[0009] In order to solve the above-mentioned problems and achieve the objectives, the fault management device of the present disclosure comprises a network information acquisition unit that acquires network status information indicating the status of a communication network from communication devices that make up the communication network, a network information control unit that controls the communication devices that make up the communication network and changes the network status information sent by the communication devices, and a fault cause analysis unit that analyzes fault causes based on the acquired network status information, and is characterized in that the fault cause analysis unit performs a two-stage analysis process of the fault causes: a first-stage analysis process that narrows down candidate fault causes from among the fault causes that may occur in the communication network, and a second-stage analysis process that determines whether a fault cause has occurred for each of the candidate fault causes narrowed down in the first-stage analysis process, and identifies the network status information to be acquired in each stage and instructs the network information control unit to acquire the identified network status information.

[0010] According to the present disclosure, it is possible to reduce the amount of information acquired from a communication network when analyzing the cause of a failure in the communication network.

[0011] FIG. 1 is a diagram showing an example of the configuration of a communication system according to a first embodiment. FIG. 1 is a diagram showing an example of the functional configuration of a fault management device according to a first embodiment.

[0012] A fault management device and a fault management method according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0013] First Embodiment. FIG. 1 illustrates an example of the configuration of a communication system 100 according to a first embodiment. The communication system 100 includes devices 1-1, 1-2, and 1-3, a communication network 2, and an operation management system 4. The communication network 2 is composed of multiple communication devices 3-1 to 3-M. The operation management system 4 is composed of multiple operation management function units 5-1 to 5-N. Note that multiple components having similar functions are distinguished from one another by assigning a common reference numeral followed by a hyphen and a reference numeral. Furthermore, when there is no need to distinguish between multiple components having similar functions, the following description may use only the common reference numeral. For example, when there is no need to distinguish between devices 1-1, 1-2, and 1-3, they will simply be referred to as device 1.

[0014] The device 1 is, for example, a sensor, an actuator, a user terminal, an application server, etc., and is a device that generates, transmits, receives, and processes data to realize applications, services, etc. Because the device 1 transmits and receives data, it can also be considered a type of communication device 3, which will be described next.

[0015] The communication network 2 is composed of multiple communication devices 3. While FIG. 1 illustrates M communication devices 3, the number of communication devices 3 is not particularly limited. The communication devices 3 receive data via wired or wireless communication between the device 1 or other communication devices 3, and transfer and transmit the data within the communication devices 3, thereby transferring application data generated and processed by the device 1 from one device 1 to another device 1. The communication devices 3 may also transmit to the operation management system 4 information acquired or measured for or during communication processing, data reception, data transfer, and data transmission processing, as well as information monitoring the hardware status of the communication devices 3 themselves or components constituting the communication devices 3, and the status of hardware resources. Hereinafter, information indicating the status of the communication network 2 as described above will be collectively referred to as network status information. The transmission of network status information may be controlled by a fault management unit (described later).

[0016] The operation management system 4 is a system for managing the operation of the communication network 2. The operation management system 4 includes a large number of operation management function units 5, and one of the operation management function units 5, the operation management function unit 5-i, is a fault management unit. Examples of operation management function units 5 other than the fault management unit include a configuration management unit, an accounting management unit, a performance management unit, and a security management unit. Another example of the operation management function unit 5 is a display unit for conveying information to the operations manager of the communication network 2 that uses the operation management system 4. However, the display unit may be located outside the operation management system 4.

[0017] 2 is a diagram illustrating an example of the functional configuration of the fault management device 10 according to the first embodiment. A device having the functions of the operation management function unit 5-i that functions as the fault management unit shown in FIG. 1 is referred to as the fault management device 10. Although the hardware configuration will be described later, the fault management device 10 does not necessarily refer to a single piece of hardware. The functions of the fault management device 10 may be realized by multiple pieces of hardware, or the functions of the fault management device 10 and the functions of the other operation management function units 5 may be realized by a single piece of hardware.

[0018] The fault management device 10 has a network information acquisition unit 11, a network information database 12, a fault cause analysis unit 13, and a network information control unit 14. In addition to the functions shown in Fig. 2, the fault management device 10 may have other functions such as a fault detection unit, a countermeasure proposal unit, a countermeasure decision unit, and a countermeasure execution unit. The network information acquisition unit 11, the network information database 12, and the network information control unit 14 may be functions that are exclusive to the fault management device 10, or may be functions that are shared with the operation management function units 5 other than the operation management function unit 5-i that functions as the fault management unit.

[0019] 2 is conceptual, and the physical connection may be a connection to one or more of the communication devices 3 that make up the communication network 2. The communication device 3 may also be connected to the communication device 3 via a communication network other than the communication network 2 managed by the operation management system 4. A communication device 3 that is not directly connected to the network information acquisition unit 11 is connected to the network information acquisition unit 11 via another communication device 3 that makes up the communication network 2, or via another communication network.

[0020] The network information acquisition unit 11 acquires network status information from the communication devices 3 that constitute the communication network 2. As described above, the communication devices 3 may include the device 1. When the network information acquisition unit 11 receives the network status information, it may add time information indicating the time when the information was acquired.

[0021] Examples of network status information include physical layer information related to communications between communication devices 3, information at the data link layer or higher, control protocol normality monitoring information, and information related to device status. Physical layer information between communication devices 3 includes, for example, information such as the inter-device distance between communication devices 3, transmission power, reception power, signal-to-noise ratio, frequency used, coding rate, and number of error corrections. Information at the data link layer or higher includes, for example, wireless resource allocation, bandwidth allocation, amount of transmitted data, amount of received data, throughput, frame loss rate, delay, and jitter. Information related to device status includes, for example, temperature, power consumption, various alarms, utilization rate of computational resources, and utilization rate of memory resources.

[0022] The network information acquisition unit 11 registers the acquired network state information in the network information database 12. At this time, the network information acquisition unit 11 may register the acquired network state information as is in the network information database 12, or may process the network state information before registering it in the network information database 12. Processing includes, for example, deleting part of the network state information, converting units, normalizing, and creating statistics from the network state information. Processed network state information is also called network state information.

[0023] The failure cause analysis unit 13 analyzes the cause of a failure based on the network status information registered in the network information database 12. The failure cause analysis unit 13 may analyze the cause of a failure using information related to the network configuration or settings of the communication device 3 held by another operation management function unit 5, such as a configuration management unit, or information external to the operation management system 4, such as temperature sensor information, image sensor information, or earthquake information.

[0024] Examples of causes of failure include wireless interference from other wireless systems, shadowing due to obstructions, fading, unexpected terminal movement, changes in antenna orientation, distortion or disconnection of communication lines, resource pressure due to other applications or other terminals, CPU (Central Processing Unit) or memory shortages, abnormal status of virtualized communication functions, overheating, power shortages, and momentary power outages.

[0025] FIG. 3 is a flowchart for explaining the operation of the fault cause analysis unit 13 according to the first embodiment.

[0026] The failure cause analysis unit 13 determines whether a failure has been detected (step S101). The failure cause analysis unit 13 can determine whether a failure has been detected based on, for example, the detection result of a failure detection unit (not shown). If no failure has been detected (step S101: No), the failure cause analysis unit 13 repeats the process of step S101.

[0027] If a failure is detected (step S101: Yes), the failure cause analysis unit 13 executes a first-stage analysis process to narrow down candidate failure causes from among failure causes that may occur in the communication network 2 (step S102). Subsequently, the failure cause analysis unit 13 executes a second-stage analysis process to determine whether each failure cause narrowed down in the first-stage failure cause analysis process has occurred (step S103). As described above, the failure cause analysis unit 13 performs failure cause analysis in two stages. Note that in FIG. 3, it is determined whether a failure has been detected, and the failure cause analysis process is performed only when a failure is detected. However, this is not limiting, and the failure cause analysis process may be started constantly at regular time intervals.

[0028] Here, fault detection can be realized by, for example, setting a threshold value for determining whether a value of network status information related to a Service Level Agreement (SLA) set for each application is a violation of the SLA or is likely to be a violation of the SLA, monitoring the value, and determining that a fault has been detected when the monitored value exceeds the threshold value. Detecting a fault by detecting an excess of the threshold is one example, and the fault detection method is not limited to this example.

[0029] The failure cause analysis unit 13 can perform at least one of the first-stage analysis process and the second-stage analysis process using a trained model trained by machine learning. For example, machine learning that performs clustering can be applied to narrow down the failure causes, which is the first-stage analysis process. For the second-stage analysis process, which determines candidates for each failure cause, a trained model constructed for each failure cause can be used. The failure cause analysis unit 13 can determine whether each failure cause has occurred using a trained model corresponding to each failure cause identified as a candidate in the first-stage analysis process. For example, an autoencoder can be applied as a machine learning method for the trained model used in the second-stage analysis process. When clustering is applied, it is necessary to know the correspondence between each cluster and a group of failure cause candidates. Furthermore, when using a trained model that determines whether a failure cause has occurred, learning using training data is required. In this case, during the trial operation phase of the communication network 2, each failure cause that may occur in the communication network 2 can be forcibly generated, and network state information at that time can be acquired and used as training data. Cluster labeling and learning model construction during the trial operation phase may be performed offline. When constructing a learning model for each failure cause, it is possible to extract features for each failure cause, i.e., network state information required for learning the learning model corresponding to each failure cause and for inference using the trained model.

[0030] Here, a case where the failure cause analysis unit 13 uses machine learning will be described. Note that, in the following description, the failure cause analysis unit 13 is described as using a trained model by machine learning in both the first-stage analysis process and the second-stage analysis process, but some of the first-stage analysis process and the second-stage analysis process may be performed without using machine learning. For example, the failure cause analysis unit 13 may perform the first-stage analysis process using a conditional expression based on a combination of values ​​of network status information without using machine learning. Alternatively, in the second-stage analysis process, the failure cause analysis unit 13 may selectively use a technique that uses machine learning and a technique that does not use machine learning depending on the failure cause.

[0031] 4 is a diagram showing an example of a learning device 6 that generates a trained model used by the fault cause analysis unit 13. The learning device 6 may be any computer that can acquire training data acquired from the communication network 2. The learning device 6 does not necessarily have to be connected to the communication network 2, and may perform learning offline using training data acquired during the trial operation phase of the communication network 2. Here, the learning device 6 is described as executing both the generation of a first-stage trained model and the generation of a second-stage trained model, but the first-stage trained model and the second-stage trained model may be generated by different learning devices 6.

[0032] The learning device 6 includes a learning data acquisition unit 61 and a model generation unit 62 .

[0033] The learning data acquisition unit 61 acquires network state information of the communication network 2 as learning data.

[0034] In the first-stage analysis process, the model generation unit 62 learns candidate failure causes that may have occurred when the network status information was acquired, based on the learning data acquired by the learning data acquisition unit 61. That is, the model generation unit 62 generates a first-stage trained model for inferring narrowed-down candidate failure causes from the network status information of the communication network 2 using clustering. This is data in which multiple types of network status information are associated with each other.

[0035] In the first stage of analysis processing, when machine learning is used to narrow down candidate failure causes, the trained model is configured as a model for classifying the cause into one of multiple cluster groups made up of network status information for each failure cause.

[0036] The learning algorithm used by the model generation unit 62 may be a known algorithm such as supervised learning, unsupervised learning, or reinforcement learning. As an example, a case will be described here in which a first-stage trained model is generated by applying the K-means method, which is unsupervised learning. Unsupervised learning refers to a technique in which training data that does not include results is provided to the learning device 6, and the features of the training data are learned.

[0037] The model generation unit 62 learns the narrowed-down candidates for the failure cause by so-called unsupervised learning, for example, in accordance with a grouping method using the K-means method. Note that "narrowed-down" here means that the number of candidates for the failure cause is smaller than the number of candidates for the failure cause that may occur in the communication network 2.

[0038] The K-means method is a non-hierarchical clustering algorithm, which uses the cluster mean to classify given clusters into k clusters.

[0039] Specifically, the K-means algorithm is processed as follows: First, a cluster is randomly assigned to each data x. Next, the center Vj of each cluster is calculated based on the assigned data. Next, the distance between each x and each vj is calculated, and x is reassigned to the cluster with the closest center. If the cluster assignment for all x remains unchanged through the above process, or if the amount of change falls below a predetermined threshold, it is determined that convergence has occurred and the process ends.

[0040] Here, narrowed-down candidates for the cause of a failure are learned by so-called unsupervised learning in accordance with learning data created based on multiple types of network status information acquired by the learning data acquisition unit 61 .

[0041] The model generation unit 62 generates a trained model by performing the above-described learning, and outputs the generated trained model as a first-stage trained model for performing the first-stage analysis processing.

[0042] The trained model storage unit 63 stores the first-stage trained model output from the model generation unit 62.

[0043] Next, for the second-stage analysis process, the learning device 6 generates a second-stage trained model that has been trained for each failure cause that may occur in the communication network 2. The second-stage trained model is a trained model for inferring, from network state information, a determination result as to whether or not a target failure cause has occurred.

[0044] The learning data acquisition unit 61 acquires, for each failure cause, network status information and information indicating whether the target failure cause has occurred when the network status information is acquired, as learning data.

[0045] The model generation unit 62 learns the determination result of whether or not a target failure cause has occurred based on learning data created based on a combination of network status information output from the learning data acquisition unit 61 and information indicating whether or not a target failure cause has occurred. Here, the learning data is data in which the network status information and the information indicating whether or not a target failure cause has occurred are associated with each other.

[0046] In generating the trained model in the second stage, the learning algorithm used by the model generation unit 62 may be a known algorithm such as supervised learning, unsupervised learning, reinforcement learning, etc. As an example, a case where a neural network is applied will be described.

[0047] The model generation unit 62 learns the determination result of whether or not a target failure factor has occurred, for example, by so-called supervised learning according to a neural network model. Here, supervised learning refers to a method of providing the learning device 6 with a set of data consisting of input and labels as results, learning the features of the learning data, and inferring the result from the input.

[0048] A neural network is composed of an input layer consisting of multiple neurons, an intermediate layer (hidden layer) consisting of multiple neurons, and an output layer consisting of multiple neurons. The intermediate layer may be one layer, or two or more layers.

[0049] Figure 5 is a diagram showing an example of a three-layer neural network. For example, in a three-layer neural network like the one shown in Figure 5, when multiple inputs are input to the input layer (X1-X3), the values ​​are multiplied by weight W1 (w11-w16) and input to the intermediate layer (Y1-Y2), and the result is further multiplied by weight W2 (w21-w26) and output from the output layer (Z1-Z3). This output result varies depending on the values ​​of weights W1 and W2.

[0050] In the present application, the neural network learns the determination result of whether or not a target fault factor has occurred by so-called supervised learning in accordance with learning data created based on a combination of network status information acquired by the learning data acquisition unit 61 and information indicating whether or not a target fault factor has occurred.

[0051] In other words, the neural network learns by inputting network state information into the input layer and adjusting the weights W1 and W2 so that the result output from the output layer approaches the determination result of whether or not the target fault factor has occurred.

[0052] The model generation unit 62 generates a trained model by executing the above-described learning and outputs it as a trained model for the second stage. The model generation unit 62 extracts the feature quantities of each failure cause from the network status information, and identifies the network status information to be acquired at each stage based on the extracted feature quantities of each failure cause.

[0053] The trained model storage unit 63 stores the second-stage trained model output from the model generation unit 62.

[0054] Next, the learning process of the learning device 6 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing the learning process of the learning device 6.

[0055] First, the generation of the first-stage trained model will be described. The training data acquisition unit 61 acquires multiple types of network state information as training data (step S201). Note that, although the training data acquisition unit 61 is configured to simultaneously acquire multiple types of network state information together in this example, it is sufficient that the multiple types of network state information can be input in an associated manner, and the data for the multiple types of network state information may be acquired at different times.

[0056] The model generation unit 62 performs a learning process to learn the narrowed-down candidates for failure causes by so-called unsupervised learning in accordance with the learning data created based on a combination of network status information acquired by the learning data acquisition unit 61, and generates a learned model (step S202).

[0057] The trained model storage unit 63 stores the trained model output from the model generation unit 62 as a first trained model (step S203).

[0058] Next, the generation of the second-stage trained model will be described using the same Figure 6 as above. The learning data acquisition unit 61 acquires multiple types of network status information and information indicating whether a target failure cause has occurred (step S201). Note that, here, multiple types of network status information and information indicating whether a target failure cause has occurred are acquired simultaneously, but it is sufficient if multiple types of network status information and information indicating whether a target failure cause has occurred are input in association with each other, and the data for multiple types of network status information and the information indicating whether a target failure cause has occurred may be acquired at different times.

[0059] The model generation unit 62 executes a learning process to learn the determination result of whether or not the target failure factor has occurred by so-called supervised learning in accordance with the learning data created based on a combination of multiple types of network status information acquired by the learning data acquisition unit 61 and information indicating whether or not the target failure factor has occurred, and generates a learned model (step S202).

[0060] The trained model storage unit 63 stores the trained model generated by the model generation unit 62 as a second trained model (step S203).

[0061] Since the second-stage trained model is generated for each failure cause, the model generation unit 62 repeats the processing of steps S201 to S203 as many times as the number of candidate failure causes that may occur in the communication network 2.

[0062] Next, an inference process using the trained model will be described. FIG. 7 is a configuration diagram of an inference device 7 related to failure causes of the communication network 2. The inference device 7 is a device that performs analysis processing of failure causes of the communication network 2 using a first-stage trained model and a second-stage trained model. The inference device 7 may be provided in the failure cause analysis unit 13 of the failure management device 10, for example, or may be a device separate from the failure management device 10 and capable of communicating with the failure cause analysis unit 13. Here, the inference device 7 is provided in the failure cause analysis unit 13. The inference device 7 has an inference data acquisition unit 71 and an inference unit 72.

[0063] First, the function of the inference device 7 when performing the first-stage analysis processing will be described. For the first-stage analysis processing, the fault cause analysis unit 13 acquires network status information for the first-stage analysis processing using the inference data acquisition unit 71.

[0064] Furthermore, the failure cause analysis unit 13 infers narrowed-down failure cause candidates obtained by the inference unit 72 using the first-stage trained model stored in the trained model storage unit 63. That is, the inference unit 72 inputs the network state information acquired by the inference data acquisition unit 71 into the first-stage trained model for narrowing down failure cause candidates from the network state information, thereby inferring to which cluster the network state information belongs, and can output the inference result as narrowed-down failure cause candidates. Here, a cluster represents one or a combination of multiple failure cause candidates. The failure cause analysis unit 13 regards the narrowed-down failure cause candidates as the result of the first-stage analysis process.

[0065] The fault cause analysis unit 13 obtains a narrowed-down list of candidate fault causes using the first-stage trained model, and then obtains a determination result as to whether or not each fault cause has occurred using the second-stage trained model corresponding to each of the narrowed-down candidate fault causes.

[0066] When the fault cause analysis unit 13 obtains narrowed-down candidate fault causes through the first-stage analysis process, the inference data acquisition unit 71 acquires network status information for the second-stage analysis process that corresponds to the narrowed-down candidate fault causes.

[0067] Furthermore, the failure cause analysis unit 13 obtains a determination result of whether each failure cause has occurred by using the second-stage trained model corresponding to the candidate failure cause that is the result of the first-stage analysis process, from among the second-stage trained models trained by the inference unit 72 for each failure cause stored in the trained model storage unit 63. That is, the inference unit 72 inputs the network state information acquired by the inference data acquisition unit 71 into the second-stage trained model corresponding to each failure cause, thereby inferring a determination result of whether each candidate failure cause narrowed down by the first-stage analysis process has occurred. The failure cause analysis unit 13 can identify a failure cause occurring in the communication network 2 based on the inference result using the second-stage trained model, and generate an analysis result of the failure cause.

[0068] 8 is a flowchart for explaining the operation when the failure cause analysis unit 13 uses machine learning. First, the failure cause analysis unit 13 acquires inference data for the first stage of analysis processing using the inference data acquisition unit 71 (step S301). Specifically, the failure cause analysis unit 13 instructs the network information control unit 14 on the type of data to acquire, and the network information acquisition unit 11 acquires the necessary data.

[0069] The fault cause analysis unit 13 inputs the inference data acquired in step S301 to the first-stage trained model by the inference unit 72 (step S302). When the inference unit 72 obtains narrowed-down fault cause candidates, which are the output from the first-stage trained model, the fault cause analysis unit 13 outputs the fault cause candidates as the inference results of the first-stage analysis process (step S303).

[0070] The fault cause analysis unit 13 acquires, via the inference data acquisition unit 71, inference data for the second stage of analysis processing in accordance with the candidate fault causes that are the inference results of the first stage of analysis processing (step S304).

[0071] The fault cause analysis unit 13 selects one candidate fault cause from the candidates for fault causes narrowed down by the first-stage analysis process, and the inference unit 72 inputs the inference data into the second-stage trained model corresponding to the selected fault cause (step S305).

[0072] The inference unit 72 outputs, as an inference result, the determination result of whether or not the target fault cause, i.e., the fault cause selected in step S305, has occurred, which is the output from the second-stage trained model (step S306).

[0073] The failure cause analysis unit 13 determines whether or not the determination results have been acquired for all the candidates (step S307). If the determination results have not been acquired for all the candidates (step S307: No), the failure cause analysis unit 13 excludes the failure cause selected in step S305 from the candidates to be selected (step S308), and repeats the process from step S305. If the determination results have been acquired for all the candidates (step S307: Yes), the failure cause analysis unit 13 ends the process of FIG. 8.

[0074] In FIG. 8, in step S304, the inference data used in the second stage analysis process is acquired all at once, but it is also possible to acquire the necessary inference data for each failure cause each time.

[0075] As described above, the failure cause analysis unit 13 can obtain the cause of a failure occurring in the communication network 2 as a result of the analysis process.

[0076] The failure cause analysis unit 13 can output the analysis results using the display unit 20. The failure cause analysis unit 13 can output the failure cause that is determined to be currently occurring to the display unit 20. The display unit 20 displays the failure cause that is determined to be currently occurring. The failure cause analysis unit 13 may also output the failure cause to a failure countermeasure proposal unit (not shown) separately.

[0077] In this embodiment, the inference process is performed using a trained model trained using network state information acquired from the communication network 2 to be inferred, but it is also possible to obtain from the outside a trained model trained using network state information acquired from another communication network, and perform the analysis process based on this trained model.

[0078] In this embodiment, the learning algorithms used by the model generation unit 62 and the inference unit 72 are described as applying unsupervised learning to the first-stage trained model and supervised learning to the second-stage trained model, but this is not limitative. Unsupervised learning, reinforcement learning, supervised learning, semi-supervised learning, etc. can also be applied to the learning algorithm.

[0079] Furthermore, the learning algorithm used in the model generation unit 62 can be deep learning, which learns to extract the features themselves, or machine learning can be performed according to other known methods, such as genetic programming, functional logic programming, or support vector machines.

[0080] Furthermore, when implementing unsupervised learning, the method is not limited to the non-hierarchical clustering using the K-means method described above, and any other known method capable of clustering may be used. For example, hierarchical clustering such as the shortest distance method may be used.

[0081] The learning device 6 and the inference device 7 may be connected to the fault management device 10 via a network, for example, and may be separate devices from the fault management device 10. Alternatively, the learning device 6 and the inference device 7 may be built into the fault management device 10, as described above. Furthermore, the learning device 6 and the inference device 7 may exist on a cloud server.

[0082] The model generation unit 62 may also perform the learning process in accordance with learning data created for multiple communication networks 2. It is also possible to add or remove communication networks 2 from which learning data is collected during the process. Furthermore, a learning device 6 that has learned about a certain communication network 2 may be applied to another communication network 2, and re-learning may be performed for the other communication network 2 to update the model.

[0083] The failure cause analysis unit 13 identifies network status information required for the failure cause analysis process at each stage depending on the stage of the failure cause analysis process and which candidate failure causes were narrowed down in the first stage, and instructs the network information control unit 14 to acquire the identified network status information. The required network status information includes the types of network status information and the frequency of each type. The method for selecting network status information will be described later.

[0084] Returning to the explanation of Fig. 2, the network information control unit 14 controls each communication device 3 constituting the communication network 2 so that each communication device 3 transmits desired network status information based on the necessary network status information transmitted from the failure cause analysis unit 13, and changes the network status information transmitted by the communication device 3. For example, the network information control unit 14 sets the type of network status information transmitted by each communication device 3 and the frequency of transmission for each type.

[0085] 2 is conceptual, and the physical connection may be a connection to one or more of the communication devices 3 that make up the communication network 2. The connection between the network information control unit 14 and the communication network 2 may also be a connection via a communication network other than the communication network 2 managed by the operation management system 4. A communication device 3 that is not directly connected to the network information control unit 14 is connected to the network information control unit 14 via another communication device 3 that makes up the communication network 2, or via another communication network.

[0086] Here, a method for selecting network status information required at each stage of the failure cause analysis will be described. For simplicity, it is assumed that there are three types of failure causes A, B, and C that can occur in the communication network 2.

[0087] FIG. 9 is a diagram showing the relationship between the entire network status information and the network status information that is the characteristic amount of each of the failure causes A, B, and C.

[0088] Among all sets of network status information, the sets of network status information that are the feature quantities of each failure cause are shown as circles. Here, to determine whether all three types of failure causes have occurred together, (feature quantity A), U (feature quantity B), and U (feature quantity C) are required. Note that "U" here means "or" in the set, and (feature quantity A), U (feature quantity B), and U (feature quantity C) are the shaded parts in Figure 9.

[0089] FIG. 10 is a conceptual diagram showing network status information required at each stage in a first example of two-stage failure cause analysis. FIG. 11 is a conceptual diagram showing network status information required at each stage in a second example of two-stage failure cause analysis. The first example is an example in which candidates are narrowed down to two in the first stage. The second example is an example in which candidates are narrowed down to "A" or "B or C" in the first stage. Other ways of narrowing down are possible, but other examples are not shown. The network status information acquired in the first stage analysis processing is identified based on the feature amount of each failure cause and the method of narrowing down the failure cause candidates in the first stage analysis processing.

[0090] In the first example shown in Fig. 10, the following network status information is acquired in the first stage. By clustering into (1), (2), and (3) in the first stage, it is possible to narrow down any two pairs of failure factors as candidates for failure factors. (1) Network status information that has the feature amount of A and the feature amount of B, but not the feature amount of C. (2) Network status information that has the feature amount of B and the feature amount of C, but not the feature amount of A. (3) Network status information that has the feature amount of C and the feature amount of A, but not the feature amount of B.

[0091] In the second stage, feature quantities corresponding to the two narrowed-down candidate failure causes are acquired. Specifically, if the first stage narrows down the candidates to A or B, the second stage acquires network state information that is the feature quantity of A and network state information that is the feature quantity of B. By using the trained models corresponding to each of the two narrowed-down candidate failure causes in the second stage, it is possible to determine whether each failure cause has actually occurred.

[0092] In the second example shown in Fig. 11, the following network status information is acquired in the first stage. By clustering into (4) and (5) in the first stage, it is possible to narrow down the candidates for the cause of the failure to either "A" or "B or C". (4) Network status information that is a feature of A but is neither a feature of B nor a feature of C. (5) Network status information that is a feature of B and a feature of C but is not a feature of A.

[0093] In the second stage, if the network conditions have been narrowed down to "A," it is possible to obtain a determination result as to whether or not failure cause A has actually occurred by acquiring network state information that is the feature amount of A and using the trained model corresponding to A. If the network conditions have been narrowed down to "B or C," it is possible to obtain a determination result as to whether or not each of failure causes B and C has actually occurred by acquiring network state information that is the feature amount of B and network state information that is the feature amount of C and using the trained model corresponding to B and the trained model corresponding to C.

[0094] The amount of network status information acquired varies depending on the narrowing-down method. An example of a method for selecting a narrowing-down method from all the narrowing-down methods that can reduce the amount of network status information acquired is shown below. In the first-stage analysis process, the failure cause candidates are narrowed down to one of a plurality of pattern combinations, and in the second-stage analysis process, a process corresponding to the narrowed-down pattern is executed. The pattern of the second analysis process is determined depending on the narrowing-down method. For example, as shown in FIG. 10 , if the first-stage analysis process narrows down three candidates, A, B, and C, to one of three patterns of two-candidate combinations, A or B, B or C, and Cor or A, the second analysis process will have three patterns: a process for determining A and B, a process for determining B and C, and a process for determining C and A.

[0095] STEP #1: For each narrowing-down method, the amount of network status information obtained in the first stage is compared with the amount of network status information obtained for each pattern in the second stage, and the amount of network status information for the stage and pattern with the largest amount of information is extracted.

[0096] For example, in the first example shown in Fig. 10, the amount of information to be acquired for each of the four stages and patterns shown, i.e., one first stage and three patterns of second stages, is calculated, and the four amounts of information are compared to extract the largest amount of information. In the second example shown in Fig. 11, the amount of information to be acquired for each of the three stages and patterns shown, i.e., one first stage and two patterns of second stages, is calculated, and the three amounts of information are compared to extract the largest amount of information.

[0097] STEP #2: Compare the maximum amount of network status information extracted in STEP #1 among all the narrowing-down methods, and select the narrowing-down method with the smallest amount of information. In other words, compare the maximum amount of information extracted in STEP #1 for each narrowing-down method, and select the narrowing-down method with the smallest amount of information.

[0098] The selection method shown in STEPs #1 and #2 above is a so-called "minimization of maximum misfortune" strategy, which selects a narrowing-down method that minimizes the amount of network status information at the stage and pattern where the amount of network status information obtained is the largest among the possible stages and patterns.

[0099] It has been explained above that the model generation unit 62 extracts the feature quantities of each failure factor from the network status information, and specifies the network status information to be acquired at each stage based on the extracted feature quantities of each failure factor. "Extracting the feature quantities of each failure factor" refers to generating a set of feature quantities for each of failure factors A, B, and C shown in Fig. 9, and "specifying the network status information to be acquired at each stage" refers to specifying the shaded areas in Figs. 10 and 11 according to how candidate failure factors are narrowed down.

[0100] As described above, according to the first embodiment, it is possible to provide a fault management device 10 including a network information acquisition unit 11 that acquires network status information indicating the status of the communication network 2 from the communication devices 3 that constitute the communication network 2, a network information control unit 14 that controls the communication devices 3 that constitute the communication network 2 to change the network status information transmitted by the communication devices 3, and a fault cause analysis unit 13 that analyzes fault causes based on the acquired network status information. The fault cause analysis unit 13 of this fault management device 10 is characterized in that it performs a fault cause analysis process in two stages: a first stage analysis process that narrows down candidate fault causes from among fault causes that can occur in the communication network 2, and a second stage analysis process that determines whether a fault cause has occurred for each of the candidate fault causes narrowed down in the first stage analysis process, and specifies network status information to be acquired in each stage and instructs the network information control unit 14 to acquire the specified network status information.

[0101] The fault management device 10 acquires the network status information required for each stage of the fault cause analysis process and executes the fault cause analysis process in two stages. Therefore, compared to determining the occurrence of all possible fault causes in the communication network 2 at once, it is possible to reduce the amount of network status information acquired at one time and the total amount of network status information acquired. This makes it possible to reduce required resources such as the communication bandwidth for transmitting the network status information, the computational resources required for the reception process of the network status information, and the memory capacity for storing the received network status information.

[0102] The failure cause analysis unit 13 can execute at least one of the first-stage analysis processing and the second-stage analysis processing using a trained model trained by machine learning. In the first embodiment, an example has been described in which the failure cause analysis unit 13 executes both the first-stage analysis processing and the second-stage analysis processing using a trained model trained by machine learning. When the first-stage analysis processing is executed using machine learning, the failure cause analysis unit 13 can input network status information into a first-stage trained model for narrowing down candidates for failure causes from network status information, and output the candidates for failure causes as the result of the first-stage analysis processing.

[0103] Furthermore, when the second-stage analysis processing is performed using machine learning, the failure cause analysis unit 13 can perform the second-stage analysis processing using a second-stage trained model for inferring a determination result as to whether or not a target failure cause has occurred from network status information learned for each failure cause. The failure cause analysis unit 13 can identify the occurring failure cause based on the determination result output by inputting network status information corresponding to each of the candidate failure causes narrowed down in the first-stage analysis processing into the second-stage trained model corresponding to each of the candidate failure causes narrowed down in the first-stage analysis processing.

[0104] Furthermore, for each narrowing-down method for narrowing down the candidate failure factors to one of a plurality of pattern combinations in the first-stage analysis process, the failure factor analysis unit 13 can calculate and compare the maximum values ​​of the amount of network status information acquired in the first-stage analysis process and the amount of network status information acquired in each of the second-stage analysis processes corresponding to each of the above patterns, and select the narrowing-down method that produces the smallest maximum value. This is a method for selecting a narrowing-down method that minimizes the amount of information in the process with the largest amount of information, in accordance with the above-mentioned "minimization of maximum misfortune" strategy, and although it does not necessarily minimize the total amount of information, it can reduce the peak of the instantaneous value of data communication volume.

[0105] The failure cause analysis unit 13 can identify the network status information to be acquired in the first stage of analysis processing based on the characteristic quantities of each failure cause and the method of narrowing down the candidate failure causes in the first stage of analysis processing.

[0106] The fault management device 10 may further include a model generation unit 62 that generates a second-stage trained model using, as training data, network status information acquired under conditions in which each possible fault factor in the communication network is forcibly generated during a test run of the communication network 2. This model generation unit 62 may be provided in a device other than the fault management device 10. The model generation unit 62 may extract a feature value of each fault factor from the network status information, and identify the network status information to be acquired at each stage based on the extracted feature value of each fault factor.

[0107] According to the first embodiment, the method includes the steps of acquiring network status information indicating the status of the communication network from a communication device 3 constituting the communication network 2, controlling the communication device 3 constituting the communication network 2 to change the network status information transmitted by the communication device 3, and analyzing the cause of the failure based on the acquired network status information, and the step of analyzing the cause of the failure is characterized in that the analysis process of the failure cause is performed in two stages: a first stage of analysis processing for narrowing down candidate failure causes from among failure causes that can occur in the communication network 2, and a second stage of analysis processing for determining whether or not the candidate failure cause has occurred for each of the candidate failure causes narrowed down in the first stage, and specifies the network status information to be acquired in each stage and instructs the acquisition of the specified network status information.

[0108] 12 is a diagram illustrating an example of a functional configuration of a fault management device 10A according to a second embodiment. The fault management device 10A according to the second embodiment further includes a fault cause occurrence rate management unit 15 in addition to the configuration of the fault management device 10 according to the first embodiment.

[0109] The failure cause occurrence rate management unit 15 has occurrence frequency information for each failure cause that can occur in the communication network 2. The occurrence frequency information may be set by an operations manager of the communication network 2 based on past experience, or may be the frequency of occurrence during the actual operation of the communication network 2. Alternatively, both may be combined, and the occurrence frequency information may be updated based on the value set by the operations manager, taking into account the occurrence frequency in actual operation.

[0110] In the first embodiment, when selecting a method for narrowing down the candidates for the cause of occurrence, a method for selecting a narrowing down method using the "minimization of maximum unhappiness" strategy has been described as an example. In the second embodiment, a narrowing down method using the "minimization of average unhappiness" strategy can be used based on the occurrence frequency of each cause of occurrence.

[0111] For example, in the narrowing down method shown as the first example in Figure 10, the amount of information of network status information required for the jth pattern in the ith stage is Qij (i = 1, 2; j = 1 when i = 1, j = 1, 2, 3 when i = 2), and the probability of transitioning from the first stage to pattern j in the second stage, calculated based on the occurrence frequency of failure factors A, B, and C, is Pj (j = 1, 2, 3).The evaluation value of the total amount of information of network status information can be calculated using the following formula (1).

[0112] Q11+ΣPj×Q2j...(1)

[0113] The evaluation value is calculated for other narrowing-down methods in the same way, and the narrowing-down method that minimizes the evaluation value can be selected. The evaluation value decreases as the amount of acquired network status information decreases, and is calculated based on the transition probability of transitioning to each pattern in the second-stage analysis processing.

[0114] As described above, according to the second embodiment, a fault management device 10A can be provided that includes a fault cause occurrence rate management unit 15 in addition to the same functions as those of the first embodiment. For each narrowing-down method used to narrow down a combination of candidate fault causes to one of multiple patterns in the first-stage analysis process, the fault cause analysis unit 13 of the fault management device 10A calculates an evaluation value for the total amount of network status information acquired in the first-stage analysis process and the second-stage analysis process based on the transition probability from the first-stage analysis process to the corresponding second-stage analysis process, which is calculated using the occurrence rate of each fault cause. The method can then select a narrowing-down method based on the evaluation value. This method of selecting a narrowing-down method based on an evaluation value corresponds to the "minimization of average unhappiness" strategy described above. Compared to the "minimization of maximum unhappiness" strategy described in the first embodiment, the peak value of the amount of information may be larger, but it is possible to reduce the overall average amount of information.

[0115] Here, each function of the communication system 100 according to the first and second embodiments is realized by a processing circuit. The processing circuit that realizes each function may be dedicated hardware or a processor that executes a program stored in a memory.

[0116] 13 is a diagram showing a processing circuit 1000, which is dedicated hardware. The processing circuit 1000, which is dedicated hardware, is, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. The functions of each unit may be realized by a separate processing circuit 1000, or the functions of each unit may be realized together by a single processing circuit 1000.

[0117] 14 is a diagram showing the configuration of a control circuit 2000 including a processor 2001 and a memory 2002. When the control circuit 2000 is used, each function of the communication system 100 is realized by software, firmware, or a combination of software and firmware. The software and firmware are written as programs and stored in the memory 2002. The processor 2001 realizes the functions of each unit by reading and executing the programs stored in the memory 2002.

[0118] The processor 2001 is a CPU, and is also called a central processing unit, processing unit, arithmetic unit, microprocessor, microcomputer, DSP (Digital Signal Processor), etc. The memory 2002 is, for example, a non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable ROM), EEPROM (registered trademark) (Electrically EPROM), a magnetic disk, a flexible disk, an optical disk, a compact disk, a mini disk, a DVD (Digital Versatile Disk), etc.

[0119] It should be noted that some of the functions of the above-described units may be realized by dedicated hardware, and other parts may be realized by software or firmware.

[0120] In this way, the processing circuit can realize the functions of each of the above-mentioned units by hardware, software, firmware, or a combination of these.

[0121] The program may be provided in a state stored in a storage medium, or may be provided via a communication channel such as the Internet.

[0122] The configurations shown in the above embodiments are examples of the contents of the present disclosure, and may be combined with other known technologies, and parts of the configurations may be omitted or modified within the scope of the gist of the present disclosure.

[0123] 1, 1-1, 1-2, 1-3 Equipment, 2 Communication network, 3, 3-1 to 3-M Communication device, 4 Operation management system, 5, 5-1, 5-i, 5-N Operation management function unit, 6 Learning device, 7 Inference device, 10, 10A Fault management device, 11 Network information acquisition unit, 12 Network information database, 13 Fault cause analysis unit, 14 Network information control unit, 15 Fault cause occurrence rate management unit, 20 Display unit, 61 Learning data acquisition unit, 62 Model generation unit, 63 Learned model storage unit, 71 Inference data acquisition unit, 72 Inference unit, 100 Communication system, 1000 Processing circuit, 2000 Control circuit, 2001 Processor, 2002 Memory.

Claims

1. A fault management device comprising: a network information acquisition unit that acquires network status information indicating the status of a communications network from communications devices that make up the communications network; a network information control unit that controls communications devices that make up the communications network to change the network status information sent by the communications devices; and a fault cause analysis unit that analyzes fault causes based on the acquired network status information, wherein the fault cause analysis unit performs analysis of the fault causes in two stages: a first stage of analysis processing that narrows down candidate fault causes from among fault causes that can occur in the communications network; and a second stage of analysis processing that determines whether the fault cause has occurred for each of the candidate fault causes narrowed down in the first stage of analysis processing, and identifies the network status information to be acquired in each stage, and instructs the network information control unit to acquire the identified network status information.

2. The fault management device described in claim 1, characterized in that the fault cause analysis unit performs at least one of the first stage analysis processing and the second stage analysis processing using a trained model trained by machine learning.

3. The fault management device according to claim 1 or 2, characterized in that the fault cause analysis unit inputs the network status information into a first-stage trained model for narrowing down the candidates for the fault cause from the network status information, and the candidates for the fault cause output are the results of the first-stage analysis process.

4. A fault management device as described in any one of claims 1 to 3, characterized in that the fault cause analysis unit executes the second stage analysis process using a second stage learned model for inferring a determination result of whether or not the target fault cause has occurred from the network status information learned for each fault cause.

5. The fault management device described in claim 4, characterized in that the fault cause analysis unit inputs the network status information corresponding to each of the candidate fault causes narrowed down in the first-stage analysis process into the second-stage trained model corresponding to each of the candidate fault causes narrowed down in the first-stage analysis process, and identifies the occurring fault cause based on the judgment result output.

6. A fault management device as described in any one of claims 1 to 5, characterized in that the fault cause analysis unit calculates and compares the maximum value of the amount of network status information obtained in the first-stage analysis process and the amount of network status information obtained in each of the second-stage analysis processes corresponding to each of the patterns, for each narrowing-down method that narrows down candidate fault causes to one of a combination of multiple patterns in the first-stage analysis process, and selects the narrowing-down method that minimizes the maximum value.

7. A fault management device as described in any one of claims 1 to 5, characterized in that the fault cause analysis unit calculates an evaluation value of the total amount of network status information obtained in the first-stage analysis processing and the second-stage analysis processing based on the transition probability of transitioning from the first-stage analysis processing to each of the second-stage analysis processing corresponding to the pattern, which is calculated using the occurrence rate of each fault cause, for each narrowing-down method that narrows down the combinations of candidate fault causes in the first-stage analysis processing to one of a plurality of patterns, and selects the narrowing-down method based on the evaluation value.

8. A fault management device as described in any one of claims 1 to 7, characterized in that the fault cause analysis unit identifies the network status information to be obtained in the first stage of analysis processing based on the characteristic quantities of each fault cause and the method of narrowing down the candidate fault causes in the first stage of analysis processing.

9. The fault management device according to claim 2, further comprising a model generation unit that generates the second-stage trained model using, as training data, the network status information obtained when each possible fault factor that may occur in the communication network is forcibly generated during a test run of the communication network.

10. The fault management device according to claim 9, characterized in that the model generation unit extracts the characteristic quantities of each fault cause from the network status information, and identifies the network status information to be acquired at each stage based on the characteristic quantities of each extracted fault cause.

11. A fault management method comprising the steps of: acquiring network status information indicating the status of a communications network from communications devices constituting the communications network; controlling communications devices constituting the communications network to change the network status information transmitted by the communications devices; and analyzing fault causes based on the acquired network status information, wherein the step of analyzing fault causes comprises performing the fault cause analysis process in two stages: a first stage of analysis processing to narrow down candidate fault causes from among fault causes that may occur in the communications network; and a second stage of analysis processing to determine whether or not each of the candidate fault causes narrowed down in the first stage of analysis processing has occurred; specifying the network status information to be acquired in each stage, and issuing an instruction to acquire the specified network status information.

Citation Information

Patent Citations

  • System, method and computer-readable medium for anomaly detection

    JP2024039820A

  • Inference device, inference method, and inference program

    WO2022201477A1