Monitoring systems, monitoring methods, and programs
The monitoring system addresses the challenge of identifying anomaly-related state data by correlating saturation periods with anomaly occurrences, enabling efficient detection and recovery in communication services.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-04
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies struggle to identify which state data is related to anomalies in a service, making it difficult to detect and recover from anomalies in communication services, as multiple communication devices can cause saturation without a clear upper limit, leading to undetected anomalies.
A monitoring system that includes a status data acquisition unit to collect data from communication devices, determines saturation periods, and identifies correlations between anomalies and device states using a normalization and relationship determination process, enabling quicker anomaly detection and recovery.
The system effectively identifies communication devices causing anomalies by correlating saturation periods with anomaly occurrences, facilitating quicker recovery and maintenance, thus improving service reliability.
Smart Images

Figure 0007842308000001 
Figure 0007842308000002 
Figure 0007842308000003
Abstract
Description
Technical Field
[0004] , , , , , , , , , ,
[0003] , ,
[0001] The present disclosure relates to a monitoring system, a monitoring method, and a program.
Background Art
[0002] Conventionally, for various purposes such as detection, recovery, or analysis of abnormalities in a service, state data regarding the state of a device to be monitored has been analyzed. For example, Patent Document 1 describes a system for determining the presence or absence of an abnormality in a device based on a model in which the relationship between the temporal change in the state indicated by the state data of a certain device and the actual operating state of the device is learned. Patent Document 2 describes a system for estimating whether the future memory usage of a device will exceed a threshold based on the history of the device's memory usage in the past and the current memory usage of the device. For example, Patent Document 3 describes that when a failure (an example of an abnormality) occurs in a system, based on a calculation model that calculates the degree of relevance between the state data regarding the state of the devices in the system and the recovery procedure for recovering the failure that has occurred in the system, an effective recovery procedure for recovering the failure that has occurred in the system is specified.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] The technologies described above require that processing related to service anomalies be performed based on status data concerning the state of the devices being monitored within the service.
[0005] One of the purposes of this disclosure is to perform processing related to anomalies in the service based on status data regarding the status of devices monitored in the service. [Means for solving the problem]
[0006] The monitoring system relating to this disclosure includes a status data acquisition unit that acquires status data relating to the status of devices to be monitored in a service, and executes processing relating to abnormalities in the service based on the status data. [Effects of the Invention]
[0007] According to this disclosure, processing related to anomalies in the service can be performed based on status data regarding the status of devices being monitored in the service. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an example of the hardware configuration of a monitoring system. [Figure 2] This figure shows an example of the time-series changes in the communication volume of a communication device when no abnormalities occur in the communication service. [Figure 3] This figure shows an example of the time-series change in communication volume of a communication device when an anomaly is detected in a communication service. [Figure 4] This figure shows an example of a notification displayed on the administrator's terminal. [Figure 5] This figure shows an example of the functions implemented in the monitoring system. [Figure 6] This figure shows an example of a state database. [Figure 7] This figure shows an example of a period set by the period setting unit. [Figure 8]This figure shows an example of the process performed by the monitoring system. [Figure 9] This figure shows an example of the functions implemented in the monitoring system of Modification Example 1-1. [Figure 10] This figure shows an example of the time-series change in communication volume of a communication device when an anomaly is detected in a communication service. [Figure 11] This figure shows an example of a notification displayed on the administrator's terminal. [Figure 12] This figure shows an example of the functions implemented in the monitoring system. [Figure 13] This figure shows an example of a state database. [Figure 14] This figure shows an example of the results of clustering. [Figure 15] This figure shows an example of the process performed by the monitoring system. [Figure 16] This figure shows an example of the functions implemented in the modification monitoring system. [Modes for carrying out the invention]
[0009] [1. First Embodiment] A first embodiment, which is an example of an embodiment of the monitoring system relating to this disclosure, will be described. In the first embodiment, an example of processing related to anomalies in a service will be described. In the conventional technology, various state data exist for various devices, so if a service provider can identify the state data related to anomalies in the service, it can be used for various purposes such as detecting, recovering from, or analyzing anomalies in the service, which is very useful. However, the conventional technology cannot identify which state data is related to an anomaly from among the various state data related to the state of the device. Therefore, the monitoring system of the first embodiment determines whether or not the state data is related to an anomaly in the service.
[0010] [1-1. Hardware configuration of the monitoring system] FIG. 1 is a diagram showing an example of the hardware configuration of a monitoring system. For example, the monitoring system 1 includes a server 10, a communication device 20, a user terminal 30, and an administrator terminal 40. Each of the server 10, the communication device 20, the user terminal 30, and the administrator terminal 40 can be connected to the network N. For example, the network N is the Internet, a public communication line, or a LAN.
[0011] The server 10 is a server computer. In this embodiment, the case where a communication carrier manages the server 10 is taken as an example. For example, the server 10 includes a control unit 11, a storage unit 12, and a communication unit 13. The control unit 11 includes at least one processor. The storage unit 12 includes at least one of a volatile memory such as a RAM and a non-volatile memory such as a flash memory. The communication unit 13 includes at least one of a communication interface for wired communication and a communication interface for wireless communication.
[0012] The communication device 20 is a device that communicates with other devices. For example, the communication device 20 is a device managed by a communication carrier. The communication device 20 may be a device for wireless communication or a device for wired communication. The communication device 20 relays at least one of a call and data communication by a user. For example, the communication device 20 is a base station of a public communication line, a private branch exchange, an access point of a wireless LAN, a router, a hub, a repeater, a modem, or a server computer used in virtualization technology. The communication device 20 may be a device of a communication carrier that provides a fully virtualized cloud-native mobile network to users.
[0013] The user terminal 30 is a computer used by a user of the communication service. For example, the user terminal 30 may be a smartphone, tablet, personal computer, or wearable device. For example, the user terminal 30 includes a control unit 31, a storage unit 32, a communication unit 33, an operation unit 34, and a display unit 35. The hardware configurations of the control unit 31, storage unit 32, and communication unit 33 may be the same as those of the control unit 11, storage unit 12, and communication unit 13, respectively. The operation unit 34 is an input device such as a touch panel or mouse. The display unit 35 is a display such as a liquid crystal or organic EL.
[0014] The administrator terminal 40 is the administrator's computer in the communication service. For example, the administrator terminal 40 is a smartphone, tablet, personal computer, or wearable device. For example, the administrator terminal 40 includes a control unit 41, a storage unit 42, a communication unit 43, an operation unit 44, and a display unit 45. The hardware configuration of the control unit 41, storage unit 42, communication unit 43, operation unit 44, and display unit 45 may be the same as that of the control unit 11, storage unit 12, communication unit 13, operation unit 34, and display unit 35, respectively.
[0015] Furthermore, the programs stored in the memory units 12, 32, and 42 may be supplied via the network N. Each computer may also include at least one of a reading unit (e.g., a memory card slot) for reading computer-readable information storage media and an input / output unit (e.g., a USB port) for inputting and outputting data to and from external devices. For example, programs stored on the information storage media may be supplied via at least one of the reading unit and the input / output unit.
[0016] Furthermore, the monitoring system 1 may include at least one computer and is not limited to the example in Figure 1. For example, the monitoring system 1 may not include at least one of the communication device 20, user terminal 30, and administrator terminal 40. In this case, the at least one exists outside the monitoring system 1. The monitoring system 1 may include only the server 10. The monitoring system 1 may include the server 10 and other computers not shown in Figure 1. The monitoring system 1 may include only the administrator terminal 40. The monitoring system 1 may include the administrator terminal 40 and other computers not shown in Figure 1.
[0017] [1-2. Overview of the Monitoring System] In this embodiment, we take the example that the communication device 20 is a server computer having a containerized network function (CNF). Furthermore, we take the example that the user terminal 30 is a smartphone. The communication device 20 relays communications from a large number of user terminals 30. For example, if the amount of communication that the communication device 20 needs to process reaches the amount of communication that the communication device 20 can process, an abnormality may occur in the communication service. The amount of communication can be expressed by a known indicator such as bps.
[0018] For example, if an anomaly occurs in a communication service, the administrator identifies the communication device 20 that caused the anomaly and performs recovery work. However, since there are many communication devices 20 in a communication service and various anomalies can occur, it can be difficult for the administrator to identify the communication device 20 that caused the anomaly. In this case, the communication device 20 that caused the anomaly may not be able to handle any further communication, and the communication volume may become saturated.
[0019] Saturation of communication volume means that the communication volume is positive (not zero) and does not change or changes very little. No change in communication volume means the change in communication volume is zero. Nearly no change in communication volume means the change in communication volume is below a threshold. The specific method for determining saturation will be described later. While an upper limit may be set for communication volume, there may be cases where the communication device 20 can process more communication volume than the upper limit set by the administrator, or where it can process less communication volume than the upper limit set by the administrator, and furthermore, the administrator themselves may not know the upper limit of the communication device 20. Therefore, in this embodiment, no upper limit for communication volume is set. For this reason, even if the communication volume of the communication device 20 reaches a certain level, no abnormality will be detected.
[0020] Figure 2 shows an example of the time-series change in the communication volume of communication device 20 when no abnormalities occur in the communication service. The horizontal axis in Figure 2 represents time. The vertical axis in Figure 2 represents the communication volume. In the example in Figure 2, the change in the communication volume of each of the three communication devices 20A to 20C located in the telecommunications carrier's facility is shown. Hereafter, when communication devices 20A to 20C are not distinguished, they will simply be referred to as communication device 20. As shown in Figure 2, when no abnormalities occur in the communication service, the communication volume of each communication device 20 is not saturated. That is, the communication volume of each communication device 20 is, in principle, always changing.
[0021] Figure 3 shows an example of the time-series change in the communication volume of communication device 20 when an anomaly is detected in the communication service. In the example in Figure 3, the communication volume of communication device 20C is saturated during the period from time t00 to time t01. During this period, communication device 20C cannot process any further communication, so a user terminal 30 that attempts to connect to communication device 20C in a particular area may not be able to use the communication service. In such a case, the user may contact the call center of the communication service and inform them that they cannot use the communication service.
[0022] For example, the administrator becomes aware of an anomaly after receiving inquiries from multiple users. At this point, the administrator has not identified which communication device 20 is causing the anomaly. Subsequently, at time t01, the communication volume of communication device 20C decreases, and the anomaly resolves itself. Furthermore, the administrator identifies that the anomaly has been resolved when they stop receiving inquiries from users. Since the user terminal 30 becomes able to use the communication service when time t01 arrives, there is a correlation between the period during which the anomaly occurred and the communication volume of communication device 20C. In this embodiment, the period from the start of saturation to the end of saturation is treated as an example of the period during which the anomaly occurred, but the period from the start of saturation to the present time, before saturation has ended, may also be treated as an example of the period during which the anomaly occurred. For example, the period from the start of saturation to a time specified by the administrator may also be treated as an example of the period during which the anomaly occurred.
[0023] For example, after an anomaly occurs and is resolved, server 10 analyzes the communication volume of each communication device 20A to 20C. Server 10 identifies that the communication volume of communication device 20C was saturated during the period when the anomaly occurred. Server 10 stores the characteristics of the anomaly that occurred and the fact that the communication volume of communication device 20C was related to that anomaly, associating them together. If a similar anomaly occurs again, server 10 sends a notification to the administrator terminal 40 prompting it to check the status of communication device 20C.
[0024] Figure 4 shows an example of a notification displayed on the administrator terminal 40. For example, when the administrator starts a maintenance tool installed on the administrator terminal 40, the administrator screen SC, which shows the notification received from the server 10, is displayed on the display unit 45. The administrator screen SC displays a message indicating that the communication volume of the communication device 20C was saturated when a similar anomaly was detected in the past, and therefore the status of the communication device 20C should be checked again this time.
[0025] For example, suppose that in the past, when the communication device 20C's traffic became saturated, the call center for the communication service received an inquiry from a user attempting to use the communication service from a specific area X. If the call center were to receive an inquiry from the same user attempting to use the communication service from the same area X at the present time, it would be detected that an anomaly has occurred in the communication service in area X. In this case, the server 10 sends display data to the administrator terminal 40, prompting it to check the status of the communication device 20C, via the administrator screen SC.
[0026] For example, when the administrator terminal 40 receives display data from the administrator screen SC, it displays the administrator screen SC on the display unit 45, which includes a message prompting the administrator to check the status of the communication device 20C. The administrator screen SC may display a graph showing the current communication volume of the communication device 20C, or it may display a graph showing the communication volume of the communication device 20C when a similar abnormality occurred in the past. The administrator checks the administrator screen SC and performs maintenance on the communication device 20C. The maintenance of the communication device 20C itself may be performed by a well-known method.
[0027] As described above, the monitoring system 1 determines whether the communication volume of the communication device 20 was saturated when an abnormality occurs in the communication service and is restored. Based on the determination result of whether the communication volume of the communication device 20 was saturated, the monitoring system 1 determines whether the communication volume of the communication device 20 is related to the abnormality in the communication service. For example, if a similar abnormality occurs again, it becomes easier for the administrator to identify the communication device 20 that caused the abnormality, and recovery work can be carried out quickly. The details of this embodiment will be described below.
[0028] [1-3. Functions implemented by the monitoring system] Figure 5 shows an example of the functions implemented by the monitoring system 1. Figure 5 shows the functions implemented by the server 10 and the functions implemented by the administrator terminal 40. The functions of the communication device 20 and the user terminal 30 are the same as those in known communication services, and are therefore omitted in Figure 5. For example, the communication device 20 has the function of transmitting status data or a part thereof to the server 10, as described later. The user terminal 30 has the function of allowing the user to use the communication service.
[0029] [1-3-1. Functions implemented by the server] For example, server 10 includes a data storage unit 100, a state data acquisition unit 101, a period setting unit 102, a normalization execution unit 103, a saturation determination unit 104, a relationship determination unit 105, and a recovery processing execution unit 106. The data storage unit 100 is implemented by a storage unit 12. The state data acquisition unit 101, the period setting unit 102, the normalization execution unit 103, the saturation determination unit 104, the relationship determination unit 105, and the recovery processing execution unit 106 are implemented by a control unit 11.
[0030] [Data Storage Unit] The data storage unit 100 stores data related to communication services. For example, the data storage unit 100 stores a status database DB.
[0031] Figure 6 shows an example of a state database DB. The state database DB is a database that stores the state data of each of the multiple communication devices 20. For example, the state database DB stores the identification information of each of the multiple communication devices 20 and state data indicating the state of the communication device 20. Other data may be stored in the state database DB. For example, the state database DB may store information indicating the determination result of the saturation determination unit, or information indicating the determination result of the relationship determination unit.
[0032] The identification information of the communication device 20 may be any information, such as an IP address, device name, or MAC address. The status data may indicate the pinpoint state of the communication device 20 at a specific point in time, but in this embodiment, the status data is data relating to the time-series changes in the state of the communication device 20. The state of the communication device 20 can also be said to be the load on the communication device 20. The state of the communication device 20 may refer to a hardware state or a software state.
[0033] In this embodiment, the amount of communication performed by the communication device 20 is given as an example where it corresponds to the state of the communication device 20, but other states of the communication device 20 may also correspond to the state of the communication device 20. Other states may include the amount of resources consumed by the communication device 20, the processing time (response time) for requests to the service, or the number of errors returned for requests. The state data may be an indicator known as a golden signal metric, or other indicators used in well-known benchmark tests. For example, the state data may be CPU usage, memory usage, power consumption, communication speed, temperature, or a combination thereof.
[0034] The status data may also indicate the status of performance in the communication service or other services. Other services are services that users utilize via the communication service, such as e-commerce services, travel booking services, payment services, or financial services. For example, the status data may be the number of users connected to the communication device 20 (logged in to the communication service or other services), the number of payments made by users via the communication device 20, or the number of orders made by users via the communication device 20. As the number of payments, payment amounts, number of orders, and order amounts increase, the load on the communication device 20 increases, so the number of payments, payment amounts, number of orders, and order amounts also correspond to the status of the communication device 20.
[0035] In this embodiment, the status data indicates the time-series changes in the state of the communication device 20. For example, the status data includes the date and time when the state of the communication device 20 was acquired, and a numerical value indicating the state of the communication device 20 at that date and time. The state of the communication device 20 may be expressed in a format other than a numerical value, such as characters. In this embodiment, since the amount of communication of the communication device 20 corresponds to the state of the communication device 20, the status data includes the date and time when the amount of communication of the communication device 20 was acquired, and the amount of communication of the communication device 20 at that date and time.
[0036] The data stored in the data storage unit 100 is not limited to the state database DB. The data storage unit 100 can store any data related to the communication service. For example, the data storage unit 100 may store data indicating the result of the saturation determination unit 104. The data storage unit 100 may store data indicating the result of the relationship determination unit 105. The data storage unit 100 may store data necessary to display the administrator screen SC on the administrator terminal 40. The data storage unit 100 may also store various thresholds used in the processing described later. The thresholds may be specified by the administrator, or they may be determined based on machine learning methods or statistical methods.
[0037] [Status data acquisition unit] The status data acquisition unit 101 acquires status data. For example, the status data acquisition unit 101 periodically requests the status of each of the multiple communication devices 20 (for example, every 1 to 30 seconds or every 1 to 5 minutes). When a communication device 20 receives a request from the server 10 via the status data acquisition unit 101, it sends its own identification information and information indicating its status to the server 10. The status data acquisition unit 101 adds the current date and time and the information received from the communication device 20 to the status data associated with the identification information of the communication device 20.
[0038] The communication device 20 may also store status data for a certain period of time and send the status data to the server 10 all at once. In this case, the status data acquisition unit 101 acquires the status data from the communication device 20 all at once and stores it in the status database DB. Alternatively, the status data acquisition unit 101 may acquire status data from the communication device 20 irregularly rather than periodically. The status data acquisition unit 101 may also acquire status data from the communication device 20 at a timing instructed by the administrator.
[0039] Communication device 20 is an example of a device to be monitored in a communication service. The device to be monitored may be any device and is not limited to communication device 20. Where "communication device 20" is written in this embodiment, it can be read as any other device to be monitored. The other device may be any device used in the communication service, for example, server 10, other computers other than server 10, components such as CPUs, memory, and network cards included in server 10 and other computers, batteries, antennas, communication cables, or sensors.
[0040] In this embodiment, the status data acquisition unit 101 acquires status data when an abnormality is detected. The status data acquired when an abnormality is detected indicates the state during all or part of the period from when the abnormality occurs until it is resolved. Hereafter, this period will be referred to as the abnormality period. The status data may indicate not only all or part of the abnormality period, but also the state during other periods other than the abnormality period. The status data acquisition unit 101 may acquire status data during the period from when the abnormality occurs until it is resolved, or it may acquire status data retrospectively after the abnormality has been resolved.
[0041] In this embodiment, the state data does not indicate the state at a specific point in time, but rather the state data indicates the state over a period of a certain length. For example, the state data acquisition unit 101 acquires state data relating to the time-series changes in the state of the device over a period set by the period setting unit 102. In this embodiment, three periods, such as a saturation period, a reference period, and an abnormal period, are set by the period setting unit 102, which will be described later. Therefore, the state data acquisition unit 101 acquires state data indicating the state of the communication device 20 during each of the saturation period, reference period, and abnormal period.
[0042] The status data acquisition unit 101 may acquire status data indicating the state of the communication device 20 during periods other than the saturation period, reference period, and abnormal period. For example, the status data acquisition unit 101 may acquire status data indicating the state of the communication device 20 before the saturation period. The status data acquisition unit 101 may acquire status data indicating the state of the communication device 20 after the reference period and abnormal period. The status data acquisition unit 101 may acquire status data indicating the state of the communication device 20 regardless of the timing of the abnormality occurrence. The status data acquisition unit 101 can acquire status data stored in the status database DB at any time.
[0043] [Period setting section] The period setting unit 102 sets the period related to the state data. The period related to the state data is the period used in the determination of the relationship determination unit 105, which will be described later. All periods in which a state is indicated in the state data may be used in the determination of the relationship determination unit 105, which will be described later, but in this embodiment, only a portion of the periods in which a state is indicated in the state data will be used in the determination of the relationship determination unit 105. Setting the period means determining the start and end times of the period.
[0044] Figure 7 shows an example of a period set by the period setting unit 102. In this embodiment, the period setting unit 102 sets three periods: a saturation period T1, a reference period T2, and an abnormal period T3. The time window W in Figure 7 is the period used in processing by the relationship determination unit 105 and other units from the state data. In the example in Figure 7, the time window W is the sum of the reference period T2 and the abnormal period T3.
[0045] For example, the period setting unit 102 may set only one or two of the saturation period T1, the reference period T2, and the abnormal period T3. That is, the period setting unit 102 may set only the saturation period T1, only the reference period T2, only the abnormal period T3, only the saturation period T1 and the reference period T2, only the saturation period T1 and the abnormal period T3, or only the reference period T2 and the abnormal period T3. The period setting unit 102 may set only the time window W without setting the saturation period T1, the reference period T2, and the abnormal period T3.
[0046] For example, the period setting unit 102 sets the saturation period T1 for the status data based on the determination result of the saturation determination unit 104. In the example in Figure 7, the period setting unit 102 sets the saturation period T1 to be the period from the saturation start timing t1, when the state indicated by the status data begins to saturate, to the saturation end timing t2, when the state indicated by the status data ends to saturation. The saturation start timing t1 is the date and time when the saturation determination unit 104 determines that saturation has started, among the dates and times when the status is indicated in the status data. The saturation end timing t2 is the date and time when the saturation determination unit 104 determines that saturation has ended, among the dates and times when the status is indicated in the status data. The period setting unit 102 may also set the saturation period T1 to be the period from the saturation start timing t1 to the current timing when saturation has not ended (saturation is continuing). For example, the period setting unit 102 may set the saturation period T1 to be the period from the saturation start timing t1 to a time specified by the administrator.
[0047] For example, when an abnormality is detected, the period setting unit 102 sets the reference period T2 and the abnormal period T3 for the status data based on the abnormality occurrence timing t3 in which the abnormality occurred. In this embodiment, when an abnormality is detected, the administrator performs an abnormality occurrence operation from the administrator terminal 40 to indicate that an abnormality has occurred. For example, when the administrator receives an inquiry from a user indicating that an abnormality has occurred, the administrator performs an abnormality occurrence operation.
[0048] For example, server 10 determines whether an error has occurred by determining whether or not it has received error data indicating an error from administrator terminal 40. If server 10 does not receive error data from administrator terminal 40, it does not determine that an error has occurred. If it receives error data from administrator terminal 40, it determines that an error has occurred. The timing at which server 10 receives error data is the error timing t3.
[0049] Furthermore, administrators may perform abnormality detection operations based on factors other than user inquiries. For example, administrators may perform abnormality detection operations by visually checking the status data of the communication device 20 on the administrator screen SC or other screens. If the execution result of a program that estimates the occurrence of an abnormality is displayed on the administrator screen SC or other screens, administrators may perform abnormality detection operations by visually checking the execution result.
[0050] Furthermore, the method by which server 10 detects the occurrence of an anomaly can be any method and is not limited to an anomaly detection operation performed by an administrator. For example, server 10 may detect the occurrence of an anomaly based on the execution result of an anomaly detection program for detecting the occurrence of an anomaly. The anomaly detection program determines whether or not an anomaly has occurred based on status data. For example, the anomaly detection program determines that an anomaly has occurred if the state indicated by the status data exceeds a standard range.
[0051] For example, an anomaly detection program may determine that an anomaly has occurred if the state indicated by the state data exceeds the standard range for a predetermined period of time or longer. Alternatively, an anomaly detection program may determine that an anomaly has occurred if the number of times or the frequency of the state indicated by the state data exceeding the standard range exceeds a threshold. In this case, the timing at which the anomaly detection program determines an anomaly is the anomaly occurrence timing t3.
[0052] In the example shown in Figure 7, the period setting unit 102 sets the period from the reference timing t4, which serves as the basis for the reference period T2, to the abnormality occurrence timing t3, when the abnormality occurs, as the reference period T2. For example, the reference timing t4 is a predetermined time before the abnormality occurrence timing t3. The reference timing t4 may be a predetermined time before the current time, or it may be a timing specified by the administrator.
[0053] For example, the period setting unit 102 sets the period from the abnormality occurrence timing t3 to the abnormality resolution timing t5, when the abnormality is resolved, as the abnormality period T3. The abnormality resolution timing t5 is the timing at which it is determined that the abnormality has been resolved. In this embodiment, when an abnormality is resolved, the administrator performs an abnormality resolution operation from the administrator terminal 40 to indicate that the abnormality has been resolved. For example, the administrator performs the abnormality resolution operation when they stop receiving inquiries from users indicating that an abnormality has occurred.
[0054] For example, server 10 determines whether an anomaly has been resolved by determining whether or not it has received an anomaly resolution data indicating an anomaly resolution operation from administrator terminal 40. If server 10 does not receive an anomaly resolution data from administrator terminal 40, it does not determine that the anomaly has been resolved. If it receives an anomaly resolution data from administrator terminal 40, it determines that the anomaly has been resolved. The timing at which server 10 receives the anomaly resolution data is the anomaly resolution timing t5.
[0055] Furthermore, the administrator may perform error resolution operations based on factors other than user inquiries. For example, the administrator may visually check the status data of the communication device 20 on the administrator screen SC or other screens and perform error resolution operations. If the execution result of a program that estimates the resolution of the error is displayed on the administrator screen SC or other screens, the administrator may visually check the execution result and perform error resolution operations.
[0056] Furthermore, the method by which server 10 detects the resolution of an anomaly can be any method and is not limited to an anomaly resolution operation performed by an administrator. For example, server 10 may detect the resolution of an anomaly based on the execution result of an anomaly resolution program used to detect the resolution of an anomaly. The anomaly resolution program determines whether or not an anomaly has been resolved based on status data. For example, the anomaly resolution program may determine that an anomaly has been resolved when the state indicated by the status data falls within a reference range.
[0057] For example, the anomaly resolution program may determine that an anomaly has been resolved if the state indicated by the state data remains within a standard range for a predetermined period of time or longer. Alternatively, the anomaly resolution program may determine that an anomaly has been resolved if the number of times or frequency in which the state indicated by the state data remains within a standard range exceeds a threshold. In this case, the timing at which the anomaly resolution program determines that the anomaly has been resolved is the anomaly resolution timing t5.
[0058] [Normalization Execution Department] The normalization execution unit 103 performs normalization on the state data. For example, the normalization execution unit 103 performs normalization based on the state data of each of the multiple communication devices 20. The normalization execution unit 103 performs normalization so that the numerical values of the states indicated by the state data of each of the multiple communication devices 20 fall within a certain range. The normalization execution unit 103 may perform normalization based on a known normalization method. For example, the normalization method may be the z-score normalization method or the minimum-maximum normalization method.
[0059] [Saturation judgment section] The saturation determination unit 104 determines whether the time-series state of the communication device 20 is saturated based on the state data. Saturation means that the amount of change in the state indicated by the state data falls below a threshold, or that the state in which the amount of change is below the threshold continues for a predetermined time or longer. In the case of communication volume, where there is no upper limit in principle, if the communication volume increases, the amount of change mainly means the increase. Therefore, saturation occurs when, after the total amount of change increases over a certain period, the increase in the state indicated by the state data falls below a threshold, or that the state in which the increase is below the threshold continues for a predetermined time or longer. The amount of change may be the cumulative value over a certain period, or it may be the change between a certain point in time and the next point in time. The amount of change may also mean the decrease in the communication volume. In other words, the saturation determination unit 104 makes the determination of saturation based on the time-series change of the amount of change itself indicated by the state data.
[0060] For example, the saturation determination unit 104 determines that the time-series state of the communication device 20 is saturated when the increase in the state indicated by the state data changes from a state where the increase is greater than or equal to a threshold to a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer. The saturation determination unit 104 determines that the time-series state of the communication device 20 is not saturated even if the increase in the state indicated by the state data changes from a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer, if the increase was less than the threshold before that. For example, the saturation determination unit 104 may determine that the time-series state of the communication device 20 is saturated when the state indicated by the state data is greater than or equal to a threshold, and the increase in that state changes from a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer. The saturation determination unit 104 determines that the time-series state of the communication device 20 is not saturated if the increase in the state indicated by the state data is less than a threshold, or if the state in which the increase is less than a threshold continues for a predetermined time or longer, as long as the state remains below the threshold.
[0061] For example, the saturation determination unit 104 determines whether the time-series state of the communication devices 20 is saturated based on the state data of each of the multiple communication devices 20 and a predetermined threshold. In other words, for each state data of the communication device 20, the saturation determination unit 104 determines whether the time-series state indicated by the state data is saturated based on a predetermined threshold. The threshold referenced by the saturation determination unit 104 when determining saturation may be specified by the administrator or calculated from past anomaly trends. The threshold may be determined based on a model using machine learning techniques.
[0062] The saturation determination unit 104 may determine whether the state of the communication device 20 is saturated or not based on a model using machine learning techniques, without using a threshold. In this case, the model is trained with training data that takes time-series state changes shown by training state data as input and outputs a label indicating whether the state is saturated or not. The saturation determination unit 104 may input the state data to be determined into the trained model and determine whether the state of the communication device 20 is saturated or not based on the label output from the model.
[0063] For example, the saturation determination unit 104 determines whether the time-series state of a certain communication device 20 is saturated by determining whether the amount of change indicated by the state data of the communication device 20 is greater than or equal to a threshold. If the saturation determination unit 104 determines that the amount of change indicated by the state data of a certain communication device 20 is greater than or equal to a threshold, it does not determine that the time-series state of the communication device 20 is saturated. If the saturation determination unit 104 determines that the amount of change indicated by the state data of a certain communication device 20 is less than a threshold, it determines that the time-series state of the communication device 20 is saturated.
[0064] For example, the saturation determination unit 104 determines whether the time-series state of a communication device 20 is saturated by determining whether the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold continues for a predetermined time or longer. If the saturation determination unit 104 determines that the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold does not continue for a predetermined time or longer, it does not determine that the time-series state of the communication device 20 is saturated. If the saturation determination unit 104 determines that the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold continues for a predetermined time or longer, it determines that the time-series state of the communication device 20 is saturated.
[0065] In this embodiment, when an abnormality is detected, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data of the communication device 20. That is, after an abnormality occurs, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data of the communication device 20. For example, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data after normalization has been performed. The saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state indicated by the state data after normalization has been performed.
[0066] In this embodiment, the saturation determination unit 104 will describe a case in which it determines whether a state is saturated based on the state within the time window W of the state data. The saturation determination unit 104 may determine whether a state is saturated regardless of the time window W. For example, the saturation determination unit 104 may determine whether a state is saturated based on the state within the reference period T2 of the state data. The saturation determination unit 104 may determine whether a state is saturated based on the state within the abnormal period T3 of the state data. The saturation determination unit 104 may determine whether a state is saturated based on the state over the entire period of the state data.
[0067] [Relationship Determination Unit] The relationship determination unit 105 determines, based on the determination result of the saturation determination unit 104, whether or not the status data is related to an anomaly in the communication service. A relationship between an anomaly and the status data means that there is a correlation between the anomaly and the status data. If the state indicated by the status data saturates when an anomaly occurs, then the anomaly is related to the status data. It is not necessary for the state indicated by the status data to necessarily saturate when an anomaly is detected; a certain degree of correlation between the anomaly and the status data is sufficient. For example, if an anomaly occurs n times (where n is a natural number), and the number of times the state indicated by the status data is saturated is k or more (where k is a natural number less than or equal to n), then the status data is related to the anomaly. However, k is assumed to be a predetermined percentage of n (for example, 70%) or more.
[0068] In this embodiment, the relationship determination unit 105 determines whether or not the status data is related to the abnormality when an abnormality is detected. That is, the relationship determination unit 105 determines whether or not the status data is related to the abnormality after the abnormality has occurred. The relationship determination unit 105 may also determine whether or not the status data is related to the abnormality before the abnormality is resolved (i.e., during the abnormality period T3), but in this embodiment, the relationship determination unit 105 determines whether or not the status data is related to the abnormality after the abnormality has been resolved as an example.
[0069] For example, the relationship determination unit 105 determines whether the status data is abnormally related based on the saturation period T1, which is determined to be the period during which a change in the status data is indicated and the state is saturated. The period during which a change in the status data is indicated is the period of timestamps included in the status data. The saturation period T1 is the period during which the saturation determination unit 104 determines that the state of the communication device 20 is saturated. That is, the saturation period T1 is the period during which the amount of change in the state indicated by the status data is less than the threshold, or the period during which the state during which the amount of change in the state indicated by the status data is less than the threshold continues.
[0070] For example, the relationship determination unit 105 may determine whether the state data is related to an abnormality based on the degree of agreement between the abnormal period T3, during which an abnormality occurred, and the saturation period T1, during which the state was determined to be saturated, within the period in which a change in the state data was indicated. Here, the degree of agreement is the length or proportion of the overlap period in which the saturation period T1 and the abnormal period T3 overlap. The more the saturation period T1 and the abnormal period T3 match, the longer the overlap period or the higher the proportion of the overlap period. For example, the relationship determination unit 105 calculates the overlap period in which the saturation period T1 and the abnormal period T3 overlap, and determines whether the state data is related to an abnormality based on this overlap period.
[0071] For example, the relationship determination unit 105 does not determine that the status data is abnormally related if the length of the overlapping period is less than a threshold, and determines that the status data is abnormally related if the length of the overlapping period is equal to or greater than the threshold. The threshold may be predetermined, or it may be determined based on the length of at least one of the saturation period T1 and the abnormal period T3. For example, the relationship determination unit 105 may determine that the status data is not abnormally related if the overlapping period is less than a predetermined percentage (e.g., 80%) of at least one of the saturation period T1 and the abnormal period T3, and determine that the status data is abnormally related if the overlapping period is equal to or greater than the predetermined percentage.
[0072] Furthermore, the method by which the relationship determination unit 105 determines whether or not the status data is related abnormally based on the saturation period T1 is not limited to a method based on the overlap period. For example, the relationship determination unit 105 may determine whether or not the status data is related abnormally by determining whether or not the saturation period T1 includes the abnormal occurrence timing t3. In this case, the relationship determination unit 105 may determine that the status data is not related abnormally if the saturation period T1 does not include the abnormal occurrence timing t3, and determine that the status data is related abnormally if the saturation period T1 includes the abnormal occurrence timing t3.
[0073] Alternatively, for example, the relationship determination unit 105 may determine whether the state data is abnormally related by determining whether the length of the saturation period T1 is greater than or equal to a threshold. The relationship determination unit 105 may determine that the state data is not abnormally related if the length of the saturation period T1 is less than the threshold, and determine that the state data is abnormally related if the length of the saturation period T1 is greater than or equal to the threshold.
[0074] For example, the relationship determination unit 105 may determine whether the status data is related to the anomaly based on the highest value of the status in other periods prior to the anomaly period T3 in which the anomaly occurred, within the period in which a change in the status data was indicated. The other period can be any period prior to the anomaly period T3. In this embodiment, we take the case where the reference period T2 corresponds to the other period as an example. The other period can be immediately before the anomaly period T3, or it can be a period that is not immediately before the anomaly period T3 but is somewhat prior to the anomaly period T3. The other period can also be a period prior to the reference period T2. Note that the anomaly period T3 may include a timing that is somewhat distant from the anomaly occurrence timing t3. The highest value is the maximum value of the communication volume within a certain period.
[0075] For example, the relationship determination unit 105 may determine whether the state data is abnormally related based on the highest value of the state in another period (in this embodiment, the reference period T2) and the highest value of the state in a period prior to the other period. The period prior to the other period is the period prior to the reference period T2. For example, the period prior to the other period may be the entire period prior to the reference period T2, or it may be a part of the period prior to the reference period T2. The highest value is the value that represents the peak of the time-series change. The relationship determination unit 105 refers to the state database DB and obtains the highest value in the other period and the highest value of the state in a period prior to the other period.
[0076] For example, the relationship determination unit 105 determines whether the highest value of the state in another period (in this embodiment, the reference period T2) is less than the highest value of the state in a period prior to that other period. If the highest value of the state in another period (in this embodiment, the reference period T2) is greater than or equal to the highest value of the state in a period prior to that other period, the relationship determination unit 105 determines that the state data is not related in an abnormal way, and if the highest value of the state in another period (in this embodiment, the reference period T2) is less than the highest value of the state in a period prior to that other period, the relationship determination unit 105 determines that the state data is related in an abnormal way.
[0077] For example, the relationship determination unit 105 may determine whether the status data is related to the abnormality based on the highest value of the state indicated by the status data during the abnormal period T3 in which the abnormality occurred, within the period in which a change in the status data was indicated. The relationship determination unit 105 obtains the highest value of the state indicated by the status data during the abnormal period T3. If the relationship determination unit 105 determines that the highest value is less than the threshold, it does not determine that the status data is related to the abnormality, but if it determines that the highest value is equal to or greater than the threshold, it determines that the status data is related to the abnormality.
[0078] For example, the relationship determination unit 105 may determine whether the state data is related to an abnormality based on the amount of change in state during each of the periods in which the state data shows a change, specifically between the saturation period T1 when the state is saturated and other periods prior to the abnormal period T3 when the abnormality occurs. As mentioned above, the reference period T2 is an example of other periods. The relationship determination unit 105 determines whether the difference between the amount of change in state during the saturation period T1 and the amount of change in state during other periods is greater than or equal to a threshold. If the relationship determination unit 105 determines that the difference is less than the threshold, it does not determine that the state data is related to an abnormality; however, if the difference is greater than or equal to the threshold, it determines that the state data is related to an abnormality.
[0079] The relationship determination unit 105 may determine whether the state data is related to an abnormality based on a determination method other than the determination method described above. For example, if the state indicated by the state data is saturated during the abnormal period T3, but the state is also saturated during periods when no abnormality occurs, the relationship determination unit 105 may determine that the state data is not related to an abnormality because saturation occurs regularly for this state data. On the other hand, if the state indicated by the state data is not saturated during periods when no abnormality occurs, but the state indicated by the state data is saturated during the abnormal period T3, the relationship determination unit 105 may determine that the state data is not related to an abnormality because saturation occurs only during the abnormal period T3.
[0080] [Recovery Processing Execution Unit] When an abnormality is detected, the recovery processing execution unit 106 executes recovery processing related to the recovery of the abnormality based on the status data determined to be related to the abnormality. The recovery processing can be any processing related to the abnormality that occurred. In this embodiment, the processing that displays the status data related to the abnormality on the administrator screen SC corresponds to the recovery processing. The recovery processing may be any other processing, for example, the processing that executes a pre-prepared recovery program, or the processing that sends a notification such as an email to the administrator.
[0081] For example, when an abnormality is detected, the recovery processing execution unit 106 sends display data for the administrator screen SC to the administrator terminal 40, thereby displaying the administrator screen SC on the administrator terminal 40. When an abnormality is detected, the recovery processing execution unit 106 displays an image on the administrator screen SC that shows the status data related to the abnormality. In the example in Figure 4, when an abnormality occurs, the recovery processing execution unit 106 displays on the administrator screen SC the status indicated by the status data determined by the relationship determination unit 105 when the same abnormality has occurred in the past. When an abnormality occurs and the status data related to that abnormality is identified, the recovery processing execution unit 106 executes a recovery process based on the status data related to that abnormality when the same abnormality occurs again.
[0082] [1-3-2. Functions implemented on the administrator terminal] For example, the administrator terminal 40 includes a data storage unit 400, an operation reception unit 401, and a display control unit 402. The data storage unit 400 is mainly implemented as a storage unit 42. The operation reception unit 401 and the display control unit 402 are mainly implemented as a control unit 41.
[0083] [Data Storage Unit] The data storage unit 400 stores data necessary for the administrator's work. For example, the data storage unit 400 stores data necessary for displaying the administrator screen SC. The data storage unit 400 may also store maintenance tools necessary for the administrator's work. The maintenance tools themselves may be well-known tools, for example, any tool capable of monitoring the status of at least one of hardware and software.
[0084] [Operation Reception Section] The operation reception unit 401 receives various operations from the administrator. For example, the operation reception unit 401 receives operations on the administrator screen SC.
[0085] [Display Control Unit] The display control unit 402 displays various screens on the display unit 45. For example, the display control unit 402 displays the administrator screen SC on the display unit 45.
[0086] [1-4. Processes executed by the monitoring system] Figure 8 shows an example of the process performed by the monitoring system 1. The control units 11 and 41 execute the programs stored in the storage units 12 and 42, respectively, thereby executing the process shown in Figure 8.
[0087] As shown in Figure 8, the server 10 communicates with each of the multiple communication devices 20, acquires status data, and stores it in the status database DB (S1). When an abnormality occurs, the administrator terminal 40 accepts an abnormality occurrence operation from the administrator (S2). The administrator terminal 40 sends abnormality occurrence data to the server 10 indicating that an abnormality occurrence operation has been performed (S3). The server 10 receives the abnormality occurrence data from the administrator terminal 40 (S4). The server 10 records the timing at which the abnormality occurrence data was received as the abnormality occurrence timing t3 in the storage unit 12 (S5). The administrator may specify the abnormality occurrence timing t3. In this case, the server 10 records the abnormality occurrence timing t3 specified by the administrator in the storage unit 12.
[0088] When the anomaly is resolved, the administrator terminal 40 accepts the administrator's anomaly resolution operation (S6). The administrator terminal 40 sends anomaly resolution data to the server 10 indicating that the anomaly resolution operation has been performed (S7). The server 10 receives the anomaly resolution data from the administrator terminal 40 (S8). The server 10 records the timing at which it received the anomaly resolution data as the anomaly resolution timing t5 in the storage unit 12 (S9). The administrator may specify the anomaly resolution timing t5. In this case, the server 10 records the anomaly resolution timing t5 specified by the administrator in the storage unit 12.
[0089] Server 10 sets the period from the reference timing t4, which is a predetermined time before the abnormality occurrence timing t3, to the abnormality occurrence timing t3 as the reference period T2 (S10). Server 10 sets the period from the abnormality occurrence timing t3 to the abnormality resolution timing t5 as the abnormality period T3 (S11). Server 10 refers to the status database DB and obtains the status data for the period including the abnormality occurrence timing t3 as the status data to be judged (S12). The status data to be judged is the status data that will be processed from S13 onwards. Subsequent processing is executed for each piece of status data.
[0090] Server 10 normalizes the state data to be judged (S13). In S13, in addition to normalization, Server 10 may perform preprocessing such as standardization or smoothing on the state data to be judged. Alternatively, for example, if the state indicated by the state data to be judged is below a predetermined threshold, Server 10 may refrain from performing the processes from S14 onward.
[0091] Server 10 determines whether the state indicated by the normalized state data is saturated (S14). In S14, Server 10 calculates the time-series change in the state indicated by the state data. Server 10 determines that the state indicated by the state data is saturated if the time-series change is below a threshold for a continuous period. Server 10 determines that the state indicated by the state data is not saturated if the time-series change is above a threshold, or if the time-series change is not below a threshold for a continuous period.
[0092] In S14, if the state indicated by the state data to be judged is not determined to be saturated (S14:N), the server 10 determines that the state data to be judged is not related to an abnormality (S15) and proceeds to the process in S21. In S14, if the state indicated by the state data to be judged is determined to be saturated (S14:Y), the server 10 sets the period from the saturation start timing t1 to the saturation end timing t2 as the saturation period T1 (S16).
[0093] Server 10 determines whether the first maximum value, which is the highest value of the state data to be judged during the reference period T2, is greater than or equal to the second maximum value, which is the highest value of the state data in a period prior to the reference period T2 (S17). If in S17 it is determined that the first maximum value is greater than or equal to the second maximum value (S17:Y), Server 10 proceeds to S15 and determines that the state data to be judged is not related to an abnormality.
[0094] In S17, if it is determined that the first highest value is less than the second highest value (S17:N), the server 10 determines whether the third highest value during the abnormal period T3 of the state data subject to determination is greater than or equal to the first highest value or the second highest value (S18). Here, we will explain the case where the server 10 compares the third highest value with the first highest value or the second highest value (i.e., only one of the first or second highest values is the subject of comparison), but the server 10 may also compare the third highest value with the first and second highest values.
[0095] In S18, if it is determined that the highest value in the abnormal period T3 is less than the first or second highest value (S18:N), the server 10 proceeds to S15 and determines that the state data being judged is not related to an abnormality. In S18, if it is determined that the third highest value, which is the highest value in the abnormal period T3, is greater than or equal to the first or second highest value (S18:Y), the server 10 determines whether the degree of agreement between the saturation period T1 and the abnormal period T3 is greater than or equal to a threshold (S19). In S19, if it is determined that the degree of agreement between the saturation period T1 and the abnormal period T3 is less than a threshold (S19:N), the server 10 proceeds to S15 and determines that the state data being judged is not related to an abnormality.
[0096] In S19, if it is determined that the degree of agreement between the saturation period T1 and the abnormal period T3 is equal to or greater than the threshold (S19:Y), the server 10 records the state data subject to judgment in the storage unit 12 as a candidate for state data related to an abnormality (S20). The server 10 determines whether all state data has been subject to judgment (S21). In S21, if there is still state data that has not been subject to judgment (S21:N), the process returns to S12 and processing for the next state data subject to judgment is executed.
[0097] In S21, if it is determined that all state data has been evaluated (S21:Y), the server 10 determines from the state data recorded as candidates in S20 that the state data is related to an anomaly (S22), and this process ends. In S22, the server 10 calculates a feature quantity indicating the degree of relationship between the state data and the anomaly based on the amount of change in the saturation period T1 and reference period T2 for each of the state data recorded as candidates in S20, and determines a predetermined number of state data in descending order of feature quantity as state data related to an anomaly. Alternatively, the process in S22 may be omitted, and the state data recorded as candidates in S20 may be determined as state data related to an anomaly without further processing.
[0098] [1-5. Summary of Embodiments] The monitoring system 1 of this embodiment determines whether the state of the communication device 20 is saturated based on state data relating to the time-series changes in the state of the communication device 20. Based on this determination result, the monitoring system 1 determines whether the state data is related to an anomaly in the communication service. As a result, the monitoring system 1 can determine whether the state data is related to an anomaly in the communication service. For example, the monitoring system 1 can speed up the recovery from an anomaly. In addition, the monitoring system 1 can also quickly detect anomalies and analyze the causes of anomalies.
[0099] Furthermore, when an abnormality is detected, the monitoring system 1 acquires status data, determines whether the status of the communication device 20 is saturated, and determines whether the status data is related to the abnormality. When an abnormality is detected, the monitoring system 1 performs recovery processing based on the status data determined to be related to the abnormality. This allows the monitoring system 1 to expedite recovery when an abnormality is detected.
[0100] Furthermore, when an anomaly is detected, the monitoring system 1 sets a period for status data based on the anomaly occurrence timing t3, and acquires status data related to the time-series changes in the state of the communication device 20 during that period. As a result, the monitoring system 1 can determine whether or not the status data is related to the anomaly based on the status data during the period related to the anomaly, thereby improving the accuracy of the determination. By accurately identifying the status data related to the anomaly, the monitoring system 1 can expedite the recovery from the anomaly.
[0101] Furthermore, the monitoring system 1 determines whether the state is saturated based on the state data after normalization. As a result, even if the range in which the amount of communication changes for each communication device 20 changes, the monitoring system 1 can absorb that range through normalization, thereby improving the accuracy of determining whether the state data is related to an anomaly.
[0102] Furthermore, the monitoring system 1 determines whether the status data is related to an abnormality based on the saturation period T1, which is determined to be the period during which the status data shows a change. In this way, the monitoring system 1 can determine whether the status data is related to an abnormality by using a period that is highly likely to be directly related to an abnormality, such as the saturation period T1, thereby improving the accuracy of the determination of whether the status data is related to an abnormality.
[0103] Furthermore, the monitoring system 1 determines whether the status data is related to the anomaly based on the highest value of the status during other periods prior to the anomaly period T3 in which the anomaly occurred, within the period in which the status data shows a change. This allows the monitoring system 1 to determine whether the status data is related to the anomaly, for example, by considering whether the communication volume is consistently high even when no anomaly is occurring, thereby improving the accuracy of the determination of whether the status data is related to the anomaly.
[0104] Furthermore, monitoring system 1 determines whether the status data is related to an anomaly based on the highest status value in other periods prior to the anomaly period T3 and the highest status value in periods prior to those other periods. This allows monitoring system 1 to determine whether the status data is related to an anomaly, for example, by considering whether the communication volume is consistently high even when no anomalies are occurring, thereby improving the accuracy of determining whether the status data is related to an anomaly.
[0105] Furthermore, monitoring system 1 determines whether the status data is related to the anomaly based on the highest value of the state indicated by the status data during the anomaly period T3, in which the anomaly occurred, within the period in which changes were shown in the status data. Even if the state indicated by the status data is saturated, if the value is not very high, that state may not be related to the anomaly. Therefore, monitoring system 1 can improve the accuracy of determining whether the status data is related to the anomaly by considering the highest value of the state indicated by the status data during the anomaly period T3.
[0106] Furthermore, the monitoring system 1 determines whether the status data is related to an anomaly based on the degree of agreement between the anomaly period T3, during which an anomaly occurred, and the saturation period T1, during which the state was determined to be saturated, within the period in which a change in the status data was indicated. The higher the degree of agreement between the saturation period T1 and the anomaly period T3, the higher the degree to which the status data is related to the anomaly. Therefore, by utilizing the degree of agreement, the monitoring system 1 can improve the accuracy of its determination of whether or not the status data is related to an anomaly.
[0107] Furthermore, the monitoring system 1 determines whether the status data is related to an anomaly based on the amount of change in the status during each of the periods in which the status data shows a change, specifically between the saturation period T1 when the status is saturated and other periods prior to the anomaly period T3 when the anomaly occurs. This allows the monitoring system 1 to determine whether the status data is related to an anomaly even if the status indicated by a certain status data is saturated, taking into account whether it is a state that is prone to saturation on a regular basis, thereby improving the accuracy of determining whether the status data is related to an anomaly.
[0108] [1-6. Variations] This disclosure is not limited to the embodiments described above. It may be modified as appropriate without departing from the spirit of this disclosure.
[0109] [1-6-1. Variation 1-1] Figure 9 shows an example of the functions implemented in the monitoring system 1 of Modification 1-1. The monitoring system 1 of Modification 1-1 includes a detection processing execution unit 107. The detection processing execution unit 107 is implemented by the control unit 11. For example, in the embodiment, a case in which a series of processes are executed when an abnormality is detected is described, but the series of processes may be executed before an abnormality is detected. Modification 1-1 describes a case in which the judgment result of the saturation determination unit 104 is used for detecting an abnormality. In Modification 1-1, the period used in the series of processes among the state data (the period corresponding to the time window W in Figure 7) is the period from a predetermined period before the present (for example, about 15 minutes) to the present.
[0110] In Modification 1-1, the status data acquisition unit 101 acquires status data before an abnormality is detected. The method of acquiring status data by the status data acquisition unit 101 is the same as in the embodiment. The status data acquisition unit 101 acquires status data in real time while the communication service is running. The status data acquisition unit 101 acquires status data regardless of whether an abnormality occurs or not. In Modification 1-1, the saturation determination unit 104 determines whether the state is saturated based on the status data before an abnormality is detected. The method of determining saturation by the saturation determination unit 104 is also the same as in the embodiment. The saturation determination unit 104 determines whether the state indicated by the status data is saturated or not, regardless of whether an abnormality occurs or not.
[0111] In Modification 1-1, the relationship determination unit 105 determines whether the state data is related to an abnormality before an abnormality is detected. Regardless of whether an abnormality has occurred, the relationship determination unit 105 determines whether the state indicated by the state data is related to an abnormality. For example, if the relationship determination unit 105 determines that the state indicated by the state data is not saturated, it determines that the state is not related to an abnormality, and if the state indicated by the state data is saturated, it determines that the state is related to an abnormality.
[0112] The detection processing execution unit 107 executes detection processing related to the detection of an anomaly based on the state data determined to be related to the anomaly. The detection processing can be any processing that detects the occurrence of an anomaly. In the modified example 1-1, the processing that displays the occurrence of an anomaly on the administrator screen SC corresponds to the detection processing. For example, when an anomaly is detected, the detection processing execution unit 107 displays an image indicating that an anomaly has been detected on the administrator screen SC. The detection processing execution unit 107 may also execute the detection processing by sending a notification to the administrator using means such as email.
[0113] In the modified example 1-1, the monitoring system 1 performs the following before an anomaly is detected: acquiring status data, determining whether the status of the communication device 20 is saturated, and determining whether the status data is related to the anomaly. Based on the status data determined to be related to the anomaly, the monitoring system 1 performs detection processing related to anomaly detection. This allows the monitoring system 1 to quickly detect anomalies.
[0114] [1-6-2. Variation 1-2] For example, monitoring system 1 can be applied to services other than communication services. For instance, monitoring system 1 may determine whether an anomaly in another service, such as an e-commerce service, a travel booking service, a financial service, a payment service, an online flea market service, or a video streaming service, is related to the status data of the device in that other service.
[0115] For example, monitoring system 1 may determine whether an anomaly in each of several services is related to the status data of the device in that service or in other services. Monitoring system 1 can be interpreted as a tenant in a cloud infrastructure or cloud platform. Monitoring system 1 consists of multiple offering services selected from a group of offering services, including software configurations or hardware configurations.
[0116] In the modified example 1-2, the status data acquisition unit 101 acquires status data for each of the multiple services. The process by which the status data acquisition unit 101 acquires status data for each service may be the same as in the embodiment. The status data acquisition unit 101 communicates with each of the multiple services and acquires the status data for said device.
[0117] In the modified example 1-2, the saturation determination unit 104 determines whether or not the device is saturated based on the device status data for each of the multiple services. The process by which the saturation determination unit 104 determines saturation based on the device status data for each service may be the same as in the embodiment.
[0118] In the modified example 1-2, the relationship determination unit 105 determines, based on the determination result of the saturation determination unit 104, whether or not the status data is related to an abnormality in each of the multiple services. The process by which the relationship determination unit 105 determines whether or not the status data is related to an abnormality in each service may be the same as in the embodiment.
[0119] The monitoring system 1 in modified example 1-2 acquires device status data for each of the multiple services. Based on the device status data for each of the multiple services, the monitoring system 1 determines whether the device is saturated or not. Based on this determination result, the monitoring system 1 determines whether the status data is related to an abnormality in each of the multiple services. In this way, the monitoring system 1 can respond to abnormalities in each of the multiple services.
[0120] [1-6-3. Other variations] For example, the above variations may be combined.
[0121] For example, in this embodiment, we have described a case where the main processing is performed on server 10, but the processing described as being performed on server 10 may also be performed on administrator terminal 40 or other computers, or it may be shared among multiple computers.
[0122] [2. Second Embodiment] A second embodiment, which is an example of another embodiment of the monitoring system 1 related to this disclosure, will be described. In the second embodiment, an example of processing related to anomalies in a service will be described. Conventional technology can only identify recovery procedures based on the state of predetermined devices among various devices in a system. Various devices may be involved in an anomaly that occurs in a system. Even if the state data of each of multiple devices is closely related, conventional technology cannot identify these relationships. Therefore, the monitoring system 1 of the second embodiment identifies the relationships between state data related to the state of devices to be monitored in a service.
[0123] [2-1. Hardware configuration of the monitoring system in the second embodiment] The hardware configuration of the monitoring system 1 in the second embodiment may be the same as that of the first embodiment.
[0124] [2-2. Overview of the Monitoring System] In this embodiment, we take the example that the communication device 20 is a server computer having a containerized network function (CNF). Furthermore, we take the example that the user terminal 30 is a smartphone. The communication device 20 relays communications from a large number of user terminals 30. For example, if the amount of communication that the communication device 20 needs to process reaches the amount of communication that the communication device 20 can process, an abnormality may occur in the communication service. The amount of communication can be expressed by a known indicator such as bps.
[0125] For example, if an anomaly occurs in a communication service, the administrator identifies the communication device 20 involved in the anomaly and performs recovery work. However, in a communication service, not just one communication device 20, but multiple communication devices 20 may be involved in the anomaly. For example, if an anomaly occurs in a communication service in a particular area, in addition to the communication device 20 that the administrator has identified as the cause of the anomaly, other communication devices 20 related to that device 20 may also be involved in the anomaly. It is difficult for the administrator to identify all communication devices 20 involved in the anomaly on their own.
[0126] For example, it is conceivable that an anomaly could be detected when the amount of data processed by the communication device 20 exceeds a threshold. However, there are cases where the communication device 20 can process more data than the upper limit set by the administrator, or where it can process less data than the upper limit set by the administrator. Furthermore, the administrator themselves may not know the upper limit of the communication device 20. Therefore, in this embodiment, no upper limit for the amount of data to be judged as an anomaly is set. For this reason, even if the amount of data processed by the communication device 20 reaches a certain level, it does not necessarily mean that an anomaly has occurred.
[0127] In this embodiment, the server 10 identifies the communication volume of other communication devices 20 that have similar characteristics to the communication volume of the communication device 20 related to the anomaly that occurred in the communication service. For example, the communication device 20 related to the anomaly is identified by analysis by the administrator. As shown in the modified example below, the communication device 20 related to the anomaly may also be identified by the saturation of the communication volume. Since the above-mentioned other communication devices 20 may be related to the anomaly, identifying the above-mentioned other communication devices 20 is useful for detecting and recovering from the anomaly.
[0128] Figure 10 shows an example of the time-series change in communication volume of communication device 20 when an anomaly is detected in the communication service. The horizontal axis in Figure 10 is the time axis. The vertical axis in Figure 10 is the axis showing communication volume. In the example in Figure 10, the change in communication volume for each of the five communication devices 20A to 20E located in the telecommunications carrier's facility is shown. Hereafter, when communication devices 20A to 20E are not distinguished, they will simply be referred to as communication device 20.
[0129] In the example shown in Figure 10, the communication volume of the communication device 20A increases sharply at a specific point in time. Subsequently, the communication device 20A becomes unable to process communication beyond a certain amount, and the communication volume becomes saturated. In this case, a user terminal 30 that attempts to connect to the communication device 20A in a specific area will be unable to use the communication service. In other words, an abnormality occurs in the communication service in that area. The user may contact the communication service call center to inform them that they are unable to use the communication service.
[0130] For example, the administrator becomes aware of an anomaly after receiving inquiries from multiple users. At this point, the administrator has not identified which communication device 20 is causing the anomaly. Subsequently, the communication volume of communication device 20A decreases, and the anomaly resolves itself. Furthermore, the administrator becomes aware that the anomaly has resolved itself because they no longer receive inquiries from users. Once the anomaly resolves itself, the user terminal 30 becomes able to use the communication service via communication device 20A.
[0131] For example, after an anomaly occurs in a specific area and is resolved, the administrator analyzes the communication volume of each communication device 20A to 20E. The administrator identifies that during the period when the anomaly occurred, the communication volume of communication device 20A increased sharply and reached saturation. The administrator inputs information into the administrator terminal 40 indicating that the communication volume of communication device 20A is related to the anomaly that occurred in the area. The administrator terminal 40 uploads this information to the server 10. The server 10 stores the area and associates it with the fact that the communication volume of communication device 20A is related to the anomaly in that area.
[0132] For example, server 10 clusters each of the communication devices 20A to 20E to determine whether there is any state data among the state data of each of the communication devices 20B to 20E that is related to the state data of communication device 20A. In the example in Figure 10, clustering identifies that the characteristics of the communication volume of communication device 20A and the characteristics of the communication volume of communication device 20E are similar.
[0133] For example, server 10 records in storage unit 12 that an anomaly in a certain area is related not only to communication device 20A designated by the administrator, but also to communication device 20E identified through clustering. If another anomaly occurs in the same area, server 10 sends a notification to administrator terminal 40 prompting it to check the status of communication device 20E as well as communication device 20A.
[0134] Figure 11 shows an example of a notification displayed on the administrator terminal 40. For example, suppose that in the past, an abnormality was detected in the communication volume of communication device 20A, and the call center for the communication service receives an inquiry from a user who tried to use the communication service from a specific area X. At this point, if the call center receives an inquiry from a user who tried to use the communication service in the same area X, it is detected that an abnormality has occurred in the communication service in area X. In this case, the server 10 sends display data to the administrator terminal 40, prompting it to check the communication device 20A specified by the administrator and the communication device 20E identified by clustering, through the administrator screen SC.
[0135] For example, when the administrator terminal 40 receives display data from the administrator screen SC, a message is displayed prompting it to check the communication volume of communication device 20A and the communication volume of communication device 20E, which is related to the communication volume of communication device 20A. The administrator screen SC may display graphs showing the current communication volumes of communication devices 20A and 20E, or it may display graphs showing the communication volumes of communication devices 20A and 20E when similar abnormalities occurred in the past. The administrator checks the administrator screen SC and performs maintenance on communication devices 20A and 20E. The maintenance of communication devices 20A and 20E itself may be performed by known methods.
[0136] As described above, when an anomaly occurs in the communication service, the monitoring system 1 identifies the communication device 20A related to the anomaly and the related communication device 20E through clustering. For example, if a similar anomaly occurs again, the monitoring system 1 makes it easier for the administrator to identify the communication device 20 that caused the anomaly, allowing for rapid recovery work. The details of the monitoring system 1 will be described below.
[0137] [2-3. Functions implemented by the monitoring system] Figure 12 shows an example of the functions implemented in monitoring system 1. Figure 12 shows the functions implemented in server 10 and the functions implemented in administrator terminal 40. The functions of communication device 20 and user terminal 30 are the same as those in known communication services and are therefore omitted in Figure 12. For example, communication device 20 has the function of transmitting status data or a part thereof to server 10, as described later. User terminal 30 has the function of allowing users to use the communication service.
[0138] [2-3-1. Functions implemented by the server] For example, server 10 includes a data storage unit 100, a status data acquisition unit 101, a clustering execution unit 108, a related data identification unit 109, and a recovery processing execution unit 106. The data storage unit 100 is implemented by a storage unit 12. The status data acquisition unit 101, the clustering execution unit 108, the related data identification unit 109, and the recovery processing execution unit 106 are implemented by a control unit 11.
[0139] [Data Storage Unit] The data storage unit 100 stores data related to communication services. For example, the data storage unit 100 stores a status database DB.
[0140] Figure 13 shows an example of a status database DB. The status database DB is a database that stores the status data for each of the multiple communication devices 20. For example, the status database DB stores device identification information that can identify each of the multiple communication devices 20, status data that indicates the state of the communication device 20, abnormal cause identification information that can identify whether or not it corresponds to the abnormal cause data described later, and related identification information that can identify whether or not it corresponds to the related data described later. Other data may be stored in the status database DB. For example, the status database DB may store information indicating the location where the communication devices 20 are located.
[0141] The identification information of the communication device 20 may be any information, for example, an IP address, a device name, or a MAC address. In this embodiment, the status data is data relating to the time-series changes in the state of the communication device 20. The state of the communication device 20 can also be said to be the load on the communication device 20. The state of the communication device 20 may refer to the hardware state or the software state.
[0142] In this embodiment, the amount of communication performed by the communication device 20 is given as an example where it corresponds to the state of the communication device 20, but other states of the communication device 20 may also correspond to the state of the communication device 20. Other states may include the amount of resources consumed by the communication device 20, the processing time (response time) for requests to the service, or the number of errors returned for requests. The state data may be an indicator known as a golden signal metric, or other indicators used in well-known benchmark tests. For example, the state data may be CPU usage, memory usage, power consumption, communication speed, temperature, or a combination thereof.
[0143] The status data may also indicate the status of performance in the communication service or other services. Other services are services that users utilize via the communication service, such as e-commerce services, travel booking services, payment services, or financial services. For example, the status data may be the number of users connected to the communication device 20 (logged in to the communication service or other services), the number of payments made by users via the communication device 20, or the number of orders made by users via the communication device 20. As the number of payments, payment amounts, number of orders, and order amounts increase, the load on the communication device 20 increases, so the number of payments, payment amounts, number of orders, and order amounts also correspond to the status of the communication device 20.
[0144] In this embodiment, the status data for each of the multiple communication devices 20 indicates the time-series change in the state of the communication device 20. For example, the status data includes the date and time when the state of the communication device 20 was acquired, and a numerical value indicating the state of the communication device 20 at that date and time. The state of the communication device 20 may be expressed in a format other than a numerical value, such as characters. In this embodiment, since the amount of communication of the communication device 20 corresponds to the state of the communication device 20, the status data includes the date and time when the amount of communication of the communication device 20 was acquired, and the amount of communication of the communication device 20 at that date and time.
[0145] Furthermore, the status data of each of the multiple communication devices 20 may represent the state at a specific pinpoint (a single point in time) rather than the time-series change in the state of the communication device 20. For example, some status data may represent a time-series change in state, while other status data may represent the state at a specific pinpoint. By collecting a large number of status data representing the state at a specific pinpoint, the overall state data may represent a time-series change in state.
[0146] Furthermore, the data stored in the data storage unit 100 is not limited to the status database DB. The data storage unit 100 can store any data related to the communication service. For example, the data storage unit 100 may store the data necessary to display the administrator screen SC on the administrator terminal 40. The data storage unit 100 may also store various thresholds used in the processing described later. The thresholds may be specified by the administrator, or they may be determined based on machine learning or statistical methods.
[0147] [Status data acquisition unit] The status data acquisition unit 101 acquires status data for each of the multiple communication devices 20 that are monitored in the communication service. For example, the status data acquisition unit 101 periodically requests the status of each of the multiple communication devices 20 (for example, every 1 to 30 seconds or every 1 to 5 minutes). When a communication device 20 receives a request from the server 10 to the status data acquisition unit 101, it sends its device identification information and information indicating its status to the server 10. The status data acquisition unit 101 adds the current date and time and the information received from the communication device 20 to the status data associated with the device identification information of the communication device 20.
[0148] The communication device 20 may also store status data for a certain period of time and send the status data to the server 10 all at once. In this case, the status data acquisition unit 101 acquires the status data from the communication device 20 all at once and stores it in the status database DB. Alternatively, the status data acquisition unit 101 may acquire status data from the communication device 20 irregularly rather than periodically. The status data acquisition unit 101 may also acquire status data from the communication device 20 at a timing instructed by the administrator.
[0149] Communication device 20 is an example of a device to be monitored in a communication service. The device to be monitored may be any device and is not limited to communication device 20. Where "communication device 20" is written in this embodiment, it can be read as any other device to be monitored. The other device may be any device used in the communication service, for example, server 10, other computers other than server 10, components such as CPUs, memory, and network cards included in server 10 and other computers, batteries, antennas, communication cables, or sensors.
[0150] In this embodiment, the status data acquisition unit 101 acquires status data when an abnormality is detected. The status data acquired when an abnormality is detected indicates the state during all or part of the period from when the abnormality occurs until it is resolved. Hereafter, this period will be referred to as the abnormality period. The status data may indicate not only all or part of the abnormality period, but also the state during other periods other than the abnormality period. The status data acquisition unit 101 may acquire status data during the period from when the abnormality occurs until it is resolved, or it may acquire status data retrospectively after the abnormality has been resolved.
[0151] In this embodiment, the status data does not represent the state at a specific pinpoint timing, but rather represents the state over a period of a certain length. The status data acquisition unit 101 may acquire status data that represents the state over the entire past period, but in this embodiment, it acquires status data that represents the state over a portion of the past period. For example, the status data acquisition unit 101 acquires status data relating to the time-series changes in the state of the device over a period that includes the time when an abnormality was detected.
[0152] [Clustering execution unit] The clustering execution unit 108 performs clustering on the state data of each of the multiple communication devices 20. Clustering is the process of identifying state data that have similar characteristics to each other. Clustering is sometimes called grouping. The clustering execution unit 108 performs clustering so that state data with similar characteristics to each other belong to the same cluster. Each of the multiple state data belonging to a given cluster exhibits similar characteristics to each other. The number of clusters may be predetermined or may not be predetermined.
[0153] The clustering method itself may be a known method. For example, the clustering execution unit 108 may perform clustering based on k-means, hierarchical clustering, density-based clustering, Gaussian mixture models, or spectral clustering. Alternatively, the clustering execution unit 108 may perform clustering based on a machine learning model created using machine learning techniques. The machine learning model may also be a model created using a known method. For example, the clustering execution unit 108 may perform clustering by inputting the state data of each of the multiple communication devices 20 to a model such as an RNN (Recurrent Neural Network), LSTM (Long Short-Term Memory), or Transformer-based model.
[0154] Figure 14 shows an example of the results of clustering. In Figure 14, a vector space is shown in which the state data of each of the numerous communication devices 20 has been characterized. The white circles in Figure 14 are anomaly relationship data, which will be described later. The black circles in Figure 14 are other state data other than anomaly relationship data. The closeness of the circles in Figure 14 means that their features are similar. In this embodiment, since the state data shows time-series changes in state, the clustering execution unit 108 performs clustering based on the time-series changes in state shown by the state data of each of the multiple communication devices 20. For example, the clustering execution unit 108 calculates the waveform features shown by each of the state data.
[0155] For example, the clustering execution unit 108 calculates time-domain features of state data. The clustering execution unit 108 calculates the mean, variance, standard deviation, minimum, maximum, rising edge timing, mean change, maximum change, minimum change, or a combination thereof of the time-series state changes represented by a given state data as time-domain features of the state data. The clustering execution unit 108 may also calculate the period during which the state represented by a given state data is saturated, the amount of communication at that time, the frequency of saturation, or a combination thereof as time-domain features of the state data. The clustering execution unit 108 may also calculate features in the frequency domain of the state data instead of the time domain.
[0156] For example, the clustering execution unit 108 performs clustering based on the features of each of the multiple state data so that state data with similar features belong to the same cluster. When a machine learning model is used, the features of the state data are sometimes called embedding representations. The actual clustering process is performed inside the machine learning model. The clustering execution unit 108 performs clustering by inputting each of the multiple state data into the machine learning model and obtaining the output from the machine learning model. For example, the output of the machine learning model is data showing the relationship between each of the multiple state data and the cluster to which that state data belongs. Unsupervised learning methods may be used for clustering. The clustering execution unit 108 may also perform clustering based on a method that captures the waveforms represented by the state data as images and determines whether the images are similar or not.
[0157] In this embodiment, the clustering execution unit 108 performs clustering based on the status data of each of the multiple devices when an abnormality is detected. For example, the clustering execution unit 108 may automatically perform clustering when the server 10 detects an abnormality, or it may perform clustering when the administrator of the communication service performs a predetermined operation on the administrator terminal 40. The clustering execution unit 108 only needs to perform clustering after an abnormality is detected. An example in which clustering is performed before an abnormality is detected will be explained in Modification 2-1 below.
[0158] In the example in Figure 14, four anomaly-related data points are specified by the administrator, and the clustering execution unit 108 identifies four clusters C1 to C4 through clustering. For example, the white circle (anomaly-related data point) in cluster C1 is the status data of communication device 20A in the examples in Figures 10 and 11. One of the black circles in cluster C1 is the status data of communication device 20E in the examples in Figures 10 and 11. The white circles in each of the other clusters C2 to C4 are the status data of communication device 20 that caused the anomaly in an area other than area X where communication device 20A was the cause of the anomaly. For other areas as well, the status data of other communication devices 20 related to the communication device 20 is identified through clustering. The number of status data points belonging to each of the clusters C1 to C4 may be the same or different from each other.
[0159] [Related Data Identification Department] The related data identification unit 109 identifies related data, which are other state data related to the abnormality factor data, which are state data relating to the cause of the abnormality in the service, based on the clustering execution result. In this embodiment, the administrator specifies which communication device 20's state data corresponds to the abnormality factor data. When the server 10 receives information from the administrator terminal 40 indicating the state data of the communication device 20 that the administrator has designated as abnormality factor data, it updates the abnormality factor identification information in the state database DB that is associated with the device identification information of the communication device 20. In the example in Figure 13, the server 10 updates the abnormality factor identification information of communication device 20A to indicate that the state data of communication device 20A is abnormality factor data relating to the cause of the abnormality in area X.
[0160] Related data are state data identified through clustering as having similar characteristics to anomaly factor data. In other words, related data are state data belonging to the same cluster as anomaly factor data. The related data identification unit 109 identifies at least one related data for each anomaly factor data. If no state data with similar characteristics to the anomaly factor data exists, the related data identification unit 109 may identify no related data at all. The related data identification unit 109 updates the related information associated with the state data of the communication device 20 identified as related data in the state database DB. In the example in Figure 13, the related data identification unit 109 identifies the state data of the communication device 20E as related data, and therefore updates the related information associated with the state data of the communication device 20E.
[0161] Furthermore, the related data identification unit 109 may identify state data as related data even if it does not belong to the same cluster as the anomaly factor data, as long as it has somewhat similar characteristics. For example, the related data identification unit 109 may identify state data as related data even if it does not belong to the same cluster as the anomaly factor data, but the difference in features (in the example of Figure 14, the distance in the vector space) is less than a threshold. Also, when the related data identification unit 109 calculates the difference in features, it can identify related data that is closer to the more distinctive data by multiplying the data point interval with a weight (weight parameter). Conversely, the related data identification unit 109 does not have to identify state data as related data even if it belongs to the same cluster as the anomaly factor data. For example, the related data identification unit 109 may identify a predetermined number of state data as related data from among multiple state data belonging to the same cluster as the anomaly factor data, in order of the smallest difference from the features of the anomaly factor data.
[0162] In this embodiment, the related data identification unit 109 identifies related data when an abnormality is detected. For example, the related data identification unit 109 may automatically identify related data when the server 10 detects an abnormality, or it may identify related data when the administrator of the communication service performs a predetermined operation on the administrator terminal 40. The related data identification unit 109 only needs to identify the related data after an abnormality is detected. An example in which related data is identified before an abnormality is detected will be explained in Modification 2-1 below.
[0163] [Recovery Processing Execution Unit] When an anomaly is detected, the recovery processing unit 106 executes recovery processing related to the recovery of the anomaly based on the anomaly cause data and related data. The recovery processing can be any processing related to the anomaly that occurred. In this embodiment, the processing that displays the status data related to the anomaly on the administrator screen SC corresponds to the recovery processing. The recovery processing may be any other processing, for example, the processing that executes a pre-prepared recovery program, or the processing that sends a notification such as an email to the administrator.
[0164] For example, when an abnormality is detected, the recovery processing execution unit 106 sends display data for the administrator screen SC to the administrator terminal 40 based on the abnormality cause data and related data, thereby displaying the administrator screen SC on the administrator terminal 40. When an abnormality is detected, the recovery processing execution unit 106 displays images representing the abnormality cause data and related data on the administrator screen SC.
[0165] In the example shown in Figure 12, when an abnormality occurs, the recovery processing execution unit 106 displays on the administrator screen SC the state indicated by the state data determined by the related data identification unit 109 when the same abnormality occurred in the past. When an abnormality occurs and state data related to that abnormality is identified, the recovery processing execution unit 106 executes recovery processing based on the state data related to that abnormality when the same abnormality occurs again.
[0166] In the example of the status database DB in Figure 13, when an anomaly occurs in area X, the recovery processing execution unit 106 refers to the anomaly cause identification information in the status database DB and identifies that the status data of the communication device 20A is the anomaly cause data that caused the anomaly. Furthermore, the recovery processing execution unit 106 refers to the related identification information in the status database DB and identifies that the status data of the communication device 20E is related data. Based on the identified anomaly cause data and related data, the recovery processing execution unit 106 generates display data for the administrator screen SC and transmits it to the administrator terminal 40.
[0167] [2-3-2. Functions implemented on the administrator terminal] For example, the administrator terminal 40 includes a data storage unit 400, an operation reception unit 401, and a display control unit 402. The data storage unit 400 is mainly implemented as a storage unit 42. The operation reception unit 401 and the display control unit 402 are mainly implemented as a control unit 41.
[0168] [Data Storage Unit] The data storage unit 400 stores data necessary for the administrator's work. For example, the data storage unit 400 stores data necessary for displaying the administrator screen SC. The data storage unit 400 may also store maintenance tools necessary for the administrator's work. The maintenance tools themselves may be well-known tools, for example, any tool capable of monitoring the status of at least one of hardware and software.
[0169] [Operation Reception Section] The operation reception unit 401 receives various operations from the administrator. For example, the operation reception unit 401 receives operations on the administrator screen SC.
[0170] [Display Control Unit] The display control unit 402 displays various screens on the display unit 45. For example, the display control unit 402 displays the administrator screen SC on the display unit 45.
[0171] [2-4. Processes executed by the monitoring system] Figure 15 shows an example of the process performed by the monitoring system 1. The control units 11 and 41 execute the programs stored in the memory units 12 and 42, respectively, thereby executing the process shown in Figure 15.
[0172] As shown in Figure 15, the server 10 communicates with each of the multiple communication devices 20, acquires status data, and stores it in the status database DB (S100). The process in S100 is executed periodically. When an abnormality occurs, the administrator terminal 40 accepts an abnormality occurrence operation from the administrator (S101). The administrator terminal 40 sends abnormality occurrence data to the server 10 indicating that an abnormality occurrence operation has been performed (S102). The server 10 receives the abnormality occurrence data from the administrator terminal 40 (S103). The server 10 records the timing at which the abnormality occurrence data was received as the abnormality occurrence timing in the storage unit 12 (S104). The administrator may specify the abnormality occurrence timing. In this case, the server 10 records the abnormality occurrence timing specified by the administrator in the storage unit 12.
[0173] When the anomaly is resolved, the administrator terminal 40 accepts the administrator's anomaly resolution operation (S105). The administrator terminal 40 sends anomaly resolution data to the server 10 indicating that the anomaly resolution operation has been performed (S106). The server 10 receives the anomaly resolution data from the administrator terminal 40 (S107). The server 10 records the timing at which it received the anomaly resolution data as the anomaly resolution timing in the storage unit 12 (S108). The administrator may specify the anomaly resolution timing. In this case, the server 10 records the anomaly resolution timing specified by the administrator in the storage unit 12.
[0174] The administrator analyzes the status data of each of the multiple communication devices 20 and identifies the abnormal cause data. When the administrator specifies the abnormal cause data, the administrator terminal 40 transmits abnormal cause identification information that identifies the abnormal cause data to the server 10 (S109). In the example shown in Figure 10, the administrator terminal 40 transmits abnormal cause identification information to the server 10 indicating that the status data of communication device 20A is abnormal cause data.
[0175] When server 10 receives abnormal cause identification information from administrator terminal 40 (S110), it updates the status database DB (S111). In the example in Figure 13, server 10 updates the abnormal cause identification information associated with the device identification information of communication device 20A in the status database DB. Server 10 performs clustering based on the status data of each of the multiple communication devices 20 (S112). Based on the clustering results, server 10 identifies related data that is associated with the abnormal cause data (S113), and this process ends. Thereafter, if a similar abnormality occurs, server 10 performs recovery processing based on the abnormal cause data and the related data.
[0176] [2-5. Summary of Embodiments] The monitoring system 1 of this embodiment acquires status data relating to the status of each of the multiple communication devices 20. The monitoring system 1 performs clustering on the status data of each of the multiple communication devices 20. Based on the results of the clustering, the monitoring system 1 identifies related data, thereby identifying the relationships between the status data of the communication devices 20. For example, even if the administrator is unaware of the existence of related data, the monitoring system 1 can identify the related data through clustering, allowing the administrator to quickly recover in the event of an anomaly. The monitoring system 1 can support the administrator's work.
[0177] Furthermore, when an anomaly is detected, the monitoring system 1 acquires the status data of each of the multiple communication devices 20. When an anomaly is detected, the monitoring system 1 performs clustering based on the status data of each of the multiple communication devices 20. When an anomaly is detected, the monitoring system 1 identifies the relevant data. When an anomaly is detected, the monitoring system 1 performs recovery processing based on the anomaly cause data and the relevant data. Through the recovery processing, the monitoring system 1 can quickly recover from anomalies in communication services. For example, if a similar anomaly occurs again, the monitoring system 1 can have the administrator check the anomaly cause data and the relevant data, thereby effectively supporting the administrator's work.
[0178] Furthermore, the monitoring system 1 performs clustering based on the time-series state changes indicated by the state data of each of the multiple communication devices 20. The monitoring system 1 can identify relationships between state data from the time-series changes in the state of the communication devices 20. For example, by identifying state data with similar time-series characteristics to anomaly cause data as related data, the monitoring system 1 can perform a more accurate recovery process based on the time-series changes in the state data.
[0179] [2-6. Variations] This disclosure is not limited to the embodiments described above. This disclosure may be modified as appropriate without departing from the spirit of this disclosure.
[0180] Figure 16 shows an example of the functions implemented in the modified version of the monitoring system 1. The modified version of the monitoring system 1 includes a detection processing execution unit 110, a saturation determination unit 104, and an abnormality cause data identification unit 111. The detection processing execution unit 110, the saturation determination unit 104, and the abnormality cause data identification unit 111 are implemented by the control unit 11.
[0181] [2-6-1. Variation 2-1] For example, in the embodiment, a series of processes are described when an abnormality is detected, but the series of processes may be executed before an abnormality is detected. Modification 2-1 describes a case in which the related data identified by the related data identification unit 109 is used for detecting an abnormality, rather than for recovering from an abnormality as in the embodiment.
[0182] In Modification 2-1, the status data acquisition unit 101 acquires status data for each of the multiple communication devices 20 before an abnormality is detected. The method of acquiring status data by the status data acquisition unit 101 is the same as in the embodiment. The status data acquisition unit 101 acquires status data in real time while the communication service is running. The status data acquisition unit 101 acquires status data regardless of whether an abnormality has occurred or not. For example, even if an abnormality has actually occurred, if it is before the abnormality is detected by the monitoring system 1, it corresponds to the above "before an abnormality is detected".
[0183] In Modification 2-1, the clustering execution unit 108 performs clustering based on the status data of each of the multiple communication devices 20 before an anomaly is detected. The method of performing clustering is the same as in the embodiment. The clustering execution unit 108 performs clustering regardless of whether an anomaly occurs or not. In Modification 2-1, the related data identification unit 109 identifies related data before an anomaly is detected. The method of identifying related data is the same as in the embodiment. The related data identification unit 109 identifies related data regardless of whether an anomaly occurs or not.
[0184] The monitoring system 1 includes a detection processing execution unit 110. The detection processing execution unit 110 performs detection processing related to the detection of anomalies based on anomaly cause data and related data. The detection processing can be any processing that detects an anomaly that has occurred. In modified example 2-1, the processing that displays the occurrence of an anomaly on the administrator screen SC corresponds to the detection processing. For example, when an anomaly is detected, the detection processing execution unit 110 displays an image indicating that an anomaly has been detected on the administrator screen SC. The detection processing execution unit 110 may also perform detection processing by sending a notification to the administrator using means such as email.
[0185] In Modification 2-1, a criterion indicating whether or not an anomaly has occurred is associated with each of the anomaly factor data and the related data. This criterion is determined based on anomalies that have occurred in the past. For example, it may be some kind of threshold, or it may be whether or not saturation has occurred, as in Modification 2-2 described below. The detection processing execution unit 110 determines whether or not each of the anomaly factor data and the related data satisfies the criterion.
[0186] For example, the detection processing execution unit 110 executes the detection process when it determines that at least one of the abnormal cause data and the related data meets the criteria. Alternatively, the detection processing execution unit 110 may execute the detection process when it determines that both the abnormal cause data and the related data meet the criteria. Or, the detection processing execution unit 110 may execute the detection process when it determines that only one of the abnormal cause data or the related data meets the criteria.
[0187] In the modified version 2-1, the monitoring system 1 acquires status data from each of the multiple communication devices 20, performs clustering, and identifies related data before an anomaly is detected. Based on the anomaly cause data and related data, the monitoring system 1 performs detection processing related to anomaly detection. This allows the monitoring system 1 to quickly detect anomalies.
[0188] [2-6-2. Modified Example 2-2] For example, in the embodiment, an example is given where an administrator analyzes status data and identifies abnormality cause data. Status data may be analyzed by a monitoring system and classified as abnormality cause data. For example, in a communication service, there are many communication devices 20 and various abnormalities can occur, so it can be difficult for an administrator to identify the communication device 20 that caused the abnormality. In this case, the communication device 20 that caused the abnormality may not be able to process any further communications, and the communication volume may become saturated.
[0189] Communication volume saturation occurs when the communication volume is positive (not zero) and does not change or changes very little. No change in communication volume means the change in communication volume is zero. Nearly no change in communication volume means the change in communication volume is below a threshold. The specific method for determining saturation will be described later. While an upper limit may be set for communication volume, there may be cases where the communication device 20 can process more communication volume than the upper limit set by the administrator, or where it can process less communication volume than the upper limit set by the administrator, and furthermore, the administrator themselves may not know the upper limit of the communication device 20. Therefore, it will be assumed that no upper limit for communication volume is set. For this reason, even if the communication volume of the communication device 20 reaches a certain level, an abnormality will not be detected.
[0190] In the example shown in Figure 10, the communication volume of communication device 20A increases rapidly, and then the communication volume of communication device 20A becomes saturated. During this period, communication device 20A cannot process any further communication, so a user terminal 30 attempting to connect to communication device 20A in a specific area may not be able to use the communication service. In modification 2-2, the period from the start of saturation to the end of saturation is treated as an example of a period during which an anomaly occurred, but the period from the start of saturation to the present time, before saturation has ended, may also be treated as an example of a period during which an anomaly occurred. For example, the period from the start of saturation to a time specified by the administrator may also be treated as an example of a period during which an anomaly occurred.
[0191] The monitoring system 1 of Modification 2-2 includes a saturation determination unit 104 and an abnormality factor data identification unit 111. The saturation determination unit 104 determines whether the state indicated by the state data of a plurality of communication devices 20 is saturated, based on the state data of each of the communication devices 20. Saturation determined by the saturation determination unit 104 is when the amount of change in the state indicated by the state data falls below a threshold, or when the state in which the amount of change is below the threshold continues for a predetermined time or longer. In the case of communication volume, where there is no upper limit in principle, if the communication volume increases, the amount of change mainly means the increase. Therefore, saturation occurs when, after the total amount of change has increased over a certain period, the increase in the state indicated by the state data falls below a threshold, or when the state in which the increase is below the threshold continues for a predetermined time or longer. The amount of change may be an accumulated value over a certain period, or it may be the amount of change between a certain point in time and the next point in time. The amount of change may also mean the decrease in communication volume. In other words, the saturation determination unit 104 makes the determination of saturation based on the time-series change of the amount of change itself indicated by the state data.
[0192] For example, the saturation determination unit 104 determines that the time-series state of the communication device 20 is saturated when the increase in the state indicated by the state data changes from a state where the increase is greater than or equal to a threshold to a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer. The saturation determination unit 104 determines that the time-series state of the communication device 20 is not saturated even if the increase in the state indicated by the state data changes from a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer, if the increase was less than the threshold before that. For example, the saturation determination unit 104 may determine that the time-series state of the communication device 20 is saturated when the state indicated by the state data is greater than or equal to a threshold, and the increase in that state changes from a state where the increase is less than the threshold, or when the state where the increase is less than the threshold continues for a predetermined time or longer. The saturation determination unit 104 determines that the time-series state of the communication device 20 is not saturated if the increase in the state indicated by the state data is less than a threshold, or if the state in which the increase is less than a threshold continues for a predetermined time or longer, as long as the state remains below the threshold.
[0193] For example, the saturation determination unit 104 determines whether the time-series state of the communication devices 20 is saturated based on the state data of each of the multiple communication devices 20 and a predetermined threshold. In other words, for each state data of the communication device 20, the saturation determination unit 104 determines whether the time-series state indicated by the state data is saturated based on a predetermined threshold. The threshold referenced by the saturation determination unit 104 when determining saturation may be specified by the administrator or calculated from past anomaly trends. The threshold may be determined based on a model using machine learning techniques.
[0194] The saturation determination unit 104 may determine whether the state of the communication device 20 is saturated or not based on a model using machine learning techniques, without using a threshold. In this case, the model is trained with training data that takes time-series state changes shown by training state data as input and outputs a label indicating whether the state is saturated or not. The saturation determination unit 104 may input the state data to be determined into the trained model and determine whether the state of the communication device 20 is saturated or not based on the label output from the model.
[0195] For example, the saturation determination unit 104 determines whether the time-series state of a certain communication device 20 is saturated by determining whether the amount of change indicated by the state data of the communication device 20 is greater than or equal to a threshold. If the saturation determination unit 104 determines that the amount of change indicated by the state data of a certain communication device 20 is greater than or equal to a threshold, it does not determine that the time-series state of the communication device 20 is saturated. If the saturation determination unit 104 determines that the amount of change indicated by the state data of a certain communication device 20 is less than a threshold, it determines that the time-series state of the communication device 20 is saturated.
[0196] For example, the saturation determination unit 104 determines whether the time-series state of a communication device 20 is saturated by determining whether the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold continues for a predetermined time or longer. If the saturation determination unit 104 determines that the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold does not continue for a predetermined time or longer, it does not determine that the time-series state of the communication device 20 is saturated. If the saturation determination unit 104 determines that the state in which the amount of change indicated by the state data of a communication device 20 is below a threshold continues for a predetermined time or longer, it determines that the time-series state of the communication device 20 is saturated.
[0197] For example, when an abnormality is detected, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data of the communication device 20. That is, after an abnormality occurs, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data of the communication device 20. For example, the saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state data after normalization has been performed. The saturation determination unit 104 determines whether the state of the communication device 20 is saturated based on the state indicated by the state data after normalization has been performed.
[0198] The saturation determination unit 104 may determine whether a state is saturated based on the states within a predetermined time window among the states indicated by the state data. The saturation determination unit 104 may also determine whether a state indicated by the state data is saturated regardless of the time window. For example, the saturation determination unit 104 may determine whether a state is saturated based on the states within a reference period among the states indicated by the state data. The saturation determination unit 104 may also determine whether a state is saturated based on the states within an abnormal period in which an abnormality occurred among the states indicated by the state data. The saturation determination unit 104 may also determine whether a state is saturated based on the states over the entire period indicated by the state data.
[0199] The abnormal cause data identification unit 111 identifies the state data of each of the multiple communication devices 20 that is determined to indicate that the state of the communication device 20 is saturated as abnormal cause data. In the example in Figure 10, since the state of communication device 20A shows a tendency toward saturation, the abnormal cause data identification unit 111 identifies the state data of communication device 20A as abnormal cause data. The processing after the abnormal cause data has been identified is the same as in the embodiment.
[0200] The monitoring system 1 in Modification 2-2 determines whether the state of a plurality of communication devices 20 is saturated based on the state data of each of the multiple communication devices 20. The monitoring system 1 identifies the state data of each of the plurality of communication devices 20 that is determined to be saturated as abnormality cause data. This allows the monitoring system 1 to determine whether the state data is related to an abnormality in the communication service. As a result, for example, the monitoring system 1 can speed up the recovery from an abnormality. In addition, for example, the monitoring system 1 can also quickly detect abnormalities and analyze the causes of abnormalities.
[0201] [2-6-3. Modified form 2-3] For example, the status data of each of the multiple communication devices 20 is generated based on the time managed by that communication device 20. The time managed by each of the multiple communication devices 20 may be out of sync with each other. In this case, the time axes of the status data of each of the multiple communication devices 20 may be out of sync with each other. For this reason, the clustering execution unit 108 may align the time axes of the status data of each of the multiple communication devices 20 and then perform clustering based on the state data with aligned time axes.
[0202] For example, the clustering execution unit 108 obtains the current time managed by each of the multiple communication devices 20. For each communication device 20, the clustering execution unit 108 calculates the difference between the current time managed by the server 10 and the current time managed by the communication device 20. For each communication device 20, the clustering execution unit 108 aligns the time axes of the multiple communication devices 20 by shifting the time axis of the communication device 20 so that the difference becomes smaller than a threshold (for example, so that the difference becomes 0).
[0203] For example, suppose the clustering execution unit 108 determines that the time managed by the communication device 20A is 15 seconds ahead of the time managed by the server 10. In this case, the clustering execution unit 108 changes the time axis indicated by the status data of the communication device 20A to be delayed by 15 seconds overall. Alternatively, suppose the clustering execution unit 108 determines that the time managed by the communication device 20B is 20 seconds behind the time managed by the server 10. In this case, the clustering execution unit 108 changes the time axis indicated by the status data of the communication device 20B to be advanced by 20 seconds overall.
[0204] In Modification Example 2-3, the example uses the case where the time managed by server 10 is used as the reference. However, the clustering execution unit 108 may use the time managed by any of the multiple communication devices 20 as the reference. For example, the clustering execution unit 108 may use the time managed by communication device 20A as the reference. In this case, if the clustering execution unit 108 determines that the time managed by communication device 20B is 32 seconds ahead of the time managed by communication device 20A, it will change the time axis indicated by the status data of communication device 20B to be delayed by 32 seconds overall. If the clustering execution unit 108 determines that the time managed by communication device 20C is 26 seconds behind the time managed by communication device 20A, it will change the time axis indicated by the status data of communication device 20C to be advanced by 26 seconds overall. The clustering execution unit 108 does not change the time axis of communication device 20A.
[0205] The monitoring system 1 in modified version 2-3 aligns the time axes of the status data of each of the multiple communication devices 20, and then performs clustering based on the state data with aligned time axes. If the time axes of the status data of each of the multiple communication devices 20 are misaligned, even if the status data originally had similar characteristics, it may be determined that they do not have similar characteristics due to the misalignment of the time axes in the status data. In this respect, the monitoring system 1 can improve the accuracy of identifying related data by aligning the time axes of the status data of each of the multiple communication devices 20.
[0206] [2-6-4. Modified Version 2-4] For example, even if there is a relationship between the status data of communication device 20A and the status data of communication device 20B, a time lag may occur between these relationships, such as when the communication volume of communication device 20A suddenly increases 30 seconds later, followed by a sudden increase in the communication volume of communication device 20B. In this case, if the communication volumes at the same point in time are compared using clustering, they may be judged as not being similar in terms of characteristics. If the clustering execution unit 108 advances the waveform shown by the status data of communication device 20B by 30 seconds overall and then compares it with the status data of communication device 20A, it becomes easier to identify similarities in the overall waveform characteristics. For this reason, the clustering execution unit 108 may shift the time axis of the status data of each of the multiple communication devices 20 and then perform clustering based on the state data with the shifted time axis.
[0207] In Modification 2-4, the amount of shift, which indicates how much the time axis should be shifted, is predetermined. Furthermore, multiple shift amounts are defined for the time axis. For example, six shift amounts such as 10 seconds, 20 seconds, 30 seconds, 40 seconds, 50 seconds, and 60 seconds are defined. In the above example, if the status data of communication device 20A is abnormal cause data, the clustering execution unit 108 advances the time axis of the status data of communication device 20B by 10 seconds and then performs clustering. Furthermore, the clustering execution unit 108 delays the time axis of the status data of communication device 20B by 10 seconds and then performs clustering.
[0208] Similarly, the clustering execution unit 108 performs clustering after advancing the time axis of the status data of the communication device 20B by 20 seconds, 30 seconds, 40 seconds, 50 seconds, and 60 seconds, respectively. Furthermore, the clustering execution unit 108 performs clustering after delaying the time axis of the status data of the communication device 20B by 20 seconds, 30 seconds, 40 seconds, 50 seconds, and 60 seconds, respectively. The clustering execution unit 108 also performs clustering on other communication devices 20, such as the communication device 20C, after advancing or delaying the time axis of the status data by a predetermined amount.
[0209] The related data identification unit 109 in Modification 2-4 identifies related data based on the results of clustering performed after the time axis has been shifted as described above. For example, the related data identification unit 109 identifies state data classified into the same cluster as the abnormal cause data by clustering based on any of the above shift amounts as related data. The related data identification unit 109 may also identify state data classified into the same cluster as the abnormal cause data by clustering based on a criterion number or more of the above shift amounts.
[0210] In the example above, the clustering execution unit 108 determines that the state data of communication device 20A and the state data of communication device 20B are similar by shifting the time axis of the state data of communication device 20B forward by 30 seconds overall. Therefore, even if there is a time lag as described above, the clustering execution unit 108 can classify the state data of communication device 20A and the state data of communication device 20B into the same cluster. The clustering execution unit 108 classifies state data that it has determined to be similar to each other for at least one of the above multiple shift amounts into the same cluster.
[0211] The monitoring system 1 in modified example 2-4 shifts the time axis of the status data of each of the multiple communication devices 20, and then performs clustering based on the status data with the shifted time axis. Even if there is a time lag between related abnormality factor data and related data, the monitoring system 1 can absorb the time lag by deliberately shifting the time axis and identify related data that is related to the abnormality factor data.
[0212] [2-6-5. Modified Version 2-5] For example, in Modification 2-4, the clustering execution unit 108 may determine the amount of time axis shift based on a portion of the time-series state changes indicated by the state data of each of the multiple communication devices 20, and then shift the time axis of the state data of each of the multiple communication devices 20 based on the determined amount of shift. In Modification 2-5, the clustering execution unit 108 determines which of the multiple shift amounts to use by using only a portion of the period indicated by the state data.
[0213] For example, the clustering execution unit 108 determines a characteristic period indicated by the abnormality factor data as part of the above-mentioned period. The characteristics that serve as the basis for determining this period are predetermined. In the example in Figure 10, the status data of the communication device 20A corresponds to the abnormality factor data. Furthermore, if the timing of a sudden increase in the communication volume of the communication device 20A indicates an abnormality, the clustering execution unit 108 determines the period including that timing as part of the above-mentioned period. This period may also be manually specified by the administrator.
[0214] For example, the clustering execution unit 108 compares the state changes of the state data of the communication device 20A over a certain period with the state changes of the state data of the other communication devices 20 over the same period for each of the multiple shift amounts. The clustering execution unit 108 determines the shift amount with the smallest difference in state changes among the multiple shift amounts as the shift amount on the time axis. Based on the determined shift amount, the clustering execution unit 108 executes the process described in Modification 2-4 and performs clustering. Clustering based on the shift amount may be the same as in Modification 2-4.
[0215] The monitoring system 1 in modified example 2-5 determines the amount of time axis shift based on a portion of the time-series state changes shown by the state data of each of the multiple communication devices 20, and shifts the time axis of the state data of each of the multiple communication devices 20 based on the determined amount of shift. Even if there is a time lag between related abnormality factor data and related data, the monitoring system 1 can absorb the time lag by shifting the time axis with the optimal amount of shift and identify related data that is related to the abnormality factor data. The monitoring system 1 can reduce the processing load because it does not have to perform processing based on unnecessary amounts of shift.
[0216] [2-6-6. Modified Version 2-6] For example, the clustering execution unit 108 may calculate the difference in values indicated by the state data of each of the multiple communication devices 20 and perform clustering based on the calculated difference. Taking Figure 10 as an example, the clustering execution unit 108 calculates the difference between the abnormal factor data, which is the state data of communication device 20A, and the state data of each of the communication devices 20B to 20E. A known index such as the Mahalanobis distance may be used for this difference. The clustering execution unit 108 classifies state data whose change in this difference is less than a threshold into the same cluster as the abnormal factor data.
[0217] In the example shown in Figure 10, the difference between the abnormal cause data, which is the state data of the communication device 20A, and the state data of the communication device 20E remains within a certain range. Therefore, the clustering execution unit 108 classifies the abnormal cause data and the state data of the communication device 20E into the same cluster. That is, even if there is a certain difference between the abnormal cause data and the state data, if the difference falls within a certain range, the clustering execution unit 108 classifies the state data into the same cluster as the abnormal cause data. The clustering execution unit 108 may perform clustering based only on the difference, rather than the amount of change. The processing after clustering is performed is as described in the embodiment.
[0218] The monitoring system 1 in modified example 2-6 calculates the difference in values indicated by the status data of each of the multiple communication devices 20, and performs clustering based on the calculated difference. This allows the monitoring system 1 to perform clustering that focuses on the difference indicated by the status data of each of the multiple communication devices 20, thereby improving the accuracy of the clustering.
[0219] [2-6-7. Modified version 2-7] For example, in some cases, some status data may not be acquired due to so-called node down events such as forced restarts or shutdowns of communication devices 20. In such cases, the status data acquisition unit 101 may acquire status data from among the multiple communication devices 20 such that all or part of the communication devices 20 from which some status data was not acquired show a predetermined value. The predetermined value may be a value specified in advance by the administrator, or it may be calculated dynamically on the spot.
[0220] In the example shown in Figure 10, suppose that some of the status data of the communication device 20B is missing due to a forced restart of the communication device 20B. In this case, the status data acquisition unit 101 may fill in the missing portion with the average value of the portion of the status data that was not missing. Alternatively, the status data acquisition unit 101 may fill in the missing portion of the status data with a value specified by the administrator. This filling in of the missing portion is sometimes called padding. The method for filling in the missing portion can be any known padding method.
[0221] In the modified example 2-7, the monitoring system 1 acquires status data from among the multiple communication devices 20 such that all or part of the communication devices 20 from which some of the status data was not acquired show a predetermined value. As a result, the monitoring system 1 can perform clustering even if some of the status data is missing.
[0222] [2-6-8. Modification 2-8] For example, the clustering execution unit 108 may perform clustering based on the status data of each of the multiple communication devices 20 and the attributes associated with that status data. The attributes can be any information that allows the status data to be classified from some perspective, and in Modification 2-8, an example is given where each of the multiple communication devices 20 is located. The location may be a fairly broad area such as a city or town, or a fairly narrow area such as a building, room, or rack where the communication device 20 is located. The attributes are not limited to the location where the communication device 20 is located, and may be, for example, the performance, role, or type of the communication device 20. The attributes may be the meaning of the status data, or acquisition conditions such as the sampling period of the status data. The data indicating the attributes of the communication device 20 is pre-stored in the data storage unit 100.
[0223] For example, the clustering execution unit 108 performs clustering based on the state data of each of the multiple communication devices 20 and the location associated with that state data. The clustering execution unit 108 performs clustering based on the state data of each of the multiple communication devices 20 located in the same location. Even if the characteristics of the state data of the multiple communication devices 20 are similar, the clustering execution unit 108 ensures that communication devices 20 located in different locations do not belong to the same cluster. The clustering execution unit 108 only performs clustering among multiple communication devices 20 located in the same location.
[0224] Clustering when attributes other than location are used may also be performed in the same manner as described above. For example, the clustering execution unit 108 performs clustering based on the state data of each of the multiple communication devices 20 associated with the same attribute. The clustering execution unit 108 ensures that even if the characteristics of the state data of the multiple communication devices 20 are similar, if the state data of the communication devices 20 are associated with different attributes, they will not belong to the same cluster. The clustering execution unit 108 performs clustering only among the multiple communication devices 20 whose state data are associated with the same attribute. The processing after clustering is performed is the same as in the embodiment.
[0225] The monitoring system 1 in modified example 2-8 performs clustering based on the status data of each of the multiple communication devices 20 and the attributes associated with that status data. This allows the monitoring system 1 to perform clustering that takes into account the attributes of the status data, thereby improving the accuracy of the clustering. For example, by performing clustering among multiple communication devices 20 whose status data is associated with the same attribute, the monitoring system 1 can accurately identify the relationships between the status data.
[0226] Furthermore, the monitoring system 1 performs clustering based on the status data of each of the multiple communication devices 20 and the location associated with that status data. This allows the monitoring system 1 to perform clustering that takes into account the location where the communication devices 20 are located, thereby improving the accuracy of the clustering. For example, by performing clustering among multiple communication devices 20 located in the same location, the monitoring system 1 can accurately identify the relationships between status data.
[0227] [2-6-9. Modified Version 2-9] For example, in Modification 2-9, the clustering execution unit 108 may average time-series data based on the status data of each of the multiple communication devices 20 at predetermined time intervals. For example, if status data is acquired at a frequency of 1 item / 1 second, the number of status data items will be 3600 items / 60 minutes. If this status data is averaged at 1-minute intervals, the 3600 items / 60 minutes of status data will be reduced to 60 items. By averaging the status data in advance, the clustering execution unit 108 in Modification 2-9 can create lightweight time-series data similar to the time-series data before averaging, thereby reducing the processing load of clustering.
[0228] [2-6-10. Other variations] For example, the above variations may be combined.
[0229] For example, monitoring system 1 may detect anomalies in a service based on SNS posts. In this case, monitoring system 1 acquires post data related to SNS posts. Post data may include text, images, videos, hashtags, or a combination thereof. Monitoring system 1 detects anomalies based on the number of posts containing the name of the service being monitored. Monitoring system 1 detects an anomaly when the number of such posts exceeds a threshold. Related data identification unit 109 may identify the relationship between SNS post data and the status data of communication device 20. For example, related data identification unit 109 may identify a relationship between SNS post data and the status data of communication device 20 if there is a correlation between the timing of an increase in the number of SNS posts and the timing of an increase in the load indicated by the status data of communication device 20.
[0230] For example, monitoring system 1 is applicable to services other than communication services. For instance, monitoring system 1 may identify related data associated with anomaly factor data that may cause anomalies in other services such as e-commerce services, travel booking services, financial services, payment services, online flea market services, or video streaming services. The devices to be monitored are not limited to communication devices as in the embodiment, but may be any devices used in these services. For example, the devices to be monitored may be server computers, personal computers, tablets, power supplies, printers, scanners, or other devices.
[0231] For example, in this embodiment, we have described a case where the main processing is performed on server 10, but the processing described as being performed on server 10 may also be performed on administrator terminal 40 or other computers, or it may be shared among multiple computers.
[0232] [3. Addendum] For example, the monitoring system can also be configured as follows:
[0233] [3-1. Notes according to the first embodiment] (1-1) A status data acquisition unit acquires status data regarding the time-series changes in the state of the devices being monitored in the service, A saturation determination unit determines whether the state is saturated or not based on the state data, A relationship determination unit determines, based on the determination result of the saturation determination unit, whether or not the status data is related to the abnormality in the service, A monitoring system including this. (1-2) The status data acquisition unit acquires the status data when the abnormality is detected. The saturation determination unit, when the abnormality is detected, determines whether the state is saturated based on the state data, The relationship determination unit, when the abnormality is detected, determines whether or not the state data is related to the abnormality. The monitoring system further includes a recovery processing execution unit that, when an abnormality is detected, executes recovery processing related to the recovery of the abnormality based on the status data determined to be related to the abnormality. The monitoring system described in (1-1). (1-3) The status data acquisition unit acquires the status data before the abnormality is detected. The saturation determination unit determines whether the state is saturated or not based on the state data before the abnormality is detected. The relationship determination unit determines whether the state data is related to the abnormality before the abnormality is detected. The monitoring system further includes a detection processing execution unit that performs detection processing related to the detection of the anomaly based on the state data which has been determined to be related to the anomaly. The monitoring system described in (1-1) or (1-2). (1-4) The status data acquisition unit acquires the status data of the device in each of the multiple services, The saturation determination unit determines whether the device is saturated or not based on the status data of the device in each of the plurality of services. The relationship determination unit determines, based on the determination result of the saturation determination unit, whether or not the status data is related to the abnormality in each of the multiple services. A monitoring system as described in any of (1-1) to (1-3). (1-5) The monitoring system further includes a period setting unit that, when an abnormality is detected, sets a period for the status data based on the timing of the abnormality occurrence, The state data acquisition unit acquires the state data relating to the time-series changes in the state of the device during the period. A monitoring system as described in any of (1-1) to (1-4). (1-6) The monitoring system further includes a normalization execution unit that performs normalization of the state data, The saturation determination unit determines whether the state is saturated based on the state data after the normalization has been performed. A monitoring system as described in any of (1-1) to (1-5). (1-7) The relationship determination unit determines whether the state data is related to the abnormality based on the saturation period during which the state data is determined to be saturated, within the period in which the change is shown in the state data. A monitoring system as described in any of (1-1) to (1-6). (1-8) The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state in other periods prior to the abnormality period in which the abnormality occurred, within the period in which the change in the state data was shown. A monitoring system as described in any of (1-1) to (1-7). (1-9) The aforementioned other period is a period prior to the aforementioned abnormal period. The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state in the other period and the highest value of the state in a period prior to the other period. The monitoring system described in (1-8). (1-10) The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state indicated by the state data during the abnormal period in which the abnormality occurred, within the period in which the change was shown in the state data. A monitoring system as described in any of (1-1) to (1-9). (1-11) The relationship determination unit determines whether the state data is related to the abnormality based on the degree of agreement between the abnormal period during which the abnormality occurred and the saturation period during which the state was determined to be saturated, within the period during which the change was indicated in the state data. A monitoring system as described in any of (1-1) to (1-10). (1-12) The relationship determination unit determines whether the state data is related to the abnormality based on the amount of change in the state during the period in which the change is shown in the state data, specifically during the saturation period when the state is saturated and during other periods prior to the abnormal period in which the abnormality occurs. A monitoring system as described in any of (1-1) to (1-11).
[0234] [3-2. Notes according to the second embodiment] (2-1) A status data acquisition unit that acquires status data regarding the status of each of the multiple devices that are monitored in the service, A clustering execution unit that performs clustering on the state data of each of the plurality of devices, Based on the results of the clustering, a related data identification unit identifies related data, which is other state data related to the abnormality factor data, which is state data relating to the cause of the abnormality in the service, A monitoring system including this. (2-2) When the abnormality is detected, the status data acquisition unit acquires the status data of each of the plurality of devices. When the clustering execution unit detects the abnormality, it performs the clustering based on the status data of each of the plurality of devices. The related data identification unit identifies the related data when the anomaly is detected. The monitoring system further includes a recovery processing execution unit that, when an abnormality is detected, performs recovery processing related to the recovery of the abnormality based on the abnormality cause data and the related data. The monitoring system described in (2-1). (2-3) The status data acquisition unit acquires the status data of each of the plurality of devices before the abnormality is detected. The clustering execution unit performs the clustering based on the status data of each of the plurality of devices before the abnormality is detected. The related data identification unit identifies the related data before the anomaly is detected. The monitoring system further includes a detection processing execution unit that performs detection processing related to the detection of the anomaly based on the anomaly cause data and the related data. The monitoring system described in (2-1) or (2-2). (2-4) The state data of each of the aforementioned plurality of devices indicates the time-series change in the state of the device. The clustering execution unit performs the clustering based on the time-series changes in the state indicated by the state data of each of the plurality of devices. A monitoring system as described in any of (2-1) to (2-3). (2-5) The aforementioned monitoring system A saturation determination unit determines whether the state indicated by the state data is saturated, based on the state data of each of the plurality of devices, An abnormality factor data identification unit identifies the state data from among the state data of each of the plurality of devices that is determined to be saturated as the abnormality factor data, The monitoring system described in (2-4) further includes the following. (2-6) The clustering execution unit aligns the time axis of the state data of each of the plurality of devices, and then performs the clustering based on the state data with the aligned time axis. The monitoring system described in (2-4) or (2-5). (2-7) The clustering execution unit shifts the time axis of the state data of each of the plurality of devices, and then performs the clustering based on the state data with the time axis shifted. A monitoring system as described in any of (2-4) to (2-6). (2-8) The clustering execution unit determines the amount of shift of the time axis based on a partial period of the temporal change of the state indicated by the state data of each of the plurality of devices, and shifts the time axis of the state data of each of the plurality of devices based on the determined amount of shift. The monitoring system according to (2-7). (2-9) The clustering execution unit calculates the difference between the values indicated by the state data of each of the plurality of devices, and executes the clustering based on the calculated difference. The monitoring system according to any one of (2-1) to (2-8). (2-10) The state data acquisition unit acquires the state data such that all or a part of the devices among the plurality of devices for which a part of the state data has not been acquired indicates a predetermined value. The monitoring system according to any one of (2-1) to (2-9). (2-11) The clustering execution unit executes the clustering based on the state data of each of the plurality of devices and the attribute associated with the state data. The monitoring system according to any one of (2-1) to (2-10). (2-12) The attribute indicates the location where each of the plurality of devices is arranged. The clustering execution unit executes the clustering based on the state data of each of the plurality of devices and the location associated with the state data. The monitoring system according to (2-11).
Claims
1. A status data acquisition unit acquires status data regarding the time-series changes in the state of the devices being monitored in the service, A saturation determination unit determines whether the state is saturated or not based on the state data, A relationship determination unit determines, based on the determination result of the saturation determination unit, whether or not the status data is related to the abnormality in the service, A monitoring system including this.
2. The status data acquisition unit acquires the status data when the abnormality is detected. The saturation determination unit, when the abnormality is detected, determines whether the state is saturated based on the state data, The relationship determination unit, when the abnormality is detected, determines whether or not the state data is related to the abnormality. The monitoring system further includes a recovery processing execution unit that, when an abnormality is detected, executes recovery processing related to the recovery of the abnormality based on the status data determined to be related to the abnormality. The monitoring system according to claim 1.
3. The status data acquisition unit acquires the status data before the abnormality is detected. The saturation determination unit determines whether the state is saturated or not based on the state data before the abnormality is detected. The relationship determination unit determines whether the state data is related to the abnormality before the abnormality is detected. The monitoring system further includes a detection processing execution unit that performs detection processing related to the detection of the anomaly based on the state data which has been determined to be related to the anomaly. The monitoring system according to claim 1 or 2.
4. The status data acquisition unit acquires the status data of the device in each of the multiple services, The saturation determination unit determines whether the device is saturated or not based on the status data of the device in each of the plurality of services. The relationship determination unit determines, based on the determination result of the saturation determination unit, whether or not the status data is related to the abnormality in each of the multiple services. The monitoring system according to claim 1 or 2.
5. The monitoring system further includes a period setting unit that, when an abnormality is detected, sets a period for the status data based on the timing of the abnormality occurrence, The state data acquisition unit acquires the state data relating to the time-series changes in the state of the device during the period. The monitoring system according to claim 1 or 2.
6. The monitoring system further includes a normalization execution unit that performs normalization of the state data, The saturation determination unit determines whether the state is saturated based on the state data after the normalization has been performed. The monitoring system according to claim 1 or 2.
7. The relationship determination unit determines whether the state data is related to the abnormality based on the saturation period during which the state data is determined to be saturated, within the period in which the change is shown in the state data. The monitoring system according to claim 1 or 2.
8. The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state in other periods prior to the abnormal period in which the abnormality occurred, within the period in which the change in the state data was shown. The monitoring system according to claim 1 or 2.
9. The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state in the other period and the highest value of the state in a period prior to the other period. The monitoring system according to claim 8.
10. The relationship determination unit determines whether the state data is related to the abnormality based on the highest value of the state data shown in the state data during the abnormal period in which the abnormality occurred, within the period in which the change in the state data was shown. The monitoring system according to claim 1 or 2.
11. The relationship determination unit determines whether the state data is related to the abnormality based on the degree of agreement between the abnormal period during which the abnormality occurred and the saturation period during which the state was determined to be saturated, within the period during which the change was indicated in the state data. The monitoring system according to claim 1 or 2.
12. The relationship determination unit determines whether the state data is related to the abnormality based on the amount of change in the state during the period in which the change is shown in the state data, specifically during the saturation period when the state is saturated and during other periods prior to the abnormal period in which the abnormality occurs. The monitoring system according to claim 1 or 2.
13. A status data acquisition unit that acquires status data relating to the status of each of a plurality of devices that are monitored in a service, A clustering execution unit that performs clustering on the state data of each of the plurality of devices, Based on the results of the clustering, a related data identification unit identifies related data, which is other state data related to the abnormality factor data, which is state data relating to the cause of the abnormality in the service, A monitoring system including this.
14. When the abnormality is detected, the status data acquisition unit acquires the status data of each of the plurality of devices. When the clustering execution unit detects the abnormality, it performs the clustering based on the status data of each of the plurality of devices. The related data identification unit identifies the related data when the anomaly is detected. The monitoring system further includes a recovery processing execution unit that, when an abnormality is detected, performs recovery processing related to the recovery of the abnormality based on the abnormality cause data and the related data. The monitoring system according to claim 13.
15. The status data acquisition unit acquires the status data of each of the plurality of devices before the abnormality is detected. The clustering execution unit performs the clustering based on the status data of each of the plurality of devices before the abnormality is detected. The related data identification unit identifies the related data before the anomaly is detected. The monitoring system further includes a detection processing execution unit that performs detection processing related to the detection of the anomaly based on the anomaly cause data and the related data. The monitoring system according to claim 13 or 14.
16. The state data of each of the aforementioned plurality of devices indicates the time-series change in the state of the device. The clustering execution unit performs the clustering based on the time-series changes in the state indicated by the state data of each of the plurality of devices. The monitoring system according to claim 13 or 14.
17. The aforementioned monitoring system A saturation determination unit determines whether the state indicated by the state data is saturated, based on the state data of each of the plurality of devices, An abnormality factor data identification unit identifies the state data from among the state data of each of the plurality of devices that is determined to be saturated as the abnormality factor data, The monitoring system according to claim 16, further comprising:
18. The clustering execution unit aligns the time axis of the state data of each of the plurality of devices, and then performs the clustering based on the state data with the aligned time axis. The monitoring system according to claim 16.
19. The clustering execution unit shifts the time axis of the state data of each of the plurality of devices, and then performs the clustering based on the state data with the time axis shifted. The monitoring system according to claim 16.
20. The clustering execution unit determines the amount of time axis shift based on a portion of the time-series state changes indicated by the state data of each of the plurality of devices, and shifts the time axis of the state data of each of the plurality of devices based on the determined amount of shift. The monitoring system according to claim 19.
21. The clustering execution unit calculates the difference between the values indicated by the state data of each of the plurality of devices, and performs the clustering based on the calculated difference. The monitoring system according to claim 13 or 14.
22. The status data acquisition unit acquires the status data of the plurality of devices such that all or part of the devices for which some of the status data was not acquired show a predetermined value. The monitoring system according to claim 13 or 14.
23. The clustering execution unit performs the clustering based on the state data of each of the plurality of devices and the attributes associated with the state data. The monitoring system according to claim 13 or 14.
24. The aforementioned attribute indicates the location where each of the plurality of devices is placed. The clustering execution unit performs the clustering based on the state data of each of the plurality of devices and the location associated with the state data. The monitoring system according to claim 23.
25. A status data acquisition step that acquires status data regarding the time-series changes in the state of the device being monitored in the service, A saturation determination step, which determines whether the state is saturated or not based on the state data, A relationship determination step that determines whether the status data is related to an abnormality in the service based on the determination result of the saturation determination step, A monitoring method that includes this.
26. A status data acquisition unit that acquires status data regarding the time-series changes in the state of devices being monitored in a service. A saturation determination unit determines whether the state is saturated or not based on the state data. Based on the determination result of the saturation determination unit, a relationship determination unit determines whether or not the status data is related to the abnormality in the service. A program that makes a computer function.
Citation Information
Patent Citations
Traffic monitoring system
JP2009017393A
Abnormal state detection device and abnormal state detection method
JP2013011987A
Failure detection device, failure detection method, failure detection program and recording medium
JP2015028700A
Network system, network control method, and control device
JP2016192661A
Clustering program, clustering method, and information processing apparatus
JP2017068748A