Failure verification system and failure verification method
The failure verification system addresses the challenge of anticipating failures in complex information systems by collecting and analyzing logs and metrics to pre-verify system failures, facilitating timely identification and response.
Patent Information
- Application Number
- JP2024225945
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-12-23
- Publication Date
- 2025-07-24
AI Technical Summary
In information systems with distributed components across cloud, on-premises, and edge devices, failures are difficult to anticipate and verify due to complex interactions and diverse infrastructures, leading to unexpected system failures during operation.
A failure verification system that collects and associates execution logs and metrics from cloud and device applications, assigns failure causes, and stores this information for pre-verification, using agents to monitor and identify failure causes in a target system composed of cloud services and devices.
Enables pre-verification of system failures that were previously difficult to confirm, allowing for timely identification and countermeasure planning.
Smart Images

Figure 2025109182000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for verifying failures occurring in an information system.
Background Art
[0002] As techniques for dealing with failures occurring in an information system, there are known a method of predicting a failure that may spread due to the failure when the failure occurs, a method of improving the estimation accuracy of failure causes, and a method of artificially generating a failure for a reliability test (for example, Patent Documents 1 to 3).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0004] In an information system configured by dispersing and arranging each component in various environments such as the cloud, on-premises, and edge devices, each operates on a different infrastructure and cooperates in a complex manner. For this reason, a system failure that was not anticipated may occur during actual operation. An object related to one aspect of the present invention is to provide a method for verifying a failure of an information system in advance.
Means for Solving the Problems
[0005] One embodiment of the failure verification system verifies failures occurring in a target system composed of a cloud service provided by cloud computing and a device that exchanges data with the cloud service. This target system includes a cloud application that provides the cloud service, cloud resources that are the hardware for running the cloud application, a device application that provides a data exchange function in the device, and device resources that are the hardware for running the device application as components. This failure verification system includes a collection unit, an assignment unit, and a storage unit. The collection unit collects execution logs of the cloud application and the device application as log information, and collects metric logs for each of the cloud resources and the device resources as metric information. The assignment unit assigns a cause of failure to any of the components of the target system to cause a failure in the target system. The storage unit stores the log information and the metric information in association with the time when the cause is assigned to the component as failure verification information.
[0006] A fault verification system according to another embodiment verifies faults in a target system including a multi-core device having a first processor core implementing a first OS and a second processor core implementing a second OS. This fault verification system includes a fault verification unit that creates fault verification information for verifying faults in the target system, a fault identification unit that identifies the cause of a fault or a sign of a fault occurring in the target system using the fault verification information, a first agent implemented in the first processor core that collects first log information representing the execution log of an application operating within the first processor core and first metric information representing the state of the hardware of the first processor core, and a second agent implemented in the second processor core that collects second log information representing the execution log of an application operating within the second processor core and second metric information representing the state of the hardware of the second processor core. The fault verification unit receives, from the first agent, the first log information and the first metric information when the fault factor is injected into the target system for each of a plurality of predefined fault factors, and receives, from the second agent via the first agent, the second log information and the second metric information when the fault factor is injected into the target system, and stores the first log information, the first metric information, the second log information, and the second metric information when the fault factor is injected into the target system as the fault verification information. The fault identification unit identifies the cause of a fault or a sign of a fault occurring in the target system by comparing the monitor information including at least a part of the first log information and the first metric information received from the first agent and the second log information and the second metric information received from the second agent via the first agent during actual operation of the target system with the fault verification information stored for each of the plurality of fault factors.
Effect of the Invention
[0007] According to the above aspect, it becomes possible to pre-verify system failures that were difficult to confirm in single-unit tests or conventional system tests.
Brief Description of Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Mode for Carrying Out the Invention
[0009] As the system infrastructure of information systems, the utilization of cloud services provided by cloud computing has become widespread. In addition, the infrastructure environment in which information systems are built (the term "infrastructure" is an abbreviation for infrastructure) is expanding and decentralized in a form that does not stay in a single data center. Hybrid clouds that leave important core data and applications that are difficult to update on-premises and then cooperate with cloud services, and edge computing that performs data processing in real time at the site, are examples of such infrastructure environments. Thus, in an information system in which many services are linked and configured as a system architecture, it has become difficult to track each individual traffic.
[0010] In such an information system, many applications are executed on an infrastructure under the management of a cloud vendor and are intricately coordinated. For this reason, in existing system tests, it may not be possible to cover all abnormal situations (for example, abnormal loads on the network or arithmetic processing units). In addition, since multiple platforms and multiple devices on the cloud that make up the information system are distributed and operate on different infrastructures and are intricately coordinated, a failure that occurs in one application may affect another application unexpectedly.
[0011] Thus, it has become difficult to verify failures, check the status, and consider countermeasures for an information system in advance through existing system tests. It is also conceivable that it will take time to identify the cause when the information system stops due to a failure during actual operation. Therefore, an embodiment of the present invention provides a method for verifying a failure of an information system and identifying or estimating its cause.
[0012] <First Embodiment> As a first embodiment, a system for verifying a failure that occurs in a target system composed of a cloud service provided by cloud computing and a device that exchanges data with the cloud service will be described.
[0013] FIG. 1 is a diagram for explaining the outline of the failure verification system 100, and FIG. 2 is a diagram showing a configuration example of the failure verification system 100.
[0014] In FIG. 1, the cloud 10, the devices 20a and 20b, and the sub-device 21 are connected via a public line 30. Note that the device 20a, the device 20b, and the sub-device 21 are connected to each other by a local line (wired, wireless, etc.) not shown. Among these, the sub-device 21 may be connected to the public line 30 or may be connected only to the local line depending on the role of the application installed in the sub-device 21.
[0015] The cloud 10 is an environment that provides cloud computing.
[0016] The cloud resource 11 is a hardware resource that provides cloud computing in the cloud 10.
[0017] The cloud app 12 is application software that runs on the cloud resource 11 and provides cloud services utilized by the target system.
[0018] The devices 20a and 20b and the sub-device 21 all exchange data with the cloud services provided by the cloud app 12 and are, for example, IoT devices. Note that "IoT" is an abbreviation for Internet of Things. Although two devices 20a and 20b and one sub-device 21 are shown in FIG. 1, the number of these devices does not have to be the number shown in FIG. 1.
[0019] In the following description, when it is not necessary to distinguish between the devices 20a and 20b, they may be collectively referred to as "device 20".
[0020] The device app 22 is application software that runs on the respective hardware of the devices 20 and the sub-device 21 and provides a function to exchange data with the cloud services provided by the cloud app 12.
[0021] In FIG. 1, the target system to be verified for faults by the fault verification system 100 includes the cloud resource 11, the devices 20a and 20b, the sub-device 21, the cloud app 12, and the device app 22 as components. The fault verification system 100 that verifies the faults of the target system with this configuration provides an information storage function 41, a fault verification function 42, and a fault identification function 43, and uses fault verification agents 44a and 44b as necessary for providing these functions.
[0022] The failure verification agents 44a and 44b are software programs respectively installed on the cloud resource 11 and the device 20.
[0023] The failure verification agent 44a is responsible for tasks such as acquiring the metrics of the cloud resource 11, obtaining the execution logs of the cloud application 12, and mediating various processes performed by the failure verification function 42 and the failure identification function 43 on the cloud resource 11.
[0024] The failure verification agent 44b is responsible for tasks such as acquiring the metrics of the device 20, obtaining the execution logs regarding the data transfer process by the device application 22, and mediating various processes performed by the failure verification function 42 and the failure identification function 43 on the device 20. The failure verification agent 44b also performs these processes for the sub-device 21 that executes the device application 22.
[0025] In the following description, the failure verification agents 44a and 44b may be simply referred to as "agent 44a" and "agent 44b" respectively.
[0026] It is assumed that the sub-device 21 has fewer hardware resources than the device 20 and lacks the necessary hardware resources for the execution of the agent 44b. In this embodiment, for such a sub-device 21, the agent 44b running on the device 20 substitutes for the above-mentioned processes.
[0027] In the example of FIG. 1, it is assumed that the agent 44b installed on the device 20b performs the above-mentioned processes for the sub-device 21. Note that the agent 44b may substitute for the processes for multiple sub-devices 21.
[0028] The information storage function 41 is a function for storing and managing the execution logs and metrics acquired by the agents 44a and 44b.
[0029] The failure verification function 42 is a function for performing pre-verification of failures in the target system, and is a function for monitoring failures intentionally generated in any component of the target system for each component of the target system. The results of this monitoring are also stored and managed by the information storage function 41.
[0030] The failure identification function 43 is a function that uses the results of the pre-verification by the failure verification function 42 stored by the information storage function 41 to identify the cause of a failure that occurred during the actual operation of the target system, and outputs the identified cause to the administrator terminal 50 to notify the administrator of the target system.
[0031] Next, the configuration of the failure verification system 100 that provides these functions will be described.
[0032] In the configuration example of FIG. 2, the failure verification system 100 includes a collection unit 110, an imparting unit 120, a storage unit 130, a detection unit 140, a specific information output unit 150, a monitoring unit 160, a collection control unit 170, and a cause identification unit 180. Among these, the collection unit 110 has a first collection intermediary unit 111, a second collection intermediary unit 112, and a reception unit 113, and the imparting unit 120 has a first imparting intermediary unit 121, a second imparting intermediary unit 122, and an imparting instruction unit 123. Further, the cause identification unit 180 has a calculation unit 181 and a cause information output unit 182.
[0033] The information storage function 41 in FIG. 1 is provided by the reception unit 113 and the storage unit 130 among the collection unit 110. Also, the failure verification function 42 in FIG. 1 is provided by the imparting instruction unit 123, the detection unit 140, and the specific information output unit 150 among the imparting unit 120. Furthermore, the failure identification function 43 in FIG. 1 is provided by the monitoring unit 160, the collection control unit 170, and the cause identification unit 180 (the calculation unit 181 and the cause information output unit 182).
[0034] Note that the agent 44a in FIG. 1 provides the first collection mediation unit 111 of the collection unit 110 and the first granting mediation unit 121 of the granting unit 120. Further, the agent 44b in FIG. 1 provides the second collection mediation unit 112 of the collection unit 110 in FIG. 2 and the second granting mediation unit 122 of the granting unit 120.
[0035] The collection unit 110 collects log information including an execution log of a cloud service provided by cloud computing and an execution log of a process executed in each of the device 20 and the sub-device 21. Further, the collection unit 110 collects metrics information which is a log of metrics regarding the hardware as a cloud infrastructure used for providing the cloud service and the hardware constituting each of the device 20 and the sub-device 21.
[0036] The first collection mediation unit 111 included in the collection unit 110 is provided in a platform that provides a cloud service, and collects the above-described execution log of the cloud service and the above-described metrics log of the cloud service. The agent 44a in FIG. 1 that provides the first collection mediation unit 111 is arranged on the cloud resource 11, and collects an execution log of the cloud app 12 and a metrics log of the cloud resource 11. In the example of FIG. 1, as the metrics log of the cloud resource 11, the metrics logs of the arithmetic processing unit, the memory, the storage device, and the interface device prepared for the execution of the cloud app 12 are collected. More specifically, the agent 44a collects a usage rate log for the arithmetic processing unit, the memory, and the interface device, and collects a log of the free capacity ratio and the access frequency (number of accesses per unit time) for the storage device.
[0037] The second collection intermediary unit 112 included in the collection unit 110 is provided in the device 20, and collects the above-described execution logs for the device 20 and the above-described metric logs for the device 20. The agent 44b in FIG. 1 that provides the second collection intermediary unit 112 is arranged on the hardware resources (device resources) of the device 20. This agent 44b collects the execution logs for the device application 22 and the metric logs for the device resources. Note that the agent 44b that provides the second collection intermediary unit 112 in the device 20b further substitutes for the collection of the execution logs of the device application 22 executed in the sub-device 21 and the collection of the metric logs for the device resources of the sub-device 21. In the example of FIG. 1, as the metric logs for the device 20 and the sub-device 21, the metric logs of the arithmetic processing unit, memory, storage device, and interface device provided in the device 20 and the sub-device 21 respectively are collected. More specifically, the agent 44b collects the usage rate logs for the arithmetic processing unit, memory, and interface device, and collects the free capacity ratio and access frequency (number of accesses per unit time) logs for the storage device.
[0038] The receiving unit 113 included in the collection unit 110 receives from the first collection intermediary unit 111 the execution logs and metric logs for the cloud service collected by the first collection intermediary unit 111. Further, the receiving unit 113 receives from the second collection intermediary unit 112 the execution logs and metric logs for the device 20 or the sub-device 21 collected by the second collection intermediary unit 112.
[0039] The imparting unit 120 imparts the cause of the failure to the components of the target system to cause a failure in the target system.
[0040] The first imparting intermediary unit 121 included in the imparting unit 120 is arranged in the cloud resource 11. The agent 44a in FIG. 1 that provides the first imparting intermediary unit 121 imparts the cause of the failure to the cloud resource 11 or the cloud application 12.
[0041] The second imparting intermediary unit 122 included in the imparting unit 120 is arranged in the device resources of the device 20. The agent 44b in FIG. 1 that provides the second imparting intermediary unit 122 imparts a cause of failure to the device resources of the device 20 or the device application 22 executed on the device 20. Note that the agent 44b that provides the second imparting intermediary unit 122 in the device 20b may, in a predetermined case, impart a cause of failure to the device resources of the sub-device 21 or the device application 22 executed on the sub-device 21.
[0042] The imparting instruction unit 123 included in the imparting unit 120 gives an instruction to the first imparting intermediary unit 121 or the second imparting intermediary unit 122 in response to a failure occurrence command including the setting of the cause of failure and the component of the target system, and causes the cause of failure to be imparted to the component. The failure occurrence command is a command issued by the administrator, and in FIG. 1, it is input to the failure verification function 42 when the administrator operates the administrator terminal 50.
[0043] Note that there may be a case where the device resources of the sub-device 21 are set as the component setting in the failure occurrence command. In this case, the imparting instruction unit 123 gives an instruction to the second imparting intermediary unit 122 included in the device 20b, and causes the cause of failure set in the failure occurrence command to be imparted to the device resources of the sub-device 21. Also, there may be a case where the device application 22 executed on the sub-device 21 is set as the component setting in the failure occurrence command. In this case, the imparting instruction unit 123 gives an instruction to the second imparting intermediary unit 122 included in the device 20b, and causes the cause of failure set in the failure occurrence command to be imparted to the device application 22. In this way, by the second imparting intermediary unit 122 included in the device 20b imparting the cause of failure not only to the device 20b itself but also to the sub-device 21, it becomes possible to impart the cause of failure in the sub-device 21 that does not include the second imparting intermediary unit 122.
[0044] Note that the grant instruction unit 123 may be configured to check for the presence or absence of an abnormality in the second grant mediation unit 122 included in the device 20b (an example of the first device). The grant instruction unit 123 checks for the presence or absence of such an abnormality when a failure occurrence command includes a device application 22 executed by the sub-device 21 as a component setting or a device resource that executes the device application 22 in the sub-device 21. Here, when it is confirmed that there is no abnormality, the grant instruction unit 123 gives an instruction to the second grant mediation unit 122 included in the device 20b to cause the component of the sub-device 21 to be given the cause of the failure occurrence set in the failure occurrence command. On the other hand, when it is confirmed that there is an abnormality, instead of the device 20b, the grant instruction unit 123 gives an instruction to the second grant mediation unit 122 included in, for example, the device 20a (an example of the second device) to cause the component of the sub-device 21 to be given the cause. By doing so, even if an abnormality occurs in the second grant mediation unit 122 (agent 44b) included in the device 20b, the second grant mediation unit 122 (agent 44b) included in the device 20a can give the cause of the failure occurrence to the component of the sub-device 21.
[0045] The storage unit 130 stores the log information and metrics information collected by the collection unit 110 in association with the time when the grant unit 120 gives the cause of the failure occurrence to the component of the target system as failure verification information.
[0046] In FIG. 1, the information storage function 41 as the storage unit 130 stores the execution log and the metrics log received by the information storage function 41 as the reception unit 113 as failure verification information. Note that the time when the cause of the failure occurrence instructed by the failure verification function 42 as the grant instruction unit 123 is given to the component is associated with this failure verification information.
[0047] The detection unit 140 detects the occurrence of an abnormality in the components of the target system. In FIG. 1, the failure identification function 43 as the detection unit 140 detects the occurrence of an abnormality in each of the cloud resource 11 that executes the cloud app 12, the device 20 that executes the device app 22, and the sub-device 21.
[0048] When the specific information output unit 150 detects the occurrence of an abnormality in other components of the target system excluding the component to which the cause of the failure has been given by the giving unit 120 according to the giving of the cause by the detection unit 140, it outputs information for specifying the other components.
[0049] In FIG. 1, for example, the giving instruction unit 123 as the failure verification function 42 gives an instruction to the agent 44a that provides the first giving intermediary unit 121 to cause the cloud resource 11 to be given the cause of the failure such as the stop of the app or an excessive processing load. At this time, it is assumed that the failure verification function 42 as the detection unit 140 detects the occurrence of an abnormality (for example, an execution error of the device app 22) in the device 20a. The failure verification function 42 as the specific information output unit 150 outputs, at this time, information for identifying the device 20a and information indicating the execution error of the device app 22 to the administrator terminal 50. These pieces of information output to the administrator terminal 50 enable the administrator to conduct a preliminary study on countermeasures when the failure given to the cloud resource 11 occurs during actual operation.
[0050] The monitoring unit 160 monitors the target system. In FIG. 1, the failure identification function 43 as the monitoring unit 160 monitors the log information and metrics information about the target system during actual operation, which are collected by the agents 44a and 44b that provide the first collection intermediary unit 111 and the second collection intermediary unit 112.
[0051] When the monitoring unit 160 detects the occurrence of a failure in the target system through monitoring, the collection control unit 170 controls the collection unit 110 to collect log information and metrics information at the time of the detected failure. In FIG. 1, when the failure identification function 43 as the collection control unit 170 detects the occurrence of a failure in the target system during actual operation, it gives instructions to the agents 44a and 44b to send the log information and metrics information collected at the time of the occurrence of the failure.
[0052] The cause identification unit 180 uses the failure verification information stored by the storage unit 130 to identify the cause of the failure detected by the monitoring unit 160 through monitoring from the log information and metrics information at the time of the occurrence of the failure. The cause identification unit 180 performs the identification of the cause of this failure using the calculation unit 181 and the cause information output unit 182.
[0053] The calculation unit 181 of the cause identification unit 180 calculates the matching rate between the log information and metrics information for each cause of failure occurrence in the failure verification information of the storage unit 130 and the log information and metrics information at the time of the occurrence of the failure detected by the monitoring unit 160 through monitoring. The matching rate is the ratio at which the log information and metrics information for each cause of failure occurrence are similar to the log information and metrics information at the time of the occurrence of the failure during actual operation.
[0054] In this embodiment, in order to calculate the matching rate for the cause of a failure, first, the matching rates of the execution log and the logs of each metric are calculated. For example, for the log information, the ratio of the number of items with matching log contents in the respective execution logs at the time of failure verification and during actual operation to the total number of items in the execution log is calculated as the matching rate. Also, for the logs of the metrics of each of the arithmetic processing unit, the memory, and the interface device, the value obtained by subtracting the difference in the usage rates at the time of failure verification and during actual operation at the time of failure occurrence from 1.0 (100%) is calculated as the matching rate. Further, for the log of the metric of the storage device, the value obtained by subtracting the difference in the ratio of free capacity at the time of failure verification and during actual operation at the time of failure occurrence from 1.0 (100%) is calculated as the matching rate. Also, the value obtained by subtracting the difference in the access frequencies at the time of failure verification and during actual operation at the time of failure occurrence from 1.0 (100%) is calculated as the matching rate. Then, the average value of the matching rate of the execution log and the matching rates of the logs of each metric is calculated, and the calculated average value is used as the matching rate for the cause of the failure.
[0055] The cause information output unit 182 included in the cause identification unit 180 outputs the identification information for a predetermined number of the causes with the highest calculated matching rates for each cause of the failure in the failure verification information as information representing the cause of the failure detected by the monitoring unit 160.
[0056] In FIG. 1, a failure identification function 43 as the monitoring unit 160 monitors each component of the target system during actual operation. When a failure occurs in the target system is detected in this monitoring, the failure identification function 43 as the collection control unit 170 gives instructions to agents 44a and 44b to collect log information and metrics information about each component at the time when the detected failure occurred. Then, the failure identification function 43 as the calculation unit 181 included in the cause identification unit 180 calculates the matching rate between the log information and metrics information for each cause of the failure occurrence and the log information and metrics information at the time when the failure detected by monitoring occurred. After that, the failure identification function 43 as the cause information output unit 182 included in the cause identification unit 180 outputs identification information about a predetermined number of such causes in descending order of the matching rate for each cause of the failure occurrence in the failure verification information to the administrator terminal 50 as information representing the cause of the failure. This information enables the administrator to quickly identify the cause of the failure that occurred in the target system during actual operation. Also, by utilizing the prior consideration for the failure due to such a cause that was carried out for the target system before the start of actual operation, it becomes possible to quickly execute countermeasures against the occurred failure.
[0057] Next, the hardware configuration of the failure verification system 100 will be described.
[0058] FIG. 3 shows a hardware configuration example of the information processing apparatus 60. The information storage function 41, the failure verification function 42, and the failure identification function 43 provided in the cloud 10 may all be provided using the information processing apparatus 60 as a cloud server. Also, the information processing apparatus 60 may be used as the cloud resource 11 on which the cloud app 12 and the agent 44a are executed. Furthermore, the information processing apparatus 60 may be used as the device 20 on which the device app 22 and the agent 44b are executed.
[0059] The information processing apparatus 60 is a computer including components such as a CPU 61, a memory 62, an input device 63, an output device 64, an auxiliary storage device 65, and a communication I / F 66. All of these components are connected to an internal bus 67 and are configured to be able to exchange data among the components. Note that "CPU" is an abbreviation for Central Processing Unit. Also, "I / F" is an abbreviation for Interface.
[0060] The CPU 61 controls each component of the information processing apparatus 60 by executing a predetermined program using, for example, the memory 62.
[0061] The input device 63 is, for example, a keyboard or a pointing device for inputting instructions, or various sensors.
[0062] The output device 64 is used, for example, for display output of various information.
[0063] The auxiliary storage device 65 is a non-volatile storage device, for example, a flash memory.
[0064] The communication I / F 66 transmits and receives various data to and from other cloud servers or to and from the device 20 in accordance with instructions sent from the CPU 61.
[0065] Note that when using the information processing apparatus 60 as each element shown in FIG. 1, the information processing apparatus 60 does not necessarily need to include all the components shown in FIG. 3, and some components may be omitted according to the use or conditions.
[0066] Next, the processing contents of each of the agents 44a and 44b, and the processing contents of the processes performed to provide the information storage function 41, the failure verification function 42, and the failure identification function 43 will be described.
[0067] First, the flowchart of FIG. 4 will be described. FIG. 4 shows a first example of the processing contents of agents 44a and 44b, and shows the processing contents when a preliminary study is conducted on the failure of the target system before the start of actual operation.
[0068] When the processing of FIG. 4 starts, first, in S101, a process of receiving a normality confirmation request is performed, and then in S102, a process of determining whether the confirmation request has been received is performed. The confirmation request received by the process of S101 is transmitted by the failure verification function 42 to agents 44a and 44b.
[0069] In the determination process of S102, when it is determined that a normality confirmation request has been received (when the determination result is YES), the process proceeds to S103, and a process of replying to the failure verification function 42 with a response indicating normality is performed, and then the process proceeds to S104. On the other hand, in the determination process of S102, when it is determined that a normality confirmation request has not been received (when the determination result is NO), the process of S103 is skipped and the process proceeds to S104.
[0070] In S104, a process of acquiring and storing the execution log of the application and the metric log of the infrastructure is performed.
[0071] By the process of S104, in the case of agent 44a, the execution log of cloud application 12 and the metric log of cloud resource 11 are acquired and stored in the storage device of cloud resource 11. On the other hand, in the case of agent 44b, the execution log on device 20 of device application 22 and the metric log of the hardware of device 20 are acquired and stored in the storage device of device 20. In the case of agent 44b arranged on device 20b, further, the execution log on sub-device 21 of device application 22 and the metric log of the hardware of sub-device 21 are acquired and stored in the storage device of device 20b.
[0072] In S105, a process of receiving a failure occurrence command issued from the failure verification function 42 is performed, and in the subsequent S106, a process of determining whether a failure occurrence command has been received is performed.
[0073] In the determination process of S106, when it is determined that a failure occurrence command has been received (when the determination result is YES), the process proceeds to S107. Then, in S107, a process of determining whether the target to which the cause of failure occurrence set in the received failure occurrence command is to be assigned is the application or infrastructure for which the agent itself is responsible for information collection is performed.
[0074] As described above, the failure occurrence command includes settings of the cause of failure occurrence and the components of the target system to which the cause is to be assigned. In the determination process of S107, in the case of the agent 44a, when the target to which the cause of failure occurrence set in the failure occurrence command is to be assigned is the cloud resource 11 or the cloud application 12, the determination result is YES. On the other hand, in the case of the agent 44b, when the target to which the cause of failure occurrence is to be assigned is the hardware of the device 20 or the device application 22 executed on the device 20, the determination result is YES. However, in the case of the agent 44b arranged in the device 20b, furthermore, when the target to which the cause of failure occurrence is to be assigned is the hardware of the sub-device 21 or the device application 22 executed on the sub-device 21, the determination result is also YES.
[0075] When the determination result of the determination process in S107 is YES, the process proceeds to S108, and a process of assigning the cause of failure occurrence set in the failure occurrence command to the target to which the cause of failure occurrence is to be assigned is performed, and then the process proceeds to S109. On the other hand, when it is determined that the target to which the cause of failure occurrence is to be assigned is not the application or infrastructure for which the agent itself is responsible for information collection (when the determination result is NO), the process of S108 is skipped and the process proceeds to S109.
[0076] In S109, a process is performed to obtain and save the execution logs of the application and the metric logs of the infrastructure. This process is the same as the process in S104, and for each component after a cause of failure is assigned in response to a failure occurrence command, the execution logs and the metric logs are obtained.
[0077] In S110, a process is performed to transmit all of the execution logs and the metric logs saved by the process in S104 and / or the process in S109 to the information storage function 41. Then, in the subsequent S111, a process is performed to delete all of the saved execution logs and the metric logs from the storage device. This process in S111 is for relieving the pressure on the storage area of the storage device due to the storage of these data.
[0078] After the completion of the process in S111, or when it is determined in the determination process in S106 that a failure occurrence command has not been received (when the determination result is NO), the process returns to S101. Then, the process of receiving a normality confirmation request and the process of obtaining and saving the execution logs of the application and the metric logs of the infrastructure continue to be performed.
[0079] The above processes are the first example of the processing contents of the agents 44a and 44b.
[0080] Next, the flowchart in FIG. 5 will be described. FIG. 5 shows the processing contents of the information storage process performed in the cloud server to provide the information storage function 41.
[0081] When the process in FIG. 5 starts, first, in S201, a process is performed to receive a failure occurrence command issued from the failure verification function 42, and then, in S202, a process is performed to determine whether a failure occurrence command has been received.
[0082] In the determination process of S202, when it is determined that a failure occurrence command has been received (when the determination result is YES), the process proceeds to S203, and the process of receiving the execution log and the metric log is performed. Then, in the subsequent S204, the process of determining whether the execution log and the metric log have been received is performed.
[0083] The execution log and the metric log received by the process of S203 are sent from the cloud resource 11 and the device 20 respectively by the execution of the process of S110 in FIG. 4 in each of the agents 44a and 44b.
[0084] In the determination process of S204, when it is determined that the execution log and the metric log have been received (when the determination result is YES), the process proceeds to S205. On the other hand, in the determination process of S204, when it is determined that the execution log and the metric log have not been received (when the determination result is NO), the process returns to S203 and continues the process of receiving the execution log and the metric log.
[0085] In S205, when the cause of the failure occurrence based on the failure occurrence command issued from the failure verification function 42 is associated, the received execution log and metric log are stored in the storage device provided in the cloud server that executes the information storage process. Note that the information about the time when the cause of the failure occurrence was given is recorded in the execution log of the one to which the cause of the failure occurrence was given among the cloud resource 11 and the device 20 respectively.
[0086] In S206, a process is performed to determine whether execution logs and metric logs have been received from all of agents 44a and 44b by the process of S203. In this determination process, when it is determined that the logs have been received from all of agents 44a and 44b (when the determination result is YES), the process proceeds to S207. On the other hand, in this determination process, when it is determined that the execution logs and metric logs have not been received from all of agents 44a and 44b (when the determination result is NO), the process returns to S203 to continue receiving the execution logs and metric logs.
[0087] In S207, in association with the assignment of the cause of the occurrence of a failure based on the failure occurrence command issued from the failure verification function 42, a process is performed to store information about the failure occurrence command in a storage device provided in a cloud server that executes information storage processing. By this process, the information on the cause of the occurrence of the failure and the component of the target system to which the cause is assigned, which is included in the failure occurrence command received by the process of S201, is stored.
[0088] After the process of S207 described above is completed, the process returns to S201 to continue the process of receiving the execution logs and metric logs.
[0089] The above processes up to this point are information storage processes.
[0090] Next, the flowchart of FIG. 6 will be described. FIG. 6 shows the details of the failure verification process performed in the cloud server to provide the failure verification function 42.
[0091] When the process of FIG. 6 starts, first, in S301, a process is performed to confirm the normality of each of agents 44a and 44b. Subsequently, in S302, a process is performed to determine whether all of agents 44a and 44b are functioning normally.
[0092] In the process of S301, first, a normality confirmation request is sent to all of agents 44a and 44b. When agents 44a and 44b receive this confirmation request through the process of S101 (Figure 4) (the determination result of S102 is YES), they perform a process of replying with an answer indicating normality (S103). However, if they are not functioning properly, they cannot reply with this answer. Therefore, in the process of S301, a process of receiving this answer is also performed. In the subsequent determination process of S302, if the answer is received from all of agents 44a and 44b before a predetermined time elapses after sending the normality confirmation request, the determination result is set to YES and the process proceeds to S306. On the other hand, if the answer is not received from all of agents 44a and 44b even after the predetermined time has elapsed, the determination result is set to NO and the process proceeds to S303.
[0093] In S303, a process of outputting information (such as identification information) about agent 44a or 44b, which has been found to be malfunctioning abnormally through the processes of S301 and S302, to the administrator terminal 50 is performed.
[0094] In S304, a process of determining whether the agent 44a or 44b, which has been found to be malfunctioning abnormally, is responsible for collecting information regarding the sub-device 21 is performed. In the example of Figure 1, when an abnormality in the operation of agent 44b arranged in device 20b is found, the determination result of S304 is YES and the process proceeds to S305.
[0095] In S305, a process is performed to change the agent responsible for collecting information about the sub-device 21 to another agent 44b. In the example of FIG. 1, a process is performed to change the agent responsible for collecting information about the sub-device 21 from the agent 44b arranged in the device 20b to, for example, the agent 44b arranged in the device 20a. After the process of S305 is performed, the agent 44b of the device 20a also assigns a cause of failure to the device resources of the sub-device 21 or the device application 22 on the sub-device 21 based on a failure occurrence command.
[0096] After the process of S305, or when it is determined that the agent 44a or 44b, which has been found to have abnormal operation in the determination process of S304, is not responsible for collecting information about the sub-device 21 (when the determination result is NO), the process proceeds to S306.
[0097] In S306, a process is performed to obtain, from the administrator terminal 50, an input of a component that deliberately causes a failure in the target system and a cause for causing the failure, which are input by the administrator of the target system to the administrator terminal 50. Then, in the subsequent S307, a process is performed to determine whether the input of the component and the cause of failure has been obtained. When it is determined that the input has been obtained (when the determination result is YES), the process proceeds to S308. On the other hand, when it is determined in the determination process of S307 that the input of the component and the cause of failure has not been obtained (when the determination result is NO), the process returns to S306 and repeats the process of obtaining the input of the component and the cause of failure.
[0098] In S308, a process is performed to create a failure occurrence command including the setting of the components acquired by the process of S306 and the failure cause. Subsequently, in S309, a process is performed to transmit the created failure occurrence command to all of agents 44a and 44b and the information storage function 41. The failure occurrence command transmitted by this process is received by the process of S105 (FIG. 4) in agents 44a and 44b. Also, the failure occurrence command transmitted by this process is received by the process of S201 in the information storage process (FIG. 5).
[0099] In S310, a process is performed to determine whether the storage of the execution log and the metrics log received and stored in the information storage function 41 after the transmission of the failure occurrence command by the process of S309 is completed. In this determination process, when it is determined that the storage is completed (when the determination result is YES), the process proceeds to S311. When it is determined that the storage is not completed (when the determination result is NO), the determination process of S310 is repeated until the storage is completed.
[0100] In S311, a process is performed to check for the occurrence of abnormalities in components other than the components targeted to cause a failure by the failure occurrence command among the components of the target system, with reference to the execution log and the metrics log stored by the information storage function 41. Then, in S312 following that, a process is performed to determine whether an abnormality has occurred in the other components. In this determination process, when it is determined that an abnormality has occurred in the other components (when the determination result is YES), in S313, a process is performed to output information (for example, identification information) about the other components to the administrator terminal 50.
[0101] After the completion of the process of S313 described above, or when it is determined in the determination process of S312 that no abnormality has occurred in the other components (when the determination result is NO), the process returns to S306.
[0102] The above processes are the failure verification process.
[0103] Next, the flowchart of FIG. 7 will be described. FIG. 7 shows a second example of the processing contents of agents 44a and 44b, and shows the processing contents during the actual operation of the target system.
[0104] The processing from S401 to S403 that is first performed after the processing of FIG. 7 starts is the same as the processing from S101 to S103 in the first example shown in FIG. 4, and detailed description thereof will be omitted. However, the confirmation request received by the processing of S401 is transmitted to agents 44a and 44b by the failure identification function 43, and the reply destination of the reply indicating normality by the processing of S403 is also the failure identification function 43.
[0105] After the processing of S403, or when the determination result of the determination processing of S402 is NO, the processing proceeds to S404, and the processing of acquiring and storing the execution log of the application and the log of the infrastructure metrics is performed in the same manner as the processing of S104 in the first example shown in FIG. 4.
[0106] In S405, the execution log and the metrics log acquired by the processing of S404 are monitored to detect a failure. In this processing, the processing of monitoring the execution log to detect error items and the processing of monitoring the metrics log to detect values outside the preset normal value range of the metrics are performed.
[0107] In S406, based on the result of monitoring in S405, a process is performed to determine whether a failure has been detected. When it is determined that a failure has been detected (when the determination result is YES), the process proceeds to S407. Then, in S407, the execution log and the metrics log when the failure is detected, that is, the execution log and the metrics log obtained by the process of S404 executed most recently, are transmitted to the failure identification function 43. Then, in the subsequent S408, a process is performed to delete all of the stored execution log and metrics log from the storage device. The process of S408 is the same as the process of S111 in the first example shown in FIG. 4, and is for eliminating the pressure on the storage area of the storage device due to the storage of these data.
[0108] After the completion of the process of S408, the process returns to S401. Then, the process of receiving a normality confirmation request and the process of obtaining and storing the execution log of the application and the metrics log of the infrastructure are continuously performed.
[0109] By the way, in the determination process of S406, when it is determined that a failure has not been detected (when the determination result is NO), the process proceeds to S409 and a process of receiving a log transmission request is performed. Then, in the subsequent S410, a process is performed to determine whether a log transmission request has been received. The log transmission request received by the process of S409 is transmitted by the failure identification function 43 to agents 44a and 44b.
[0110] In the determination process of S410, when it is determined that a log transmission request has been received (when the determination result is YES), the process returns to S407, and the process of transmitting the execution log and the metrics log to the fault identification function 43 is performed. Then, in the subsequent S408, the process of deleting all the stored execution logs and metrics logs from the storage device is performed. Then, after that, the process returns to S401, and the process of receiving the normality confirmation request and the process of acquiring and storing the execution log of the application and the metrics log of the infrastructure continue to be performed. On the other hand, when it is determined in the determination process of S410 that a log transmission request has not been received (when the determination result is NO), the processes of S407 and S408 are skipped and the process returns to S401.
[0111] The above processes are the second example of the processing contents of the agents 44a and 44b.
[0112] Next, the flowchart of FIG. 8 will be described. FIG. 8 shows the processing contents of the fault identification process performed in the cloud server to provide the fault identification function 43.
[0113] The processes from S501 to S505 that are first performed when the process of FIG. 8 is started are the same as the processes from S301 to S305 in the fault verification process shown in FIG. 6, and detailed descriptions are omitted. However, in the agents 44a and 44b, the confirmation request transmitted by the confirmation process of S501 is received by the process of S401 (FIG. 7) (the determination result of S402 is YES), and the process of returning a reply indicating normality (S403) is performed. In the confirmation process of S501, the process of receiving this reply is also performed.
[0114] After the process of S505, or when the determination result of the determination process of S502 is YES, the process proceeds to S506. Then, in S506, a process of receiving the execution log and the metrics log transmitted in response to the detection of the occurrence of a failure by the processes of S406 and S407 (Fig. 7) of agent 44a or 44b is performed. Then, in the subsequent S507, a process of determining whether the execution log and the metrics log have been received is performed, and when it is determined that they have been received (when the determination result is YES), the process proceeds to S508. On the other hand, in this determination process, when it is determined that the execution log and the metrics log have not been received (when the determination result is NO), the process returns to S506, and the process of receiving the execution log and the metrics log is continued.
[0115] In S508, a process of transmitting a log sending request is performed for each of the remaining agents 44a and 44b excluding the source of the execution log and the metrics log received by the process of S506. Then, in the subsequent S509, a process of receiving the execution log and the metrics log sent by the processes of S409 and S410 (Fig. 7) in each of the remaining agents 44a and 44b is performed.
[0116] In S510, a process of calculating the matching rate between the execution log and the metrics log received by the process of S509 and the execution log and the metrics log stored in the information storage function 41 is performed for each cause of the occurrence of a failure.
[0117] By the processes of S205 and S207 of the information storage process (Fig. 5) described above, the cause of the occurrence of a failure and the execution log and the metrics log for each component of the target system are linked by the information when the cause of the occurrence of a failure is given. In the process of S510, according to this link, the matching rate for each cause of the occurrence of a failure is calculated.
[0118] In S511, a process of extracting a predetermined number (3 in the example of Fig. 8) of factors with the highest matching rates calculated by the process of S510 from among the causes of the occurrence of a failure is performed.
[0119] In S512, a process is performed to output information (e.g., identification information) on the cause of the failure extracted by the process of S511 to the administrator terminal 50.
[0120] After the completion of the process of S512 described above, the process returns to S506, and the process of receiving the execution log and the metrics log transmitted in response to the detection of the failure is continued.
[0121] The above processes are the failure identification processes.
[0122] Next, a specific example of the failure verification of the target system performed using the failure verification system 100 will be described with reference to FIGS. 9 and 10.
[0123] The relationship between each element shown in FIG. 9 and each element shown in FIG. 1 will be described. "C001", "C002", "C003", "C004", "C005", and "C006" on the cloud 10 are examples of cloud resources 11. Also, the data collection application [S003], the data processing application [S004], the data storage application [S005], the data display application [S006], the file transmission application [S007], and the file storage application [S008] are examples of cloud applications 12. These cloud applications 12 are each executed on the cloud resources 11 described above. And the "agent" superimposed on "C001", "C002", "C003", "C004", "C005", "C006" in FIG. 9 is the agent 44a.
[0124] On the other hand, the data transmission application [S001] and the file acquisition application [S002] provided in the device 20 are specific examples of the device application 22. Also, "D001" is a specific example of the hardware of the device 20 that executes the data transmission application [S001] and the file acquisition application [S002]. And the "agent" superimposed on "D001" in FIG. 9 is the agent 44b.
[0125] During the actual operation of the target system of FIG. 9 configured as described above, here, a case where a failure occurs in the downstream communication for sending various data and files from the cloud 10 to the device 20 will be described. Note that this failure is caused by an abnormality in the communication network on the cloud 10 between the file storage application [S008] responsible for file storage and the file transmission application [S007] responsible for file transmission to the device 20.
[0126] For such a target system, the administrator of the target system performs pre-verification of failures in the target system using the failure verification function 42 of the failure verification system 100. In this pre-verification, the administrator operates the administrator terminal 50 to input the components that deliberately cause a failure and the factors causing the failure in the target system. The failure verification function 42 checks the normality of each of the agents 44a and 44b (S301 to S305 in FIG. 6, S101 to S103 in FIG. 4), and then creates and transmits a failure occurrence command according to this input (S306 to S309 in FIG. 6).
[0127] When each of the agents 44a and 44b receives this failure occurrence command from the failure verification function 42, it assigns the cause of the failure set in the failure occurrence command to the components of the target system set in the failure occurrence command (S105 to S108 in FIG. 4). Then, each of the agents 44a and 44b collects the execution log and the metric log after the cause of the failure is assigned for each component of the target system and transmits them to the information storage function 41 (S109 to S110 in FIG. 4).
[0128] On the one hand, when the information storage function 41 receives a failure occurrence command from the failure verification function 42, it also receives the execution log and the metrics log for each component of the target system from each of the agents 44a and 44b (S201 to S204 in FIG. 5). The information storage function 41 associates and stores the information about the received failure occurrence command, the execution log, and the metrics log when the cause of the failure occurrence set in the failure occurrence command is assigned (S205 to S207 in FIG. 5). Among the various tables shown in FIG. 10, the failure occurrence command list 70, the application execution log 71, the cloud resource metrics log 72, and the device resource metrics log 73 show the data examples of the target system in FIG. 9 saved at this time.
[0129] The failure occurrence command list 70 is a list of information about the failure occurrence command. In the failure occurrence command list 70, the "failure verification ID" is information for identifying the failure occurrence command, and the "time" is information indicating the time when the cause of the failure occurrence set in the failure occurrence command is assigned. Also, the "target" is information indicating the component of the target system to which the cause of the failure occurrence is assigned in the failure occurrence command. Furthermore, the "content" is information indicating the cause of the failure occurrence assigned to the component of the target system set in the failure occurrence command.
[0130] The application execution log 71 is the execution log for each application. In the application execution log 71, the "application ID" is information for identifying the application (cloud application 12 and device application 22), and the "infrastructure ID" is information for identifying the infrastructure on which the application identified by the "application ID" is executed. Also, the "time" is information indicating the time when the cause of the failure occurrence set in the failure occurrence command is assigned, and the "log information" is the content of the execution log for the application identified by the "application ID". In the example of the application execution log 71 in FIG. 10, it is shown that an error "Error NoX" has occurred in the content of the execution log.
[0131] The cloud resource metrics log 72 is a log of metrics for each cloud resource 11. In the cloud resource metrics log 72, the "cloud resource ID" is information for identifying the cloud resource 11, and the "time" is information indicating the time when the cause of the failure occurrence set in the failure occurrence command was given. Also, "CPU", "memory", and "network" are information indicating the utilization rate, which is a metric for each of the CPU, memory, and interface device that the cloud resource 11 identified by the "cloud resource ID" has.
[0132] The device resource metrics log 73 is a log of metrics for each device 20 (or each sub-device 21). In the device resource metrics log 73, the "device ID" is information for identifying the device 20 and the sub-device 21, and the "time" is information indicating the time when the cause of the failure occurrence set in the failure occurrence command was given. Also, "CPU", "memory", and "network" are information indicating the utilization rate, which is a metric for each of the CPU, memory, and interface device that the device 20 or sub-device 21 identified by the "device ID" has.
[0133] In the pre-verification of failures in the target system, information as described above is stored by the information storage function 41. Next, a description will be given of the state of identifying the cause of a failure that occurred during the actual operation of the target system, which the failure verification system 100 performs using this information.
[0134] When the failure identification function 43 of the failure verification system 100 is started in response to the start of actual operation in the target system, the failure identification function 43 checks the normality of each of the agents 44a and 44b (S501 to S505 in FIG. 8, S401 to S403 in FIG. 7). Thereafter, each of the agents 44a and 44b collects and monitors the execution log and the metrics log after the cause of failure is assigned for each component of the target system (S404 to S405 in FIG. 7). When an abnormality is detected from the execution log or the metrics log in this monitoring, each of the agents 44a and 44b transmits the collected execution log and metrics log to the failure identification function 43 (S406 to S410 in FIG. 7). Among the various tables shown in FIG. 10, the application execution log 81, the cloud resource metrics log 82, and the device resource metrics log 83 show data examples of the target system in FIG. 9 that were transmitted to the failure identification function 43 at this time.
[0135] Each item of the application execution log 81, the cloud resource metrics log 82, and the device resource metrics log 83 is the same as each item of the application execution log 71, the cloud resource metrics log 72, and the device resource metrics log 73, respectively. Here, the description of each item of the application execution log 81, the cloud resource metrics log 82, and the device resource metrics log 83 is omitted.
[0136] Here, the agent 44b that collects and monitors the execution log of the file acquisition application [S002] on the device 20 obtains error information "Error NoX" from the application execution log 81. Then, the agent 44b determines that a failure has occurred in the file acquisition application [S002], and transmits the application execution log 81 and the device resource metrics log 83 for the device 20 to the failure identification function 43 (S406 to S407 in FIG. 7).
[0137] When the failure identification function 43 receives the application execution log 81 and the device resource metrics log 83 from the agent 44b of the device 20, it sends a log transmission request to each of the remaining agents 44a (S506 to S508 in FIG. 8). When each agent 44a receives this log transmission request, it sends the stored application execution log 81 and the cloud resource metrics log 82 to the failure identification function 43 (S409 to S410 and S407 in FIG. 7). The failure identification function 43 receives these application execution log 81 and cloud resource metrics log 82 (S509 in FIG. 8).
[0138] Here, the failure identification function 43 calculates the matching rate between the execution log and the metrics log and the execution log and the metrics log stored in the information storage function 41 for each cause of the failure occurrence (S510 in FIG. 8). For this calculation, the application execution log 81, the cloud resource metrics log 82, and the device resource metrics log 83, and the application execution log 71, the cloud resource metrics log 72, and the device resource metrics log 73 of the information storage function 41 are used. Then, the failure identification function 43 extracts a predetermined number (3 in this example) of causes of failure occurrence with the highest calculated matching rate and outputs them to the administrator terminal 50. Among the various tables shown in FIG. 10, the cause information output 90 shows an example of the cause of failure occurrence output by the failure identification function 43 at this time.
[0139] In the cause information output 90, the "actual failure ID" is information for identifying a failure that occurred during the actual operation of the target system, the "time" is information indicating the time when the failure occurred, and the "target" is information for identifying the component of the target system where the failure occurred. Also, the "matching rate 1", "matching rate 2", and "matching rate 3" are information showing three causes of failure occurrence and their matching rates in order from the highest matching rate, and are information indicating the cause of the failure identified by the failure identification function 43.
[0140] In the cause information output 90 of FIG. 10, it is shown that the failure identified by the actual failure ID "I001" occurred at the time "05:00:00" in the file acquisition application [S002], i.e., [S002]. Also, it is shown that among the failures pre-verified by the failure verification function 42, the failure causes identified by the failure verification IDs "G002", "G001", and "G003" were identified by the failure identification function 43 as three candidates for the cause of this failure. Furthermore, it is shown that the matching rates of the failure causes identified by the failure verification IDs "G002", "G001", and "G003" are "90%", "50%", and "10%", respectively.
[0141] By referring to such cause information output 90 output from the administrator terminal 50, the administrator of the target system can promptly implement various countermeasures (such as contacting the cloud vendor or starting an alternative server) according to the examination results in the pre-verification of the failure cause.
[0142] <Second Embodiment> As described above, in the first embodiment, failure verification information representing logs / metrics when a failure is intentionally generated before actual operation is created. Then, when a failure occurs during actual operation, the cause of the failure is estimated by comparing the logs / metrics related to the failure with the failure verification information.
[0143] In the second embodiment, not only can the cause of a failure that occurred during actual operation be estimated, but also the cause of a sign of a failure can be estimated. Also, in the second embodiment, the verification target system includes a multi-core device in which a plurality of processor cores operate using different operating systems (OSs). Furthermore, in the second embodiment, a fail-safe function is provided when a failure occurs or a sign of a failure is detected.
[0144] FIG. 11 shows an example of a system to be verified and a failure verification system in the second embodiment. In the example shown in FIG. 11, the system to be verified includes a plurality of devices connected to network 200. Specifically, the system to be verified includes one or more devices 210, one or more sub-devices 220, and one or more multi-core devices 230. In FIG. 11, one device 210, one sub-device 220, and one multi-core device 230 are depicted. Note that the system to be verified may not include device 210. Also, the system to be verified may not include sub-device 220. Furthermore, the system to be verified may include server 250 provided on cloud 240.
[0145] Network 200 corresponds to, for example, public line 30 shown in FIG. 1, but may include a private network. Also, network 200 may be a wireless network, a wired network, or may include both a wireless network and a wired network. Cloud 240 corresponds to cloud 10 shown in FIG. 1 and is realized by a plurality of computers. The plurality of computers may be connected to each other via a network.
[0146] The failure verification system in the second embodiment includes an information storage unit 301, a failure verification unit 302, a failure identification unit 303, a fail-safe control unit 304, and failure verification agents 310 to 312. In this example, the information storage unit 301, the failure verification unit 302, the failure identification unit 303, and the fail-safe control unit 304 are provided on cloud 240. Also, failure verification agent 310 is installed in device 210, and failure verification agents 311 to 312 are installed in multi-core device 230.
[0147] Device 210 corresponds to device 20 shown in FIG. 1 and can communicate with server 250 via network 200. Further, device 210 includes a processor and a memory, and a general-purpose OS is implemented. The general-purpose OS includes a network stack, file I / O, etc. in this embodiment and is directly operable from cloud 240 (e.g., fault verification unit 302, fault identification unit 303, fail-safe control unit 304). The general-purpose OS is, for example, Linux (registered trademark). Furthermore, application program A1 (hereinafter, app A1) and fault verification agent 310 are implemented in device 210.
[0148] Sub-device 220 corresponds to sub-device 21 shown in FIG. 1. That is, sub-device 220 has less resources (CPU capacity, memory capacity) compared to device 210. Therefore, it is assumed that no fault verification agent is implemented in sub-device 220.
[0149] Multi-core device 230 includes a plurality of processor cores. Specifically, multi-core device 230 includes A-core 231 and M-core 232. A-core 231 is a processor core for executing general-purpose processing, and a general-purpose OS such as Linux is implemented. Further, app A2 and fault verification agent 311 are implemented in A-core 231. M-core 232 is a processor core for executing specific processing, and an embedded OS such as RTOS (real-time OS) is implemented. It is assumed that the embedded OS does not support communication via a network. Further, app A3 and fault verification agent 312 are implemented in M-core 232. In this embodiment, app A3 controls peripheral device 260. Note that peripheral device 260 includes devices mounted on the board (such as LEDs) and devices controlled through peripherals such as USB / Ethernet (registered trademark) ports.
[0150] The multi-core device 230 includes a shared memory 233 shared by the A core 231 and the M core 232. The shared memory 233 includes an inter-core communication area 233a and an embedded app information storage area 233b. The inter-core communication area 233a is used for communication between the A core 231 and the M core 232. In the embedded app information storage area 233b, operation state information representing the operation state of the embedded app (i.e., app A3) implemented in the M core 232 is stored. The operation state information includes, for example, the values of the parameters used by app A3.
[0151] Note that when the M core 232 stops, the app A3 implemented in the M core 232 is executed on the A core 231 by a fail-safe function described later. Therefore, in this embodiment, the app A3 is implemented not only in the M core 232 but also in the A core 231. Normally, the app A3 implemented in the M core 232 is executed, and when a failure occurs in the M core 232, the app A3 implemented in the A core 231 is executed.
[0152] The information storage unit 301 stores the log / metric information collected by each of the failure verification agents 310 to 312. The log information represents, for example, the operations (i.e., execution logs) of apps A1 to A3. The metric information represents, for example, the states of the hardware resources (CPU usage rate, memory usage rate, etc.) of each device (210, 220, 230).
[0153] The failure verification unit 302 can intentionally cause a failure in the devices (210, 220, 230) by giving a failure occurrence command to the failure verification agents 310 to 312 in response to an instruction from the administrator of the failure verification system. Further, the failure verification unit 302 stores the log / metric information collected by the failure verification agents 310 to 312 in the information storage unit 301. Furthermore, the failure verification unit 302 can verify the operation of the system under verification based on the log / metric information when the failure is intentionally caused. For example, when the communication of a predetermined path in the system under verification is disconnected, when a predetermined application is stopped, when an excessive load occurs on the CPU of a predetermined device, or when the usage rate of the memory of a predetermined device becomes high, it can detect what log / metric information can be obtained. In the following description, the log / metric information obtained when the failure verification unit 302 intentionally causes a failure may be referred to as "failure verification information (or reference log / metric information)".
[0154] The failure identification unit 303 estimates the cause of a failure or its sign that has occurred in the system under verification based on the log / metric information collected by the failure verification agents 310 to 312 during the actual operation of the system under verification. At this time, the failure identification unit 303 estimates the cause of the failure or its sign by comparing the newly collected log / metric information with the failure verification information stored in the information storage unit 301. In the following description, the log / metric information collected by the failure verification agents 310 to 312 during the actual operation of the system under verification may be referred to as "monitoring information". However, the monitoring information may be a part of the log / metric information collected by each of the failure verification agents 310 to 312.
[0155] When a failure or its sign is detected in the multi-core device 230, the fail-safe control unit 304 gives a fail-safe instruction to the failure verification agent 311 so that the operation of the multi-core device 230 continues. In this embodiment, when a failure or its sign of the M core 232 is detected, the fail-safe control unit 304 gives an instruction to the failure verification agent 311 to execute a fail-safe process for operating the application A3 that was operating on the M core 232 on the A core 231. Thereby, the fail-safe process in the multi-core device 230 is executed.
[0156] The failure verification agent 310 implemented in the device 210 generates a specified failure in the device 210 or the sub-device 220 in response to a failure occurrence command given from the failure verification unit 302. At this time, the failure verification agent 310 collects log / metric information from the device 210 and the sub-device 220 and transmits it to the failure verification unit 302. Also, during the actual operation of the system to be verified, the failure verification agent 310 periodically collects log / metric information from the device 210 and the sub-device 220 and transmits it to the failure identification unit 303.
[0157] The failure verification agent 311 implemented in the A core 231 within the multi-core device 230 has the same functions as the failure verification agent 310. That is, the failure verification agent 311 generates a specified failure in the A core 231 in response to a failure occurrence command given from the failure verification unit 302. At this time, the failure verification agent 311 collects log / metric information from the A core 231 and transmits it to the failure verification unit 302. Also, during the actual operation of the system to be verified, the failure verification agent 311 periodically collects log / metric information from the A core 231 and transmits it to the failure identification unit 303. Note that since the failure verification agent 311 operates on a general-purpose OS, it may be called the "general-purpose OS failure verification agent (311)".
[0158] The general-purpose OS failure verification agent 311 can cause a specified failure in the M core 232 by giving an instruction to the failure verification agent 312 implemented in the M core 232 in response to a failure occurrence command given from the failure verification unit 302. At this time, the general-purpose OS failure verification agent 311 transmits the log / metric information of the M core 232 received from the failure verification agent 312 to the failure verification unit 302. Also, during the actual operation of the system under verification, the general-purpose OS failure verification agent 311 transmits the log / metric information of the M core 232 periodically received from the failure verification agent 312 to the failure identification unit 303.
[0159] When the general-purpose OS failure verification agent 311 receives an execution command for fail-safe processing from the fail-safe control unit 304, it executes the fail-safe processing in cooperation with the failure verification agent 312. In this embodiment, in response to the execution command for fail-safe processing, the multi-core device 230 is reconfigured so that the application A3 operating on the M core 232 operates on the A core 231.
[0160] The failure verification agent 312 can cause a failure in the M core 232 in response to an instruction given from the general-purpose OS failure verification agent 311. At this time, the failure verification agent 312 passes the log / metric information of the M core 232 to the general-purpose OS failure verification agent 311. Also, during the actual operation of the system under verification, the failure verification agent 312 periodically collects the log / metric information of the M core 232 and passes it to the general-purpose OS failure verification agent 311. Communication between the general-purpose OS failure verification agent 311 and the failure verification agent 312 is performed using the inter-core communication area 233a. Note that since the failure verification agent 312 operates on the embedded OS, it may be called the "embedded OS failure verification agent (312)".
[0161] The administrator who manages the above-described failure verification system can input necessary instructions using the administrator terminal 320. Also, information related to a failure detected in the failure verification system is transmitted to the administrator terminal 320. The information related to the failure includes information indicating the cause of the occurred failure or the sign of the failure. Thereby, the administrator can recognize the cause of the failure or the sign of the failure that has occurred in the failure verification system. Therefore, when a failure occurs, the administrator can immediately take measures against the failure. Also, when a sign of a failure is detected, the administrator can take necessary measures against the cause before the failure occurs.
[0162] In the failure verification system having the above configuration, the information storage unit 301 is realized by using a storage device that constitutes the cloud 240. The failure verification unit 302, the failure identification unit 303, and the fail-safe control unit 304 are realized by using one or more computers that constitute the cloud 240. The failure verification agents 310 to 312 are application programs installed in the corresponding devices, and their functions are realized by being executed by the processors of the devices.
[0163] The failure verification by the above-described failure verification system includes the following phases. (1) Preparation phase (pre-verification phase) (2) Failure identification phase (3) Fail-safe phase Note that hereinafter, the failure verification for the multi-core device 230 will be described.
[0164] (1) Preparation phase (pre-verification phase) The administrator uses the administrator terminal 320 to set the failure verification content for the failure verification unit 302. For example, for each failure cause, it is set which failure cause is to be injected into which component of which device. Also, a corresponding failure verification agent is installed for each device.
[0165] FIG. 12 is a flowchart showing an example of the processing on the cloud 240 side in the preparation phase. FIG. 13 is a flowchart showing an example of the processing of the failure verification agent in the preparation phase.
[0166] In S601 shown in FIG. 12, the failure verification unit 302 transmits a log / metric request to the general-purpose OS failure verification agent 311. The log / metric request requests execution logs of applications executed on the multi-core device 230 and metric information representing the operating states of various hardware components of the multi-core device 230.
[0167] In S611 shown in FIG. 13, the general-purpose OS failure verification agent 311 waits for a log / metric request. Then, when the general-purpose OS failure verification agent 311 receives the log / metric request, in S612, it collects the log / metric information of core A 231. Subsequently, in S613, the general-purpose OS failure verification agent 311 instructs the embedded OS failure verification agent 312 to collect the log / metric information of core M 232. Then, the embedded OS failure verification agent 312 collects the log / metric information of core M 232 and passes it to the general-purpose OS failure verification agent 311. Thereby, the general-purpose OS failure verification agent 311 obtains the log / metric information of core M 232. After that, in S614, the general-purpose OS failure verification agent 311 transmits the log / metric information of core A 231 and the log / metric information of core M 232 to the failure verification unit 302.
[0168] In S602 shown in FIG. 12, the failure verification unit 302 stores the log / metric information received from the general-purpose OS failure verification agent 311 (i.e., the log / metric information transmitted in S614) in the information storage unit 301 as "normal-time failure verification information". The "normal-time failure verification information" is stored in association with the time when the log / metric request was transmitted.
[0169] In S603, the failure verification unit 302 transmits a failure occurrence command to the general-purpose OS failure verification agent 311. The failure occurrence command represents the component to which a failure cause should be injected and the content of the failure cause. An example of the failure occurrence command is "Component: Core A 231, Content of failure cause: Stop Application A2".
[0170] In S615 shown in FIG. 13, the general-purpose OS failure verification agent 311 waits for a failure occurrence command. Then, when the general-purpose OS failure verification agent 311 receives the failure occurrence command, in S616, it injects the specified failure cause into the specified component. That is, the general-purpose OS failure verification agent 311 causes the specified failure to occur. When the specified component is Core M 232, the general-purpose OS failure verification agent 311 instructs the embedded OS failure verification agent 312 to inject the specified failure cause. Then, the embedded OS failure verification agent 312 causes the specified failure to occur in Core M 232.
[0171] S617 to S619 are substantially the same as S612 to S614. That is, the general-purpose OS failure verification agent 311 collects the log / metric information of Core A 231, and the embedded OS failure verification agent 312 collects the log / metric information of Core M 232. Then, the general-purpose OS failure verification agent 311 transmits the log / metric information of Core A 231 and the log / metric information of Core M 232 to the failure verification unit 302.
[0172] In S604 shown in FIG. 12, the failure verification unit 302 stores the log / metric information received from the general-purpose OS failure verification agent 311 (that is, the log / metric information transmitted in S619) in the information storage unit 301 as "failure verification information at the time of failure occurrence". The "failure verification information at the time of failure occurrence" is stored in association with the time when the failure occurrence command was transmitted.
[0173] Note that the processes of S603 - S604 and S615 - S619 are performed for each factor of the intentionally generated failure. For example, for failure factors such as "Stop App A2", "Increase the CPU load of Core A231", and "Increase the memory usage of Core M232", "failure verification information at the time of failure" is saved respectively.
[0174] Also, in the example shown in FIGS. 12 - 13, a log / metric request is sent before the failure occurrence command, but the second embodiment is not limited to this procedure. For example, the processes of S601 - S602 and S611 - S614 may be executed due to the failure occurrence command.
[0175] In S605 shown in FIG. 12, the failure verification unit 302 verifies the operation of the system under verification based on the "failure verification information at the time of failure". For example, it is verified what kind of operating state occurs for what kind of failure factor. At this time, the failure verification unit 302 may verify the operation of the system under verification based on the "failure verification information in normal times" and the "failure verification information at the time of failure". Then, in S606, the failure verification unit 302 transmits the verification result to the administrator terminal 320.
[0176] FIG. 14 shows an example of various information stored in the information storage unit 301 in the preparation phase. FIG. 14A shows an example of a failure factor list. The failure factor list represents the failure factors to be injected in S616 shown in FIG. 13 and is set by the administrator. The failure factor ID identifies the failure factor to be injected. The time represents the time when the corresponding failure occurrence command was sent (or the time when the failure factor was injected). The target represents the component into which the failure factor is to be injected. The content represents the content of the failure factor.
[0177] FIG. 14B shows an example of application management information. The application management information includes information related to applications implemented in the system under verification. The application ID identifies each application implemented in the system under verification. The FS availability indicates whether fail-safe processing can be performed on the application. The FS deadline represents the period during which fail-safe processing can be performed on the application. The FS flag indicates whether fail-safe processing is being performed on the application. The control device represents the peripheral devices controlled by the application.
[0178] FIG. 14C shows an example of failure verification information. The failure verification information includes information related to the state of the device and information related to the state of the application. Here, the failure verification information corresponds to the log / metric information transmitted from the general-purpose OS failure verification agent 311 in S619 shown in FIG. 13. Further, the failure verification information may include the log / metric information transmitted from the general-purpose OS failure verification agent 311 in S614 shown in FIG. 13. Note that in FIG. 14C, the failure verification information is stored for each cause of the failure. However, the failure verification information does not necessarily have to be stored for each cause of the failure and may be stored in the order in which the failure verification unit 302 receives it.
[0179] The device ID identifies each device within the system under verification. The OS represents the OS implemented on the device. Note that since the device identified by "D002" (i.e., the multi-core device 230 shown in FIG. 11) has two processor cores, two records are allocated to the device. The time represents the time when the corresponding failure occurrence command was transmitted (or the time when the cause of the failure was injected). The CPU represents the CPU usage rate of the device. The memory represents the memory usage rate of the device. The NW represents the load on the network to which the device is connected. The application ID identifies the applications executed in the system under verification. The infrastructure ID represents the hardware on which the application is executed. The log information represents the execution log of the application.
[0180] In this way, in the preparation phase, a failure occurrence command is sent from the failure verification unit 302 to the corresponding failure verification agent, so that the specified failure factor is injected into the specified component in the system under verification. The failure verification agent collects log / metric information corresponding to each failure factor and sends it to the failure verification unit 302. Then, the log / metric information received from the failure verification agent is stored as failure verification information as shown in FIG. 14C.
[0181] Here, the general-purpose OS failure verification agent 311 operating on the general-purpose OS and the embedded OS failure verification agent 312 operating on the embedded OS transmit information to each other via the shared memory 233. Therefore, the failure verification unit 302 can inject a desired failure factor into a component that cannot communicate via the network (i.e., the M core 232), and can also acquire the log / metric information of such a component.
[0182] (2) Failure identification phase FIG. 15 is a flowchart showing an example of the processing of the failure verification agent in the failure identification phase. The processing of this flowchart is repeatedly executed at predetermined time intervals, for example, during the actual operation of the system under verification. Alternatively, the processing of this flowchart may be executed in response to an instruction from the failure identification unit 303.
[0183] S621 to S622 are substantially the same as S612 to S613 or S617 to S618 shown in FIG. 13 executed in the preparation phase. That is, the general-purpose OS failure verification agent 311 collects the log / metric information of the A core 231, and the embedded OS failure verification agent 312 collects the log / metric information of the M core 232. However, in S623, the general-purpose OS failure verification agent 311 sends the log / metric information of the A core 231 and the log / metric information of the M core 232 to the failure identification unit 303.
[0184] In S624, the general-purpose OS failure verification agent 311 and / or the embedded OS failure verification agent 312 writes the operation state information indicating the operation state of the fail-safe target application to the embedded application information storage area 233b. In the example shown in FIG. 11, the embedded OS failure verification agent 312 writes the operation state information of the application A3 to the embedded application information storage area 233b. After this, in order to reduce the used area of the shared memory 233, it is preferable to delete the information transmitted to the failure identification unit 303 in S623.
[0185] FIG. 16 is a flowchart showing an example of the processing of the failure identification unit 303 in the failure identification phase. The processing of this flowchart is periodically executed by the failure identification unit 303, for example. Note that, as described above, the failure verification agents implemented in each device periodically transmit the collected log / metric information to the failure identification unit 303. Then, the failure identification unit 303 stores the log / metric information received from each failure verification agent in the information storage unit 301. That is, the log / metric information related to each device is stored as "monitoring information" in the information storage unit 301.
[0186] In S631, the failure identification unit 303 acquires the latest log / metric information from the information storage unit 301. At this time, it is assumed that the failure identification unit 303 acquires the log / metric information shown in FIG. 17A.
[0187] In S632, the failure identification unit 303 compares the latest log / metric information acquired in S631 with the failure verification information stored in the information storage unit 301. Here, it is assumed that failure verification information is stored in the information storage unit 301 for each failure factor registered in the failure factor list shown in FIG. 14A. In this case, the failure identification unit 303 compares the latest log / metric information with each failure verification information. Specifically, the failure identification unit 303 calculates the matching rate between the latest log / metric information and each failure verification information in the same manner as the processing of S510 shown in FIG. 8. The matching rate may be calculated by the same method as in the first embodiment described above.
[0188] The matching rate is calculated based on the comparison between the latest log / metric information and the failure verification information of each failure factor (in FIG. 14, G001 to G003). At this time, the matching rate may be calculated by comparing only the metric information (one or more of CPU usage rate, memory usage rate, network usage rate), or the matching rate may be calculated by comparing both the log information and the metric information. Also, the matching rate may be calculated for each device. For example, when comparing the failure verification information shown in FIG. 14C and the log / metric information shown in FIG. 17A, for each failure factor G001 to G003, the matching rate for device D001 and the matching rate for device D002 may be calculated respectively, or the matching rate for device D001, the matching rate for A core 231 (device D002_general-purpose OS), and the matching rate for M core 232 (device D002_embedded OS) may be calculated respectively. Alternatively, the failure verification information of the entire system to be verified may be compared with the log / metric information of the entire system to be verified. As an example, the matching rate for the third record in FIGS. 14C and 17A may be calculated as follows. Matching rate = (90÷95)×(89÷90)×100 = 93.7%
[0189] However, the method for calculating the matching rate is not particularly limited. That is, the matching rate is a value representing the degree of coincidence or similarity between the log / metric information collected during the actual operation of the system to be verified and the failure verification information stored for each failure factor, and the calculation method is not particularly limited. As an example, the following calculation results can be obtained. (1) Matching rate with failure factor G001: 50 percent (2) Matching rate with failure factor G002: 90 percent (3) Matching rate with failure factor G003: 10 percent
[0190] In S633, the failure identification unit 303 determines whether the maximum value of the matching ratio calculated in S632 exceeds a predetermined threshold. The predetermined threshold is determined by, for example, the administrator of the failure verification system and is set in the failure identification unit 303. As an example, the threshold is 80 percent.
[0191] When a matching ratio exceeding the threshold is detected, the failure identification unit 303 determines that a failure has occurred or a sign of a failure has occurred in the system under verification. At this time, the failure identification unit 303 identifies the failure factor having a matching ratio exceeding the threshold. In the above example, the matching ratio with the failure factor G002 exceeds the threshold. Then, the failure identification unit 303 identifies the target of the identified failure factor (that is, the component into which the failure factor was injected in the preparation phase). In the example shown in FIG. 14A, as the failure factor G002, the application A3 is stopped. Therefore, in this case, "application A3" is obtained as the target of the identified failure factor. In the following description, a component into which a failure factor having a matching ratio exceeding the threshold was injected in the preparation phase may be referred to as a "failure candidate component".
[0192] In S634, the failure identification unit 303 refers to the application management information shown in FIG. 14B and determines whether it is possible to perform a fail-safe process on the failure candidate component. In this example, it is set that a fail-safe process can be performed on the application A3. When a fail-safe can be performed on the failure candidate component, the failure identification unit 303 changes the FS flag of the component from "0" to "1" in S635. "FS flag = 0" represents a state in which the fail-safe process is not being performed, and "FS flag = 1" represents a state in which the fail-safe process is being performed. When a fail-safe cannot be performed on the failure candidate component, the process of S635 is skipped.
[0193] In S636, the failure identification unit 303 transmits information representing a failure candidate component (i.e., a component into which a failure factor having a match rate exceeding a threshold value is injected in the preparation phase) to the administrator terminal 320. In the example shown in FIG. 17B, the match rate between the log / metric information collected at 15:00:00 and the failure verification information obtained when injecting the failure factor G002 exceeds the threshold value. Here, referring to the failure factor list shown in FIG. 14A, the failure factor G002 represents "Application A3 has stopped". Therefore, the failure identification unit 303 transmits information representing that "a failure or sign of a failure has occurred in Application A3" to the administrator terminal 320. Thereby, the administrator can recognize the component in which a failure or sign of a failure has occurred.
[0194] When a match rate exceeding the threshold value is not detected (S633: No), in S637, the failure identification unit 303 checks whether a failure has actually occurred in the system to be verified. The occurrence of a failure can be recognized, for example, from application logs, an alarm signal output from the system to be verified, or a report from a user of the system to be verified.
[0195] When a failure has actually occurred in the system to be verified, in S638, the failure identification unit 303 acquires the log / metric information collected at the time when the failure occurred from the information storage unit 301. Then, the failure identification unit 303 calculates the match rate between the log / metric information at the time of failure occurrence and each failure verification information stored in the information storage unit 301.
[0196] The processes of S639 to S641 are substantially the same as those of S634 to S636 described above. However, in S639 to S641, the failure diagnosis unit 303 may determine whether it is possible to perform fail-safe on the component that is outputting the alarm signal or the component reported by the user. Further, the failure diagnosis unit 303 may determine whether it is possible to perform fail-safe processing on a component related to any one of a predetermined number of failure factors having a high matching rate with the log / metric information at the time of failure occurrence. In this case, it is preferable that the failure diagnosis unit 303 determines whether it is possible to perform fail-safe processing on the component related to the failure factor having the highest matching rate with the log / metric information at the time of failure occurrence. Further, the failure diagnosis unit 303 extracts a plurality (for example, three) of failure factors having a high matching rate with the log / metric information at the time of failure occurrence, and transmits information representing these plurality of failure factors to the administrator terminal 320.
[0197] In the example shown in FIG. 17C, a failure occurred at 16:00:00, and the failure diagnosis unit 303 calculated the matching rate between the log / metric information collected at 16:00:00 and each failure verification information. The matching rate with the failure factor G002 is 70%, the matching rate with the failure factor G003 is 50%, and the matching rate with the failure factor G001 is 30%. Here, referring to the failure factor list shown in FIG. 14A, the failure factor G002 represents "Application A3 has stopped", the failure factor G003 represents "High CPU load of M core 232", and the failure factor G001 represents "Application A2 has stopped". Therefore, the failure diagnosis unit 303 transmits information representing that "the possibility that Application A3 has stopped is the highest", "the possibility that the CPU load of M core 232 is high is the second highest", and "the possibility that Application A2 has stopped is the third highest" as the cause of the occurred failure to the administrator terminal 320. Thereby, the administrator can obtain information for analyzing the cause of the occurred failure.
[0198] If no failure factor with a match rate exceeding the threshold is detected (S633: No), and if no failure has occurred in the system under verification (S637: No), it is determined that there is no failure or sign of failure, and the process of the failure identification unit 303 proceeds to S642. In S642, the failure identification unit 303 determines whether fail-safe processing is being performed. Whether fail-safe processing is being performed is indicated by the FS flag. When fail-safe processing is being performed, the failure identification unit 303 changes the FS flag from "1" to "0" in S643. Note that when fail-safe processing is not being performed, S643 is skipped. After that, in S644, the failure identification unit 303 transmits information indicating that the abnormal state of the system under verification has been resolved to the administrator terminal 320.
[0199] In this way, the failure identification unit 303 constantly or periodically monitors the state of the system under verification based on the log / metric information received from each failure verification agent. When failure verification information with a match rate higher than the threshold with the received log / metric information is detected, the failure identification unit 303 presumes that a failure or its sign has occurred and notifies the administrator of the failure factor corresponding to the detected failure verification information. Therefore, the administrator can quickly recognize the cause of the failure or its sign, and the burden of the recovery work is reduced. At this time, log / metric information related to components that cannot communicate directly with the cloud via the network (for example, a processor core in which an embedded OS is implemented within a multi-core device) is transmitted to the failure verification unit 302 and / or the failure identification unit 303 using a failure verification agent on an OS that can communicate directly with the cloud via the network. Therefore, it is also possible to appropriately presume a failure or its sign of a component that cannot communicate directly with the cloud via the network (in FIG. 11, M core 232 or application A3 operating on M core 232).
[0200] (3) Fail-Safe Phase FIG. 18 is a flowchart showing an example of the processing of the fail-safe control unit 304 in the fail-safe phase. The fail-safe control unit 304 controls the fail-safe process according to the value of the FS flag. The FS flag is set according to the procedure shown in FIG. 16.
[0201] In S651, the fail-safe control unit 304 monitors the value of the FS flag. If the value of the FS flag is "0", the fail-safe control unit 304 determines that there is no need to perform the fail-safe process, and continues the process of monitoring the value of the FS flag. On the other hand, if the value of the FS flag is "1", the fail-safe control unit 304 determines that it is necessary to perform the fail-safe process, and in S652, determines whether it is possible to perform the fail-safe process on the application in which the FS flag is set to "1" in S635 or S640 shown in FIG. 16. Whether the fail-safe process can be performed on the application is represented by the application management information shown in FIG. 14B. When the fail-safe process cannot be performed on the application, the fail-safe control unit 304 notifies the administrator in S658 that the fail-safe process cannot be performed on the application.
[0202] When the fail-safe process can be performed on the application, the fail-safe control unit 304 requests the fail-safe verification agent corresponding to the application to perform the fail-safe process in S653. The fail-safe verification agent corresponding to the application means a fail-safe verification agent that operates on a general-purpose OS and is implemented in the multi-core device that executes the application. For example, in the case of performing the fail-safe process on the application A3 shown in FIG. 11, it corresponds to the general-purpose OS fail-safe verification agent 311 that operates on the general-purpose OS implemented in the A core 231 of the multi-core device 230. The fail-safe verification agent that has received the request for execution performs the fail-safe process and notifies the fail-safe control unit 304 to that effect, as will be described later.
[0203] When the fail-safe process for the application is executed, the fail-safe control unit 304 transmits, in S654, information indicating the application for which the fail-safe process is being executed to the administrator terminal 320. At this time, it is preferable that the fail-safe control unit 304 also transmits information indicating the execution period of the fail-safe process to the administrator terminal 320. The execution period of the fail-safe process is set as the application management information shown in FIG. 14B. As a result, the administrator can recognize which application the fail-safe process is being executed for and the maximum period during which the fail-safe process can be executed.
[0204] When the fail-safe process is being executed, the fail-safe control unit 304 monitors the value of the FS flag in S655. If the value of the FS flag is "1", the fail-safe control unit 304 determines that the fail-safe state continues and continues the process of monitoring the value of the FS flag. On the other hand, if the value of the FS flag is "0", the fail-safe control unit 304 determines that it is necessary to return from the fail-safe state to the normal state, and in S656, requests the failure verification agent corresponding to the application to stop the fail-safe process. The failure verification agent that has received the stop request returns the application from the fail-safe state to the normal state and notifies the fail-safe control unit 304 to that effect, as will be described later.
[0205] When the application returns from the fail-safe state to the normal state, the fail-safe control unit 304 transmits, in S657, information indicating that the application has returned to the normal state to the administrator terminal 320. As a result, the administrator can recognize that the application that was operating in the fail-safe state has returned to the normal state.
[0206] FIG. 19 is a flowchart showing an example of the process of the failure verification agent in the fail-safe phase. The failure verification agent executes the fail-safe process in response to a request received from the fail-safe control unit 304.
[0207] In S661, the general-purpose OS failure verification agent 311 waits for a request to execute fail-safe processing sent from the fail-safe control unit 304. This execution request is sent in S653 shown in FIG. 18.
[0208] Upon receiving the request to execute fail-safe processing, the general-purpose OS failure verification agent 311, in S662, acquires the operating state information of the specified application (hereinafter referred to as the target application) from the embedded application information storage area 233b and sets it for the target application operating on the general-purpose OS. Each failure verification agent (311, 312) implemented in the multi-core device 230 detects the operating state of the application implemented on the same OS as itself when the application is in the execution state and writes it to the embedded application information storage area 233b. For example, in the multi-core device 230 shown in FIG. 11, when application A3 is running on the embedded OS implemented in M core 232, the embedded OS failure verification agent 312 detects the operating state of application A3 and writes it to the embedded application information storage area 233b. When application A3 is running on the general-purpose OS implemented in A core 231, the general-purpose OS failure verification agent 311 detects the operating state of application A3 and writes it to the embedded application information storage area 233b.
[0209] In S663, the general-purpose OS failure verification agent 311 changes the access right of the peripheral device controlled by the target application to the general-purpose OS. Then, in S664, the general-purpose OS failure verification agent 311 starts the target application on the general-purpose OS. At this time, the general-purpose OS failure verification agent 311 may instruct the embedded OS failure verification agent 312 to stop the target application that was operating on the embedded OS. After that, the general-purpose OS failure verification agent 311 notifies the fail-safe control unit 304 that the target application has shifted to the fail-safe state.
[0210] In the example shown in FIG. 11, application A3 is operating on the embedded OS in M core 232. Here, it is assumed that the general-purpose OS failure verification agent 311 receives a request from the fail-safe control unit 304 to perform fail-safe processing on application A3. In this case, the general-purpose OS failure verification agent 311 acquires the operating state information of application A3 from the embedded application information storage area 233b and sets it in application A3 implemented in A core 231. Also, the general-purpose OS failure verification agent 311 changes the access right to the peripheral device 260 to the general-purpose OS. Then, the general-purpose OS failure verification agent 311 starts application A3 implemented in A core 231. Thereby, the fail-safe state shown in FIG. 20 is configured. That is, application A3 implemented in A core 231 controls the peripheral device 260.
[0211] When the fail-safe processing is being performed, the general-purpose OS failure verification agent 311 waits for a request to stop the fail-safe processing at S665. This stop request is transmitted at S656 shown in FIG. 18.
[0212] When the general-purpose OS failure verification agent 311 receives a request to stop the fail-safe processing, at S666, the embedded OS failure verification agent 312 acquires the operating state information about the target application from the embedded application information storage area 233b and sets it in the target application (i.e., application A3) operating on the embedded OS.
[0213] At S667, the general-purpose OS failure verification agent 311 changes the access right to the peripheral device controlled by the target application to the embedded OS. Then, at S668, the embedded OS failure verification agent 312 starts the target application on the embedded OS. At this time, the general-purpose OS failure verification agent 311 stops the target application that was operating on the general-purpose OS. After that, the general-purpose OS failure verification agent 311 notifies the fail-safe control unit 304 that it has returned from the fail-safe state to the normal state.
[0214] Thus, in a multi-core device on which a general-purpose OS and an embedded OS are implemented, when an app running on the embedded OS is presumed to be the cause of a failure or a sign of a failure, the general-purpose OS failure verification agent and the embedded OS failure verification agent cooperate to perform fail-safe processing, enabling the app to run on the general-purpose OS. This makes it possible to continue the operation of the app.
[0215] As described above in detail regarding the disclosed embodiments and their advantages, those skilled in the art will be able to make various changes, additions, and omissions without departing from the scope of the invention clearly described in the claims.
Explanation of Signs
[0216] 10 Cloud 11 Cloud Resource 12 Cloud App 20, 20a, 20b Device 21 Sub-device 22 Device App 30 Public Line 41 Information Storage Function 42 Failure Verification Function 43 Failure Identification Function 44a, 44b Failure Verification Agent (Agent) 50 Administrator Terminal 60 Information Processing Device 61 CPU 62 Memory 63 Input Device 64 Output Device 65 Auxiliary Storage Device 66 Communication I / F 67 Internal Bus 100 Failure Verification System 110 Collection Unit 111 First Collection Intermediary Unit 112 Second Collection Intermediary Unit 113 Reception Unit 120 Granting Unit 121 First Granting Intermediary Unit 122 Second Grant Intermediary Section 123 Grant Instruction Section 130 Storage Section 140 Detection Section 150 Specific Information Output Section 160 Monitoring Section 170 Collection Control Section 180 Cause Identification Section 181 Calculation Section 182 Cause Information Output Section 230 Multicore Device 302 Fault Verification Section 303 Fault Identification Section 304 Fail-Safe Control Section 311 General-Purpose OS Fault Verification Agent 312 Embedded OS Fault Verification Agent
Claims
1. A target system composed of a cloud service provided by cloud computing and a device that exchanges data with the cloud service, the target system comprising: a cloud application that provides the cloud service, cloud resources that are the hardware for running the cloud application, a device application that provides the function of data exchange in the device, and device resources that are the hardware for running the device application, and a failure verification system for verifying a failure occurring in the target system, a collection unit that collects execution logs of the cloud application and the device application as log information, and collects metric logs regarding each of the cloud resources and the device resources as metric information, an application unit that assigns a cause of a failure to any of the components of the target system to cause a failure in the target system, a storage unit that associates and stores the log information and the metric information when the cause is assigned to the component as failure verification information, characterized by comprising the above.
2. a detection unit that detects the occurrence of an abnormality in the components of the target system, a specific information output unit that outputs information for specifying other components when the occurrence of an abnormality in other components excluding the component to which the cause has been assigned in the target system is detected in response to the assignment of the cause, The failure verification system according to claim 1, further comprising the above.
3. a monitoring unit that monitors the target system, a collection control unit that controls the collection unit when a failure in the target system is detected by the monitoring, and causes the collection unit to collect the log information and the metric information at the time point when the detected failure occurs, a cause identification unit that identifies the cause of the failure detected by the monitoring from the log information and the metric information at the time point when the failure occurs, using the failure verification information, The failure verification system according to claim 1, further comprising the above.
4. The cause identification unit a calculation unit that calculates a matching rate between the log information and the metric information for each cause in the failure verification information for each cause and the log information and the metric information at the time point when the failure occurs, A cause information output unit that outputs identification information of a predetermined number of the factors in descending order of the matching rate as information representing the cause of the failure detected by the monitoring; The failure verification system according to claim 3, characterized by comprising the above.
5. The imparting unit is A first imparting intermediary unit arranged in the cloud resource, which imparts the factor to the cloud resource or the cloud application; A second imparting intermediary unit arranged in the device resource, which imparts the factor to the device resource or the device application; An imparting instruction unit that gives an instruction to the first imparting intermediary unit or the second imparting intermediary unit according to a failure occurrence command including the setting of the factor and the component, and causes the factor set in the failure occurrence command to be imparted to the component set in the failure occurrence command; The failure verification system according to claim 1, characterized by comprising the above.
6. The collection unit is A first collection intermediary unit arranged in the cloud resource, which collects the execution log of the cloud application and the metric log of the cloud resource; A second collection intermediary unit arranged in the device resource, which collects the execution log of the device application and the metric log of the device resource; A receiving unit that receives the execution log and the metric log collected by the first collection intermediary unit from the first collection intermediary unit, and receives the execution log and the metric log collected by the second collection intermediary unit from the second collection intermediary unit; The failure verification system according to claim 1 or 5, characterized by comprising the above.
7. The target system is composed of the cloud service, the device, and a sub-device that exchanges data with the cloud service; The second imparting intermediary unit arranged in the device resource further imparts the factor to the device application executed by the sub-device or the hardware resource that executes the device application by the sub-device according to the setting of the component in the failure occurrence command; The failure verification system according to claim 5, characterized by the above.
8. There are a plurality of devices constituting the target system. The granting instruction unit checks for the presence or absence of an abnormality in the second granting intermediary unit provided in the first device among the plurality of devices, in the failure occurrence command, when, as the setting of the component, the device application executed by the sub-device or the hardware resource for executing the device application in the sub-device is set, when it is confirmed that there is no such abnormality, an instruction is given to the second granting intermediary unit provided in the first device to cause the factor set in the failure occurrence command to be granted to the component for the sub-device set in the failure occurrence command, when it is confirmed that there is such an abnormality, an instruction is given to the second granting intermediary unit provided in a second device different from the first device among the plurality of devices to cause the factor set in the failure occurrence command to be granted to the component for the sub-device set in the failure occurrence command The failure verification system according to claim 7, characterized in that.
9. The target system is composed of the cloud service, the device, and a sub-device that processes data transfer between the cloud service and the device, The second collection intermediary unit further collects an execution log of the device application executed by the sub-device and a metrics log of the hardware resource for executing the device application in the sub-device, The receiving unit further receives the execution log and the metrics log of the sub-device collected by the second collection intermediary unit from the second collection intermediary unit, The failure verification system according to claim 6, characterized in that.
10. A failure verification method performed by a failure verification system for verifying a failure occurring in a target system including a cloud service provided by cloud computing and a device that exchanges data with the cloud service, the target system including, as components, a cloud application that provides the cloud service, cloud resources that are the hardware for executing the cloud application, a device application that provides the data exchange function in the device, and device resources that are the hardware for executing the device application Collect the execution logs of the cloud application and the device application as log information, Collect the metric logs for each of the cloud resources and the device resources as metric information, Impose a cause of failure on any of the components of the target system to cause a failure in the target system, Link the log information and the metric information when the cause is imposed on the component, and save them as failure verification information for each cause, A failure verification method characterized by the above.
11. A failure verification system for verifying a failure of a target system including a multi-core device having a first processor core on which a first OS is implemented and a second processor core on which a second OS is implemented, A failure verification unit that creates failure verification information for verifying a failure of the target system, A failure identification unit that identifies the cause of a failure or a sign of a failure occurring in the target system using the failure verification information, A first agent that is implemented on the first processor core and collects first log information representing the execution log of an application operating within the first processor core and first metric information representing the state of the hardware of the first processor core, A second agent that is implemented on the second processor core and collects second log information representing the execution log of an application operating within the second processor core and second metric information representing the state of the hardware of the second processor core, The failure verification unit, for each of a plurality of preset failure factors, Receives the first log information and the first metric information from the first agent when the failure factor is injected into the target system, and receives the second log information and the second metric information from the second agent via the first agent when the failure factor is injected into the target system, Saves the first log information, the first metric information, the second log information, and the second metric information as the failure verification information when the failure factor is injected into the target system, The failure identification unit, By comparing at least a part of the monitoring information including the first log information and the first metrics information received from the first agent during the actual operation of the target system, and the second log information and the second metrics information received from the second agent via the first agent, with the failure verification information stored for each of the plurality of failure factors, identify the cause of the failure or the sign of a failure occurring in the target system A failure verification system characterized by the above.
12. When the matching rate between the monitoring information and the failure verification information stored for the first failure factor among the plurality of failure factors exceeds a predetermined threshold, the failure identification unit transmits information representing the first failure factor to the administrator terminal The failure verification system according to claim 11, characterized by the above.
13. When the corresponding application corresponding to the first failure factor is operating on the second OS within the second processor core, further comprising a fail-safe control unit that transmits a request to perform fail-safe processing to the first agent The first agent operates the corresponding application on the first OS within the first processor core in response to the implementation request The failure verification system according to claim 12, characterized by the above.
14. The failure identification unit extracts a predetermined number of failure factors from the plurality of failure factors in order from the one with the highest matching rate between the monitoring information and the failure verification information, and transmits information representing the extracted predetermined number of failure factors to the administrator terminal The failure verification system according to claim 11, characterized by the above.
15. When the corresponding application corresponding to any of the predetermined number of failure factors extracted by the failure identification unit is operating on the second OS within the second processor core, further comprising a fail-safe control unit that transmits a request to perform fail-safe processing to the first agent The first agent operates the corresponding application on the first OS within the first processor core in response to the implementation request The failure verification system according to claim 14, characterized by the above.
Citation Information
Patent Citations
Device, method and program for generating simulated fault, and test system
JP2011123783A
Failure contact efficiency system
JP2013222313A
Failure factor estimating device and failure factor estimating method
JP2021128538A