Design support system and design support method

The design support system addresses complex failures in distributed information systems by using fault verification agents to monitor and modify resources, ensuring continuous service provision.

JP2026018108APending Publication Date: 2026-02-05FUJI ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024119182
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

In information systems with distributed components across cloud, on-premise, and edge devices, complex interactions lead to unexpected failures due to different infrastructures, making it difficult to verify and address system failures effectively.

Method used

A design support system that includes fault verification agents to monitor and identify failure causes, modify cloud resources, and output requirements to maintain service provision, using a configuration of cloud services, devices, and fault verification functions.

Benefits of technology

Supports the design of systems to maintain service provision even when failures occur by identifying and addressing potential failure causes, reducing the risk of system downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026018108000001_ABST
    Figure 2026018108000001_ABST
Patent Text Reader

Abstract

To support design work of a system aiming at maintenance of provision of service even when a failure occurs.SOLUTION: In a design support system, an imparting part 120 imparts a factor that may cause a failure in a target system composed of a cloud service provided by cloud computing and a device for exchanging data with the cloud service to any of components of the target system. The detection unit 140 detects the occurrence of a failure in the target system. The changing unit 160 changes a cloud resource for executing a cloud application that provides a cloud service. The requirement identification section 170 identifies, on the basis of a result of detection by the detection section in response to application of a factor by the application section each time when the changing section is caused to change a cloud resource by a plural number of times, a requirement for a cloud resource with which a failure does not occur even if the factor is applied. The output unit 180 outputs information indicating the requirement specified by the requirement specifying unit.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for supporting the design work of an information system. [Background technology]

[0002] Several technologies are known for improving the ability to deal with failures that occur in information systems (see, for example, Patent Documents 1 and 2). For example, there is known a technology that supports a failure analyst in efficiently identifying possible failure locations when a failure occurs in a network system. There is also known a control cloud server technology that enables information obtained within a control system to be utilized beyond a single control system. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-36809 [Patent Document 2] Japanese Patent Publication No. 2023-45180 Summary of the Invention [Problem to be solved by the invention]

[0004] In systems where components are distributed across a variety of environments, such as the cloud, on-premise and edge devices, each component operates on a different infrastructure and interacts in a complex manner, which can lead to unexpected system failures during actual operation. [Means for solving the problem]

[0005] A design support system according to one embodiment supports the design of a target system including a cloud service provided by cloud computing and a device that exchanges data with the cloud service. The target system includes, as its components, a cloud app, a cloud resource, a device app, and a device resource. The cloud app is an app that provides a cloud service, and the cloud resource is hardware that executes the cloud app. The device app is an app that provides a function for exchanging data on the device, and the device resource is hardware that executes the device app. The design support system includes an assigning unit, a detecting unit, a modifying unit, a requirement identifying unit, and an outputting unit. The assigning unit assigns a factor that may cause a failure in the target system to any of the components. The detecting unit detects the occurrence of a failure. The modifying unit modifies the cloud resource. The requirement identifying unit identifies cloud resource requirements that will not cause a failure even when the factor is assigned, based on the results of the detection unit's detection of the factor assigned by the assigning unit each time the modifying unit modifies the cloud resource multiple times. The outputting unit outputs information indicating the requirements identified by the requirement identifying unit. [Effects of the Invention]

[0006] According to the above aspect, it is possible to support the design work of a system that aims to maintain the provision of services even when a failure occurs. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating an overview of a design support system. [Figure 2] FIG. 1 illustrates an example of the configuration of a design support system. [Figure 3] FIG. 2 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 4] FIG. 1 illustrates an example of a target system. [Figure 5] 10 is a flowchart illustrating an example of processing content of a fault verification agent. [Figure 6] 10 is a flowchart showing an example of processing contents of information storage processing. [Figure 7] FIG. 10 is a diagram illustrating an example of data stored by the information storage process. [Figure 8] 10 is a flowchart illustrating an example of the processing content of a fault verification process. [Figure 9] FIG. 10 is a diagram illustrating an example of data in a fault verification list table. [Figure 10] 10 is a flowchart illustrating an example of processing content of a design proposal process. [Figure 11] 10 is a flowchart illustrating an example of a process content of a condition specifying process. [Figure 12] FIG. 10 is a diagram illustrating an example of data in a detailed information list table. [Figure 13] 10 is a flowchart illustrating an example of a process content of a requirement identification process. [Figure 14] FIG. 10 is a diagram illustrating an example of fail-safe rate data. DETAILED DESCRIPTION OF THE INVENTION

[0008] Cloud services provided by cloud computing are increasingly being used as the system infrastructure for information systems. In addition, the infrastructure environment ("infrastructure" is an abbreviation for "infrastructure") on which information systems are built is expanding and becoming decentralized, no longer confined to a single data center. Examples of such infrastructure environments include hybrid clouds, which keep important mission-critical data and applications that are difficult to update on-premises and link them with cloud services, and edge computing, which processes data in real time on-site. In information systems where many services are linked together as part of a system architecture, it is becoming difficult to track each individual traffic.

[0009] Furthermore, in such information systems, many applications run on infrastructure managed by cloud vendors and are interconnected in complex ways. For this reason, existing system testing may not be able to cover all abnormalities (such as abnormalities in the network or processing unit load). Furthermore, because the multiple cloud platforms and devices that make up an information system are distributed and run on different infrastructures, interconnecting in complex ways, an unexpected failure in one application may affect other applications.

[0010] As such, it is becoming difficult to verify such information system failures, confirm the situation, and consider countermeasures in advance using existing system tests.

[0011] In the following embodiment, a design support system is described that supports the design of a target system that is configured with a cloud service provided by cloud computing and a device that exchanges data with the cloud service. This design support system provides information for preventing failures from occurring, derived from the status of failures that occur due to factors such as excessive access to components of the target system or forced shutdown of components.

[0012] FIG. 1 is a diagram for explaining an outline of a design support system 100, and FIG. 2 is a diagram showing an example of the configuration of the design support system 100.

[0013] 1, the cloud 10, the devices 20a and 20b, and the sub-device 21 are connected via a public line 30. The devices 20a, 20b, and the sub-device 21 are connected to each other by a local line (wired, wireless, etc.) not shown. Of these, the sub-device 21 may be connected to the public line 30 or may be connected only to the local line, depending on the role of the application installed in the sub-device 21.

[0014] Cloud 10 is an environment that provides cloud computing.

[0015] The cloud resources 11 are hardware resources that provide cloud computing in the cloud 10 .

[0016] Cloud applications 12 are application software executed on cloud resources 11 and provide cloud services used by the target system. Although two cloud applications 12 are shown in FIG. 1, the number of cloud applications 12 does not have to be the number shown in FIG. 1.

[0017] The devices 20a and 20b and the sub-device 21 all exchange data with a cloud service provided by at least one cloud application 12, and are, for example, IoT devices. Note that "IoT" is an abbreviation for Internet of Things. While two devices 20a and 20b and one sub-device 21 are shown in FIG. 1, the number of these devices does not have to be the number shown in FIG. 1.

[0018] In the following description, when there is no need to distinguish between the devices 20a and 20b, they may be collectively referred to as "device 20."

[0019] The device application 22 is application software that runs on the hardware of each of the device 20 and the sub-device 21, and provides a function for exchanging data with the cloud service provided by the cloud application 12.

[0020] 1, a target system that is the target of design work support by design support system 100 includes, as its components, cloud resource 11, devices 20a and 20b, a sub-device 21, a cloud application 12, and a device application 22. Design support system 100, which verifies faults in the target system with this configuration, provides the functions of information storage function 41, fault verification function 42, fault identification function 43, and design proposal function 44, and uses fault verification agents 45a and 45b as necessary to provide each function.

[0021] The fault verification agents 45a and 45b are software that are deployed in the cloud resource 11 and the device 20, respectively.

[0022] The fault verification agent 45 a performs processes such as acquiring metrics of the cloud resources 11 , acquiring execution logs of the cloud applications 12 , and acting as an intermediary between various processes that each function of the design support system 100 performs on the cloud resources 11 .

[0023] The fault verification agent 45b performs processes such as obtaining metrics of the device 20, obtaining an execution log of data transfer processing by the device application 22, and mediating various processes that each function of the design support system 100 performs on the device 20. The fault verification agent 45b also performs these processes on the sub-device 21 that executes the device application 22.

[0024] In the following description, the fault verification agent 45a and the fault verification agent 45b may be simply referred to as "agent 45a" and "agent 45b," respectively.

[0025] It is assumed that the sub-device 21 has fewer hardware resources than the device 20 and lacks the hardware resources necessary to execute the agent 45b. In this embodiment, for such a sub-device 21, the agent 45b executed on the device 20 performs the above-mentioned processing on behalf of the sub-device 21.

[0026] 1, the agent 45b arranged in the device 20b performs the above-described processing for the sub-device 21. Note that the agent 45b may also perform processing for a plurality of sub-devices 21.

[0027] The information storage function 41 is a function that stores and manages the execution logs and metrics acquired by the agents 45a and 45b.

[0028] The fault verification function 42 is a function for verifying faults in the target system in advance, assigning a factor that will cause a fault to any of the components of the target system, and monitoring each component of the target system for the occurrence of a fault caused by the assignment of this factor. The results of this monitoring are also stored and managed by the information storage function 41.

[0029] The fault identification function 43 is a function that uses the results of advance verification by the fault verification function 42, which are stored in the information storage function 41, to identify the cause of a fault that occurs during actual operation of the target system, and outputs the identified cause to the administrator terminal 50 to notify the administrator of the target system. Note that a detailed description of the fault identification function 43 will be omitted from the following explanation.

[0030] The design proposal function 44 monitors the target system by having the fault verification function 42 apply a factor that will cause a failure multiple times while changing the conditions for the application. Based on the results of this monitoring, the design proposal function 44 identifies the conditions for the application that will become critical and cause a failure in the target system. The design proposal function 44 also monitors the target system by having the fault verification function 42 apply the factor to the target system multiple times while changing the cloud resources 11 under conditions that satisfy the conditions. Based on the results of this monitoring, the design proposal function 44 identifies requirements for the cloud resources 11 that will not cause a failure even when the factor is applied and satisfies the conditions, and outputs the requirements to the administrator terminal 50 as design support information for the target system.

[0031] Next, the configuration of the design support system 100 that provides these functions will be described.

[0032] In the configuration example of FIG. 2, the design support system 100 includes a collection unit 110, an assignment unit 120, a storage unit 130, a detection unit 140, a condition specification unit 150, a change unit 160, a requirement specification unit 170, and an output unit 180.

[0033] 1 is provided by a collection unit 110 and a saving unit 130. Also, a fault verification function 42 in Fig. 1 is provided by an assignment unit 120. Furthermore, a design proposal function 44 in Fig. 1 is provided by a detection unit 140, a condition specification unit 150, a change unit 160, a requirement specification unit 170, and an output unit 180.

[0034] 1 provides the first collection intermediation unit 111 of the collection unit 110 and the first assignment intermediation unit 121 of the assignment unit 120. Also, the agent 45b in FIG. 1 provides the second collection intermediation unit 112 of the collection unit 110 and the second assignment intermediation unit 122 of the assignment unit 120 in FIG.

[0035] The collection unit 110 collects log information including an execution log for cloud services provided by cloud computing and an execution log for processes executed in each of the device 20 and the sub-device 21. The collection unit 110 also collects metrics information, which is a log of metrics related to hardware as a cloud infrastructure used to provide the cloud services and the hardware constituting each of the device 20 and the sub-device 21.

[0036] The first collection intermediation unit 111 included in the collection unit 110 is provided on a platform that provides cloud services and collects the execution logs and metrics logs described above for the cloud services. The agent 45a in FIG. 1 that provides the first collection intermediation unit 111 is located on the cloud resources 11 and collects execution logs and metrics logs for the cloud resources 11. In the example of FIG. 1, the agent 45a collects metrics logs for the processing unit, memory, storage device, and interface device provided for executing the cloud applications 12 as metrics logs for the cloud resources 11. More specifically, the agent 45a collects usage rate logs for the processing unit and interface device, and usage amount logs for the memory. Furthermore, the agent 45a collects free space percentages and access frequencies (number of accesses per unit time) for the storage devices.

[0037] The second collection intermediation unit 112 included in the collection unit 110 is provided in the device 20 and collects the above-described execution logs for the device 20 and the above-described metrics logs for the device 20. The agent 45b in FIG. 1 that provides the second collection intermediation unit 112 is located on the hardware resources (device resources) of the device 20. This agent 45b collects execution logs for the device apps 22 and collects metrics logs for the device resources. Note that the agent 45b that provides the second collection intermediation unit 112 in the device 20b also acts as an agent for collecting execution logs for the device apps 22 executed in the sub-device 21 and collecting metrics logs for the device resources of the sub-device 21. In the example of FIG. 1, as metrics logs for the device 20 and the sub-device 21, metrics logs for the arithmetic processing unit, memory, storage device, and interface device that the device 20 and the sub-device 21 respectively include are collected. More specifically, the agent 45b collects logs of the utilization rate of the arithmetic processing unit and the interface unit, logs of the memory usage, and logs of the free capacity ratio and the access frequency (number of accesses per unit time) of the storage device.

[0038] The receiving unit 113 included in the collecting unit 110 receives the execution logs and metrics logs for the cloud services collected by the first collection intermediation unit 111 from the first collection intermediation unit 111. The receiving unit 113 also receives the execution logs and metrics logs for the devices 20 or sub-devices 21 collected by the second collection intermediation unit 112 from the second collection intermediation unit 112.

[0039] The assigning unit 120 assigns a cause that may cause a failure in the target system (a failure occurrence cause) to any of the components of the target system.

[0040] The first grant mediation unit 121 included in the grant unit 120 is disposed in the cloud resource 11. The agent 45a in FIG.

[0041] The second assignment intermediation unit 122 included in the assignment unit 120 is disposed in the device resources of the device 20. The agent 45b in FIG. 1 that provides the second assignment intermediation unit 122 assigns a failure cause to the device resources of the device 20 or the device application 22 executed by the device 20. Note that the agent 45b that provides the second assignment intermediation unit 122 in the device 20b may also assign a failure cause to the device resources of the sub-device 21 or the device application 22 executed by the sub-device 21 in certain cases.

[0042] The assignment instruction unit 123 included in the assignment unit 120 creates a fault occurrence command including settings for the fault occurrence cause and the components of the target system, and sends it to the first assignment mediation unit 121 or the second assignment mediation unit 122, instructing the component to be assigned the fault occurrence cause.

[0043] In some cases, the failure occurrence command sets the device resources of the sub-device 21 as a configuration element setting. In this case, the assignment instruction unit 123 instructs the second assignment intermediation unit 122 included in the device 20b to assign the failure cause set in the failure occurrence command to the device resources of the sub-device 21. In other cases, the failure occurrence command sets the device app 22 to be executed by the sub-device 21 as a configuration element setting. In this case, the assignment instruction unit 123 instructs the second assignment intermediation unit 122 included in the device 20b to assign the failure cause set in the failure occurrence command to the device app 22. In this way, the second assignment intermediation unit 122 included in the device 20b assigns the failure cause not only to the device 20b itself but also to the sub-device 21, thereby making it possible to assign the failure cause to a sub-device 21 that does not include the second assignment intermediation unit 122.

[0044] The assignment instruction unit 123 may be configured to check whether or not there is an abnormality in the second assignment intermediation unit 122 included in the device 20b. When the failure occurrence command includes, as the configuration of a component, a device app 22 executed by the sub-device 21 or a device resource that executes the device app 22 in the sub-device 21, the assignment instruction unit 123 checks whether or not there is an abnormality. If the assignment instruction unit 123 confirms that there is no abnormality, it instructs the second assignment intermediation unit 122 included in the device 20b to assign the failure occurrence cause set in the failure occurrence command to the component of the sub-device 21. On the other hand, if the assignment instruction unit 123 confirms that there is an abnormality, it instructs the second assignment intermediation unit 122 included in, for example, the device 20a (an example of a second device) instead of the device 20b to assign the cause to the component of the sub-device 21. By doing this, even if an abnormality occurs in the second assignment intermediation unit 122 (agent 45b) provided in device 20b, the second assignment intermediation unit 122 (agent 45b) provided in device 20a can assign the cause of the failure to the components of sub-device 21.

[0045] The storage unit 130 associates the log information and metrics information collected by the collection unit 110 with the cause of a failure that the assignment unit 120 assigns to the component of the target system, and stores the information as failure verification information.

[0046] 1, the information storage function 41 serving as the storage unit 130 stores, as fault verification information, the execution log and the metrics log received by the information storage function 41 serving as the receiving unit 113. Note that this fault verification information is linked to the time when the fault occurrence cause instructed by the fault verification function 42 serving as the assignment instruction unit 123 was assigned to the component.

[0047] The detection unit 140 detects the occurrence of an abnormality in a component of the target system as the occurrence of a failure in the target system. In Fig. 1, the design proposal function 44 as the detection unit 140 detects the occurrence of an abnormality in each of the cloud resource 11 that executes the cloud application 12, and the device 20 and sub-device 21 that execute the device application 22.

[0048] The condition specification unit 150 specifies the conditions for the assignment of a fault occurrence factor to the target system that will become critical and cause a fault to occur in the target system as a result of the assignment by the assignment unit 120. This condition specification is performed based on the results of detection of the occurrence of a fault in the target system by the detection unit 140 in response to the assignment of a fault occurrence factor each time the assignment unit 120 assigns a fault occurrence factor multiple times while changing the assignment conditions.

[0049] 1, the design proposal function 44 as the condition identification unit 150 sends instructions multiple times to the fault verification function 42 as the assignment unit 120 to assign a fault cause to the components of the target system. The design proposal function 44 changes the conditions for assigning the fault cause in the instructions each time it sends the instructions.

[0050] For example, if a failure cause is a predetermined amount of access load on cloud application 12, the condition for assigning the failure cause is the access volume in the access load. In this case, design proposal function 44 sends instructions multiple times to fault verification function 42 to assign an access load to cloud application 12 multiple times while changing the access volume. At this time, design proposal function 44 identifies the critical access volume at which a failure occurs in the target system based on the state of detection by detection unit 140 of the occurrence of a failure in the target system in response to the assigned access load for each of the multiple times.

[0051] Furthermore, for example, if the cause of a failure is the random stopping of multiple cloud applications 12, the condition for assigning the cause of a failure is the selection of the cloud application 12 to be stopped. In this case, the design proposal function 44 sends a command to the fault verification function 42 multiple times, causing the fault verification function 42 to send an execution stop command multiple times while changing the cloud application 12 to which the command is sent. At this time, the design proposal function 44 identifies the cloud application 12 that will cause the target system to stop providing services, based on the state of detection by the detection unit 140 of the occurrence of a failure in the target system in response to the stopping of the cloud application 12 for each of the multiple times.

[0052] The change unit 160 changes the configuration of the target system, more specifically, changes the cloud resources 11 that are components of the target system.

[0053] The requirement identification unit 170 identifies requirements of the cloud resource 11 that will not cause a failure even when a failure causing factor is assigned to the target system by the assignment unit 120. This requirement is identified based on the result of detection of the occurrence of a failure in the target system by the detection unit 140 in response to the assignment of a failure causing factor by the assignment unit 120, each time the change unit 160 changes the cloud resource 11 multiple times.

[0054] 1 , for example, if the cause of a failure is a predetermined amount of access load on cloud application 12, design proposal function 44 as change unit 160 changes, for example, the processing capacity of cloud resource 11. For example, the processing capacity of cloud resource 11 is changed by changing or reducing the hardware assigned as cloud resource 11 that executes cloud application 12, which is a component of the target system. At this time, design proposal function 44 as requirement identification unit 170 identifies a processing capacity of cloud resource 11 that will not cause a failure even when the cause of a failure is assigned by fault verification function 42. This identification is performed based on the status of detection of the occurrence of a failure in the target system by detection unit 140 in response to the assignment of an access load by fault verification function 42, each time the processing capacity of cloud resource 11 is changed multiple times.

[0055] 1, for example, if the cause of a failure is the random stopping of multiple cloud applications 12, the design proposal function 44 serving as the change unit 160 changes the configuration of the target system by, for example, multiplexing the cloud resources 11. At this time, the design proposal function 44 serving as the requirement identification unit 170 selects a region in which the cloud resources 11 newly used for multiplexing are installed in the cloud service that receives the cloud resources 11. In this selection, a region is selected that will allow the target system to continue providing services even if the cloud applications 12 running on the cloud resources 11 before multiplexing are stopped. This selection is made based on the detection status of a failure in the target system in response to the region selection by the fault verification function 42 each time the selection of the region in which the cloud resources 11 are installed is changed multiple times.

[0056] In this embodiment, the failure cause that the requirement specifying unit 170 causes the assigning unit 120 to assign to the target system in order to specify the requirements of the cloud resource 11 is assumed to satisfy the critical condition that causes a failure in the target system specified by the condition specifying unit 150. By doing so, it is expected that the amount of change to the configuration of the target system to improve the availability of the target system can be reduced.

[0057] The output unit 180 outputs information indicating the requirements of the cloud resource 11 that will not cause a failure even when the assigning unit 120 assigns a failure cause to the target system, as identified by the requirement identifying unit 170, and sends the information to the administrator terminal 50. By operating the administrator terminal 50 and referring to this information, the administrator of the target system can easily design a target system that allows the provision of services to be maintained even when a failure occurs.

[0058] Next, the hardware configuration of the design support system 100 will be described.

[0059] 3 shows an example of the hardware configuration of the information processing device 60. All of the functions shown in FIG. 1 provided in the cloud 10 may be provided by using the information processing device 60 as a cloud server. The information processing device 60 may also be used as the cloud resource 11 on which the cloud application 12 and the agent 45a are executed. Furthermore, the information processing device 60 may also be used as the device 20 on which the device application 22 and the agent 45b are executed.

[0060] The information processing device 60 is a computer equipped with the following components: a CPU 61, a memory 62, an input device 63, an output device 64, an auxiliary storage device 65, and a communication I / F 66. All of these components are connected to an internal bus 67, and are configured to enable data exchange between the components. Note that "CPU" is an abbreviation for Central Processing Unit. Also, "I / F" is an abbreviation for Interface.

[0061] The CPU 61 controls each component of the information processing device 60 by, for example, using the memory 62 to execute a predetermined program.

[0062] The input device 63 is, for example, a keyboard or pointing device for inputting instructions, or various sensors.

[0063] The output device 64 is used, for example, to display and output various types of information.

[0064] The auxiliary storage device 65 is a non-volatile storage device, such as a flash memory.

[0065] The communication I / F 66 transmits and receives various data to and from other cloud servers or the device 20 in accordance with instructions sent from the CPU 61 .

[0066] When an information processing device 60 is used as each element shown in FIG. 1, the information processing device 60 does not need to include all of the elements shown in FIG. 3, and some elements may be omitted depending on the application or conditions.

[0067] Next, the processing contents of each of the agents 45a and 45b in FIG. 1, and the processing contents performed to provide the information storage function 41, the fault verification function 42, and the design proposal function 44 will be described.

[0068] In the following description, an example of design support for a target system shown in Fig. 4 will be described as a specific example of design support performed using the design support system 100. The correspondence between each element of the target system shown in Fig. 4 and each element shown in Fig. 1 will be described first.

[0069] The target system in Fig. 4 has four cloud resources 11 on a cloud 10 as its components. To identify these cloud resources 11, resource IDs "C001," "C002," "C003," and "C004" are assigned to each of them. "ID" is an abbreviation for "identifier."

[0070] Furthermore, four cloud applications 12, which are components of the target system, are executed on each of these cloud resources 11. These cloud applications 12 provide functions of data collection, data processing, data storage, and data display, respectively. These cloud applications 12 are assigned application IDs [S002], [S003], [S004], and [S005], respectively. User terminals 51 are terminals used by users of the target system.

[0071] On the other hand, the device 20 has a device resource 23 as a component of the target system. The device resource 23 is hardware of the device 20 and is assigned a resource ID of "D001." The device resource 23 executes a device application 22, which is a component of the target system. The device application 22 provides a data transmission function and is assigned an application ID of [S001].

[0072] 4 further shows a resource information table 70. The resource information table 70 shows information about specific resources held by the cloud resource 11 or device resource 23 identified by the resource ID. The resource information table 70 is assumed to be stored in advance in a storage device provided in a cloud server that executes the information storage process described below to provide the information storage function 41.

[0073] It is assumed that the resources held by the cloud resource 11 are changeable. The resource information table 70 also shows information on the range of change in the resources held by the cloud resource 11.

[0074] First, the processing contents of the agents 45a and 45b will be described with reference to Fig. 5. Fig. 5 is a flowchart showing an example of the processing contents of the agents 45a and 45b.

[0075] 5 starts, first, in S101, a process of receiving a normality confirmation request is performed, and then, in S102, a process of determining whether or not the confirmation request has been received is performed. The confirmation request received in the process of S101 is transmitted to the agents 45a and 45b by the fault verification function 42.

[0076] In the determination process of S102, if it is determined that a normality confirmation request has been received (if the determination result is YES), the process proceeds to S103, where a response indicating normality is returned to the fault verification function 42, and then the process proceeds to S104. On the other hand, in the determination process of S102, if it is determined that a normality confirmation request has not been received (if the determination result is NO), the process of S103 is skipped and the process proceeds to S104.

[0077] In S104, a process is performed to acquire and store the application execution log and the infrastructure metrics log.

[0078] By the processing of S104, in the case of agent 45a, an execution log of cloud application 12 and a log of metrics of cloud resource 11 are acquired and stored in a storage device of cloud resource 11. On the other hand, in the case of agent 45b, an execution log of device application 22 on device 20 and a log of hardware metrics of device 20 are acquired and stored in a storage device of device 20. Note that in the case of agent 45b located in device 20b, an execution log of device application 22 on sub-device 21 and a log of hardware metrics of sub-device 21 are also acquired and stored in a storage device of device 20b.

[0079] In S105, a process of receiving a fault occurrence command issued from the fault verification function 42 is performed, and in the following S106, a process of determining whether or not the fault occurrence command has been received is performed.

[0080] In the determination process of S106, when it is determined that a failure occurrence command has been received (when the determination result is YES), the process proceeds to S107. Then, in S107, a process is performed to determine whether the target to which the failure occurrence cause set in the received failure occurrence command is assigned is an application or infrastructure for which the device is responsible for collecting information.

[0081] As described above, the failure occurrence command includes settings for the failure cause and the component of the target system to which the cause is to be assigned. In the determination process of S107, in the case of agent 45a, the determination result is YES if the target to which the failure cause set in the failure occurrence command is to be assigned is the cloud resource 11 or the cloud application 12. On the other hand, in the case of agent 45b, the determination result is YES if the target to which the failure cause is to be assigned is the hardware of the device 20 or the device application 22 executed by the device 20. However, in the case of agent 45b located in device 20b, the determination result is also YES if the target to which the failure cause is to be assigned is the hardware of the sub-device 21 or the device application 22 executed by the sub-device 21.

[0082] If the determination result of the determination process of S107 is YES, the process proceeds to S108, where a process is performed to assign the failure cause set in the failure command to the target to which the failure cause set in the failure command is to be assigned, and then the process proceeds to S109. On the other hand, if it is determined that the target to which the failure cause is to be assigned is not an application or infrastructure for which the device is responsible for collecting information (if the determination result is NO), the process of S108 is skipped and the process proceeds to S109.

[0083] In S109, a process is performed in which an application execution log and an infrastructure metrics log are acquired and saved. This process is similar to the process in S104, and an execution log and a metrics log are acquired for each component after a failure cause has been assigned in response to a failure occurrence command.

[0084] In S110, all of the execution logs and metrics logs saved by the process of S104 or S109 are sent to the information storage function 41. Then, in the following S111, all of the saved execution logs and metrics logs are deleted from the storage device. This process of S111 is intended to relieve pressure on the storage area of ​​the storage device due to the saving of this data.

[0085] After the process of S110 is completed, or if it is determined in the determination process of S106 that a failure occurrence command has not been received (the determination result is NO), the process returns to S101. Then, the process of receiving a normality confirmation request and the process of acquiring and saving application execution logs and infrastructure metrics logs are continued.

[0086] The above processing is performed by the agents 45a and 45b.

[0087] Next, the information storage process performed in the cloud server to provide the information storage function 41 will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the processing contents of the information storage process.

[0088] When the process of FIG. 6 starts, first in S201, a process of receiving a fault occurrence command issued from the fault verification function 42 is performed, and then in S202, a process of determining whether or not the fault occurrence command has been received is performed.

[0089] In the determination process of S202, if it is determined that a failure occurrence command has been received (if the determination result is YES), the process proceeds to S203, where a process is performed to receive the execution log and metrics log. Then, in the following S204, a process is performed to determine whether the execution log and metrics log have been received.

[0090] The execution log and metrics log received by the process of S203 are sent from the cloud resource 11 and the device 20, respectively, by the execution of the process of S110 in FIG. 5 in the agents 45a and 45b, respectively.

[0091] If it is determined in the determination process of S204 that the execution log and metrics log have been received (if the determination result is YES), the process proceeds to S205. On the other hand, if it is determined in the determination process of S204 that the execution log and metrics log have not been received (if the determination result is NO), the process returns to S203, and the process of receiving the execution log and metrics log continues.

[0092] In S205, the received execution log and metrics log are stored in a storage device of the cloud server that executes the information storage process, linked to the time when a failure cause was assigned based on the failure command issued from the failure verification function 42. Note that information about when a failure cause was assigned is recorded in the execution log of each of the cloud resources 11 and the device 20 to which a failure cause was assigned.

[0093] In S206, a process is performed to determine whether or not execution logs and metrics logs have been received from all of the agents 45a and 45b by the process of S203. If it is determined in this determination process that execution logs and metrics logs have been received from all of the agents 45a and 45b (if the determination result is YES), the process proceeds to S207. On the other hand, if it is determined in this determination process that execution logs and metrics logs have not been received from all of the agents 45a and 45b (if the determination result is NO), the process returns to S203, and reception of execution logs and metrics logs continues.

[0094] In S207, a process is performed in which information about the failure occurrence command is stored in a storage device provided in the cloud server that executes the information storage process, linked to when a failure occurrence cause is assigned based on the failure occurrence command issued from the failure verification function 42. This process stores information about the failure occurrence cause and the component of the target system that assigns the cause, which is included in the failure occurrence command received in the process of S201.

[0095] After the above-mentioned process of S207 is completed, the process returns to S201, and the process of receiving the execution log and metrics log continues.

[0096] The above process is the information storage process. Figure 7 shows an example of data stored by this information storage process for the target system shown in Figure 4.

[0097] The application execution log 71 is an execution log for each application, including the cloud application 12 and the device application 22. In the application execution log 71, an "application ID" is information for identifying the application. An "infrastructure ID" is information for identifying the infrastructure on which the application identified by the "application ID" is executed, and indicates the resource ID of the cloud resource 11 or the device resource 23. A "time" is information indicating the time when the cause of the failure set in the failure command was assigned, and "log information" is the content of the execution log for the application identified by the "application ID." In the example of the application execution log 71 in FIG. 7, the content of the execution log for the cloud application 12 of "data storage" in FIG. 4, whose "application ID" is S004, indicates that an error "Error No. X" has occurred.

[0098] The cloud resource metrics log 72 is a log of metrics for each cloud resource 11. In the cloud resource metrics log 72, "resource ID" is information that identifies the cloud resource 11, and indicates the resource ID of the cloud resource 11. "Time" is information that indicates the time when the failure cause set in the failure occurrence command was assigned. Furthermore, "CPU," "memory," and "network" are metric values ​​for the CPU, memory, and interface device, respectively, of the cloud resource 11 identified by the "cloud resource ID." In the example of the cloud resource metrics log 72 in FIG. 7, utilization rates are used as metric values ​​for "CPU" and "network," and usage amounts are used as metric values ​​for "memory."

[0099] The device resource metrics log 73 is a log of metrics for each device 20 (or each sub-device 21). In the device resource metrics log 73, "device ID" is information for identifying the device 20 and sub-device 21, and indicates the resource ID of the device resource 23. "Time" is information indicating the time when the failure cause set in the failure occurrence command was applied. Furthermore, "CPU," "memory," and "network" are information indicating the utilization rates, which are metrics for the CPU, memory, and interface device, respectively, of the device 20 or sub-device 21 identified by the "device ID." In the example of the device resource metrics log 73 in FIG. 7, utilization rates are used as metric values ​​for "CPU" and "network," and usage amount is used as metric value for "memory."

[0100] The fault occurrence command list 74 is a list of information about fault occurrence commands. In the fault occurrence command list 74, "fault verification ID" is information that identifies the fault occurrence command, and "time" is information that indicates the time when the fault occurrence cause set in the fault occurrence command was assigned. Furthermore, "target" is information that indicates the component of the target system to which the fault occurrence cause set in the fault occurrence command is assigned. Furthermore, "fault occurrence cause" is information that indicates the fault occurrence cause set in the fault occurrence command that is assigned to the component of the target system.

[0101] Next, a fault verification process performed in the cloud server to provide the fault verification function 42 will be described with reference to Fig. 8. Fig. 8 is a flowchart showing an example of the fault verification process.

[0102] When the processing of FIG. 8 starts, first, in S301, a process is performed to check the normality of each of the agents 45a and 45b, and then, in S302, a process is performed to determine whether all of the agents 45a and 45b are functioning normally.

[0103] In the process of S301, first, a normality confirmation request is sent to all of the agents 45a and 45b. When the agents 45a and 45b receive this confirmation request through the process of S101 (FIG. 5) (the determination result of S102 is YES), they perform a process (S103) of returning a response indicating that the agents are normal. However, if the agents are not functioning normally, they cannot return the response. Therefore, the process of S301 also includes a process of receiving this response. In the subsequent determination process of S302, if the responses are received from all of the agents 45a and 45b within a predetermined time after the transmission of the normality confirmation request, the determination result is YES, and the process proceeds to S306. On the other hand, if the responses are not received from all of the agents 45a and 45b even after the predetermined time has passed, the determination result is NO, and the process proceeds to S303.

[0104] In S303, a process is performed in which information (for example, identification information) about the agent 45a or 45b that is found to be malfunctioning and operating abnormally through the processes of S301 and S302 is output to the administrator terminal 50.

[0105] In S304, a process is performed to determine whether the agent 45a or 45b found to be operating abnormally is the agent responsible for collecting information about the sub-device 21. In the example of Fig. 1, if an abnormality in the behavior of the agent 45b located in the device 20b is found, the determination result in S304 is YES, and the process proceeds to S305.

[0106] In S305, a process is performed to change the agent responsible for collecting information about the sub-device 21 to another agent 45b. In the example of Fig. 1, a process is performed to change the agent responsible for collecting information about the sub-device 21 from the agent 45b located in the device 20b to, for example, the agent 45b located in the device 20a. After the process of S305 is performed, the agent 45b of the device 20a also assigns a fault cause to the device resource of the sub-device 21 or the device application 22 on the sub-device 21 based on the fault occurrence command.

[0107] After the processing of S305, or when it is determined that the agent 45a or 45b whose operation is found to be abnormal in the judgment processing of S304 is not in charge of collecting information of the sub-device 21 (when the judgment result is NO), the processing proceeds to S306.

[0108] In S306, a process is performed to acquire the failure cause and the conditions for assigning the failure cause from the design proposal function 44. The conditions for assigning the failure cause include information indicating the components of the target system to which the failure cause is assigned and more detailed information about the failure cause.

[0109] In S307, a process is performed to create a failure occurrence command including the failure cause and the conditions for granting acquired in the process of S306, and in the following S308, a process is performed to send the created failure occurrence command to all of the agents 45a and 45b and the information storage function 41. The failure occurrence command sent by this process is received by the process of S105 (FIG. 5) in the agents 45a and 45b. In addition, the failure occurrence command sent by this process is received by the process of S201 in the information storage process (FIG. 6).

[0110] In S309, a process is performed to determine whether or not the saving of the execution log and metrics log received and saved in the information saving function 41 after the failure occurrence command is sent in the process of S308 has been completed. In this determination process, if it is determined that the saving has been completed (if the determination result is YES), the process returns to S306. On the other hand, in this determination process, If it is determined that the saving is not complete (if the determination result is NO), the determination process of S309 is repeated until the saving is complete.

[0111] The above processing is the fault verification processing.

[0112] Next, a design proposal process performed in the cloud server to provide the design proposal function 44 will be described.

[0113] First, we will explain the fault verification list table 80. Fig. 9 shows an example of data in the fault verification list table 80. This fault verification list table 80 is created in advance by an administrator who operates the administrator terminal 50, and is saved in the information storage function 41.

[0114] The “fault verification ID” is information for identifying a fault occurrence command created by the fault verification function 42 , and is the same as the information indicated as “fault verification ID” in the fault occurrence command list 74 .

[0115] The "time" is information indicating the time (scheduled time) when the cause of a failure will be assigned in accordance with the failure command identified by the "failure verification ID."

[0116] "Target" is information indicating the component of the target system to which a fault cause is assigned in accordance with the fault generation command identified by the "Fault Verification ID", and the component is indicated using the resource ID or application ID described above. Note that if all applications among the components of the target system are to be subject to the assignment of a fault cause, a special ID "Sxxx" is indicated as the "Target".

[0117] The "Failure Cause" is information indicating the failure cause assigned to the component of the target system in accordance with the failure command identified by the "Failure Verification ID." More detailed information about this failure cause is shown in the "Details of Failure Cause" section.

[0118] The "probable failure" is information indicating a failure that is probable to occur in the target system due to the assignment of a failure cause performed in accordance with a failure occurrence command identified by the "failure verification ID."

[0119] The "type of response to failure" is information indicating whether the design support system 100 will take the lead in considering countermeasures to prevent failures indicated by the "anticipated failures." Here, type "A" is assigned to a failure for which there are countermeasures that the design support system 100 itself can take the lead in considering. On the other hand, type "B" is assigned to a failure for which there are no countermeasures that the design support system 100 itself can take the lead in considering, and for which the administrator of the target system (the administrator using the administrator terminal 50) will take the lead in considering. Note that if it is impossible to respond to the failure or if the "anticipated failures" are "unknown," type "C" is assigned.

[0120] "Effectiveness of countermeasures" is information indicating whether the countermeasures shown in the fault verification list table 80 are effective as countermeasures for the faults indicated by "anticipated faults." In the example of FIG. 9, three countermeasures, "resource up," "multiplexing," and "service change," are shown as design changes to the target system. In addition, in the example of FIG. 9, for each of these countermeasures, a circle indicating that it is an effective countermeasure or a penalty mark indicating that it is an ineffective countermeasure (not a countermeasure) is shown for each "anticipated fault."

[0121] "Resource up" is a countermeasure to increase the processing capacity of a component (cloud resource 11). "Multiplexing" is a countermeasure to multiplex the cloud resource 11 that executes the cloud application 12, which is a component. "Resource up" and "multiplexing" are both countermeasures to change the cloud resource 11, which is a component of the target system. In this embodiment, "resource up" and "multiplexing" can be considered by the design support system 100 itself when the "fault response type" for the "expected fault" is type "A."

[0122] On the other hand, "service change" is a countermeasure to change the service provided by the target system. In this embodiment, "service change" is a countermeasure that is considered mainly by the administrator of the target system in any case, including when the "failure response type" for the "anticipated failure" is type "A."

[0123] In this embodiment, priority information is attached to these countermeasures. This information indicates the priority of the countermeasures to be adopted when multiple countermeasures for the same failure indicated by the "probable failure" are indicated.

[0124] 9, for example, when "access load" is the "cause of failure," the countermeasures for the "anticipated failure" of "[3] Application stop" are shown to be "resource up," "multiplexing," and "service change." Therefore, "resource up," which has the highest priority, is selected as the countermeasure for the failure that occurs due to the assignment of a failure cause that is performed in accordance with a failure generation command whose "fault verification ID" is "G001."

[0125] The design proposal process performed to provide the design proposal function 44 is a process performed in accordance with the data shown in the fault verification list table 80. This design proposal process will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the processing contents of the design proposal process.

[0126] 10 is started each time the time indicated in the "Time" field in the fault verification list table 80 is reached. First, in S401, a process is performed to obtain the fault to be verified that is associated with the time indicated in the "Time" field from the fault verification list table 80. By this process, various data associated with the time indicated in the "Time" field are obtained from the fault verification list table 80.

[0127] In S402, the fault occurrence cause and the conditions for assigning the fault occurrence cause are sent to the fault verification function 42, and fault verification is performed. By this process, information indicated in the items "Fault verification ID", "Fault occurrence cause", "Target", and "Details of fault occurrence cause" out of the data acquired from the fault verification list table 80 by the process of S401 is sent to the fault verification function 42. The information sent to the fault verification function 42 by this process is received by the process of S306 in the fault verification process (FIG. 8).

[0128] In S403, the execution log and metrics log stored in the storage device of the cloud server that executes the information storage process are referenced to detect the occurrence of a failure in the target system. The execution log and metrics log referenced in this process were saved by the process of S205 in the information storage process (FIG. 6). For example, the application execution log 71 shown in FIG. 7 indicates that an error called "Error No. X" has occurred, as described above, and in the process of S403, a process is performed to find such an error indication from the execution log.

[0129] In S404, a process is performed to determine whether a fault has occurred as a result of the detection in the process of S403. If it is determined in this determination process that a fault has occurred (if the determination result is YES), the process proceeds to S406. On the other hand, if it is determined in this determination process that a fault has not occurred (if the determination result is NO), the process proceeds to S405. Then, in S405, a process is performed to output information indicating that no fault has occurred in the target system due to the fault cause identified by the "fault verification ID" to the administrator terminal 50, and then this design proposal process ends.

[0130] In S406, a condition identification process is performed. This process is a process for providing the function of the condition identification unit 150, and is a process for identifying the conditions for adding a critical failure cause that causes a failure in the target system. The details of this process will be described later.

[0131] In S407, a requirement identification process is performed. This process is a process for providing the function of the requirement identification unit 170, and is a process for identifying requirements of the cloud resource 11 that will not cause a failure in the target system even when a failure cause is added. In this embodiment, a requirement of the cloud resource 11 (contents of design changes to the target system) that will not cause a failure in the target system even when a failure cause that satisfies the conditions for adding criticality identified by the condition identification process is identified. Details of this process will be described later.

[0132] When the process of S407 is completed, this design proposal process ends.

[0133] The above processing is the design proposal processing.

[0134] Next, the condition specification process, which is the process of S406 in the design proposal process described above, will be described. Fig. 11 is a flowchart showing an example of the process contents of the condition specification process.

[0135] 11, first, in S411, a process for changing the conditions for assigning a fault occurrence cause is performed. When the process of S411 is performed for the first time after the start of the condition identification process, a process for changing the conditions for assignment acquired from the fault verification list table 80 by the process of S401 of the design proposal process (FIG. 10) is performed. On the other hand, when the process of S411 is performed during a repetition of the process in this condition identification process, a process for further changing the conditions for assignment after the change by the most recent process of S411 is performed is performed.

[0136] In S412, the fault occurrence cause acquired from the fault verification list table 80 in the process of S401 of the design proposal process (FIG. 10) and the conditions assigned after the change in the process of S411 are sent to the fault verification function 42 to perform fault verification. The information sent to the fault verification function 42 in this process is received in the process of S306 in the fault verification process (FIG. 8).

[0137] In S413, the execution log and metrics log stored in the storage device of the cloud server that executes the information storage process are referenced to detect the occurrence of a fault in the target system. This process is the same as the process in S403 in the fault verification process (FIG. 8).

[0138] In S414, a process of adding information indicating the result of the detection in the process of S413 to the detailed information list table 81 is performed.

[0139] Here, we will explain the detailed information list table 81. Fig. 12 shows examples of data in the detailed information list table 81, where [A] shows an example when the cause of the failure is "access load" and [B] shows an example when the cause of the failure is "random stop of the application."

[0140] The "fault verification ID" is information for identifying a fault occurrence command created by the fault verification function 42. This item stores the ID acquired from the fault verification list table 80 by the processing of S401 of the design proposal process (FIG. 10), plus a sub-number corresponding to the number of times that the fault verification function 42 has been made to perform fault verification by the processing of S412.

[0141] The items "time" and "cause of fault occurrence" store the data of the corresponding items acquired from the fault verification list table 80 by the process of S401 of the design proposal process (FIG. 10).

[0142] The items "Details of the cause of the failure" and "Target" store data indicating the conditions for assignment after the change made in the process of S411.

[0143] 12, in the data example [A], the "Details of the cause of the fault," which is a condition for assigning the "Access Load" obtained from the fault verification list table 80, has been changed from "1000 accesses / s" to "700 accesses / s," "800 accesses / s," "900 accesses / s," etc. In addition, in the data example [B], the "Target," which is a condition for assigning the "Random Stop of Application" obtained from the fault verification list table 80, has been changed from "Sxxx" to "S004."

[0144] In addition, when changing the conditions for assigning a failure cause, the target of the condition change and the change content of the condition are pre-determined depending on the failure cause. For example, it is pre-determined that the condition for assigning "access load" is changed by increasing the access volume indicated in "Details of Failure Cause" from 70 percent by 10 percent. Also, it is pre-determined that the condition for assigning "random app stop" is changed by selecting cloud apps 12 to be assigned a failure cause one by one.

[0145] The "Whether or not a fault has occurred" field stores information indicating the result of the detection performed by the process of S414. If this field is "Yes," it indicates that a fault has been detected in the target system, and if this field is "No," it indicates that a fault has not been detected in the target system.

[0146] The "Occurred Fault" item indicates the details of the fault in the target system detected by the processing of S414 using the number assigned to each piece of information on "Probable fault" in the fault verification list table 80. The details of this fault are identified based on the execution log and metrics log in which the occurrence of the fault was detected in the processing of S413.

[0147] In the data example of [A] in Figure 12, in the fault verification with the "Fault Verification ID" of "G001-3", "Occurred Fault" is marked with [3], indicating that a fault of "application stopped" occurred. Also, in the data example of [B], in the fault verification with the "Fault Verification ID" of "G003-1", "Occurred Fault" is marked with [1] and [3], indicating that a "fault in the target resource" and a "fault in the region" occurred.

[0148] By the processing of S414, data is stored in each item of the detailed information list table 81 described above.

[0149] Continuing the explanation of the condition identification process, in S415, a process is performed to determine whether or not a failure has occurred as a result of the detection in the process of S413. In this determination process, if it is determined that a failure has occurred (if the determination result is YES), the process proceeds to S416. On the other hand, in this determination process, if it is determined that no failure has occurred (if the determination result is NO), the process returns to S411, and the process of changing the condition for assigning the failure cause is performed again, and thereafter the process of S412 is performed again.

[0150] In S416, the conditions for assigning a fault cause when a fault is detected, i.e., the conditions for assignment after the most recent change in the processing of S411, are identified. The conditions identified by this processing become the conditions for assigning a fault cause that is critical to cause a fault.

[0151] In S417, information on the failure that occurred due to the assignment under the conditions identified in the processing of S416 and information on countermeasures for the occurred failure (contents of design changes to the target system) are output to the administrator terminal 50. In this processing, information on each countermeasure that is associated with the "failure that is expected to occur" that matches the occurred failure and for which the "effectiveness of the countermeasure" is effective is output from the information acquired from the failure verification list table 80 in the processing of S401 of the design proposal processing (FIG. 10).

[0152] When the process of S417 is completed, the condition specification process ends, and the process returns to the design proposal process of FIG.

[0153] The above-described processing is the condition specification processing.

[0154] Next, a description will be given of the requirement specification process, which is the process of S407 in the design proposal process of Fig. 10. Fig. 13 is a flowchart showing an example of the process contents of the requirement specification process.

[0155] 13, first, in S421, a process is performed to identify a countermeasure for the "occurred fault" from the information added to the detailed information list table 81 by the process of S414 of the condition identification process (FIG. 11). In this process, first, a process is performed to select an effective countermeasure from among the countermeasures associated with the "expected fault" that matches the "occurred fault" from the data acquired from the fault verification list table 80 by the process of S401 of the design proposal process (FIG. 10). Then, a process is performed to identify the countermeasure with the highest priority from among the selected countermeasures.

[0156] In S422, a process is performed to determine whether the countermeasure identified in the process of S421 is a countermeasure that can be considered mainly by the design support system 100. In this determination process, if it is determined that the identified countermeasure is a countermeasure that can be considered mainly by the design support system 100 (if the determination result is YES), the process proceeds to S423. In this case, the process from S423 onwards is performed to identify requirements for the cloud resource 11 that will not cause a failure. On the other hand, if it is determined that the identified countermeasure is not a countermeasure that can be considered mainly by the design support system 100 (if the determination result is NO), the requirements identification process ends. When the requirements identification process ends, the process returns to the design proposal process of FIG. 10.

[0157] In the fault verification list table 80 illustrated in FIG. 9, when the "fault response type" for the "anticipated fault" is type "A," only the countermeasures of "resource up" and "multiplexing" are considered as countermeasures that can be considered primarily by the design support system 100.

[0158] In S423, a process is performed to implement design changes, which are the countermeasures identified in the process of S421, on the target system.

[0159] As an example, assume that the "occurred failure" in the detailed information list table 81 is "[3]," i.e., "application stopped," as shown in [A] in FIG. 12. According to the failure verification list table 80 in FIG. 9, this failure occurred due to the imposition of an "access load" on the cloud application 12 with the application ID "S004," and the highest priority countermeasure is "resource up." Therefore, in this case, a design change is made in S423 to increase the processing capacity of the cloud resource 11 with the resource ID "C003" that executes the cloud application 12 with "S004" (see FIG. 4). In this embodiment, as a specific example of this design change, a design change is made to increase the number of CPU cores allocated to the cloud resource 11 with "C003."

[0160] As another example, assume that the "Occurred Fault" in the detailed information list table 81 is "[1], [3]," i.e., "Fault in the target resource" and "Fault in the region," as shown in [B] in FIG. 12. According to the fault verification list table 80 in FIG. 9, this fault occurred due to the assignment of "random application stop" to the application ID "Sxxx," i.e., all cloud applications 12, and the highest priority countermeasure is "multiplexing." Therefore, in this case, a design change to multiplex the cloud resources 11 is performed by the processing of S423 (see FIG. 4).

[0161] In S424, the fault occurrence cause acquired in the process of S401 of the design proposal process (FIG. 10) and the conditions for applying the fault occurrence cause that become critical and cause a fault, which were identified in the process of S416 of the condition identification process (FIG. 11), are sent to the fault verification function 42. The information sent to the fault verification function 42 by this process is received by the process of S306 in the fault verification process (FIG. 8), thereby performing fault verification. Note that in S424, this sending process is performed a predetermined number of times, and therefore fault verification is performed multiple times.

[0162] In S425, during each of the fault verifications performed multiple times by the processing of S424, the execution log and metrics log stored in the storage device of the cloud server that executes the information storage processing are referenced to detect the occurrence of a fault in the target system. This processing is the same as the processing of S403 in the fault verification processing (FIG. 8).

[0163] In S426, a process is performed to calculate the fail-safe rate for the target system after the design change based on the detection result obtained in the process of S425.

[0164] The fail-safe rate is the percentage of the number of times that a failure in the target system is not detected by the processing of S425 compared to the number of times that failure verification is performed multiple times by the processing of S424. This fail-safe rate is an example of an index that indicates the possibility that a failure will not occur in the target system even if a failure cause is added.

[0165] In S427, a process is performed to determine whether the process of S423 for implementing the design change has been performed a predetermined number of times, and if it is determined that it has been performed (if the determination result is YES), the process proceeds to S429. On the other hand, if it is determined in this determination process that the number of times the process of S423 has been performed is less than the predetermined number of times (if the determination result is NO), the process proceeds to S428.

[0166] In S428, a process is performed to change the contents of the design change most recently implemented in the target system, and then the process returns to S423, where the design change after the content change is implemented in the target system.

[0167] In S429, the fail-safe rate data 82 is output to the manager terminal 50.

[0168] Here, we will explain the fail-safe rate data 82. Fig. 14 shows examples of the fail-safe rate data 82. Of these, [A] shows an example when the cause of the failure is "access load," and [B] shows an example when the cause of the failure is "random stop of the application."

[0169] As described above, if the "occurred failure" in the detailed information list table 81 is "3" ("application stopped") as shown in [A] of FIG. 12, in this embodiment, a design change is made to increase the number of CPU cores allocated to the cloud resource 11. The fail-safe rate data 82 in [A] of FIG. 14 shows an example of this case. This example shows the fail-safe rate for each time a failure verification is performed while changing the number of CPU cores of the cloud resource 11 of "C003" that executes the cloud application 12 of "S004" from "4" to "8" and then to "12." Note that this example shows both the usage of the processing unit and the usage of memory for each time a failure verification is performed, and these data are obtained from the execution log and the metrics log referenced by the processing of S425.

[0170] As described above, if the "occurred failure" in the detailed information list table 81 is [1] ("failure in the target resource") or [3] ("failure in the region") as shown in [B] of FIG. 12, in this embodiment, a design change is made to multiplex the cloud resource 11. The fail-safe rate data 82 in [B] of FIG. 14 shows an example of this case. In this example, the fail-safe rate is shown for each time a failure verification is performed while changing the region of the new cloud resource 11 used to multiplex the execution environment of the cloud application 12 of "S004" from "JP1" to "JP2" and "US1." In this example, the new cloud resources 11 are assigned resource IDs of "C003-1," "C003-2," and "C003-3," respectively. Here, the region "JP 1" of the cloud resource 11 of "C003-1" is assumed to be the same region as the cloud resource 11 of "C003" that was running the cloud application 12 of "S004" in the target system before the design change.

[0171] In addition, when changing the content of the design change for the target system, the target to be changed and the content of the change to the target are assumed to be predetermined depending on the countermeasure. That is, for example, it is assumed that the change to the design change for the countermeasure "resource increase" is predetermined to increase the number of CPU cores allocated to the cloud resource 11 by a predetermined number. Also, for example, it is assumed that the change to the design change for the countermeasure "multiplexing" is predetermined to change the region of the new cloud resource 11 used for multiplexing the cloud resource 11.

[0172] Continuing the explanation of the requirements identification process, in S430, a process is performed to determine whether there has been a design change that has achieved a 100 percent fail-safe rate calculated in the process of S426, and if it is determined that such a design change has occurred (if the determination result is YES), the process proceeds to S431. On the other hand, if it is determined in this determination process that there has not been a design change that has achieved a 100 percent fail-safe rate (if the determination result is NO), the requirements identification process ends and the process returns to the design proposal process of FIG. 10.

[0173] In S431, information indicating the contents of the design change that achieved a fail-safe rate of 100 percent is output to the administrator terminal 50 as the result of identifying the requirements of the cloud resource 11 that will not cause a failure in the target system even when a failure cause is added.

[0174] For example, when the fail-safe rate data 82 is [A] in Fig. 14, information is output by the processing of S431 indicating that the number of CPU cores to be allocated to the cloud resource 11 of "C003" is 12. When the fail-safe rate data 82 is [B] in Fig. 14, information is output indicating that the cloud resource 11 of "C003" is multiplexed using the cloud resource 11 of "C003-2" or "C003-3". The region of the cloud resource 11 of "C003-2" is "JP 2nd", and the region of the cloud resource 11 of "C003-3" is "US 1st".

[0175] When the process of S431 is completed, this requirement specification process is completed and the process returns to the design proposal process of FIG.

[0176] The above processing is the requirement identification processing.

[0177] 13, if it is determined in the determination process of S430 that there is no design change that has achieved a 100 percent fail-safe rate, the requirements identification process is terminated. Alternatively, the process may return to S421, and a process may be performed to identify a countermeasure with the next highest priority, following the countermeasure identified in the most recent process of S421, from among the effective countermeasures associated with the "anticipated failure" that matches the "occurring failure." In this case, it is preferable to proceed with the process from S422 onwards for the identified countermeasure.

[0178] Through the design proposal process described above, information on failures that occur due to the addition of critical failure factors that cause failures and information on countermeasures for the failures are provided to the administrator of the target system (see S417 in FIG. 11). Furthermore, if the design support system 100 is able to take the lead in examining countermeasures, fail-safe rate data 82 and information on requirements for cloud resources 11 that will not cause failures even when failure factors are added are provided to the administrator (see S429 and S431 in FIG. 13). Therefore, the administrator can use this provided information as design support material to proceed with the design of the target system.

[0179] Although the disclosed embodiments and their advantages have been described in detail above, it will be appreciated that those skilled in the art may make various modifications, additions, and omissions without departing from the scope of the invention as clearly set forth in the claims. [Explanation of symbols]

[0180] 10. Cloud 11 Cloud Resources 12. Cloud Apps 20, 20a, 20b devices 21 Sub-Devices 22 Device Apps 30 Public lines 41 Information storage function 42 Fault Verification Function 43 Fault Identification Function 44 Design proposal function 45a, 45b Fault Verification Agent (Agent) 50 Administrator terminal 60 Information processing equipment 61 CPU 62 memory 63 Input Device 64 Output Devices 65 Auxiliary storage device 66 Communication I / F 67 Internal Bus 70 Resource Information Table 71 App execution log 72 Cloud Resource Metrics Log 73 Device Resource Metrics Log 74 List of failure instructions 80 Fault Verification List Table 81 Detailed Information List Table 82 Fail-safe rate data 100 Design Support System 110 Collection Department 111 First Collection and Brokerage Department 112 2nd Collection and Brokerage Department 113 Receiving Department 120 Granting Department 121 First Granting Intermediary Department 122 Second Granting Intermediary Department 123 Assignment Instruction Section 130 Preservation Department 140 Detection unit 150 Condition specification part 160 Changes 170 Requirements Specification Department 180 Output section

Claims

1. A design support system for supporting the design of a target system that is configured of a cloud service provided by cloud computing and a device that exchanges data with the cloud service, the target system having, as components, a cloud app that provides the cloud service, cloud resources that are hardware that executes the cloud app, a device app that provides a function of exchanging the data in the device, and device resources that are hardware that executes the device app, an assigning unit that assigns a factor that may cause a failure in the target system to any of the components; a detection unit that detects the occurrence of the failure; a change unit that changes the cloud resources; a requirement specifying unit that specifies a requirement of the cloud resource that will not cause the failure even when the cause is assigned, based on a result of the detection by the detection unit in response to the assignment of the cause by the assignment unit, each time the change unit changes the cloud resource multiple times; an output unit that outputs information indicating the requirement identified by the requirement identification unit; A design support system comprising:

2. a condition specifying unit that specifies a condition for the assignment that becomes critical and causes the fault to occur based on a result of the detection by the detection unit with respect to the assignment of the cause each time the assignment unit assigns the cause while changing the condition for the assignment of the cause; The requirement specifying unit specifies a requirement of the cloud resource that will not cause the failure even when the factor that satisfies the condition for the assignment is assigned.

2. The design support system according to claim 1.

3. The design support system described in claim 1 or 2, characterized in that the requirement identification unit calculates an index representing the possibility that the failure will not occur even if the factor is assigned based on the results of the detection by the detection unit in response to the assignment of the factor by the assignment unit each time the change unit makes changes to the cloud resources multiple times, and identifies the requirement based on the index.

4. a collection unit that collects execution logs of the cloud applications as log information and also collects logs of metrics related to the cloud resources as metrics information; a storage unit that stores the log information and the metrics information as fault verification information; Further provided with The detection unit detects the occurrence of the failure using the failure verification information.

3. The design support system according to claim 1 or 2.

5. The detection unit further detects a type of the detected failure, The requirement specification unit specifies the requirement when the type matches a predetermined type.

3. The design support system according to claim 1 or 2.

6. the factor is a predetermined amount of access load on the cloud application, The change unit changes the processing capacity of the cloud resource, The requirement identification unit identifies the processing capacity of the cloud resource at which the failure does not occur even when the predetermined amount of access load is applied, based on a result of the detection by the detection unit in response to the application of the predetermined amount of access load by the application unit, each time the change unit changes the processing capacity of the cloud resource multiple times.

3. The design support system according to claim 1 or 2.

7. the cause is a stop of a cloud application running on the cloud resource; the change unit multiplexes the cloud resources that execute the cloud application and changes the region of the new cloud resources used for the multiplexing; The requirement identification unit identifies the region in which the failure will not occur even when the granting of the stop of the cloud application executed in the cloud resource is granted, based on the result of the detection by the detection unit of the stop of the cloud application by the granting unit each time the change unit multiplexes the cloud resource and changes the region.

3. The design support system according to claim 1 or 2.

8. A design support method performed by a design support system that supports the design of a target system that is configured to include a cloud service provided by cloud computing and a device that exchanges data with the cloud service, the target system having, as components, a cloud app that provides the cloud service, cloud resources that are hardware that executes the cloud app, a device app that provides a function of exchanging the data in the device, and device resources that are hardware that executes the device app, The cloud resource is changed a plurality of times, and each time the change is made a plurality of times, a factor that may cause a failure in the target system is assigned to any of the components, and the occurrence of the failure in response to the assignment of the factor is detected; Identifying requirements for the cloud resource that will not cause the failure even when the cause is applied, based on a result of detecting the occurrence of the failure in response to the application of the cause each time the change is applied multiple times; Output information indicating the identified requirements A design support method comprising:

Citation Information

Patent Citations

  • Failure analysis support device, failure analysis support method, and failure analysis support program

    JP2019036809A

  • Control cloud server

    JP2023045180A