Business influence range presentation device and business influence range presentation method
The business impact scope presentation device accurately identifies operations affected by microservice abnormalities by analyzing trace data and redundant paths, enhancing operational management in complex systems.
Patent Information
- Application Number
- JP2023209670
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-24
AI Technical Summary
Existing operation support methods struggle to accurately identify the business impact scope when abnormalities occur in microservices due to performance degradation or failure, as they rely on pre-defined associations and fail to consider dynamic changes in request paths.
A business impact scope presentation device and method that collects test logs and monitoring data to analyze trace data, identifies redundant request paths, and determines the affected operations by correlating microservices and business functions using use case association information.
Enables precise identification of business operations impacted by abnormalities in microservices, facilitating proactive measures to maintain system performance and achieve business goals.
Smart Images

Figure 2025093794000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a business impact scope presentation device and a business impact scope presentation method, and is suitable for application to, for example, a business impact scope presentation device related to a technology for presenting the impact scope on a business.
Background Art
[0002] In recent years, with the penetration of microservices, the monitoring of systems has become more complex. Since stakeholders are seeking decision-making materials in business, operation managers need to perform operation management considering the relationship between the operation target and the business on a daily basis. Therefore, a technology for specifying the relationship between a system and a business from the movement of the system is required.
[0003] To introduce such an operation support method for supporting operation management, it is necessary for the operation manager to define in advance the relationship between the business and the system, and it is a prerequisite that the operation manager understands the relationship. Therefore, the operation manager needs to cooperate with developers to prepare for appropriately operating the system. By introducing such an operation support method, the operation manager can always confirm the relationship between the business and the system.
[0004] Patent Document 1 describes an operation support method that holds association information between an IT (Information Technology) system and a business, holds information on systems important in business, recovery levels for business continuity, and recovery procedures, and enables appropriate selection of countermeasures when an incident occurs. According to the method described in Patent Document 1, by associating and holding the business functions that make up the business support system, the target recovery level of the business, the achievement judgment criteria, and the implementation items, it is possible to present countermeasures for achieving the target recovery level when a failure occurs in the business support system.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the operation support method described in Patent Document 1, the operation support method is retrieved from the association information between the already grasped business content and the business functions constituting the business support system. Therefore, when the operation support method described in Patent Document 1 is applied to microservices, the following problems may occur when an abnormality such as performance degradation or failure occurs in the request path including this microservice and API (Application Programming Interface). That is, in the operation support method described in Patent Document 1, since the operation support method is retrieved from the association information between the already grasped business content and the business functions, it is impossible to retrieve the operation support method in consideration of the above abnormality, and it is difficult to specify the content of the business affected by the abnormality from the request path.
[0007] The present invention has been made in consideration of the above points, and when an abnormality occurs in the request path including microservices in the production environment, it is intended to propose a business impact range presentation device and a business impact range presentation method capable of more accurately specifying the content of the business affected by the abnormality of the request path.
Means for Solving the Problems
[0008] In order to solve such problems, in the present invention, test log collection means is provided for collecting test record information indicating the execution result of a test on monitored software in which one or more operation steps indicating the constituent elements of the content of each of a plurality of operations and a plurality of microservices for executing processing by the operation of the one or more operation steps are connected via one or more request paths, and for collecting test case information indicating the relationship between the content of each operation and the one or more operation steps; a use case association unit for generating use case association information by associating the request path corresponding to each operation step among the one or more request paths with the content of each operation based on the test record information and the test case information collected by the test log collection means; a monitoring information acquisition unit for acquiring trace data indicating the execution result of each microservice as the execution result in the production environment for the monitored software; a trace data shaping unit for determining whether there is a redundant request path in normal operation in the request path, and creating the trace data excluding the redundant request path when the redundant request path exists; and a business impact scope analysis unit for determining whether an abnormality has occurred in any one of the one or more request paths based on the trace data acquired by the monitoring information acquisition unit. When the business impact scope analysis unit determines that the abnormality has occurred in any one of the request paths, the business impact scope analysis unit refers to the use case association information based on the one of the request paths excluding the redundant request path, and specifies the content of the operations affected by the request path in which the abnormality has occurred among the content of each operation.
[0009] In the present invention, a test log collection unit collects test record information indicating an execution result of a test on monitored software in which one or more operation steps indicating components of the content of each of a plurality of operations and a plurality of microservices that execute processing by the operations of the one or more operation steps are connected via one or more request paths, and collects test case information indicating a relationship between the content of each operation and the one or more operation steps, a use case association unit generates use case association information by associating a request path corresponding to each operation step among the one or more request paths with the content of each operation based on the test record information and the test case information collected in the test log collection step, a monitoring information acquisition unit acquires trace data indicating an execution result of each microservice as an execution result in a production environment for the monitored software, a trace data shaping unit determines whether there is a redundant request path in normal operation in the request path, and creates the trace data excluding the redundant request path when the redundant request path exists, and a business impact scope analysis unit determines whether an abnormality has occurred in any of the one or more request paths based on the trace data acquired in the monitoring information acquisition step. In the business impact scope analysis step, when the business impact scope analysis unit determines that the abnormality has occurred in any of the request paths, the business impact scope analysis unit refers to the use case association information based on any of the request paths excluding the redundant request path, and identifies the content of the operations affected by the request path in which the abnormality has occurred among the content of each operation.
Effect of the Invention
[0010] According to the present invention, when an abnormality occurs in a request path including microservices, it is possible to identify the content of the operations affected by the abnormality of the request path.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The following description and drawings are examples for explaining the present invention, and for the sake of clarity of explanation, appropriate omissions and simplifications are made. The present invention can be implemented in various other forms. Unless otherwise limited, each component may be singular or plural. In the drawings, the positions, sizes, shapes, ranges, etc. of the components shown may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings. In the following description, various information may be described using expressions such as "table" and "list", but the various information may be represented by other data structures. In order to indicate independence from the data structure, "XX table", "XX list", etc. may be referred to as "XX information". When explaining identification information, when expressions such as "identification information", "identifier", "name", "ID", "number", etc. are used, these can be replaced with each other.
[0013] (1) Example This example is based on the execution results in the test environment and the production environment of a monitoring target system that manages monitoring target software including a plurality of microservices that define the processing content of one or more operation steps belonging to a plurality of operations. When an abnormality occurs in the request path including the microservice, it identifies the content of the operations affected by the abnormality in the request path.
[0014] FIG. 1 is a block diagram showing a configuration example of a computer system including a business impact range presentation device 100 according to an embodiment of the present invention. In FIG. 1, the computer system includes a monitoring target system 1, a test device 2, a monitoring device 3, a network 4, a business impact range presentation device 100, and a display device 5. Note that the business impact range presentation device 100 is communicably connected to the test device 2, the monitoring device 3, and the display device 5 via the network 4. The network 4 is, for example, a public network such as the Internet, a LAN (Local Area Network), a WAN (Wide Area Network), or the like.
[0015] The monitored system 1 is an IT system for executing operations, and is composed of, for example, a computer (not shown) including a processor, a storage device, an input device, an output device, and a communication device. The processor is composed of, for example, a CPU (Central Processing Unit) or an MPU (Micro-processing unit). The processor executes a monitored program for executing processing according to the content of the operations of the IT system (hereinafter sometimes referred to as a use case) in the production environment. At this time, in the database 10 belonging to the storage device, data such as an access history and an error history generated by the processor during the execution of the monitored program are stored as system logs. Note that the means for storing the system logs is not limited.
[0016] The test device 2 is a device for confirming whether the monitored system 1 operates normally in the test environment before being introduced into the production environment, and is composed of, for example, a computer (not shown) including a processor, a storage device, an input device, an output device, and a communication device. In the database 20 belonging to the storage device of the test device 2, for example, information (test record information) indicating the test results when the monitored system 1 executes the monitored program in the test environment is stored as test logs. The test logs are, for example, information for confirming whether a specific use case can be realized normally, and are used as information on what operations were performed and whether the expected processing results and the like were obtained. Note that the means for storing the test logs is not limited.
[0017] The monitoring device 3 is a device that collects various monitoring data from the monitoring target system 1 in a test environment and a production environment in order to check whether the monitoring target system 1 is operating normally. For example, it is composed of a computer (not shown) equipped with a processor, a storage device, an input device, an output device, and a communication device. In the database 30 belonging to the storage device of the monitoring device 3, as monitoring data, for example, metric data such as CPU utilization rate and memory utilization rate, trace data including the path of API (Application Programming Interface) requests (hereinafter referred to as "API request path"), and data logs such as access history and error history are stored. This API request path is, for example, a path regarding in what order the APIs are operating. Note that the means for storing the monitoring data is not limited.
[0018] The display device 5 is, for example, a device belonging to a server computer which is physical computer hardware owned by the operation management department of the IT system. The display device 5 has a function of displaying the information output from the business impact scope presentation device 100 via the network 4 and the visualization chart attached to this information.
[0019] FIG. 2 is a configuration diagram showing a configuration example of monitoring target software managed by a monitoring target system according to an embodiment of the present invention. In FIG. 2, the monitoring target system 1 includes use cases 11, 12, ···, operation steps 21, 22, 23, ···, microservices 31, 32, ···, microservices 41, 42, ···, microservice 51, and API endpoints 31A, 32A, 41A, 42A, 51A as monitoring target software for executing the business of the IT system. The use cases 11, 12 are the contents of the business that the IT system wants to achieve. For example, the use case 11 indicates that the content of the business of the IT system is the settlement of orders of registered users, and the use case 12 indicates that the content of the business of the IT system is the settlement of orders of guests. The operation steps 21, 22, 23 For example, they are components belonging to use case 11, and are components when use case 11 is divided into a plurality according to its content. Operation step 21 is a step for operating the processing of the order settlement of the registered user. Operation step 21 is, for example, checkout. Operation step 22 is a step for operating the processing of the order settlement of the registered user, and is, for example, order creation. Operation step 23 is a step for operating the processing of the order settlement of the registered user. Operation step 23 is, for example, payment.
[0020] Micro services 31, 32, 41, 42, 51 are programs for providing service functions, and are independent programs from each other. Micro service 31 is a program having a function of executing the processing of checkout. Micro service 32 is a program having a function of executing the processing of order creation. Micro service 41 is a program having a function of executing the processing of calculation. Micro service 42 is a program having a function of executing the processing of inventory confirmation. Micro service 51 is a program having a function of executing the processing of coupon acquisition. API endpoints 31A, 32A, 41A, 42A, 51A are reception ports for execution triggers of the functions of the respective micro services 31, 32, 41, 42, 51. One use case includes a plurality of operation steps. One operation step is executed via one or more micro services and API endpoints.
[0021] That is, in order to execute the processing of one operation step, one operation step is connected via the input / output interface of a request (access) belonging to any microservice or API endpoint. For example, operation step 21 connected to use case 11 is connected to API endpoint 31A of microservice 31. Operation step 22 connected to use case 11 is connected to endpoint 51A of microservice 51 via API endpoint 32A of microservice 32 and API endpoint 41A of microservice 41, and is also connected to API endpoint 32A of microservice 32 and API endpoint 42A of microservice 42. At this time, each microservice and each API endpoint constitute an API request path indicating the transmission path of a request (access) from any microservice (client) to another microservice (server).
[0022] FIG. 3 is a functional block diagram showing a configuration example of a business impact scope presentation device 100 according to an embodiment of the present invention. In FIG. 3, the business impact scope presentation device 100 includes, for example, as software resources, an input unit 110, an output unit 120, a storage unit 130, an arithmetic unit 140, and a communication unit 150. Note that the business impact scope presentation device 100 is configured by a computer (not shown) including, as hardware resources, a processor, a main memory device, an auxiliary storage device, an input device, an output device, and a communication device.
[0023] The main memory device is a device that stores computer programs and data, and is, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), a non-volatile semiconductor memory, or the like.
[0024] The auxiliary storage device is, for example, a hard disk drive, an SSD (Solid State Drive), an optical storage medium (i.e., CD (Compact Disc), DVD (Digital Versatile Disc), etc.), a storage system, an IC card (Integrated Circuit Card), an SD (Secure Digital) memory card, etc., a reading / writing device for a recording medium, and a storage area of a cloud server, etc. The computer programs and data stored in the auxiliary storage device are read into the main storage device at any time.
[0025] The input device is, for example, a keyboard, a mouse, a touch panel, a card reader, a voice input device, etc. The output device (display device) is a user interface that provides various information such as the progress of processing and the processing results to the user. The output device is, for example, a screen display device as a display (i.e., a liquid crystal monitor, an LCD (Liquid Crystal Display), or a graphics card, etc.), a voice output device (i.e., a speaker, etc.), or a printing device, etc.
[0026] The communication device is a wired or wireless communication interface that realizes communication with other devices via communication means such as a LAN or the Internet. The communication device is, for example, a NIC (Network Interface Card), a wireless communication module, a USB (Universal Serial Bus) module, or a serial communication module, etc.
[0027] Here, the processor is configured using, for example, a CPU or an MPU. At this time, for example, when the processor operates according to the input program loaded into the main memory device, the function of the input unit 110 is realized, and when the processor operates according to the output program loaded into the main memory device, the function of the output unit 120 is realized. Also, when the processor operates according to the arithmetic program loaded into the main memory device, the function of the arithmetic unit 140 is realized, and when the processor operates according to the communication program loaded into the main memory device, the function of the communication unit 150 is realized. Furthermore, when the processor operates with the main memory device as the target for storing data and information, the function of the storage unit 130 is realized.
[0028] Specifically, the input unit 110 receives information when the user operates a keyboard or a mouse via the input device, accepts the input information as input information, and outputs the input information to the arithmetic unit 140.
[0029] The output unit 120 generates screen information to be displayed on the display (display unit) of the business impact scope presentation device 100 or the display device 5, and outputs the generated image information to the display or the display device 5.
[0030] The storage unit 130 is a database that stores various types of information. The storage unit 130 stores test case information 131, test record information 132, use case association information 133, target value information 134 of monitoring items, trace data 135, and update information 136. The information stored in the storage unit 130 will be described later.
[0031] The arithmetic unit 140 includes, for example, a test log collection unit 141, a use case association unit 142, a monitoring information acquisition unit 143, a business impact scope analysis unit 144, and a trace data formatting unit 146.
[0032] The test log collection unit 141 is a test log collection program that collects test case information 131 and the execution results of tests from the test device 2, and also acquires monitoring data collected when the monitoring device 3 executes tests. The test log collection unit 141 collects test record information indicating the execution results of tests on monitoring target software in which one or more operation steps indicating the constituent elements of the content of each of a plurality of operations and a plurality of microservices that execute processing by operating the one or more operation steps are connected via one or more request paths, and also collects test case information indicating the relationship between the content of each operation and each operation step.
[0033] Specifically, the test log collection unit 141 reads and collects, as test case information indicating the relationship between each use case and each test case (operation step), test case information (see FIG. 4) described later from the test device 2 via the network, and stores the collected test case information in the storage unit 130. Also, when the test device 2 executes a test on the monitoring target system 1 and the monitoring device 3 collects monitoring data from the monitoring target system 1, the test log collection unit 141 reads, as monitoring data from the monitoring device 3, for example, trace data 135, target value information 134 of monitoring items, the response time (processing time) of each request, etc., generates test record information (see FIG. 5) described later with reference to the read monitoring data and test case information 131, and stores the generated test record information 132 in the storage unit 130.
[0034] Here, the method for acquiring the test case information 131 is not limited to the exemplified method. For example, it may be acquired from a CI / CD (Continuous Integration / Continuous Delivery) tool (such as tools like Jenkins, CircleCI, etc.) that manages the test device 2. Also, specific names and test contents may be acquired from another data source using the test case ID (for example, UC-1) and the use case ID (for example, UC-1) as keys. The method for acquiring the test log is not limited to the exemplified method. A method of acquiring it in real time using the API of the monitoring device 3 may also be used.
[0035] The use case association unit 142 is a use case association program that patterns API requests and extracts the relationship between API requests and use cases based on the test record information generated by the test log collection unit 141. The use case association unit 142 generates use case association information by associating the request path corresponding to each operation step among one or more request paths with the content of each business based on the test record information 132 collected by the test log collection unit 141 and the test case information. Specifically, the use case association unit 142 acquires the test execution results (for example, test record information) and monitoring data (for example, trace data), and extracts a series of microservices, API endpoints, and request contents of the same operation (all API requests belonging under the same TraceID) from the acquired information and data to generate nested structure information. At this time, the use case association unit 142 generates the nested structure information as the use case association information 133 shown in FIG. 6, and stores the generated use case association information 133 in the storage unit 130. That is, the use case association unit 142 generates the use case association information 133 by associating the request path corresponding to each operation step among one or more request paths with each use case. Note that the process of extracting the relationship between API requests and use cases will be described later.
[0036] The monitoring information acquisition unit 143 is a monitoring information acquisition program that acquires, from the monitoring device 3, the target value information 134 of the monitoring items and monitoring data, for example, trace data indicating the execution results of each microservice as the execution results in the production environment for the monitored software managed by the monitored system 1. Specifically, the monitoring information acquisition unit 143 acquires, from the configuration file of the monitoring device 3, information regarding the target values (measurement period, threshold, percentage target) that the monitoring items of each monitored API should achieve, and acquires trace data (URI, parameters, method, request body, status code, response body, start time, end time, TraceID, SpanID, ParentID) belonging to the monitoring data of the monitoring device 3. At this time, the monitoring information acquisition unit 143 stores the information regarding the target values that the monitoring items of each monitored API should achieve in the storage unit 130 as the target value information of the monitoring items (see FIG. 7) described later, and stores the acquired trace data in the storage unit 130 as the trace data (see FIG. 8) described later.
[0037] The business impact scope analysis unit 144 is a business impact scope analysis program that determines whether an abnormality including performance degradation or occurrence of a failure has occurred in any of the plurality of API request paths based on the trace data, and compares the use case association information (see FIG. 6) described later with the trace data (see FIG. 8) described later to associate the API request paths that do not reach the target value with the use cases. Specifically, the business impact scope analysis unit 144 compares the characteristics of an API request that does not reach the target value, for example, an API request whose request processing time exceeds "500 ms", with the pattern of the API request path recorded in the use case association information, and searches for a use case in the use case association information in which there is an API request path that matches the API request that does not reach the target value.
[0038] The business impact scope analysis unit 144 determines whether an abnormality has occurred in any of the one or more request paths based on the trace data acquired by the monitoring information acquisition unit 143.
[0039] The trace data shaping unit 146 is a monitoring information shaping program that excludes redundant request paths from the trace data 135 used in the search for use cases in the business impact scope analysis unit 144. The trace data shaping unit 146 determines whether there is a redundant request path in the normal operation of the request path, and if there is a redundant request path, creates the trace data 135 with the redundant request path excluded. Specifically, the trace data shaping unit 146 detects a redundant request path in the trace data 135 acquired by the monitoring information acquisition unit 143 compared to a normal request path such as a retransmission process (retry) caused by an error during a request from the client to the server, and excludes it from the trace data. The details of the process of excluding this redundant request path will be described later.
[0040] When the above-described business impact scope analysis unit 144 determines that an abnormality has occurred in any of the request paths, it refers to the use case association information based on any of the request paths with the redundant request path excluded, and specifies the content of the business that is affected by the request path where the abnormality has occurred among the contents of each business.
[0041] The communication unit 150 is a communication program that transmits and receives information to and from an external device. Specifically, the communication unit 150 acquires necessary data from the test device 2 and the monitoring device 3. In addition, the communication unit 150 transmits an instruction to generate a UI (User Interface) for information input as screen information to the display device 5.
[0042] FIG. 4 is a configuration diagram showing a configuration example of the test case information 131 according to an embodiment of the present invention. In FIG. 4, the test case information 131 is information generated in advance in the test device 2 as information necessary for the test device 2 to execute a test on the monitoring target system 1, and includes a test case ID 131A, a test case name 131B, a related use case ID, a related use case name 131D, and a test content 131E.
[0043] Test case ID 131A is an identifier that uniquely identifies the number of a test case. For this test case ID 131A, for example, as the first test case, information such as "TC-1" is recorded. Test case name 131B is an identification name for identifying the name of a test case. For test case name 131B, for example, when performing a test on the checkout of a registered user (a test showing the processing content of operation step 21 in Figure 2), "Checkout of registered user" is recorded. The related use case ID is an identifier that uniquely identifies the number of the use case to be verified in a test case. For the related use case ID, for example, when the number of the use case to be verified in a test case (order settlement of a registered user) is "UC-1", information such as "UC-1" is recorded. Related use case name 131D is an identification name for the use case to be verified in a test case. For related use case name 131D, for example, when the use case to be verified in a test case (TC-1) is the order settlement of a registered user (use case 11 in Figure 2), information such as "Order settlement of registered user" is recorded. Test content 131E is the content of the test by test device 2. For test content 131E, for example, information such as "Add multiple items to the cart. Check whether the addition was successful." is recorded as the operation on the monitored system 1 that is the object of the test, the expected movement, input, and output.
[0044] Here, for example, in the case of test case (TC-1), it will be recorded that "The test case of test case TC-1 verifies the use case with use case number UC-1. The test is completed if multiple items can be added to the cart and the success screen can be confirmed."
[0045] FIG. 5 is a configuration diagram showing a configuration example of test record information 132 according to an embodiment of the present invention. In FIG. 5, the test record information 132 is information collected and recorded as test results when the test device 2 executes a test on the monitoring target system 1. The test record information 132 includes a test case ID 132A, a URI (Uniform Resource Identifier) 132B, a parameter 132C, a method 132D, a request body 132E, a status code 132F, a response body 132G, a start time 132H, an end time 132I, a TraceID 132J, a SpanID 132K, and a ParentID 132L.
[0046] The test case ID 132A is the number of the test case, similar to the test case ID 131A. The URI 132B is the path of the specified endpoint when the monitoring target system 1 executes a test. For example, the information of " / userCheckout" is recorded in the URI 132B as the path of the endpoint (API endpoint 31A) when the monitoring target system 1 executes a test by a test case (TC-1). The parameter 132C is additional information added in a specific format at the end of the path specifying the destination of the request (access) when the monitoring target system 1 executes a test. For example, when the path of the endpoint (API endpoint 31A) is " / userCheckout?productID=ESPC7Z", the information of "productID=ESPC7Z" is recorded in the parameter 132C.
[0047] The method 132D represents the type of request from the client (source of the request) to the server (destination of the request) using the communication protocol HTTP (Hypertext Transfer Protocol). For example, the information of "GET" is recorded in the method 132D as a request from the client to the server.
[0048] The request body 132E is the information sent from the client to the server when the client requests an API endpoint. Note that when the method 132D is "GET" and there is no information to be sent from the client to the server, no information is recorded in the request body 132E.
[0049] The status code 132F is the status code in the communication protocol HTTP, that is, the code returned when the server responds to the client when the client requests an API endpoint. For example, the information of "200", which is the status code indicating a normal response from the server to the client in HTTP, is recorded in the status code 132F.
[0050] The response body 132G is the response information sent from the server to the client when the client requests the server's API endpoint. For example, the information of "{"result":"0"}", which is sent from the server to the client as the return value of the endpoint call, is recorded in the response body 132G. Note that when there is no information to be sent from the server to the client, no information is recorded in the response body 132G. In this embodiment, at least one of the status code 132F and the response body 132G may be used to determine whether there is a redundant request path as described later.
[0051] The start time 132H is the time when a specific request is sent from the client to the server. For example, when the start time is 1:42:08 on November 4, 2022, at least the information of "2011-11-04 01:42:08.5653016" is recorded in the start time 132H.
[0052] The end time 132I is the time when the server completes the processing and sends a response to the client. For the end time 132I, when the end time is 1:42:09 on November 4, 2022, at least the information of "2022-11-04 01;42:09.1837765" is recorded. At this time, the time from the start time 132H to the end time 132I is the processing time required for the client to process the request. Note that the start time 132H and the end time 132I can also record the time in the order of ms as the time less than a second.
[0053] TraceID132J is an identifier that identifies a series of requests when the client requests an API endpoint. For TraceID132J, for example, in the case of a test case (TC-1), the information of "e8bdc860" is recorded as the identifier of the request when the client requests the API endpoint 31A.
[0054] SpanID132K is an identifier that identifies the processing content of the endpoint when a request is sent from the client to the API endpoint. For SpanID132K, for example, the information of "8b8d57f6" is recorded as the identifier that identifies the processing of the API endpoint 31A when a request is sent from the client to the API endpoint 31A.
[0055] ParentID132L is an identifier that identifies the processing of the request immediately preceding the request directly sent from the client to the API endpoint among the requests when a request is sent from the client to the API endpoint. For ParentID132L, for example, in the case of test case (TC-1), among the requests when the client requests API endpoint 31A, since there is no preceding request, no information is recorded. In the case of test case (TC-2), among ParentID132L, in the record where " / calculatePrice" is recorded in URI132B, information "91b8a50e" is recorded as an identifier (an identifier that identifies the processing of API endpoint 32A) for identifying the processing of the immediately preceding request.
[0056] Here, for example, in the case of use case 11 indicating "settlement of registered user's order", the processing of "Checkout", which is the operation in operation step 21 of test case TC-1, is recorded as being executed using API endpoint 31A of microservice 31 indicating "Checkout serivce". Also, the processing of "order creation", which is the operation in operation step 22 of test case TC-2, is recorded as being executed using API endpoint 32A of microservice 32 indicating "Order service", API endpoint 41A of microservice 41 indicating "Calculate service", API endpoint 42A of microservice 42 indicating "Inventory service", and API endpoint 51A of microservice 51 indicating "Discount service".
[0057] Figure 6 is a configuration diagram showing a configuration example of use case association information according to an embodiment of the present invention. In Figure 6, use case association information 133 is information generated by test device 2 to associate a use case with an API request path. The use case association information 133 includes a use case ID 133A, an API request path 133B, and a request content 133C.
[0058] Use case ID 133A is an identification number that uniquely identifies a use case. For example, in the case of use case 11 indicating "settlement of an order for a registered user", information such as "UC-1" is recorded in Use case ID 133A. API request path 133B is the path of the request sent from the client to the server when the use case is executed. In API request path 133B, when the request is processed via one or more API endpoints, information on the microservices passed through, the order of the API endpoints, and the HTTP method is recorded in, for example, JSON (JavaScript (registered trademark) Object Notation) format. Request content 133C is the content of the request sent from the client to the server. In request content 133C, on the API request path, the parameters of the first request or the structural information of the request body are recorded in, for example, JSON format.
[0059] Here, for example, when Use case ID 133A is "UC-1", it is shown that the first operation related to the use case is "call the / userCheckout API endpoint of the Checkout microservice and execute it with the productID as a parameter using the GET method".
[0060] FIG. 7 is a configuration diagram showing a configuration example of the target value information of the monitoring items according to an embodiment of the present invention. In FIG. 7, the target value information 134 of the monitoring item is information generated by the monitoring device 3 as information including the availability target and the performance target that the API should achieve, and includes a target value ID 134A, an API endpoint 134B, a microservice 134C, a measurement period 134D, a threshold value 134E, and a percentage target 134F.
[0061] The target value ID 134A is an identification number for identifying the target value and measurement method of the monitoring indicators of the API (API to be monitored) that is the monitoring target of the monitoring device 3. For example, information such as "SLO-1" is recorded in the target value ID 134A as the first of the SLO (Service Level Objective) "service level target".
[0062] The API endpoint 134B is the API endpoint of the API to be monitored by the monitoring device 3. For example, in the case of the API endpoint 31A, information such as " / userCheckout" is recorded in the API endpoint 134B.
[0063] The microservice 134C is the name of the microservice to which the API to be monitored belongs. For example, in the case where the microservice to which the API to be monitored belongs is microservice 31, information such as "Checkout" is recorded in the microservice 134C.
[0064] The measurement period 134D is the calculation period of the monitoring indicators of the API to be monitored (for example, the processing time required for request processing) necessary for calculating the percentage target 134F. For example, in the case where the calculation period is one month, information such as "for one month" is recorded in the measurement period 134D.
[0065] The threshold value 134E is a reference value for determining whether the monitoring indicators of the API to be monitored are normal or abnormal. For example, in the case where the processing time required for request processing is used as the monitoring indicator of the API to be monitored, information such as "500ms" is recorded in the threshold value 134E. At this time, if the processing time required for request processing is "500ms" or less, it is determined to be normal, and if it exceeds "500ms", it is determined to be abnormal.
[0066] The percentage target 134F is a target value indicating the ratio of the measured values determined to be normal among all the measured values obtained during a certain period as the measured values of the API to be monitored. For the percentage target 134F, for example, when the target is that 99% of all the measured values are normal, information of "99%" is recorded.
[0067] Here, for example, when the target value ID134A is "SLO-1", as the target value, it is shown that "among all the requests received by the / userCheckout API endpoint of the Checkout microservice within one month, requests that ended within 500 ms or less account for 99% or more".
[0068] Figure 8 is a configuration diagram showing a configuration example of trace data according to an embodiment of the present invention. In Figure 8, the trace data 135 includes data collected by the monitoring device 3 from the monitoring target system 1 during the operation of the monitoring target system 1. The trace data 135 manages a URI 135A, a parameter 135B, a method 135C, a request body 135D, a status code 135E, a response body 135F, a start time 135G, an end time 135H, a TraceID 135I, a SpanID 135J, and a ParentID 135K.
[0069] The URI 135A is the path of the specified endpoint when executing a test case. The URI 135A corresponds to the URI 132B in Figure 5. The parameter 135B is additional information expressed in a specific format at the end of the path specifying the destination of the request, and corresponds to the parameter 132C in Figure 5.
[0070] The method 135C represents the type of request from the client to the server using the communication protocol HTTP. The method 135C corresponds to the method 132D in Figure 5. The request body 135D is the information sent from the client to the server. The request body 135D corresponds to the request body 132E in Figure 5.
[0071] Status code 135E is a code returned to the client upon response from the server in the communication protocol HTTP. Status code 135E corresponds to status code 132F in FIG. 5. Status code 135E is, for example, a status code indicating a normal response when the response from the server to the client in HTTP is normal, and information of "200" is recorded. On the other hand, when the response is not normal, information other than "200" (for example, "400"), which is a status code indicating an abnormal response, is recorded.
[0072] Response body 135F is response information sent from the server to the client. Response body 135F corresponds to response body 132G in FIG. 5. In response body 135F, for example, information of "{"result":"0"}" sent from the server to the client as a return value of an endpoint call is recorded.
[0073] Start time 135G is the time when a specific request is sent to the server. Start time 135G corresponds to start time 132H in FIG. 5. Note that information of "2022-11-07 01:30:08.3226328" is recorded in the start time 135G of the first record.
[0074] End time 135H is the time when the server completes the process and sends a response to the client. End time 135H corresponds to end time 132I in FIG. 5. Note that information of "2022-11-07 01:30:08.8826696" is recorded in the end time 135H of the first record. The time from start time 135G to end time 135H is the processing time required for the client (operation step or microservice) to process the request. Also, start time 135G and end time 135H can record time on the order of ms as time less than a second.
[0075] TraceID135I is an identifier that identifies a series of requests when a client requests an API endpoint. TraceID135I corresponds to TraceID132J in Figure 5. Note that the information "4kzkn1y" is recorded in TraceID135I of the first record.
[0076] SpanID135J is an identifier that identifies the processing content of an endpoint when a request is sent from a client to an API endpoint. SpanID135J corresponds to SpanID132K in Figure 5. Note that the information "aj07npa" is recorded in SpanID135J of the first record.
[0077] ParentID135K is an identifier that identifies the processing of the previous request that directly requested the API endpoint from the client among the requests when a request is sent from the client to the API endpoint. ParentID135K corresponds to ParentID132L in Figure 5.
[0078] Here, for example, in the case of the first record, it is shown that "the operation on the / userCheck API endpoint starting from 2022-11-07 01:30:08.3226328 was executed with the GET method with the parameter proeudctID=ESPC7Z and ended at 2022-11-07 01:30:08.8826696. The TraceID of this operation is 4kzkn1y and the SpanID is aj07npa".
[0079] Figure 9 is a flowchart showing an example of the procedure of the business impact range presentation process according to an embodiment of the present invention. The business impact range presentation process is executed by the business impact range presentation device 100, for example, when an execution instruction for the business impact range presentation process is received from the display device 5 via the input unit 110. Here, first, an overview of the business impact range presentation method using the business impact range presentation device 100 will be described.
[0080] When the business impact scope indication method is used, the test log collection unit 141 collects test record information indicating the execution results of tests on monitored software in which one or more operation steps indicating the constituent elements of the content of each of a plurality of operations and a plurality of microservices that execute processing by the operations of the one or more operation steps are connected via one or more request paths, and also collects test case information indicating the relationship between the content of each operation and each operation step, which is a test log collection step; the use case association unit 142 generates use case association information 133 by associating the request path corresponding to each operation step among the one or more request paths with the content of each operation based on the test record information 132 and the test case information collected in the test log collection step, which is a use case association step; the monitoring information acquisition unit 143 acquires trace data indicating the execution results of the respective microservices as the execution results in the production environment for the monitored software, which is a monitoring information acquisition step; the trace data shaping unit 146 determines whether there is a redundant request path in the normal operation in the above-mentioned request path, and if such a redundant request path exists, creates trace data 135 excluding the redundant request path, which is a trace data shaping step; the business impact scope analysis unit 144 determines whether an abnormality has occurred in any one of the one or more request paths based on the trace data 135 acquired in the monitoring information acquisition step, which is a business impact scope analysis step. In the business impact scope analysis step, when the business impact scope analysis unit 144 determines that an abnormality has occurred in any one of the request paths, it refers to the use case association information 133 based on the one of the request paths excluding the redundant request path, and identifies the content of the operations affected by the request path in which the abnormality has occurred among the content of each operation.
[0081] Next, the business impact scope presentation process will be specifically described. When the process of identifying the impact scope by the business impact scope presentation device 100 is started, the business impact scope presentation device 100 collects test case information 131 from the test device 2 (step S101). Specifically, the test log collection unit 141 collects the test case information 131 from the test device 2 in order to collect the information defining the relationship between the test case and the use case, and stores the collected test case information 131 in the storage unit 130. At this time, for example, the test content is recorded for each of the three test cases ("TC-1", "TC-2", "TC-3") included in the use case of "UC-1" in the test case information 131.
[0082] Next, the business impact scope presentation device 100 collects the execution records of the tests performed by the test device 2 on the monitored system 1 (step S102). Specifically, the test log collection unit 141 collects the records of the monitoring data (trace data, etc.) collected from the monitoring device 3 during the test execution period to generate test record information 132. At this time, the test log collection unit 141 includes the test case information (test case ID) in the test record information 132 including the test log in order to match the test execution content with the monitoring data (trace data) by time (timestamp). For example, in the case of the test case of "TC-2" in the test record information 132, when a request is made to the / createOrder API endpoint of the microservice Order service, the request body includes "productID" and "userID", and a feature (API request path) indicating that the request is executed via a plurality of microservices and API endpoints is recorded.
[0083] Here, the method of matching the test case with the monitoring data (trace data) of the test execution result is not limited to the exemplified method. For example, the test case and the monitoring data (trace data) of the test execution result may be specified by the URI of the monitored API endpoint.
[0084] Next, the business impact scope presentation device 100 associates the test execution records recorded in the test record information 132 with the use cases (step S103). Specifically, the use case association unit 142 groups all API requests related to a specific use case by TraceID from the information recorded in the test case information 131 and the information (test log) recorded in the test record information 132, extracts the characteristics of the request group for each TraceID, and generates use case association information 133. For example, based on the API request (TraceID: 55c94b0d) of the test case "TC-2" related to the use case "UC-1", the use case association unit 142 extracts the API request path (the path that completes the process via the API / createOrder of the order service, the API / calculatePrice of the calculate service, the API / getDiscount of the discount service, and the / inventoryCheck of the inventory service) and the request content (the payload contains the productID and userID), and generates use case association information 133 in the format shown in FIG. 6 from the extracted API request path and request content. ate service's API / calculatePrice, discount service's API / getDiscount, inventory service's / inventoryCheck to complete the process) and the request content (the payload contains productID and userID), and generates use case association information 133 in the format shown in FIG. 6 from the extracted API request path and request content.
[0085] Here, the data items and extraction methods that are the characteristics of the request are not limited to the exemplified methods. For example, if the request can be uniquely identified, feature quantities other than the API request path and request content may be adopted.
[0086] Next, the business impact scope presentation device 100 collects monitoring data of the production environment from the monitoring device 3 (step S104). Specifically, the monitoring information acquisition unit 143 collects monitoring data including the characteristics (URI, parameters, method, request body, status code, response body, start time, end time, TraceID, SpanID, ParentID) of each API request from the monitoring device 3, generates trace data 135 in the format shown in FIG. 8, and stores the generated trace data 135 in the storage unit 130.
[0087] Next, the business impact scope presentation device 100 identifies the APIs that do not meet the target values (step S105). Specifically, the business impact scope analysis unit 144 determines whether an API fails to meet the target value based on the ratio of requests whose expected performance is below the threshold during a certain period. For example, when the business impact scope analysis unit 144 monitors the movement of the monitored system 1 and discovers that the performance of the API " / calculatePrice" is below the target value, it analyzes the alert information or monitoring data from the monitoring device 3 to identify " / calculatePrice".
[0088] Here, the method for identifying the APIs that do not meet the monitoring target values is not limited to the exemplified method. For example, depending on the nature of the monitored system 1, each development project may use its own judgment criteria.
[0089] Next, the business impact scope presentation device 100 extracts the requests that passed through the APIs that do not meet the target values (step S106). Specifically, the business impact scope analysis unit 144 extracts all the requests that passed through the API endpoints targeted by the alerts based on the trace data, and among all the extracted requests, identifies the requests affected by the threshold of the target value information 134 of the monitoring items. For example, from the trace data, all requests whose URIs passed through the API endpoint of / calculatePrice are searched. Among all the requests whose URIs passed through the API endpoint of / calculatePrice, the requests whose processing time of / calculatePrice exceeds the threshold value of 500 ms (threshold value 134E recorded in the target value information 134 of the monitoring items) are identified and recorded as the affected requests (requests whose processing time exceeds the threshold value).
[0090] Next, the business impact scope presentation device 100 identifies the use cases affected by the characteristics of the request (API request) (step S107). Specifically, the business impact scope analysis unit 144 compares the characteristics of the API request (trace data 135) with the characteristics of the API request recorded in the use case association information 133 (API request path 133B and request content 133C), and determines that the use case in which the contents of both exactly match is "affected". For example, referring to the trace data 135 which is the characteristic of the API request, if the processing time of / calculatePrice exceeds the threshold value of 500 ms in the request record with TraceID "ythy6f0" in the trace data 135, it can be seen that this request is a request affected.
[0091] Here, the method of comparing the characteristics of API requests is not limited to the exemplified method. For example, clustering using the feature quantities of API requests and displaying them in the order of distance may be used. Also, since a strict judgment criterion may underestimate the impact range, the search method may be determined according to the required detection accuracy.
[0092] By the way, in the API request specified by TraceID "ythy6f0" in the trace data 135, an error has occurred in the request recorded in the fourth record, and a retry operation of the request (the request recorded in the fifth record) has occurred. If the characteristics of the API request in the trace data 135 are compared with the characteristics of the API request recorded in the use case association information 133 as it is, since the trace data contains redundant requests, the use case that should originally match cannot be detected. Therefore, in this embodiment, before comparing the trace data with the use case association information, redundant requests for normal requests such as retries caused by error occurrences during requests from the client to the server are detected, and a process of excluding the redundant requests from the trace data (hereinafter referred to as "trace data shaping process") is performed. Details of the trace data shaping process will be described later.
[0093] In step S107 described above, the information on the use cases affected by the identified request is sent to the display (display unit) of the business impact scope presentation device 100 and also sent to the display device 5 via the communication unit 150. The display (display unit) displays the information on the use cases affected by the identified request as will be described later. Thereafter, the business impact scope presentation device 100 ends the processing in this routine.
[0094] FIG. 10 is a flowchart showing an example of the procedure of the trace data shaping process. The trace data shaping process is executed by the trace data shaping unit 146. As described above, the trace data shaping unit 146 determines whether there is a redundant request path in the request path during normal operation. If there is a redundant request path, the trace data 135 excluding the redundant request path is created. This will be specifically described below.
[0095] First, the trace data shaping unit 146 extracts API requests that have not responded normally from the API request paths recorded in the trace data 135 (step S201). Specifically, the trace data shaping unit 146 searches for and extracts API requests with an error in the request result among the API requests recorded in the trace data 135.
[0096] Here, as a method for determining that an API request is an error, examples include (A) a method in which an error is considered when the status code 135E of the HTTP request recorded in the status code 135E of the trace data 135 is other than the normal value (code 200), and (B) a method in which an error is considered when the status code 135E is an error code defined by the monitored software.
[0097] Also, as a method for determining that an API request is an error, a method of determining based on the content of the response body 135F of the (C) trace data 135 can be exemplified. In this method, for example, there are two methods. First, as the first method, it is determined using the character string (word) included in the response body 135F. For example, if it is OK / True, it is determined to be normal, while if it is NG / ERROR / false, it is determined to be abnormal. As the second method, the response body 135F is analyzed, and it is determined from the values corresponding to keys such as result / code / status. For example, if result = “true” / code = “OK” / status = “OK”, it is determined to be normal. Also, if result = “false” / code = “NG” / status = “ERR”, it is determined to be abnormal.
[0098] Furthermore, as a method for determining that an API request is an error, a method can be exemplified in which (D) when the status code 132F of the API request recorded in the test record information 132 or the response body 132G is regarded as information of a normal request and the corresponding information (status code 135E, response body 135F) of the API request in the trace data 135 does not match, it is determined as an error.
[0099] In addition, as a method for determining that an API request is an error, a method can be exemplified in which (E) the characteristics of a normal request are extracted from the information of the API request from past trace data, and an API request that does not match is determined as an error. For example, in a call from [API-A1] of microservice A to [API-B1] of microservice B, if the frequency of the response body 135F of the request being “result = 0” is high in the past trace data 135, it is determined that the response body 135F of a normal request has the characteristic of “result = 0”.
[0100] In addition, as a method for determining that an API request is an error, for example, (F) a method of extracting a pattern of the API request path from the past trace data 135, comparing it with a similar path, and excluding an API request that does not match as an error can be exemplified. For example, API [A1] of microservice A → API [B1] of microservice B → API [C1] of microservice C, API [A1] of the microservice A → API [B1] of the microservice B, API [A1] of the microservice A → API [B1] of the microservice B → API [C1] of the microservice C and API [B1] of the microservice B → API [C1] of the microservice C are called in this order. The path of the API request does not exist in the past API requests. Among the similar (called from [A1] of microservice A to [B1] of microservice B and from [B1] of microservice B to [C1] of microservice C) paths, if the pattern called by the path of API [A1] of microservice A → API [B1] of microservice B → API [C1] of microservice C is frequent, the non-matching calls (API [A1] of microservice A → API [B1] of microservice B → API [C1] of microservice C, API [A1] of the microservice A → API [B1] of the microservice B, and API [B1] of the microservice B → API [C1] of microservice C) are determined as errors and excluded (aligned with the pattern of past requests).
[0101] Using any one or all of these methods, it is determined whether the API request recorded in the trace data 135 is an error, and in the case of an error, it is extracted as an API request that has not had a normal response. For example, in the trace data 135, the API request recorded in the fourth record (the record where URI135A is " / inventoryCheck") is not a code 200 indicating a normal response (code 400) for the status code 135E according to the method of (A) above, and thus is extracted as an API request that has not had a normal response.
[0102] Next, the trace data shaping unit 146 excludes the API requests extracted as not having a normal response from the trace data 135 (step S202). Specifically, the API requests extracted in step S201 and the paths of the API requests further called from the corresponding API requests are deleted from the record of the trace data 135. For example, in the trace data 135, the API request recorded in the fourth record (the record where URI135A is " / inventoryCheck") was extracted as an API request that did not have a normal response in step S201, so it is deleted from the trace data 135. Here, in the monitored software, there are no API requests to other microservices from " / inventoryCheck", that is, there is no record in which the SpanID "4jb2tnu" of " / inventoryCheck" is recorded in ParentID135K. Therefore, the process is completed by deleting the record of " / inventoryCheck" from the trace data 135. However, if there is an API request path to other microservices, the parent-child relationship of the API requests is traced from SpanID135J and ParentID135K, and all API requests included in the path starting from the API request that did not have a normal response are deleted from the trace data 135.
[0103] Through the above processing, it is possible to detect redundant requests for normal requests, such as retries caused by errors when making requests from the client to the server, and exclude them from the trace data. By doing so, when an abnormality occurs in the request path including microservices in the production environment, it is possible to more accurately identify the content of the business affected by the abnormality of the request path.
[0104] FIG. 11 is a configuration diagram showing an example of a display screen of a display device according to an embodiment of the present invention. In FIG. 11, the display screen 500 of the display device 5 is a display screen starting from a use case. A plurality of use cases 501, 502,... are displayed on the display screen 500. A use case list is displayed for each use case. At this time, the use case 501 that requires particular attention is highlighted. In the area adjacent to the use case 501, microservices 511, 512,..., 521, 522,..., 531 are displayed, and API endpoints 511A, 512A,..., 521A, 522A,..., 531A belonging to each microservice are displayed. The use case 501, each of the microservices 511 to 531, and each of the API endpoints 511A to 531A are displayed in a tree structure based on use case association information 133.
[0105] Here, the display of the display device 5 or the business impact range presentation device 100 functions as a display unit that displays the components of the monitored software managed by the monitored system 1 and adjusts the displayed content based on the analysis result of the business impact range analysis unit 144. This display unit, for example, highlights and displays the request path that fails to reach the target value and the use case affected by the request path where an abnormality has occurred among the components of the monitored software. That is, the API endpoint 521A including the API that fails to reach the target value is highlighted and displayed together with the use case 501.
[0106] Also, the API request path including the use case 501, the microservice, and the API endpoint is displayed in a tree structure via arrows. The content of the API request is displayed in a tooltip on each API point. When the highlighted API endpoint 521A is clicked, the related information 541 is displayed. As the related information 541, for example, the end time, alert, HTTP method, request content, etc. are displayed. Note that the API monitoring threshold, target value information, and trace data of requests related to a specific API can also be displayed in the related information 541. Further, the content of the image information displayed on the display screen 500 is provided to the operation administrator as investigation information.
[0107] In this embodiment, the display device 5 displays the components of the software to be monitored including, for example, the redundant request paths excluded as described above. By doing so, it is possible to visually recognize which part the redundant request path was.
[0108] Also, in this embodiment, the display device 5 displays, for example, the excluded redundant request path in a visually prominent manner, for example, highlighted. By doing so, it is possible to make it easier to more visually recognize which part the redundant request path was.
[0109] FIG. 12 is a configuration diagram showing an example of another display screen of the display device according to an embodiment of the present invention. In FIG. 12, the display screen 550 of the display device 5 is a display screen starting from microservices. A plurality of microservices 551, 552, 553, 554, 555 are displayed on the display screen 550. A list of microservices of the monitoring target system 1 is displayed for each of the microservices 551 to 555. In the area adjacent to each of the microservices 551 to 555, a plurality of API endpoints 551A, 551B, ···, 552A, 552B, ···, 553A, ···, 554A, 554B, 554C, ···, 555A, 555B, ··· are displayed, and a plurality of use cases 561, 562, ··· are displayed. Each of the microservices 551 to 555 and each API endpoint are displayed in a tree structure based on the use case association information 133. The API request path related to the use case is displayed, for example, by an arrow 571 connecting between the API endpoints.
[0110] The API endpoint 553A that does not reach the target value is highlighted and displayed. Also, the use case 561 affected by the API endpoint 553A that does not reach the target value is also highlighted and displayed. When the highlighted API endpoint 553A is clicked, related information 581 is displayed. For the related information 581, for example, the end time, alert, HTTP method, request content, etc. are displayed. Note that it is also possible to display the monitoring threshold value of the API, the target value information, and the trace data of the request related to a specific API in the related information 581. In addition, the content of the image information displayed on the display screen 550 is provided to the operation administrator as investigation information.
[0111] According to this embodiment, when an abnormality occurs in the API request path including microservices, it is possible to identify the use cases affected by the abnormality of the API request path. Further, according to this embodiment, even when the release of microservices is fast, it is possible to easily identify the use cases affected by the abnormality of the API request path. Furthermore, according to this embodiment, it is possible to accurately grasp the characteristics of the API request from the test log (test record information), and based on the grasped content, associate the API endpoint that fails to reach the target value with the use case. As a result, by visualizing the API endpoints that fail to reach the target value related to important use cases, proactive measures can be taken. For example, when the performance of the API endpoint associated with the use case deteriorates, by improving or optimizing the API endpoint, the important use case can be operated normally. Since measures can be actively taken to achieve the business goals, as a result, it can contribute to an improvement in the contribution degree of the IT system to the business.
[0112] Moreover, according to this embodiment, when a retry operation due to an error or the like occurs in the API request path in the production environment, the redundant API request path is surely excluded from the trace data 135 as described above. Therefore, even when the API request path is extracted from the trace data 135 in the production environment when a retry operation occurs, the trace data 135 will not erroneously fail to match the test record information 132 that should originally match, so that the use case corresponding to the API request path can be correctly extracted.
[0113] In this embodiment, the trace data 135 includes the processing time required for processing each request by each operation step or each microservice, including the processing time of each request in each request path. The business impact scope analysis unit 144 refers to the trace data 135 to compare the processing time of each request with a set threshold value. When the processing time of any request exceeds the threshold value, it is determined that an abnormality has occurred, and the request path of the request whose processing time exceeds the threshold value is analyzed as a request path where an abnormality has occurred and the target value has not been reached.
[0114] In this embodiment, the business impact scope analysis unit 144 refers to the use case association information 133 based on the request path where the target value has not been reached, and extracts all request paths including the request path where the target value has not been reached from among one or more request paths as the request paths to be analyzed.
[0115] In this embodiment, the business impact scope analysis unit 144 determines whether the request path to be analyzed exists in the trace data. If it is determined that the request path to be analyzed does not exist in the trace data 135, the content of the business related to the microservice connected to the request path to be analyzed among the contents of each business is analyzed as the content of the business that may be affected by the request path to be analyzed. If it is determined that the request path to be analyzed exists in the trace data 135, the content of the business related to the microservice connected to the request path to be analyzed among the contents of each business is analyzed as the content of the business affected by the request path to be analyzed.
[0116] The business impact scope presentation device 100 according to this embodiment further includes a display unit that displays the components of the software to be monitored and adjusts the displayed content based on the analysis result of the business impact scope analysis unit 144. The display unit emphasizes and displays the components of the software to be monitored and the content of the business affected by the request path where the target value has not been reached and the request path where an abnormality has occurred.
[0117] The business impact scope presentation device 100 according to this embodiment further includes a display unit that displays the components of the software to be monitored and adjusts the displayed content based on the analysis results of the business impact scope analysis unit. The display unit emphasizes and displays the content of the business that may be affected by the request path to be analyzed among the components of the software to be monitored and the content of the business affected by the request path to be analyzed.
[0118] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, each of the elements described in parallel in this embodiment may be in a mode in which at least one of the elements is connected in series to another element.
Industrial Applicability
[0119] The present invention can be applied to, for example, a business impact scope presentation device related to a technique for presenting the scope of impact on a business.
Explanation of Reference Numerals
[0120] 1... System to be monitored, 2... Test device, 3... Monitoring device, 4... Network, 5... Display device, 100... Business impact scope presentation device, 110... Input unit, 120... Output unit, 130... Storage unit, 131... Test case information, 132... Test record information, 133... Use case association information, 134... Monitoring item and target value information, 135... Trace data, 136... Update information, 140... Arithmetic unit, 141... Test log collection unit, 142... Use case association unit, 143... Monitoring information acquisition unit, 144... Business impact scope analysis unit, 146... Trace data formatting unit, 150... Communication unit
Claims
1. A test log collection unit that collects test record information indicating the execution results of tests on monitored software in which one or more operation steps indicating the constituent elements of the content of each of a plurality of operations and a plurality of microservices that execute processing by the operations of the one or more operation steps are connected via one or more request paths, and collects test case information indicating the relationship between the content of each operation and the one or more operation steps, A use case association unit that generates use case association information by associating the request paths corresponding to the respective operation steps among the one or more request paths with the content of each operation based on the test record information and the test case information collected by the test log collection unit, A monitoring information acquisition unit that acquires trace data indicating the execution results of the respective microservices as the execution results in the production environment for the monitored software, A trace data shaping unit that determines whether there is a redundant request path in normal operation in the request path, and creates the trace data excluding the redundant request path if the redundant request path exists, A business impact scope analysis unit that determines whether an abnormality has occurred in any of the one or more request paths based on the trace data acquired by the monitoring information acquisition unit, Comprising, The business impact scope analysis unit, When it is determined that the abnormality has occurred in any of the request paths, referring to the use case association information based on any of the request paths excluding the redundant request path, and specifying the content of the operations affected by the request path in which the abnormality has occurred among the content of each operation A business impact scope presentation device characterized by the above.
2. The business impact scope presentation device according to Claim 1, wherein The trace data, Includes the processing time of each request in each request path as the processing time required for processing each request by each operation step or each microservice, The business impact scope analysis unit, Referring to the trace data, compare the processing time of each request with a set threshold value. When the processing time of any request exceeds the threshold value, it is determined that the abnormality has occurred, and the request path of the request whose processing time exceeds the threshold value is analyzed as the request path where the abnormality has occurred and the request path that fails to reach the target value, which is a business impact scope presentation device.
3. The business impact scope presentation device according to claim 2, wherein the business impact scope analysis unit refers to the use case association information based on the request path that fails to reach the target value, and extracts all the request paths including the request path that fails to reach the target value from the one or more request paths as the request paths to be analyzed, which is a business impact scope presentation device.
4. The business impact scope presentation device according to claim 3, wherein the business impact scope analysis unit determines whether the request path to be analyzed exists in the trace data. When it is determined that the request path to be analyzed does not exist in the trace data, analyze the content of the business related to the microservice connected to the request path to be analyzed among the contents of each business as the content of the business that may be affected by the request path to be analyzed. When it is determined that the request path to be analyzed exists in the trace data, analyze the content of the business related to the microservice connected to the request path to be analyzed among the contents of each business as the content of the business affected by the request path to be analyzed, which is a business impact scope presentation device.
5. The business impact scope presentation device according to claim 1, further comprising a display unit that displays the components of the software to be monitored and adjusts the displayed content based on the analysis result of the business impact scope analysis unit, wherein the display unit displays the components of the software to be monitored including the redundant request paths excluded, which is a business impact scope presentation device.
6. The business impact scope presentation device according to claim 2, further comprising a display unit that displays the components of the software to be monitored and adjusts the displayed content based on the analysis result of the business impact scope analysis unit, wherein the display unit A business impact scope presentation device characterized by emphasizing and presenting the content of the business affected by the request path that fails to meet the target value and the request path where the abnormality has occurred among the components of the software to be monitored.
7. The business impact scope presentation device according to claim 4, further comprising a display unit that displays the components of the software to be monitored and adjusts the displayed content based on the analysis result of the business impact scope analysis unit, wherein the display unit is a business impact scope presentation device characterized by emphasizing and presenting the content of the business that may be affected by the request path to be analyzed and the content of the business affected by the request path to be analyzed among the components of the software to be monitored.
8. A test log collection step in which a test log collection unit collects test record information indicating the execution result of a test on software to be monitored in which one or more operation steps indicating the components of the content of each of a plurality of operations and a plurality of microservices that execute processing by the operations of the one or more operation steps are connected via one or more request paths, and also collects test case information indicating the relationship between the content of each operation and the one or more operation steps, a use case association step in which a use case association unit generates use case association information by associating the request path corresponding to each operation step among the one or more request paths with the content of each operation based on the test record information and the test case information collected in the test log collection step, a monitoring information acquisition step in which a monitoring information acquisition unit acquires trace data indicating the execution result of each microservice as the execution result in the production environment for the software to be monitored, a trace data shaping step in which a trace data shaping unit determines whether there is a redundant request path in normal operation in the request path, and creates the trace data with the redundant request path excluded if the redundant request path exists, a business impact scope analysis step in which a business impact scope analysis unit determines whether an abnormality has occurred in any of the one or more request paths based on the trace data acquired in the monitoring information acquisition step, and in the business impact scope analysis step, When the business impact scope analysis unit determines that the abnormality has occurred in any of the request paths, it refers to the use case association information based on any of the request paths excluding the redundant request path, and identifies the content of the business affected by the request path where the abnormality has occurred among the contents of each business. A method for presenting a business impact scope, characterized by the above.
9. The method for presenting a business impact scope according to Claim 8, wherein The trace data includes the processing time required for processing each request by each operation step or each microservice, including the processing time of each request in each request path. In the business impact scope analysis step, the business impact scope analysis unit refers to the trace data, compares the processing time of each request with a set threshold value, and determines that the abnormality has occurred when the processing time of any request exceeds the threshold value. The request path of the request whose processing time exceeds the threshold value is analyzed as the request path where the abnormality has occurred and the target value has not been reached. A method for presenting a business impact scope, characterized by the above.
10. The method for presenting a business impact scope according to Claim 9, wherein In the business impact scope analysis step, the business impact scope analysis unit refers to the use case association information based on the request path where the target value has not been reached, and extracts all the request paths including the request path where the target value has not been reached from among the one or more request paths as the request paths to be analyzed. A method for presenting a business impact scope, characterized by the above.
11. The method for presenting a business impact scope according to Claim 10, wherein In the business impact scope analysis step, The business impact scope analysis unit determines whether the request path of the analysis target exists in the trace data. When it is determined that the request path of the analysis target does not exist in the trace data, the content of the business related to the microservice connected to the request path of the analysis target among the contents of each business is analyzed as the content of the business that may be affected by the request path of the analysis target. When it is determined that the request path of the analysis target exists in the trace data, the content of the business related to the microservice connected to the request path of the analysis target among the contents of each business is analyzed as the content of the business affected by the request path of the analysis target. A method for presenting the business impact scope is characterized by the above.
12. The method for presenting the business impact scope according to claim 8, while displaying the components of the monitored software, the display unit that adjusts the displayed content based on the analysis result of the business impact scope analysis unit displays the components of the monitored software including the excluded redundant request path. A method for presenting the business impact scope is characterized by the above.
13. The method for presenting the business impact scope according to claim 9, while displaying the components of the monitored software, the display unit that adjusts the displayed content based on the analysis result of the business impact scope analysis unit emphasizes and displays the content of the business affected by the request path that fails to reach the target value and the request path where an abnormality has occurred among the components of the monitored software. A method for presenting the business impact scope is characterized by the above.
14. The method for presenting the business impact scope according to claim 11, while displaying the components of the monitored software, the display unit that adjusts the displayed content based on the analysis result of the business impact scope analysis unit emphasizes and displays the content of the business that may be affected by the request path of the analysis target and the content of the business affected by the request path of the analysis target among the components of the monitored software. A method for presenting the business impact scope is characterized by the above.
Citation Information
Patent Citations
Presentation device, presentation method and presentation program
JP2020160567A