Service monitoring system, and service monitoring method

The service monitoring system addresses the challenge of tracing services across cloud and edge environments by using tracing units to determine the end point of processing, ensuring reliable and centralized management of service traces.

JP2025104133APending Publication Date: 2025-07-09HITACHI BUILDING SYST CO LTD

Patent Information

Application Number
JP2023222001
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-09

Smart Images

  • Figure 2025104133000001_ABST
    Figure 2025104133000001_ABST
Patent Text Reader

Abstract

To provide a service monitoring system which can determine an end point of a service trace.SOLUTION: The present invention is directed to a service monitoring system having a first machine in a first dispersed tracing environment and a second machine in a second dispersed tracing environment, and for monitoring a service provided by the first machine and the second machine. A second cooperation unit discriminates a trace from trace information included in a request for a service of a monitoring target received from a first cooperation unit, and, if a second tracing unit determines an end point of processing relating to the service of the second machine, notifies the first cooperation unit of end point determination information indicating that the end point of the trace is determined.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to a technique for monitoring services.

Background Art

[0002] The development infrastructure of microservice architecture is expected to have effects such as facilitating cooperation between services through loose coupling and improving the maintainability of the system. However, there is a problem that the traceability of the system decreases due to microservice transformation.

[0003] In recent years, distributed tracing technology has been applied to improve the traceability of microservices, but it is not easy to uniquely associate complex request propagation information with the services and functions to be monitored. In particular, to trace the services provided by processing spanning cloud environments and edge environments, a unified distributed tracing environment between the cloud and the edge is required.

[0004] Regarding this problem, a technique has been disclosed in which complete request units and semi-open request units are generated based on the collected kernel events, and an accurate processing path is discovered by analyzing the generated request units based on the causal relationship set between the kernel events (see Patent Document 1).

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the technology described in Patent Document 1, there is a problem that it is impossible to determine the end point of a trace for a service that spans separated tracing environments such as the cloud and the edge.

[0007] The present invention has been made in consideration of the above points, and intends to propose a service monitoring system or the like that can determine the end point of a service trace.

Means for Solving the Problem

[0008] In order to solve such a problem, in the present invention, there is provided a service monitoring system including a first machine in a first distributed tracing environment and a second machine in a second distributed tracing environment, for monitoring a service provided by the first machine and the second machine. The first machine includes a first execution unit that executes processing related to the service to be monitored, a first tracing unit that traces the processing related to the service executed by the first execution unit and determines the end point of the processing in the first machine, and a first cooperation unit that relays a request for the service including trace information that can identify the trace of the service to the second machine. The second machine includes a second cooperation unit that receives a request for the service to be monitored from the first cooperation unit, a second execution unit that executes processing related to the service, and a second tracing unit that traces the processing related to the service executed by the second execution unit and determines the end point of the processing in the second machine. The second cooperation unit identifies the trace from the trace information included in the request for the service to be monitored received from the first cooperation unit, and when the end point of the processing related to the service in the second machine is determined by the second tracing unit, notifies the first cooperation unit of end point determination information indicating that the end point of the trace has been determined.

[0009] In the above configuration, when a service request is relayed from the first machine to the second machine and the end point of the trace of the service is determined in the second machine, the end point determination information is notified to the first machine. According to the above configuration, for example, the first machine and the second machine can grasp the end point of the trace of the service spanning the first machine and the second machine.

Advantages of the Invention

[0010] According to the present invention, a highly reliable service monitoring system can be realized. Other problems, configurations, and effects than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Embodiments for Carrying Out the Invention

[0012] (I) First Embodiment Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. This embodiment is merely an example for realizing the present invention and does not limit the technical scope of the present invention. Not all of the elements and their combinations described in this embodiment are essential for the solution means of the invention.

[0013] In the following description, the processing performed by the program may be described. The computer performs the processing defined by the program using the memory of the main storage device and the like by a processor (for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit)). Therefore, the subject of the processing performed by executing the program may be the processor. By the processor executing the program, a functional unit for performing the processing is realized.

[0014] Similarly, the subject of the processing performed by executing the program may be a controller, a device, a system, a computer, or a node including a processor. The subject of the processing performed by executing the program only needs to be an arithmetic unit and may include a dedicated circuit for performing a specific process. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or the like.

[0015] In the following description, the program may be installed in a computer from a program source. The program source may be, for example, a program distribution server or a non-transitory computer-readable storage medium. When the program source is a program distribution server, the program distribution server includes a processor and a storage resource (storage) for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Also, in an embodiment, two or more programs may be realized as one program, or one program may be realized as two or more programs.

[0016] In the following description, various data will be described in tabular form. However, the data format is not limited to the tabular form, and other data formats such as queues, lists, and CSV (Comma Separated Value) may be used.

[0017] The notations such as "first", "second", "third", etc. in this specification and the like are attached to identify components, and do not necessarily limit numbers or orders. Also, the numbers for identifying components are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Also, it does not prevent a component identified by a certain number from also having the functions of a component identified by another number.

[0018] Note that in the following description, for the same elements in the drawings, the same numbers are assigned and the description is omitted as appropriate. Also, when describing elements of the same type without distinction, the common part (excluding the branch number) of the reference signs including the branch number is used, and when describing elements of the same type separately, reference signs including the branch number may be used. For example, when describing the application execution unit without particular distinction, it may be described as "application execution unit 230", and when describing each application execution unit separately, it may be described as "application execution unit 230A", "application execution unit 230B", etc.

[0019] In FIG. 1, reference numeral 100 indicates a cloud and edge cooperation trace monitoring system according to the first embodiment as a whole. The cloud and edge cooperation trace monitoring system 100 includes a cloud platform 110 and an edge terminal 120. Communication between the cloud platform 110 and the edge terminal 120 is performed via a network such as the Internet 101.

[0020] The cloud platform 110 is connected to client terminals 130 such as user terminals and monitoring terminals via networks such as the Internet 102 and a communication network 103. The edge terminal 120 is connected to field devices 140 such as elevators, sensors, and access control systems via a network such as a local network 104.

[0021] In the cloud and edge cooperation trace monitoring system 100, a first distributed tracing environment is provided in the cloud platform 110, a second distributed tracing environment different from the first distributed tracing environment is provided in the edge terminal 120, and the operating state of the service is monitored. Such a service is a service that processes a request transmitted from the client terminal 130 (for example, a user terminal) in the cloud platform 110 and then processes it in the edge terminal 120 to operate and control the field device 140.

[0022] More specifically, in the cloud and edge cooperation trace monitoring system 100, an identifier (trace ID) unique within the system is issued at the first point where a request is received from the client terminal 130, and a trace log including the trace ID and additional information (annotation) such as a processing result in the cloud platform 110 is stored in a storage device. Also, in the cloud and edge cooperation trace monitoring system 100, the trace ID is propagated when a request is issued from the cloud platform 110 to the edge terminal 120, and a trace log including the trace ID and the annotation in the edge terminal 120 is stored in the storage device.

[0023] According to the above processing, in the service provided across the cloud platform 110 and the edge terminal 120, the trace log in the cloud platform 110 and the trace log in the edge terminal 120 are associated by the trace ID.

[0024] FIG. 2 is a diagram showing an example of the cloud platform 110. The cloud platform 110 includes a common database 210, a cloud-edge cooperation control unit 220, an application execution unit 230, a tracing function unit 240, a log database 250, a service request propagation information management unit 260, and a service operation determination unit 270.

[0025] FIG. 3 is a diagram showing an example of the edge terminal 120. The edge terminal 120 includes a common database 310, a cloud-edge cooperation control unit 320, an application execution unit 330, and a tracing function unit 340.

[0026] FIG. 4 is a diagram showing an example of the trace cooperation information tables 211 and 311. The trace cooperation information tables 211 and 311 are stored in the common databases 210 and 310. The trace cooperation information tables 211 and 311 include information on the trace ID issued when a service is requested, the edge terminal ID for identifying the edge terminal 120 that performs the processing related to the service, the trace end point determination flag indicating the end point of the trace of the service, and the deletion determination flag information for deleting the trace log associated with the trace ID.

[0027] FIG. 5 is a diagram showing an example of a trace log stored in the log database 250. The log database 250 stores trace logs 243 and 343 generated by the tracing function units 240 and 340. The trace log 243 includes information such as a trace ID issued when a service is requested, a span ID identifying a span indicating a process (segment) in the cloud platform 110 related to the service, a parent ID identifying the span that called the span, the start time of the span, and the end time of the span. The trace log 343 includes information such as a trace ID issued when a service is requested, a span ID identifying a span indicating a process (segment) in the edge terminal 120 related to the service, a parent ID identifying the span that called the span, the start time of the span, and the end time of the span.

[0028] In this embodiment, a span is a process for each endpoint that executes a process related to a service, for example, a process for each API (Application Programming Interface). An example in which one endpoint is included in one APP (Application Software) will be described.

[0029] FIG. 6 is a diagram showing an example of the service request propagation information table 262. The service request propagation information table 262 includes information such as a service name indicating the name of a service to be monitored, an endpoint identifier array indicating endpoints (for example, the application execution units 230 and 330) that execute processes related to the service, and an operation determination threshold value that is a threshold value for determining whether the service is operating normally.

[0030] In the following description, as a first embodiment, a service operation monitoring process (FIG. 7) will be described, which traces the processing related to a service (request) that processes a service request transmitted from a client terminal 130 (for example, a user terminal) by an application execution unit 230 in a cloud platform 110 and an application execution unit 330 in an edge terminal 120, and monitors the operating state of the service to operate and control field devices 140.

[0031] (Service operation monitoring process (FIG. 7) according to the first embodiment) FIG. 7 is a flowchart showing an example of a service operation monitoring process according to the first embodiment. As a premise of the service operation monitoring process, service request propagation information that uniquely associates a service to be monitored with request propagation information is registered and managed in a service request propagation information management unit 260. Further, the processing related to the services executed by each of the application execution units 230 and 330 is traced by tracing function units 240 and 340.

[0032] First, in step S710, a pair edge cooperation control unit 220 receives a request relay request for the edge terminal 120.

[0033] Next, in step S720, a pair edge request relay process (FIG. 8) is performed.

[0034] Next, in step S730, a pair cloud request relay process (FIG. 9) is performed.

[0035] Next, in step S740, the trace end determination units 242 and 342 in the tracing function units 240 and 340 determine the end point of the trace based on the fact that the trace context has been closed. Note that, for the technology of determining that the trace context has been closed, a known technology (for example, the function of OSS (Open Source Software)) may be used. Additionally, the trace end determination units 242 and 342 may determine the end point of the trace based on the fact that the proxy units 232 and 332 have sent instructions to other machines (such as the edge terminal 120 and the field device 140).

[0036] Next, in step S750, the trace end process (Figure 10) is performed.

[0037] Next, in step S760, the service operation determination process (Figure 11) is performed.

[0038] Next, in step S770, the trace link information registration / reference unit 222 in the edge connection control unit 220 deletes the trace link information of the service (trace) whose operation state has been determined in the trace link information table 211 stored in the common database 210. The trace link information registration / reference unit 322 in the cloud connection control unit 320 deletes the trace link information of the said service (trace) in the trace link information table 311 stored in the common database 310.

[0039] In the present embodiment, at least one application execution unit 230 operates in the cloud platform 110, and at least one application execution unit 330 operates in the edge terminal 120. In the present embodiment, three application execution units 230A, 230B, and 230C are described in the cloud platform 110, and three application execution units 330D, 330E, and 330F are described in the edge terminal 120. However, the number of the application execution units 230 and 330 may be other numbers.

[0040] The application execution units 230 and 330 include an application function unit 231 and 331 that perform actual processing of application functions, and a proxy unit 232 and 332 that mediates communication between the plurality of application execution units 230 and 330 and other system components etc.

[0041] The proxy units 232 and 332 cooperate with the tracing function units 240 and 340 respectively to perform tracing of requests processed through the application execution units 230 and 330. The tracing function units 240 and 340 include a trace information management unit 241 and 341 that manages trace information such as a trace ID which is an identifier uniquely indicating each trace, and a trace end point determination unit 242 and 342 that determines the end point of each trace, and generate a trace log 243 and 343 indicating the tracing result. The trace log 243 of the trace whose end point is determined by the trace end point determination unit 242 is stored from the tracing function unit 240 to the log database 250. The trace log 343 of the trace whose end point is determined by the trace end point determination unit 342 is stored to the log database 250 by the trace log transfer units 224 and 324 described later.

[0042] In the first embodiment, the edge cooperation control unit 220 and the cloud cooperation control unit 320 play a role of relaying requests from the application execution unit 230C in the cloud platform 110 to the application execution unit 330D in the edge terminal 120.

[0043] The edge cooperation control unit 220 includes a request relay unit 221, a trace cooperation information registration reference unit 222, a common database synchronization processing unit 223, and a trace log transfer unit 224. The request relay unit 221 receives a request from the application execution unit 230C, and the edge request relay process (Fig. 8) described later is performed, and the request is relayed to the edge terminal 120.

[0044] (Edge Request Relay Process (Fig. 8) According to the First Embodiment) FIG. 8 is a flowchart showing an example of the edge request relay process according to the first embodiment.

[0045] First, in step S810, the request relay unit 221 extracts an identifier (for example, edge terminal ID) that uniquely indicates the destination edge terminal 120 from the destination information of the request.

[0046] Next, in step S820, the request relay unit 221 extracts trace information (for example, trace ID) that uniquely identifies the trace from the header of the request.

[0047] Next, in step S830, the trace link information registration reference unit 222 registers the information extracted in steps S810 and S820 in the trace link information table 211 in the common database 210.

[0048] In step S840, the request relay unit 221 inserts the trace information extracted in step S820 into the request body of a new request to be sent to the edge terminal 120 after converting it into a format of a predetermined data format such as JSON. Since the trace information is inserted into the request body of the request, the trace information can be propagated even when the information in the header of the request is rewritten when passing through the network (IoT device).

[0049] In step S850, the request relay unit 221 sends the new request generated in step S840 to the edge terminal 120 via a network such as the Internet 101.

[0050] The cloud connection control unit 320 includes a request relay unit 321, a trace connection information registration / reference unit 322, a common database synchronization processing unit 323, and a trace log transfer unit 324. The request relay unit 321 receives a request transmitted from the request relay unit 221 of the edge connection control unit 220 in the cloud platform 110, performs a cloud request relay process (Fig. 9) described later, and relays the request to the application execution unit 330D in the edge terminal 120.

[0051] (Cloud request relay process (Fig. 9) according to the first embodiment) Fig. 9 is a flowchart showing an example of the cloud request relay process according to the first embodiment.

[0052] First, in step S910, the request relay unit 321 extracts trace information from the request body of the received request.

[0053] Next, in step S920, the trace connection information registration / reference unit 322 registers the trace information extracted in step S910 and an identifier uniquely indicating the edge terminal 120, such as the edge terminal ID, in the trace connection information table 311 in the common database 310.

[0054] In step S930, the request relay unit 321 inserts the trace information extracted in step S910 into the header of a new request to be transmitted to the application execution unit 330D.

[0055] In step S940, the request relay unit 321 transmits the new request generated in step S930 to the application execution unit 330D.

[0056] Upon receiving the new request, the application execution units 330D, 330E, and 330F execute processing and perform operations and controls on the field device 140. As a result, the trace end point determination unit 342 determines that the trace (target trace) of the service corresponding to the processing executed by the application execution units 330D, 330E, and 330F has reached the end point (step S740), and the trace end point processing (FIG. 10) described later is performed.

[0057] (Trace End Point Processing (FIG. 10) According to the First Embodiment) FIG. 10 is a flowchart showing an example of the trace end point processing according to the first embodiment.

[0058] First, in step S1010, the trace link information registration reference unit 322 registers the trace end point determination flag for the target trace in the trace link information table 311 as True.

[0059] Next, in step S1020, the trace log transfer unit 324 that has confirmed that the trace end point determination flag of the target trace is registered as True transfers the trace log 343 of the target trace to the log database 250.

[0060] Since the common databases 210 and 310 are periodically synchronized by the common database synchronization processing units 223 and 323, the trace end point determination flag of the target trace in the trace link information table 211 is also registered as True by the service request propagation information registration unit 261.

[0061] Therefore, the service operation determination unit 270 that has confirmed that the trace end point determination flag of the target trace in the trace link information table 211 is registered as True starts the service operation determination processing (FIG. 11) described later.

[0062] The service operation determination unit 270 includes a trace log reference unit 271, a service request propagation information reference unit 272, and a service operation state determination notification unit 273.

[0063] (Service operation determination process according to the first embodiment (FIG. 11)) FIG. 11 is a flowchart showing an example of a service operation determination process according to the first embodiment.

[0064] First, in step S1110, the trace log reference unit 271 refers to the trace log of the target trace stored in the log database 250, and extracts the identifiers (span IDs) of each endpoint that becomes the request path in the target trace as an endpoint array.

[0065] Next, in step S1120, the service request propagation information reference unit 272 refers to the service request propagation information table 262 of the service request propagation information management unit 260, and identifies the service corresponding to the target trace based on the endpoint array extracted in step S1110.

[0066] Next, in step S1130, the service operation state determination notification unit 273 determines the operation state of the service based on the operation determination threshold value of the service in the service request propagation information table 262. For example, if the time required for the service is within the operation determination threshold value of "10 seconds", the service operation state determination notification unit 273 determines that it is "normal", and if it exceeds "10 seconds", it determines that it is "abnormal".

[0067] In step S1140, the service operation state determination notification unit 273 notifies the operation state of the service determined in step S1130 to the client terminal 130 (for example, the monitoring terminal). An example of the screen displayed on the client terminal 130 (service operation monitoring information browsing screen 1200) is shown in FIG. 12.

[0068] In step S1150, the trace link information registration reference unit 222 registers the deletion determination flag of the target trace in the trace link information table 211 as True.

[0069] Since the common databases 210 and 310 are periodically synchronized by the common database synchronization processing units 223 and 323, the deletion determination flag for the target trace in the trace link information table 311 is also registered as True.

[0070] The trace link information registration reference units 222 and 322 periodically refer to the trace link information tables 211 and 311 and delete the trace link information for which the deletion determination flag is registered as True.

[0071] FIG. 12 is a diagram showing an example of a screen (service operation monitoring information viewing screen 1200) capable of viewing the operation status of the service to be monitored.

[0072] On the service operation monitoring information viewing screen 1200, for each service, information such as the operation status, trace ID, trace start time, trace end time, the endpoint that executed the process related to the service on the first machine, and the endpoint that executed the process related to the service on the second machine is displayed. In the present embodiment, the first machine is a cloud platform 110, a computer, etc. provided with a first distributed tracing environment that receives and processes requests related to the service to be monitored, and the second machine is an edge terminal 120, a computer, etc. provided with a second distributed tracing environment that receives and processes requests for the service from the first machine.

[0073] (Effect of the First Embodiment) In the first embodiment, the trace end determination unit 342 determines the end point of the trace within the edge terminal 120, and the service request propagation information registration unit 261 registers the trace end determination flag of the trace. Then, the common database synchronization processing units 223 and 323 synchronize the trace end determination flag of the trace, and the trace end determination flag of the trace can also be confirmed in the trace cooperation information table 211 within the cloud platform 110. Therefore, according to the first embodiment, all the machines that cooperate in determining the end point of the trace of the service spanning multiple machines can make the determination.

[0074] Also, in the first embodiment, when the trace log transfer units 224 and 324 confirm that the trace end determination flag of the trace is registered in the trace cooperation information tables 211 and 311, they transfer all the trace logs associated with the trace to the log database 250. Therefore, according to the first embodiment, it becomes possible to quickly collect and centrally manage the trace logs of the service spanning multiple machines.

[0075] Also, in the first embodiment, the service operation state determination notification unit 273 determines the operation state of the service by combining the trace log (log information) of the trace referred to by the trace log reference unit 271 and the service request propagation information (service information) specified by the service request propagation information reference unit 272, and notifies the result. Therefore, according to the first embodiment, it becomes possible to uniquely associate the service to be monitored with the request propagation information and perform service operation monitoring using the trace log.

[0076] Also, in the first embodiment, when the service operation determination unit 270 has completed the determination of the operation state of the service corresponding to the trace and the notification of the result, the deletion determination flag of the trace is registered in the trace link information table 211, and the common database synchronization processing units 223 and 323 synchronize the common databases 210 and 310 between machines. Therefore, according to the first embodiment, the trace link information of the trace that spans a plurality of machines and has been used for service operation monitoring can be deleted in all the machines that are linked.

[0077] (Hardware Configuration of Edge Terminal 120) FIG. 13 is a diagram showing an example of the hardware configuration of the edge terminal 120.

[0078] The edge terminal 120 includes a processor 1310 including a CPU interconnected via an internal communication line 1301 such as a bus, a main memory device 1320, an auxiliary storage device 1330, a network interface 1340, an input device 1350, and an output device 1360.

[0079] The processor 1310 controls the overall operation of the edge terminal 120. The main memory device 1320 is composed of, for example, a volatile semiconductor memory and is used as a work memory of the processor 1310. The auxiliary storage device 1330 is composed of a large-capacity non-volatile storage device such as a hard disk device, an SSD (Solid State Drive), or a flash memory, and is used to hold various programs and data for a long time.

[0080] The executable program 1331 stored in the auxiliary storage device 1330 is loaded into the main memory device 1320 when the edge terminal 120 is started or when necessary, and is executed by the processor 1310, thereby realizing each system that executes various processes.

[0081] Note that the executable program 1331 may be recorded on a non - transitory recording medium, read from the non - transitory recording medium by a medium reading device, and loaded into the main memory device 1320. Alternatively, the executable program 1331 may be acquired from an external computer via a network and loaded into the main memory device 1320.

[0082] The network interface 1340 is an interface device for connecting the edge terminal 120 to each network in the system or communicating with other computers. The network interface 1340 is composed of, for example, a NIC (Network Interface Card) such as a wired LAN (Local Area Network) or a wireless LAN.

[0083] The input device 1350 is composed of a keyboard, a pointing device such as a mouse, etc., and is used for the user to input various instructions and information to the edge terminal 120. The output device 1360 is composed of, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display, and an audio output device such as a speaker, and is used to present necessary information to the user when necessary.

[0084] (II) Second Embodiment As a second embodiment, there may be a case where a plurality of edge terminals 120 are placed and the edge environment becomes multi - stage. In this case, in order to efficiently manage the trace cooperation information between the edge terminals 120, the table configuration of the trace cooperation information tables 211, 311 is changed to that of the trace cooperation information tables 211A, 311A shown in FIG. 14, and identifiers that uniquely identify the edge terminals 120, such as edge terminal IDs, are held in multiple stages. The identifier of the edge terminal 120 that communicates with the cloud platform 110 is registered as the parent edge terminal ID, and the identifiers of the subsequent - stage edge terminals 120 are held as an array of child edge terminal IDs.

[0085] (Effect of the Second Embodiment) In the second embodiment, the trace link information of a plurality of edge terminals 120 is hierarchically stored in the trace link information tables 211A and 311A. Therefore, according to the second embodiment, the effects of the first embodiment can be obtained even when the edge environment has multiple levels.

[0086] The present invention is not limited to the above-described embodiments, and can be implemented using any components without departing from the gist thereof. The above-described embodiments and modifications are merely examples, and the present invention is not limited to these contents as long as the features of the invention are not impaired. Further, although various embodiments and modifications have been described above, the present invention is not limited to these contents. Other aspects conceivable within the scope of the technical idea of the present invention are also included in the scope of the present invention.

[0087] For example, a part of each function provided in each device of the above-described embodiment may be provided in another device, or a function provided in a separate device may be provided in the same device.

[0088] Also, the configuration of the program described in the above embodiment is an example. For example, a part of the program may be incorporated into another program, or a plurality of programs may be configured as one program.

[0089] (III) Supplementary Note The above-described embodiments include, for example, the following contents.

[0090] In the above-described embodiment, the case where the present invention is applied to a service monitoring system has been described. However, the present invention is not limited thereto, and can be widely applied to various other systems, devices, methods, and programs.

[0091] Also, in the above-described embodiment, the configuration of each table is an example. One table may be divided into two or more tables, or all or part of two or more tables may be one table.

[0092] In addition, in the above-described embodiments, the screens illustrated and described are merely examples, and any design may be used as long as the information presented is the same.

[0093] In addition, in the above-described embodiments, the output of information is not limited to display on a display. The output of information may be audio output by a speaker, output to a file, printing on a paper medium or the like by a printing device, projection on a screen or the like by a projector, or any other mode.

[0094] The above-described embodiments include, for example, the following characteristic configurations.

[0095] (1) A service monitoring system (e.g., cloud and edge cooperation trace monitoring system 100, service monitoring system) that includes a first machine (e.g., cloud platform 110, server device, mainframe, computer, system, virtual machine) of a first distributed tracing environment and a second machine (e.g., edge terminal 120, computer, system, virtual machine) of a second distributed tracing environment, and monitors services provided by the first machine and the second machine. The first machine includes a first execution unit (e.g., application execution unit 230) that executes processing related to a service to be monitored, a first tracing unit (e.g., tracing function unit 240) that traces the processing related to the service executed by the first execution unit and determines an end point of processing in the first machine, and a first cooperation unit (e.g., edge cooperation control unit 220) that relays a request for the service including trace information (e.g., trace ID) that can identify the trace of the service to the second machine. The second machine includes a second cooperation unit (e.g., cloud cooperation control unit 320) that receives a request for a service to be monitored from the first cooperation unit, a second execution unit (e.g., application execution unit 330) that executes processing related to the service, and a second tracing unit (e.g., tracing function unit 340) that traces the processing related to the service executed by the second execution unit and determines an end point of processing in the second machine. The second cooperation unit identifies a trace from the trace information included in the request for the service to be monitored received from the first cooperation unit, and when the end point of the processing related to the service in the second machine is determined by the second tracing unit, notifies the first cooperation unit of end point determination information (e.g., trace end point determination flag "True") indicating that the end point of the trace has been determined (e.g., refer to trace end point processing).

[0096] In the above configuration, when a service request is relayed from the first machine to the second machine and the end point of the trace of the service is determined in the second machine, the end point determination information is notified to the first machine. According to the above configuration, for example, the first machine and the second machine can grasp the end point of the trace of the service spanning the first machine and the second machine.

[0097] (2) A storage unit (for example, log database 250) that stores a trace log (for example, trace log 243) generated by tracing by the first tracing unit and a trace log (for example, trace log 343) generated by tracing by the second tracing unit is provided. When the first machine receives a service request from a client terminal (for example, client terminal 130, user terminal), the first machine issues trace information that can uniquely identify the trace of the service. The first tracing unit traces the process related to the service executed by the first execution unit to generate a trace log including the trace information, and stores the generated trace log in the storage unit. The second tracing unit traces the process related to the service executed by the second execution unit to generate a trace log including the trace information. When the second cooperation unit determines that the end point of the process in the second machine has been determined by the second tracing unit, the second cooperation unit transfers the trace log of the service generated by the second tracing unit to the storage unit (for example, refer to step S1020).

[0098] In the above configuration, in the first machine, trace information is issued, a trace log in the first machine including the trace information is generated, the trace information is linked to the second machine, and a trace log in the second machine including the trace information is generated in the second machine. In the above configuration, when the end point of the process in the second machine is determined for the service to be monitored, the trace log generated in the first machine and the trace log generated in the second machine are aggregated in the storage unit, so that the trace log of the end-to-end process related to the service spanning the first machine and the second machine can be centrally managed.

[0099] (3) The trace log generated by the first tracing unit includes the time when the first execution unit starts the process related to the service (for example, the start time of the span) and the time when the first execution unit ends the process related to the service (for example, the end time of the span). The trace log generated by the second tracing unit includes the time when the second execution unit starts the process related to the service (for example, the start time of the span) and the time when the second execution unit ends the process related to the service (for example, the end time of the span). A management unit (for example, the service request propagation information management unit 260) that manages service information (for example, the service request propagation information table 262) including information (for example, service name) that can identify the service to be monitored and threshold information (for example, the operation determination threshold) indicating the time threshold indicating that the service is operating normally, and reads out from the storage unit the trace log of the service to be monitored generated by the first tracing unit and the trace log of the service generated by the second tracing unit, calculates the time required for the service from the read trace log, determines whether the service is operating normally based on the threshold information included in the service information of the service, and outputs the determined result. A determination unit (for example, the service operation determination unit 270).

[0100] In the above service monitoring system, for example, the trace log generated by the first tracing unit includes endpoint information (e.g., span ID) that can identify the first execution unit, and the trace log generated by the second tracing unit includes endpoint information (e.g., span ID) that can identify the second execution unit. The service information of the service to be monitored managed by the management unit includes endpoint information (e.g., endpoint identification, span ID) that can identify the execution unit that executes the process related to the service. The determination unit reads from the storage unit the trace log of the service to be monitored generated by the first tracing unit and the trace log of the service generated by the second tracing unit, and identifies the service information that matches the endpoint information of the read trace log (see, for example, step S1120).

[0101] According to the above configuration, for example, the user can grasp whether the service spanning the first machine and the second machine is operating normally or not.

[0102] (4) The first machine includes a first storage unit (e.g., common database 210), the second machine includes a second storage unit (e.g., common database 310), the first machine includes a first processing unit that instructs the update of corresponding information (e.g., deletion determination flag) in the second storage unit when the information (e.g., deletion determination flag) in the first storage unit is updated, the second machine includes a second processing unit that instructs the update of corresponding information (e.g., trace end point determination flag) in the first storage unit when the information (e.g., trace end point determination flag) in the second storage unit is updated, the first storage unit stores cooperation information (e.g., trace cooperation information table 211) including trace information that can identify the trace of the service to be monitored, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined, the second storage unit stores cooperation information (e.g., trace cooperation information table 311) including trace information that can identify the trace, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined, the second cooperation unit updates the end point determination flag of the trace cooperation information to information (e.g., "True") indicating that the end point of the trace has been determined when the end point of the process in the second machine is determined by the second tracing unit (e.g., refer to step S1010), the second processing unit notifies the updated information to the first processing unit when there is an update to the trace cooperation information stored in the second storage unit, the first processing unit updates the trace cooperation information stored in the first storage unit to the notified information (e.g., "True"), the first cooperation unit updates the deletion determination flag of the trace cooperation information to information (e.g., "True") indicating that it is a deletion target when it is determined by the determination unit whether the service is operating normally (e.g., refer to step S1150), the first processing unit notifies the updated information to the second processing unit when there is an update to the trace cooperation information stored in the first storage unit, and the second processing unitUpdate the information (e.g., "True") notified of the linkage information of the trace stored in the first storage unit.

[0103] According to the above configuration, when the operating state of the service is determined, deletion information is notified from the first machine to the second machine, so that the unnecessary linkage information can be deleted on all the machines that perform linkage.

[0104] (5) The first linkage unit extracts machine information (e.g., edge terminal ID) that can identify the second machine from the destination of the request for the service to be monitored (see, for example, step S810), extracts trace information (e.g., trace ID) that can identify the trace of the service from the header of the request (see, for example, step S820), stores linkage information including the trace information, the machine information, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined in the first storage unit (see, for example, step S830), inserts the trace information into the body of the request for relaying to the second machine (see, for example, step S840), and the second linkage unit extracts the trace information from the body of the request received from the first linkage unit (see, for example, step S910), and stores linkage information including the trace information, machine information that can identify the second machine, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined in the second storage unit (see, for example, step S920).

[0105] (6) A plurality of the second machines are provided. One second machine (for example, edge terminal 120 "devulise_1") calls another second machine (for example, edge terminal 120 "devulise_2") to perform processing related to the service to be monitored (for example, refer to the second embodiment). In the cooperation information stored in the first storage unit (for example, trace cooperation information table 211A) and the cooperation information stored in the second storage unit (for example, trace cooperation information table 311A), there are trace information (for example, trace ID) that can identify the trace of the service to be monitored, machine information (for example, edge terminal ID (parent)) indicating the second machine that is the calling source, machine information (for example, edge terminal ID (child)) indicating the second machine that is the call destination, an end point determination flag (for example, trace end point determination flag) indicating whether the end point of the trace has been determined, and a deletion determination flag (for example, deletion determination flag) indicating whether the operating state of the service has been determined.

[0106] According to the above configuration, for example, even when the second machines are configured in multiple stages, all the machines can grasp the end points of the traces of the services spanning the first machine and the second machines.

[0107] It should be understood that the items included in the list in the form of "at least one of A, B, and C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Similarly, the items listed in the form of "at least one of A, B, or C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).

Description of Reference Numerals

[0108] 100... Cloud and Edge Cooperation Trace Monitoring System, 110... Cloud Platform, 120... Edge Terminal.

Claims

1. A service monitoring system comprising a first machine in a first distributed tracing environment and a second machine in a second distributed tracing environment, and monitoring services provided by the first machine and the second machine, wherein the first machine comprises a first execution unit that executes processing related to a service to be monitored, a first tracing unit that traces the processing related to the service executed by the first execution unit and determines an end point of the processing in the first machine, and a first cooperation unit that relays a request for the service including trace information that can identify a trace of the service to the second machine, wherein the second machine comprises a second cooperation unit that receives a request for a service to be monitored from the first cooperation unit, a second execution unit that executes processing related to the service, and a second tracing unit that traces the processing related to the service executed by the second execution unit and determines an end point of the processing in the second machine, wherein the second cooperation unit identifies a trace from the trace information included in the request for the service to be monitored received from the first cooperation unit, and when an end point of the processing related to the service in the second machine is determined by the second tracing unit, notifies the first cooperation unit of end point determination information indicating that the end point of the trace has been determined, a service monitoring system.

2. comprises a storage unit that stores a trace log generated by tracing by the first tracing unit and a trace log generated by tracing by the second tracing unit, wherein when the first machine receives a request for a service from a client terminal, it issues trace information that can uniquely identify a trace of the service, the first tracing unit traces the processing related to the service executed by the first execution unit to generate a trace log including the trace information, and stores the generated trace log in the storage unit, the second tracing unit traces the processing related to the service executed by the second execution unit to generate a trace log including the trace information, and when an end point of the processing in the second machine is determined by the second tracing unit, the second cooperation unit transfers the trace log of the service generated by the second tracing unit to the storage unit, The service monitoring system according to claim 1.

3. The trace log generated by the first tracing unit includes the time when the first execution unit starts processing related to the service and the time when the first execution unit ends processing related to the service. The trace log generated by the second tracing unit includes the time when the second execution unit starts processing related to the service and the time when the second execution unit ends processing related to the service. A management unit that manages service information including information capable of identifying the service to be monitored and threshold information indicating a time threshold indicating that the service is operating normally. A determination unit that reads out the trace log of the service to be monitored generated by the first tracing unit and the trace log of the service generated by the second tracing unit from the storage unit, calculates the time required for the service from the read trace log, determines whether the service is operating normally based on the threshold information included in the service information of the service, and outputs the determined result. The service monitoring system according to claim 2, comprising:

4. The first machine includes a first storage unit. The second machine includes a second storage unit. The first machine includes a first processing unit that instructs the update of the corresponding information in the second storage unit when the information in the first storage unit is updated. The second machine includes a second processing unit that instructs the update of the corresponding information in the first storage unit when the information in the second storage unit is updated. The first storage unit stores cooperation information including trace information capable of identifying the trace of the service to be monitored, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined. The second storage unit stores cooperation information including trace information capable of identifying the trace, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined. When the end point of the process in the second machine is determined by the second tracing unit, the second cooperation unit updates the end point determination flag of the cooperation information of the trace to information indicating that the end point of the trace has been determined. When there is an update to the trace linkage information stored in the second storage unit, the second processing unit notifies the updated information to the first processing unit, and the first processing unit updates the trace linkage information stored in the first storage unit with the notified information. When it is determined by the determination unit whether the service is operating normally, the first cooperation unit updates the deletion determination flag of the trace linkage information to information indicating that it is a deletion target. When there is an update to the trace linkage information stored in the first storage unit, the first processing unit notifies the updated information to the second processing unit, and the second processing unit updates the trace linkage information stored in the first storage unit with the notified information. The service monitoring system according to claim 3.

5. The first cooperation unit extracts machine information capable of identifying the second machine from the destination of the request of the service to be monitored, extracts trace information capable of identifying the trace of the service from the header of the request, and includes the trace information, the machine information, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined. The linkage information is stored in the first storage unit, and the trace information is inserted into the body of the request for relaying to the second machine. The second cooperation unit extracts the trace information from the body of the request received from the first cooperation unit, and stores the linkage information including the trace information, machine information capable of identifying the second machine, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined in the second storage unit. The service monitoring system according to claim 4.

6. A plurality of the second machines are provided, and one second machine calls another second machine to perform processing related to the service to be monitored. The linkage information stored in the first storage unit and the linkage information stored in the second storage unit include trace information capable of identifying a trace of a service to be monitored, machine information indicating a second machine that is a call source, machine information indicating a second machine that is a call destination, an end point determination flag indicating whether the end point of the trace has been determined, and a deletion determination flag indicating whether the operating state of the service has been determined. The service monitoring system according to claim 4.

7. A service monitoring method for monitoring a service provided by a first machine in a first distributed tracing environment and a second machine in a second distributed tracing environment, The first machine includes a first execution unit, a first tracing unit, and a first cooperation unit. The first execution unit executes processing related to a service to be monitored. The first tracing unit traces the processing related to the service executed by the first execution unit and determines an end point of the processing in the first machine. The first cooperation unit relays a request for the service including trace information capable of identifying a trace of the service to the second machine. The second machine includes a second execution unit, a second tracing unit, and a second cooperation unit. The second cooperation unit receives a request for a service to be monitored from the first cooperation unit. The second execution unit executes processing related to the service. The second tracing unit traces the processing related to the service executed by the second execution unit and determines an end point of the processing in the second machine. The second cooperation unit identifies a trace from the trace information included in the request for the service to be monitored received from the first cooperation unit, and when the end point of the processing related to the service in the second machine is determined by the second tracing unit, notifies the first cooperation unit of end point determination information indicating that the end point of the trace has been determined. Service monitoring method.

Citation Information

Patent Citations

  • Porous inorganic material

    JP1978008613A

Cited By

  • Tracking events of interest in a live production environment

    US12585572B1