Intelligent Monitoring Method, Device, Electronic Device and Readable Storage Medium

By loading the link data acquisition plug-in in the monitoring platform to collect and aggregate monitoring data, the problems of single monitoring dimensions and missing link data on the existing monitoring platform are solved, and more accurate alarms and faster abnormal positioning are achieved.

CN112527599BActive Publication Date: 2025-05-27KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011483121.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-15
Publication Date
2025-05-27
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

The existing monitoring platform has a single monitoring dimension and cannot provide effective monitoring data for different user groups, resulting in inaccurate alarm services and missing link data, resulting in long-term problem location.

Method used

Load the link data acquisition plug-in in the application to be monitored to collect link data when executing transactions, and obtain service indicators, application indicators and physical resource indicators, perform aggregation and analysis to generate complete monitoring data.

Benefits of technology

Provides complete monitoring data, improves abnormal positioning efficiency, and can quickly locate the abnormal root cause based on link data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112527599B_ABST
    Figure CN112527599B_ABST
Patent Text Reader

Abstract

The present invention relates to data processing, and discloses an intelligent monitoring method, including: loading a link data acquisition plug-in in an application to be monitored, and collecting link data generated on each node when executing a transaction of the application to be monitored based on the acquisition plug-in; obtaining a set of service metric values of the application to be monitored based on the link data, and obtaining a set of application metric values of the application to be monitored every first preset time; obtaining a set of physical resource metric values of the server to which each node belongs every second preset time; after the transaction is executed, updating the set of business metric values corresponding to the application to be monitored, and aggregating the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain target monitoring data corresponding to the link ID. The present invention also provides an intelligent monitoring device, an electronic device, and a readable storage medium. The present invention provides complete monitoring data and improves the efficiency of anomaly location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to an intelligent monitoring method, device, electronic equipment and readable storage medium. Background Art

[0002] With the development of information technology, monitoring plays a vital role in ensuring the stability of the system platform. Monitoring the system can timely learn about system problems and solve them. However, the existing monitoring platform has the following shortcomings:

[0003] 1. The monitoring dimension is relatively single, and it cannot provide effective monitoring data for different user groups, and thus cannot provide accurate alarm services;

[0004] 2. The link data is missing, which takes a long time to locate the problem and makes it impossible to solve the problem quickly.

[0005] Therefore, there is an urgent need for an intelligent monitoring method to provide complete monitoring data and improve the efficiency of abnormal location. Summary of the invention

[0006] In view of the above, it is necessary to provide an intelligent monitoring method that aims to provide complete monitoring data and improve the efficiency of abnormal location.

[0007] The intelligent monitoring method provided by the present invention comprises:

[0008] Loading a link data collection plug-in in the application to be monitored, when receiving a transaction request for the application to be monitored, assigning a link ID to the transaction, and collecting link data generated on each node of the transaction when executing the transaction based on the collection plug-in;

[0009] Acquire a service indicator value set of the application to be monitored based on the link data, acquire application indicator items of the application to be monitored from a preset database, and acquire the application indicator value set of the application to be monitored every first preset time;

[0010] At every second preset time, obtaining a set of physical resource indicator values ​​of the server to which each node belongs;

[0011] After the transaction is executed, the business indicator value set corresponding to the application to be monitored is updated, the service indicator value set, application indicator value set, physical resource indicator value set and business indicator value set are aggregated to obtain the target monitoring data corresponding to the link ID, and the target monitoring data is sent to the preset client.

[0012] Optionally, loading a link data collection plug-in in the application to be monitored includes:

[0013] Load the link data acquisition plug-in into the code of the application to be monitored, and generate plug-in definition information based on the acquisition plug-in;

[0014] Generate link data tracing code based on the plug-in definition information.

[0015] Optionally, after collecting the link data generated at each node of the transaction when executing the transaction based on the acquisition plug-in, the method further includes:

[0016] Store the link data into fixed-length arrays of different channels in the storage buffer in sequence. When the storage amount in the storage buffer exceeds the storage threshold, batch-store the data in the storage buffer into the log corresponding to the application to be monitored.

[0017] Optionally, aggregating the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set to obtain the target monitoring data corresponding to the link ID includes:

[0018] Perform trend analysis, year-on-year analysis, and month-on-month analysis on each metric value in the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set to obtain the first monitoring data;

[0019] Based on the mapping relationship between the metric set and the aggregation operation formula, respectively obtain the aggregation operation formulas corresponding to the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set and perform the aggregation operation to obtain the second monitoring data;

[0020] Summarize the first monitoring data and the second monitoring data to obtain the target monitoring data corresponding to the link ID.

[0021] Optionally, the method further includes:

[0022] Judge whether each metric value in the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set is abnormal;

[0023] If a certain metric value is abnormal, determine the target user group corresponding to the abnormal metric value according to the mapping relationship between the metric set and the user group, and send a warning message to the target user group.

[0024] Optionally, the service metric value set includes the numerical value sets corresponding to the request type, request parameters, response information, request duration, and request result;

[0025] The application metric value set includes the numerical value sets corresponding to the number of blocked threads, the number of deadlocked threads, the number of waiting threads, the number of garbage collections, service throughput, and service exception information;

[0026] The set of physical resource index values includes the value sets corresponding to CPU, memory, disk, network IO, and bandwidth.

[0027] Optionally, the link ID is composed of a global ID, an application ID, and a transaction ID.

[0028] To solve the above problems, the present invention also provides an intelligent monitoring device, which includes:

[0029] An acquisition module, configured to load a link data acquisition plug-in in the application to be monitored. When receiving a transaction request for the application to be monitored, assign a link ID to the transaction, and collect link data generated on each node of the transaction based on the acquisition plug-in when executing the transaction;

[0030] A first acquisition module, configured to obtain the set of service index values of the application to be monitored based on the link data, obtain the application index items of the application to be monitored from a preset database, and obtain the set of application index values of the application to be monitored every first preset time;

[0031] A second acquisition module, configured to obtain the set of physical resource index values of the server to which each node belongs every second preset time;

[0032] An aggregation module, configured to update the set of business index values corresponding to the application to be monitored after the transaction is executed, aggregate the set of service index values, the set of application index values, the set of physical resource index values, and the set of business index values to obtain the target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client.

[0033] To solve the above problems, the present invention also provides an electronic device, which includes:

[0034] At least one processor; and,

[0035] A memory communicatively connected to the at least one processor; wherein,

[0036] The memory stores an intelligent monitoring program executable by the at least one processor. The intelligent monitoring program is executed by the at least one processor so that the at least one processor can execute the above intelligent monitoring method.

[0037] To solve the above problems, the present invention also provides a computer-readable storage medium, on which an intelligent monitoring program is stored. The intelligent monitoring program can be executed by one or more processors to implement the above intelligent monitoring method.

[0038] Compared with the prior art, the present invention first loads a link data acquisition plug-in into the application to be monitored, and acquires link data generated on each node of the transaction when executing the transaction of the application to be monitored based on the acquisition plug-in. This step can achieve the acquisition purpose by loading the acquisition plug-in without changing the business layer code, making the acquisition of link data extensible and maintainable. Then, based on the link data, a set of service metric values of the application to be monitored is obtained. Every first preset time, a set of application metric values of the application to be monitored is obtained. Every second preset time, a set of physical resource metric values of the server to which each node belongs is obtained. After the transaction is completed, the set of business metric values corresponding to the application to be monitored is updated. Finally, the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values are aggregated, making the obtained target monitoring data more complete. And because the link data is captured, when an index is abnormal, the root cause of the abnormality can be quickly located according to the link data. Therefore, the present invention provides complete monitoring data and improves the efficiency of anomaly location. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of an intelligent monitoring method provided by an embodiment of the present invention;

[0040] Figure 2 It is a schematic diagram of modules of an intelligent monitoring device provided by an embodiment of the present invention;

[0041] Figure 3 It is a schematic structural diagram of an electronic device for implementing the intelligent monitoring method provided by an embodiment of the present invention;

[0042] The implementation, functional features, and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] In order to make the object, technical solution, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] It should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0045] The present invention provides an intelligent monitoring method. Referring to Figure 1 As shown, it is a schematic flowchart of the intelligent monitoring method provided by an embodiment of the present invention. This method can be executed by an electronic device, and the electronic device can be implemented by software and / or hardware.

[0046] In this embodiment, the intelligent monitoring method includes:

[0047] S1. Load a link data acquisition plugin in the application to be monitored. When a transaction request for the application to be monitored is received, assign a link ID to the transaction, and collect the link data generated on each node of the transaction when executing the transaction based on the acquisition plugin.

[0048] This embodiment adopts a plug-in and pluggable acquisition scheme to collect link data. The libs of various plugins are predefined and integrated into the kernel through the Java SPI method. When it is necessary to collect the link data of a certain application, only need to load the acquisition plugin into the code of the application, without changing the business layer code, making the acquisition of link data scalable and maintainable.

[0049] The loading of the link data acquisition plugin in the application to be monitored includes:

[0050] A11. Load the link data acquisition plugin into the code of the application to be monitored, and generate plugin definition information based on the acquisition plugin;

[0051] A12. Generate link data buried point code based on the plugin definition information.

[0052] The plugin definition information includes the target class, target method, and interceptor for bytecode enhancement. The subsequent buried point information will be stored in the log corresponding to the application to be monitored.

[0053] The complete link data corresponding to a transaction is composed of the call chain data of each node participating in the transaction. The call chain data includes data generated by various operations such as interface calls, cache calls, MQ message sending and receiving, and database access. The generated data includes request parameters, response information, request duration, and request results (normal or abnormal).

[0054] A link data is a directed acyclic graph composed of one or more Spans (spans). A Span represents a logical unit in the link with a start time and an execution duration (one Span corresponds to one node). Logical causal relationships are established between Spans through nesting or sequential sorting.

[0055] The Span includes: operation name (e.g., the RPC service accessed, the URL address accessed, etc.), start and end times, and the reference relationship of the Span.

[0056] The link data is in the open-tracing data structure. The link data ID (Global trace ID-segment trace ID-span ID) is composed of a global ID, an application ID, and a transaction ID. Among them, the Global trace is the globally unique link ID, the segment trace ID is the unique ID in units of services (or applications), and the span ID is the ID at the transaction level under the service (or application). There may be multiple segment traces under one Global trace, and there may be multiple spans under one segment trace.

[0057] In this embodiment, after collecting the link data generated on each node of the transaction when executing the transaction based on the collection plugin, the method further includes:

[0058] Storing the link data into the fixed-length arrays of different channels in the storage buffer in sequence. When the storage capacity of the storage buffer exceeds the storage threshold, the data in the storage buffer is batch-stored into the log corresponding to the application to be monitored.

[0059] In this embodiment, the storage buffer is a lightweight implementation library of the producer-consumer mode. There is one or more channels in this implementation library. Each channel manages one or more fixed-length arrays. This implementation library provides a selector for determining which fixed-length array in the storage buffer a data element is written into, which can effectively reduce the spin lock waiting time caused by concurrency.

[0060] Meanwhile, there is a circular pointer in each fixed-length array, which can specify that the value field (Atomic Integer type) in the fixed-length array starts incrementing from the start value. When the value increments to the end value (int type), the value field will be reset to the start value, thus realizing the function of storing through the circular pointer. In this way, the fixed-length array can achieve the mode of circular overwrite writing.

[0061] S2. Obtain the set of service metric values of the application to be monitored based on the link data, obtain the application metric items of the application to be monitored from the preset database, and obtain the set of application metric values of the application to be monitored every first preset time.

[0062] The set of service metric values includes the numerical sets corresponding to the request type (interface call, cache call, MQ message sending and receiving, and database access), request parameters, response information, request duration, and request result (normal or abnormal). Through the set of service metric values, the concurrency and response time of various operation events of the application to be monitored can be determined, and the overall health status of the service can be determined.

[0063] The application metric items include the number of blocked threads, the number of deadlock threads, the number of waiting threads, the number of garbage collections, service throughput, and service exception information. Since the applications in this embodiment are all within the Java scope, the metric values of each application metric item can be collected through MBeans (manageable Java objects).

[0064] S3. Obtain the set of physical resource metric values of the servers to which the respective nodes belong every second preset time.

[0065] In this embodiment, the link data of each node further includes the IP information of the server to which each node belongs. Through the IP information, the set of physical resource metric values of the server to which each node belongs can be obtained.

[0066] The set of physical resource metric values includes the sets of numerical values corresponding to CPU, memory, disk, network IO, and bandwidth. In this embodiment, by deploying a polling program once every 60 seconds on each server, the corresponding kernel instructions of Linux are run regularly to collect the physical resource metric data of each server.

[0067] The physical resource metric data can reflect the current performance of the servers to which the respective nodes of the transaction belong, and the current performance of the server affects the execution of the transaction.

[0068] S4. After the execution of the transaction is completed, update the set of business metric values corresponding to the application to be monitored, aggregate the service metric value set, application metric value set, physical resource metric value set, and business metric value set to obtain the target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client.

[0069] In this embodiment, by simulating the implementation of the mysql dump data transfer protocol, the change events such as addition, deletion, and modification of business data in the mysql database are monitored, the data change records are obtained in real time and preliminary data cleaning is performed, and then aggregation operations are performed according to the preset time window and aggregation rules to obtain each business metric value.

[0070] For example, when it is necessary to monitor user order data, relevant order reports can be aggregated and generated according to the user ID, transaction amount, etc. of the order.

[0071] The aggregating the service metric value set, application metric value set, physical resource metric value set, and business metric value set to obtain the target monitoring data corresponding to the link ID includes:

[0072] B11. Perform trend analysis, year-on-year analysis, and month-on-month analysis on each metric value in the service metric value set, application metric value set, physical resource metric value set, and business metric value set to obtain the first monitoring data;

[0073] B12. Based on the mapping relationship between the metric set and the aggregation operation formula, respectively obtain the aggregation operation formulas corresponding to the service metric value set, application metric value set, physical resource metric value set, and business metric value set and perform the aggregation operation to obtain the second monitoring data;

[0074] B13. Summarize the first monitoring data and the second monitoring data to obtain the target monitoring data corresponding to the link ID.

[0075] In this embodiment, the method further includes:

[0076] C11. Determine whether each metric value in the service metric value set, application metric value set, physical resource metric value set, and business metric value set is abnormal;

[0077] C12. If a certain metric value is abnormal, determine the target user group corresponding to the abnormal metric value according to the mapping relationship between the metric set and the user group, and send a warning message to the target user group.

[0078] In this embodiment, the service metric set is for underlying R & D personnel, the application metric set is for application developers, the physical resource metric set is for operation and maintenance personnel, and the business metric set is for business personnel. When any metric value in any metric set is abnormal (by comparing with the corresponding threshold value to determine whether each metric value is abnormal), an alarm message can be sent to the corresponding user group. For the abnormal event of the metric value in the service metric set, the link call information can be reported simultaneously to facilitate quickly locating the root cause of the abnormality.

[0079] As can be seen from the above embodiments, for the intelligent monitoring method proposed by the present invention, first, a link data acquisition plug-in is loaded into the application program to be monitored, and based on the acquisition plug-in, the link data generated on each node of the transaction when executing the transaction of the application program to be monitored is acquired. This step can achieve the acquisition purpose by loading the acquisition plug-in without changing the business layer code, making the acquisition of link data have scalability and maintainability. Then, based on the link data, the service metric value set of the application program to be monitored is obtained. Every first preset time, the application metric value set of the application program to be monitored is obtained. Every second preset time, the physical resource metric value set of the server to which each node belongs is obtained. After the transaction is completed, the business metric value set corresponding to the application program to be monitored is updated. Finally, the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set are aggregated, making the obtained target monitoring data more complete. And because the link data is captured, when a metric is abnormal, the root cause of the abnormality can be quickly located according to the link data. Therefore, the present invention provides complete monitoring data and improves the efficiency of abnormal location.

[0080] As Figure 2 shown, it is a schematic diagram of the modules of an intelligent monitoring device provided by an embodiment of the present invention.

[0081] The intelligent monitoring device 100 of the present invention can be installed in an electronic device. According to the functions achieved, the intelligent monitoring device 100 can include an acquisition module 110, a first acquisition module 120, a second acquisition module 130, and an aggregation module 140. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0082] In this embodiment, the functions of each module / unit are as follows:

[0083] The acquisition module 110 is used to load a link data acquisition plug-in into the application program to be monitored. When receiving a transaction request for the application program to be monitored, a link ID is assigned to the transaction, and based on the acquisition plug-in, the link data generated on each node of the transaction when executing the transaction is acquired.

[0084] In this embodiment, a plug-in and pluggable data acquisition scheme is adopted to acquire link data. Libs of various plug-ins are predefined and integrated into the kernel through the Java SPI method. When it is necessary to acquire the link data of a certain application, only the acquisition plug-in needs to be loaded into the code of the application, without changing the business layer code, making the acquisition of link data scalable and maintainable.

[0085] Loading the link data acquisition plug-in in the application to be monitored includes:

[0086] A21. Loading the link data acquisition plug-in into the code of the application to be monitored, and generating plug-in definition information based on the acquisition plug-in;

[0087] A22. Generating link data tracing code based on the plug-in definition information.

[0088] The plug-in definition information includes the target class, target method, and interceptor for bytecode enhancement. Subsequent tracing information will be stored in the log corresponding to the application to be monitored.

[0089] The complete link data corresponding to a transaction is composed of the call chain data of each node participating in the transaction. The call chain data includes data generated by various operations such as interface calls, cache calls, MQ message sending and receiving, and database access. The generated data includes request parameters, response information, request duration, and request results (normal or abnormal).

[0090] A link data is a directed acyclic graph composed of one or more Spans. A Span represents a logical unit in the link with a start time and an execution duration (one Span corresponds to one node). Logical causal relationships are established between Spans through nesting or sequential sorting.

[0091] A Span includes: operation name (for example, the RPC service accessed, the URL address accessed, etc.), start and end times, and reference relationships of the Span.

[0092] The link data is in the open-tracing data structure. The link data ID (Global trace ID - segment trace ID - span ID) consists of a global ID, an application ID, and a transaction ID. Among them, the Global trace is a globally unique link ID, the segment trace ID is a unique ID per service (or application), and the span ID is an ID at the transaction level under the service (or application). There may be multiple segment traces under one Global trace, and there may be multiple spans under one segment trace.

[0093] In this embodiment, after collecting the link data generated on each node of the transaction when executing the transaction based on the collection plugin, the collection module 110 is further configured to:

[0094] Store the link data into fixed-length arrays of different channels in the storage buffer in sequence. When the storage amount of the storage buffer exceeds the storage threshold, batch-store the data in the storage buffer into the log corresponding to the application to be monitored.

[0095] In this embodiment, the storage buffer is a lightweight implementation library of the producer - consumer mode. There is one or more channels in this implementation library. Each channel manages one or more fixed-length arrays. This implementation library provides a selector for determining which fixed-length array in the storage buffer a data element is written into, which can effectively reduce the spin lock waiting time caused by concurrency.

[0096] At the same time, there is a circular pointer in each fixed-length array, which can specify that the value field (Atomic Integer type) in the fixed-length array starts to increment from the start value. When the value increments to the end value (int type), the value field will be reset to the start value, thus realizing the function of storing through the circular pointer. In this way, the fixed-length array can achieve the mode of circular overwrite writing.

[0097] The first acquisition module 120 is configured to obtain the set of service metric values of the application to be monitored based on the link data, obtain the application metric items of the application to be monitored from a preset database, and obtain the set of application metric values of the application to be monitored every first preset time.

[0098] The set of service metric values includes the numerical sets corresponding to request types (interface calls, cache calls, MQ message sending and receiving, and database access), request parameters, response information, request duration, and request results (normal or abnormal). Through the set of service metric values, the concurrency and response time of various operation events of the application to be monitored can be determined, and the overall health status of the service can be determined.

[0099] The application metric items include the number of blocked threads, the number of deadlock threads, the number of waiting threads, the number of garbage collections, service throughput, and service exception information. Since the applications in this embodiment are all within the Java scope, the metric values of each application metric item can be collected through MBeans (manageable Java objects).

[0100] The second acquisition module 130 is used to acquire the set of physical resource metric values of the servers to which the respective nodes belong every second preset time.

[0101] In this embodiment, the link data of each node further includes the IP information of the server to which each node belongs. Through the IP information, the set of physical resource metric values of the server to which each node belongs can be acquired.

[0102] The set of physical resource metric values includes the sets of numerical values corresponding to CPU, memory, disk, network IO, and bandwidth. In this embodiment, by deploying a polling program once every 60 seconds on each server, the corresponding kernel instructions of Linux are run regularly to collect the physical resource metric data of each server.

[0103] The physical resource metric data can reflect the current performance of the servers to which the respective nodes of the transaction belong, and the current performance of the server affects the execution of the transaction.

[0104] The aggregation module 140 is used to update the set of business metric values corresponding to the application to be monitored after the transaction is completed, aggregate the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain the target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client.

[0105] In this embodiment, by simulating the implementation of the mysql dump data transfer protocol, the change events such as addition, deletion, and modification of business data in the mysql database are monitored, the data change records are obtained in real time and preliminary data cleaning is performed, and then aggregation operations are performed according to the preset time window and aggregation rules to obtain each business metric value.

[0106] For example, when it is necessary to monitor user order data, relevant order reports can be aggregated and generated according to the user ID, transaction amount, etc. of the order.

[0107] Aggregating the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set includes:

[0108] B21. Performing trend analysis, year-on-year analysis, and month-on-month analysis on each metric value in the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set to obtain first monitoring data;

[0109] B22. Based on the mapping relationship between the metric set and the aggregation operation formula, respectively obtaining the aggregation operation formulas corresponding to the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set and performing aggregation operations to obtain second monitoring data;

[0110] B13. Summarizing the first monitoring data and the second monitoring data to obtain the target monitoring data corresponding to the link ID.

[0111] In this embodiment, the aggregation module 140 is further configured to:

[0112] C21. Judging whether each metric value in the service metric value set, the application metric value set, the physical resource metric value set, and the business metric value set is abnormal;

[0113] C22. If a certain metric value is abnormal, determining the target user group corresponding to the abnormal metric value according to the mapping relationship between the metric set and the user group, and sending a warning message to the target user group.

[0114] In this embodiment, the service metric set is for underlying R & D personnel, the application metric set is for application developers, the physical resource metric set is for operation and maintenance personnel, and the business metric set is for business personnel. When any metric value in any metric set is abnormal (by comparing with the corresponding threshold to determine whether each metric value is abnormal), an alarm message can be sent to the corresponding user group. For the abnormal event of the metric value in the service metric set, the link call information can be reported simultaneously to facilitate quickly locating the root cause of the abnormality.

[0115] As can be seen from the above embodiments, for the intelligent monitoring device 100 proposed by the present invention, first, a link data acquisition plug-in is loaded in the application to be monitored, and link data generated on each node of the transaction when executing the application to be monitored is collected based on the acquisition plug-in. This step can achieve the acquisition purpose by loading the acquisition plug-in without changing the business layer code, making the acquisition of link data scalable and maintainable. Then, a set of service metric values of the application to be monitored is obtained based on the link data. Every first preset time, a set of application metric values of the application to be monitored is obtained. Every second preset time, a set of physical resource metric values of the server to which each node belongs is obtained. After the transaction is completed, the set of business metric values corresponding to the application to be monitored is updated. Finally, the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values are aggregated, making the obtained target monitoring data more complete. And because the link data is captured, when an index is abnormal, the root cause of the abnormality can be quickly located according to the link data. Therefore, the present invention provides complete monitoring data and improves the efficiency of abnormal location.

[0116] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the intelligent monitoring method provided by an embodiment of the present invention.

[0117] The electronic device 1 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The electronic device 1 can be a computer, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing and is composed of a group of loosely coupled computers to form a super virtual computer.

[0118] In this embodiment, the electronic device 1 includes, but is not limited to, a memory 11, a processor 12, and a network interface 13 that can communicate with each other through a system bus. The memory 11 stores an intelligent monitoring program 10, and the intelligent monitoring program 10 can be executed by the processor 12. Figure 3 Only the electronic device 1 with components 11-13 and the intelligent monitoring program 10 is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine some components, or have different component arrangements.

[0119] Among them, the memory 11 includes a memory and at least one type of readable storage medium. The memory provides a cache for the operation of the electronic device 1; the readable storage medium can be a non-volatile storage medium such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the readable storage medium can be an internal storage unit of the electronic device 1, such as the hard disk of the electronic device 1; in other embodiments, the non-volatile storage medium can also be an external storage device of the electronic device 1, such as a plug-in hard disk equipped on the electronic device 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. In this embodiment, the readable storage medium of the memory 11 is generally used to store the operating system and various application software installed in the electronic device 1, such as storing the code of the intelligent monitoring program 10 in an embodiment of the present invention. In addition, the memory 11 can also be used to temporarily store various types of data that have been output or will be output.

[0120] In some embodiments, the processor 12 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with other devices. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the intelligent monitoring program 10, etc.

[0121] The network interface 13 can include a wireless network interface or a wired network interface, and the network interface 13 is used to establish a communication connection between the electronic device 1 and a client (not shown in the figure).

[0122] Optionally, the electronic device 1 can further include a user interface. The user interface can include a display (Display), an input unit such as a keyboard (Keyboard). Optionally, the user interface can further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display can also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0123] It should be understood that the above embodiments are for illustrative purposes only and the scope of the patent application is not limited by this structure.

[0124] The intelligent monitoring program 10 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 12, it can achieve:

[0125] Load a link data acquisition plugin in the application to be monitored. When a transaction request for the application to be monitored is received, assign a link ID to the transaction, and collect link data generated on each node of the transaction when executing the transaction based on the acquisition plugin;

[0126] Obtain a set of service metric values of the application to be monitored based on the link data, obtain application metric items of the application to be monitored from a preset database, and obtain a set of application metric values of the application to be monitored every first preset time;

[0127] Obtain a set of physical resource metric values of the servers to which the respective nodes belong every second preset time;

[0128] After the transaction is completed, update the set of business metric values corresponding to the application to be monitored, aggregate the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client.

[0129] Specifically, for the specific implementation method of the above intelligent monitoring program 10 by the processor 12, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here. It should be emphasized that to further ensure the privacy and security of the above sets of metric values, the above sets of metric values can also be stored in a node of a blockchain.

[0130] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or non-volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0131] The intelligent monitoring program 10 is stored on the computer-readable storage medium. The intelligent monitoring program 10 can be executed by one or more processors. The specific implementation manner of the computer-readable storage medium of the present invention is basically the same as that of the above-mentioned embodiments of the intelligent monitoring method, and will not be elaborated here.

[0132] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0133] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0134] In addition, in each embodiment of the present invention, the functional modules can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0135] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0136] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0137] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, essentially a decentralized database, is a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0138] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular form does not exclude the plural form. A plurality of units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.

[0139] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An intelligent monitoring method, characterized in that, the method includes: loading a link data acquisition plugin into the application to be monitored. When a transaction request for the application to be monitored is received, assign a link ID to the transaction, and based on the acquisition plugin, collect link data generated on each node of the transaction when the transaction is executed. The link ID is composed of a global ID, an application ID, and a transaction ID; obtain a set of service metric values of the application to be monitored based on the link data, obtain application metric items of the application to be monitored from a preset database, and obtain a set of application metric values of the application to be monitored every first preset time; obtain a set of physical resource metric values of the server to which each node belongs every second preset time; after the transaction is executed, update the set of business metric values corresponding to the application to be monitored, aggregate the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client; the loading the link data acquisition plugin into the application to be monitored includes: loading the link data acquisition plugin into the code of the application to be monitored, and generating plugin definition information based on the acquisition plugin; generating link data instrumentation code based on the plugin definition information.

2. The intelligent monitoring method according to claim 1, characterized in that, after collecting the link data generated on each node of the transaction when the transaction is executed based on the acquisition plugin, the method further includes: sequentially storing the link data into fixed-length arrays of different channels in a storage buffer. When the storage amount of the storage buffer exceeds a storage threshold, batch-store the data in the storage buffer into the log corresponding to the application to be monitored.

3. The intelligent monitoring method according to claim 1, characterized in that, the aggregating the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain target monitoring data corresponding to the link ID includes: performing trend analysis, year-on-year analysis, and month-on-month analysis on each metric value in the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values to obtain first monitoring data; respectively obtaining aggregation operation formulas corresponding to the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values based on the mapping relationship between the metric set and the aggregation operation formula, and performing aggregation operations to obtain second monitoring data; summarizing the first monitoring data and the second monitoring data to obtain target monitoring data corresponding to the link ID.

4. The intelligent monitoring method according to any one of claims 1-3, characterized in that, the method further includes: judging whether each metric value in the set of service metric values, the set of application metric values, the set of physical resource metric values, and the set of business metric values is abnormal; If a certain index value is abnormal, according to the mapping relationship between the index set and the user group, determine the target user group corresponding to the abnormal index value, and send a warning message to the target user group.

5. The intelligent monitoring method according to claim 1, characterized in that the service index value set includes the numerical sets corresponding to the request type, request parameters, response information, request duration, and request result; the application index value set includes the numerical sets corresponding to the number of blocked threads, number of deadlocked threads, number of waiting threads, number of garbage collections, service throughput, and service exception information; the physical resource index value set includes the numerical sets corresponding to CPU, memory, disk, network IO, and bandwidth.

6. The intelligent monitoring method according to claim 1, characterized in that the link ID is composed of a global ID, an application ID, and a transaction ID.

7. An intelligent monitoring device for implementing the intelligent monitoring method according to any one of claims 1-6, characterized in that the device includes: a collection module, configured to load a link data collection plug-in in the application program to be monitored, and when receiving a transaction request for the application program to be monitored, assign a link ID to the transaction, and collect link data generated on each node of the transaction when executing the transaction based on the collection plug-in; a first acquisition module, configured to obtain the service index value set of the application program to be monitored based on the link data, obtain the application index items of the application program to be monitored from a preset database, and obtain the application index value set of the application program to be monitored every first preset time; a second acquisition module, configured to obtain the physical resource index value set of the server to which each node belongs every second preset time; an aggregation module, configured to update the business index value set corresponding to the application program to be monitored after the transaction is executed, aggregate the service index value set, application index value set, physical resource index value set, and business index value set to obtain the target monitoring data corresponding to the link ID, and send the target monitoring data to a preset client.

8. An electronic device, characterized in that the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores an intelligent monitoring program executable by the at least one processor, and the intelligent monitoring program is executed by the at least one processor so that the at least one processor can execute the intelligent monitoring method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that an intelligent monitoring program is stored on the computer-readable storage medium, and the intelligent monitoring program can be executed by one or more processors to implement the intelligent monitoring method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Monitoring data processing method, equipment, server and storage medium

    CN107704360A

  • Monitored and management servers, data acquisition and analysis method and management system

    CN109032904A