Method, electronic device, and storage medium for batch calculating application healthiness
By decomposing the applied health calculation task into multi-stage sub-tasks and processing it in batches, the problems of high resource requirements for stand-alone computing and slow distributed computing speed are solved, and more efficient and accurate application health calculation is achieved.
Patent Information
- Application Number
- CN202210074374.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-01-21
AI Technical Summary
In the prior art, due to the huge amount of data, the CPU computing power, memory capacity, I/O speed and network stability are very high during single-machine computing, and the distributed computing speed is relatively slow.
The applied health calculation task is broken down into multi-stage subtasks, including basic monitoring and inspection tasks, resource utilization analysis tasks and score calculation tasks, and batch processing is carried out through time-series monitoring item data, jitter identification and uptrend identification to reduce calculation density and improve calculation speed.
By decomposing tasks and batch processing, the requirements for stand-alone computing are reduced, the calculation speed and application health accuracy are improved, and the problem of slow distributed computing is solved.
Smart Images

Figure CN114416487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method, an electronic device, and a storage medium for batch computing application health. Background Art
[0002] With the rapid development of computer technology, network technology has been widely applied, and various software applications emerge in an endless stream, providing many services for users. In order for applications to provide services stably, in the operation and maintenance scenario, it is usually necessary to calculate the application health to reflect the health degree of the application's running state within an observation period.
[0003] Generally, the application health can be directly reflected by monitoring data. However, such a simple basis cannot accurately reflect the health degree of the application. In addition, with the increasing size of the user group and the huge increase in data volume, the computing density is high, and the requirements for CPU computing power, memory capacity, I / O speed, and network stability are very high during single-machine computing; moreover, due to the huge data volume, the distributed computing speed is relatively slow. Summary of the Invention
[0004] Embodiments of this application provide a method, an electronic device, and a storage medium for batch computing application health, so as to at least solve the problem in the related art that due to the huge data volume, the requirements for CPU computing power, memory capacity, I / O speed, and network stability are very high during single-machine computing, and the distributed computing speed is relatively slow.
[0005] In a first aspect, embodiments of this application provide a method for batch computing application health, including: decomposing an application health computing task into multi-stage subtasks, where the multi-stage subtasks include a basic monitoring and inspection task, a resource utilization rate analysis task, and a score calculation task; when executing the basic monitoring and inspection task, collecting time-series monitoring item data; when executing the resource utilization rate analysis task, calculating the jitter identification and the rising trend identification of each monitoring item according to the time-series monitoring item data; when executing the score calculation task, scoring each of the monitoring items according to the time-series monitoring item data, the jitter identification, and the rising trend identification, and calculating a score for representing the application health.
[0006] In some of these embodiments, calculating the jitter identification and the rising trend identification of each monitoring item according to the time-series monitoring item data includes: grouping the time-series monitoring item data and calculating the jitter identification and the rising trend identification of each monitoring item in batches.
[0007] In some of these embodiments, the grouping of the timing monitoring item data includes: obtaining a pre-constructed E-R relationship model, where the E-R relationship model is used to express the subordination relationships among applications, clusters, computer rooms, hosts, and monitoring items; according to the E-R relationship model, counting the number of clusters where the applications are located, and the number of hosts included in each of the clusters; grouping the clusters by a fixed number of hosts to obtain the timing monitoring item data of each group of clusters.
[0008] In some of these embodiments, the batch calculation of the jitter flag for each monitoring item includes: based on the grouped timing monitoring item data, within a preset time period, determining whether there is jitter. If the number of jitter times is greater than a preset number threshold, or the maximum value of the jitter duration is greater than a preset time threshold, then the jitter flag indicates jitter anomaly.
[0009] In some of these embodiments, the determination of whether there is jitter includes: subtracting the monitoring item data at adjacent times and taking the absolute value to obtain the amplitude; in the case where the amplitude is greater than the amplitude threshold, determining whether the change trends of the monitoring item data at the adjacent times are opposite. If so, it is determined that there is jitter.
[0010] In some of these embodiments, the determination of whether the change trends of the monitoring item values at adjacent times are opposite includes: subtracting the monitoring item value at time t from the monitoring item value at time t - 1 to obtain a first difference; subtracting the monitoring item value at time t + 1 from the monitoring item value at time t to obtain a second difference; if the product of the first difference and the second difference is negative, it is determined that the change trends of the monitoring item values at time t and time t - 1 are opposite.
[0011] In some of these embodiments, the batch calculation of the upward trend flag for each monitoring item includes: performing a continuity detection on the grouped timing monitoring item data. In the case of no discontinuous data, comparing the monitoring item at the previous time with the monitoring item data at the current time. If it is an upward trend, detecting the length of the upward sequence. If the length of the upward sequence is greater than the length threshold, then the upward trend flag of the monitoring item corresponding to the upward sequence indicates an upward trend anomaly, until the following situations occur: in the case of a stable trend, recording the number of times of the stable appearance. If it is higher than the number threshold, it is determined that the upward trend ends; in the case of a downward trend, it is determined that the upward trend ends.
[0012] In some of these embodiments, the monitoring items include CPU utilization rate, memory utilization rate, disk occupancy rate, number of garbage collections, application alarm volume, and alarm waiting response duration, and the scoring rules for each of the monitoring items based on the timing monitoring item data, the jitter flag, and the upward trend flag are different.
[0013] Second aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, characterized in that a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method described in any one of the above.
[0014] Third aspect, an embodiment of the present application provides a storage medium, characterized in that a computer program is stored in the storage medium, wherein the computer program is configured to execute the method described in any one of the above when running.
[0015] Compared with the related art, the method for batch calculating the health of applications provided by the embodiments of the present application, considering that calculating the health of a large number of applications in batch will involve a large amount of data storage and calculation, so adopts a top-down design idea, decomposes the application health calculation task into multi-stage subtasks, namely basic monitoring and inspection tasks, resource utilization rate analysis tasks and score calculation tasks, which reduces the implementation difficulty as a whole. Moreover, when calculating the application health in the embodiments of the present application, not only based on the time-series monitoring item data, but also based on the jitter identifier and the rising trend identifier, it can more accurately detect the changes of the application over time during operation. At the same time, considering the periodic and trend characteristics of the time-series data, thus, the accuracy of the application health can be improved. In addition, for the two-stage subtasks of resource utilization rate analysis task and score calculation task, the data is grouped, and then the data is processed in batches, which reduces the calculation intensity, and at the same time reduces the requirements for CPU computing power, memory capacity, I / O speed and network stability during single-machine calculation; moreover, by batch calculation to reduce the calculation intensity, compared with the complex distributed calculation method, the calculation speed can be improved. Description of the Drawings
[0016] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 is a flowchart of the method for batch calculating the health of applications according to an embodiment of the present application;
[0018] Figure 2 is a schematic diagram of a piece of information recorded in the hardware monitoring historical data table according to an embodiment of the present application;
[0019] Figure 3 is a schematic diagram of a piece of information recorded in the computer room monitoring historical data table according to an embodiment of the present application;
[0020] Figure 4 is a schematic diagram of a piece of information recorded in the anomaly detection historical data table according to an embodiment of the present application;
[0021] Figure 5 is a schematic diagram of a piece of information recorded in the application health information table according to an embodiment of the present application;
[0022] Figure 6 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the present application. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided in the present application without making creative efforts fall within the scope of protection of the present application. In addition, it can also be understood that although the efforts made in such a development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as insufficient disclosure of the content of the present application.
[0024] Referring to "embodiment" in the present application means that the specific features, structures or characteristics described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0025] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. The words such as "a", "an", "one", "the" and the like involved in this application do not indicate a limitation in quantity and may represent a singular or plural number. The terms "including", "comprising", "having" and any variations thereof involved in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may further include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products or devices. The similar words such as "connected", "coupled" and "linked" involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The "plurality" involved in this application means greater than or equal to two. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The terms "first", "second", "third" and the like involved in this application are only used to distinguish similar objects and do not represent a specific order of the objects.
[0026] Figure 1 is a flowchart of a method for calculating the health degree of a batch calculation application according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0027] S101: Decompose the application health degree calculation task into multi-stage subtasks, where the multi-stage subtasks include a basic monitoring and inspection task, a resource utilization rate analysis task, and a score calculation task;
[0028] S102: When executing the basic monitoring and inspection task, collect time series monitoring item data;
[0029] S103: When executing the resource utilization rate analysis task, calculate the jitter flag and the rising trend flag of each monitoring item according to the time series monitoring item data;
[0030] S104: When executing the score calculation task, score each monitoring item according to the time series monitoring item data, the jitter flag, and the rising trend flag, and calculate the score representing the application health degree.
[0031] Based on the above content, considering that calculating the healthiness of a large number of applications in batch will involve a large amount of data storage and calculation, the embodiments of this application adopt a top-down design idea, decomposing the application healthiness calculation task into subtasks in multiple stages, namely, the basic monitoring and inspection task, the resource utilization rate analysis task, and the score calculation task, which reduces the implementation difficulty as a whole. Moreover, when calculating the application healthiness, the embodiments of this application not only rely on the time-series monitoring item data, but also rely on the jitter flag and the rising trend flag, which can more accurately detect the changes of the application over time during operation. At the same time, considering the periodic and trend characteristics of the time-series data, the accuracy of the application healthiness can be improved.
[0032] To describe the embodiments of this application more clearly, the above steps will be described in detail below.
[0033] S101. Adopt a top-down design idea and decompose the application healthiness calculation task into subtasks in three stages, namely, the basic monitoring and inspection task, the resource utilization rate analysis task, and the score calculation task.
[0034] When performing the basic monitoring and inspection task in S102, collect the time-series monitoring item data. For example, collect the time-series data of six monitoring items. Among them, four monitoring items are resource utilization rates, including CPU utilization rate, memory utilization rate, disk occupancy rate, and garbage collection times; the other two monitoring items are the alarm handling situations, including the application alarm volume and the alarm waiting response duration. And the time-series data of the above six monitors are all collected within the observation period.
[0035] In some embodiments, it is necessary to collect the monitoring item data every hour. Considering that the data is collected once every hour and the time interval is relatively large, the obtained time-series monitoring item data is relatively accidental, which will affect the accuracy of the subsequent calculated jitter flag and rising trend flag. Therefore, in the embodiments of this application, the monitoring item data is first collected at a preset interval time (such as 60s) to ensure the sampling accuracy, and then the average value (or the empirical value is the p85 percentile value) of the collected monitoring item data within a preset duration (such as one hour) is calculated, and this average value is used as the above time-series monitoring item data, which can more accurately reflect the status of the monitoring items within this hour. Thus, the accuracy of the subsequent calculation of the jitter flag and the rising trend flag results is improved, and finally the purpose of improving the accuracy of the application healthiness calculation result is achieved.
[0036] As an example, data of each monitoring item is collected at a fixed time interval (such as 60s), and the collected data is stored in the hardware monitoring history data table. This hardware monitoring history data table is a Hive table, and the detailed information such as the affiliation of the host, each monitoring item of the host and the corresponding values, and time information is stored in the Hive table. Figure 2It is a schematic diagram of a piece of information recorded in the hardware monitoring historical data table according to an embodiment of the present application. As Figure 2 shown, its meaning is as follows: An application named aerospike runs in multiple clusters, including the aerospike-fp-data cluster. This cluster includes a host with the hardware identifier "iface": "em1". This host is located in the production computer room, and the value of a certain monitoring item of this host at a certain timestamp on November 17, 2021 is 0.0.
[0037] S103. When performing the resource utilization analysis task, calculate the jitter identification and rising trend identification of each monitoring item according to the time-series monitoring item data. Preferably, first group the time-series monitoring item data, and then calculate the jitter identification and rising trend identification of each monitoring item in batches.
[0038] As an example, an E-R relationship model is pre-constructed to represent the subordination relationship between applications, clusters, computer rooms, hosts, and monitoring items. For example:
[0039] An application is maintained by one person in charge, but one person in charge can release multiple applications;
[0040] An application may run on multiple clusters, but one cluster can only provide services for one application; one cluster may be deployed in multiple computer rooms, and one computer room can deploy multiple clusters;
[0041] A computer room houses multiple hosts, and one host is only placed in one computer room;
[0042] One host only provides services for one application, and one application can run on multiple hosts;
[0043] There may be multiple pieces of the same type of hardware in one host. For example, a certain host may have multiple CPU processors;
[0044] One piece of hardware or service may be monitored by multiple monitoring items, and one monitoring item can only monitor one piece of hardware or service. For example, a "CPU utilization" monitoring item can only monitor one CPU processor.
[0045] Therefore, according to the above E-R relationship model, the number of clusters where the application is located can be counted, as well as the number of hosts included in each cluster. With reference to the memory size and the total number of hosts, the maximum number of hosts that can be processed in each batch is determined, and the clusters are grouped by a fixed number of hosts to obtain the time series monitoring item data of each group of clusters. Here, it can be implemented through the existing greedy algorithm of "change-making at the supermarket checkout". For example, if the fixed number of input hosts is 11, in reality, cluster A contains 1 host, cluster B contains 3 hosts, and cluster C contains 5 hosts. Then, for one batch after grouping, the time series monitoring item data in 1 cluster A, 0 cluster B, and 2 cluster C are processed. Therefore, through batch processing of data, each batch only processes the time series monitoring item data of a fixed number of hosts, reducing the computational intensity and also reducing the requirements for CPU computing power, memory capacity, I / O speed, and network stability during single-machine computing.
[0046] Optionally, according to the E-R relationship model, layer-by-layer screening is performed to obtain the time series monitoring item data at the computer room or even cluster level, which can be stored in the computer room monitoring history data table. For example, Figure 3 is a schematic diagram of a piece of information recorded in the computer room monitoring history data table according to the embodiment of the present application. As Figure 3 shown, its meaning is: The application named aerospike runs in multiple clusters, including the aerospike-fp-Indonesia cluster, which is located in the IDNU computer room, and the CPU utilization rate of this host is relatively high at 12:00 on November 9, 2021, with an average value of approximately 1.67. Compared with the Figure 2 hardware monitoring history data table shown above, Figure 3 the computer room monitoring history data table shown improves the production level of data from the host to the computer room, making it more convenient to analyze the utilization rate of various resources.
[0047] Optionally, for the hardware monitoring history data table, the time series monitoring item data collected due to reasons such as disk I / O and network latency are cleaned to reduce interference and make the accuracy of the application health degree calculated subsequently higher.
[0048] In addition, in the embodiment of the present application, the intermediate data produced by the subtasks in the latter two stages can be saved to an external hive table, which enables the data to also provide data support for the development of other services.
[0049] In some of these embodiments, according to the grouped time series monitoring item data, the jitter flag and the rising trend flag of each monitoring item are calculated in batches, and the jitter flag and the rising trend flag can be saved to the anomaly detection history data table. The specific calculation methods of the jitter flag and the rising trend flag are as follows:
[0050] First, calculate the jitter flag of the monitored item: Determine whether there is jitter within each preset time period. The manifestation of jitter is that the curve representing the data of the monitored item in the image shows a fluctuating pattern of being alternately high and low, including "abnormal upward spike", "abnormal downward spike", etc. The categories of jitter can be divided into slight, small, relatively large, and large in terms of the jitter amplitude; and can be divided into single-time and continuous in terms of the jitter duration. The embodiments of the present application mainly judge large jitters, that is, only when the large jitter level is reached is it considered that the meaning of jitter recorded in the embodiments of the present application is reached.
[0051] Here, it is necessary to judge whether the number of jitters is greater than the preset number threshold (the empirical value is 10 times), or whether the maximum value of the jitter duration is greater than the preset time threshold (the empirical value is 30 minutes). If so, the jitter flag is represented as jitter anomaly. Figure 4 It is a schematic diagram of a piece of information recorded in the anomaly detection historical data table according to the embodiments of the present application, as Figure 4 shown, and its meaning is: The application named elves runs in the elves cluster, and this cluster is located in the HZ computer room. The CPU utilization rate of the monitored item in this computer room had a jitter anomaly at 17:00 on December 2, 2021, that is, the jitter flag is shake.
[0052] As for the method of judging whether there is jitter, for example, subtract the data of the monitored item at adjacent times (time t and time t-1) and take the absolute value to obtain the amplitude. If the amplitude is not greater than the amplitude threshold, no further judgment is made; otherwise, judge whether the change trends of the monitored item data at adjacent times are opposite. If so, it is judged as jitter.
[0053] As for the method of judging whether the change trends of the monitored item data at adjacent times are opposite, for example, subtract the value of the monitored item at time t from the value of the monitored item at time t-1 to obtain the first difference; subtract the value of the monitored item at time t+1 from the value of the monitored item at time t to obtain the second difference; if the product of the first difference and the second difference is negative, it is judged that the change trends of the monitored item values at time t and time t-1 are opposite.
[0054] Next, calculate the upward trend flag of the monitored item: For example, perform a continuity detection on the time-series monitored item data. In the case of uninterrupted data, compare the monitored item at the previous moment with the monitored item data at the current moment. If it is an upward trend, detect the length of the upward sequence. If the length of the upward sequence is greater than the length threshold, the upward trend flag of the monitored item corresponding to the upward sequence is represented as an upward trend anomaly until the following situations occur: ① In the case of a stable trend, record the number of times of the stable appearance. If it is higher than the number threshold, it is determined that the upward trend ends; ② In the case of a downward trend, it is determined that the upward trend ends, and then the upward trend flag of the corresponding monitored item is represented as a normal upward trend.
[0055] In some of these embodiments, continuity detection is performed on the time-series monitoring item data. Assuming the detection period is 1 hour, that is, data is collected every minute. Therefore, the amount of data for one detection period of a monitoring item of a certain host is 60. If the amount of data for one detection period is no more than 60, the detection is carried out normally; if the amount of data for one detection period is more than 60, the detection is stopped and this part of the data is considered problematic because in reality, changes in business requirements or computer room maintenance may occur, new applications are released or the original applications are taken off the shelf (or suspended), resulting in insufficient data. However, under normal circumstances, the amount of data for one detection period will not be more than 60.
[0056] As an example, time-series monitoring item data for 2 days, that is, 48 hours, can be collected. Assuming that continuous detection has been passed and there is no discontinuous data, 48 data will be obtained. Next, the value at the previous moment (t - 1 moment) is compared with the value at the current moment (t moment) in turn. If it is an upward trend, these two values and their corresponding time periods are recorded. Repeat this step until one of the following three situations is encountered:
[0057] Situation 1: It is found that the curve (i.e., the trend) is stable. Record the number of times of stability. If the number of times is higher than 5, it is judged that the upward trend ends; detect the length of the previous sequence. If the length of the continuously rising sequence is greater than 12, it is considered that an "upward trend" anomaly exists in this monitoring item of the host. If the subsequent data volume meets the minimum detection volume, continue the detection; if the number of stable times is not higher than 5, only record the number of stable times and do not end the current detection.
[0058] Situation 2: It is found that the curve drops. At this time, it is determined that the continuous upward trend ends. Still detect the length of the previous sequence. If the sequence length is greater than 12, it is considered that an "upward trend" anomaly is detected in this monitoring item of the host; if the subsequent data volume meets the minimum detection volume, continue the detection.
[0059] Situation 3: All the data in this time period have been detected, and the last continuously rising sequence is detected. If the length of the continuously rising sequence is greater than 12, it is considered that an "upward trend" anomaly is detected in this monitoring item of the host.
[0060] S104, when performing the score calculation task, score each monitoring item according to the time-series monitoring item data, the jitter flag, and the upward trend flag, and calculate the score representing the application health.
[0061] As an example, each monitoring item is scored according to different rules based on the time-series monitoring item data, the jitter flag, and the upward trend flag to obtain the scores of each monitoring item; then the scores of each monitoring item are weighted and summed according to the specified weights to obtain the comprehensive anomaly score; finally, the score representing the application health is calculated according to the comprehensive anomaly score.
[0062] For example, each monitoring item is scored according to different rules based on the time-series monitoring item data, jitter flag, and upward trend flag to obtain the score of each monitoring item, as follows:
[0063] (1) For CPU utilization, the monitoring item is expressed as "cpu.busy", the full score of the anomaly score is 100 points, the default is 0 points, and the scoring is as follows:
[0064] Scoring item 1: If p85 ∈ (90, +∞), score 40 points; if p85 ∈ [80, 90], score 30 points; if p85 ∈ (0, 80), score 0 points;
[0065] Scoring item 2: If there is a jitter anomaly, score 30 points; otherwise, score 0 points;
[0066] Scoring item 3: If there is an upward trend anomaly, score 30 points; otherwise, score 0 points.
[0067] (2) For memory utilization, the monitoring item is expressed as "mem.memuse.percent", the full score of the anomaly score is 100 points, the default is 0 points, and the scoring is as follows:
[0068] Scoring item 1: If p85 ∈ (90, +∞), score 50 points; if p85 ∈ [80, 90], score 40 points; if p85 ∈ (0, 80), score 0 points;
[0069] Scoring item 2: If there is an upward trend anomaly, score 50 points; otherwise, score 0 points.
[0070] (3) For disk occupancy rate, the monitoring item is expressed as "df.bytes.used.percent", the full score of the anomaly score is 100 points, the default is 0 points, and the scoring is as follows:
[0071] Scoring item 1: If p85 ∈ (90, +∞), score 50 points; if p85 ∈ [80, 90], score 40 points; if p85 ∈ (0, 80), score 0 points;
[0072] Scoring item 2: If there is an upward trend anomaly, score 50 points; otherwise, score 0 points.
[0073] (4) For the number of garbage collections, the monitoring item is expressed as "old.gc.count", the full score of the anomaly score is 100 points, the default is 0 points, and the scoring is as follows:
[0074] Scoring item 1: If the maximum value max_val > 1, score 30 points;
[0075] Scoring item 2: If there is a jitter anomaly, score 30 points; otherwise, score 0 points;
[0076] Scoring item 3: If there is an upward trend anomaly, score 40 points; otherwise, score 0 points.
[0077] (5) The applied alarm volume, expressed as "falcon, prometheus", if greater than 5, is scored 100 points; otherwise, it is scored 0 points.
[0078] (6) The alarm waiting response duration, if greater than 24 hours, is scored 100 points; otherwise, it is scored 0 points.
[0079] It should be noted that for some monitoring items, if the scoring item does not involve jitter anomalies or upward trend anomalies, it is considered that the scores corresponding to the jitter identifier or upward trend identifier are 0.
[0080] According to the above content, adding up the scores of each scoring item can obtain the scores of each monitoring item. Then, the scores of each monitoring item can be weighted and summed according to the specified weights to obtain the comprehensive anomaly score, and the application health score can be calculated based on the comprehensive anomaly score. For example, the application health score = 100 - comprehensive anomaly score.
[0081] Optionally, the above comprehensive anomaly score can be saved in the application health information table, and the detailed scores of each monitoring item can also be saved in this application health information table. Figure 5 It is a schematic diagram of a piece of information recorded in the application health information table according to an embodiment of the present application, as Figure 5 shown, which means that the comprehensive anomaly score of the application named qiming-td on December 1, 2021 is 6.0 points. For example, the score of CPU utilization is 30.0 points, and the score of memory utilization is 0.0 points.
[0082] Optionally, according to the application health score, evaluate the four states of the application: excellent, good, medium, and poor. The specific method is as follows:
[0083] Excellent: α3 <= 100 - comprehensive anomaly score;
[0084] Good: α2 <= 100 - comprehensive anomaly score < α3;
[0085] Medium: α1 <= 100 - comprehensive anomaly score < α2;
[0086] Poor: 100 - application comprehensive anomaly score < α1;
[0087] That is: Poor < α1 <= Medium < α2 <= Good < α3 <= Excellent, where α1 = 40, α2 = 80, and α3 = 90.
[0088] Based on the above, embodiments of the present application can start from the perspective of operation and maintenance, analyze the time-series monitoring item data of clusters, computer rooms, and hosts that support the daily operation of "applications", so as to reflect the hardware working status, and then batch calculate the application health. According to the application health, information is notified to the corresponding operation and maintenance personnel, which helps to quickly troubleshoot and solve problems that may exist in the application, such as serious resource utilization jitter, approaching resource exhaustion, and untimely alarm handling, continuously optimize the application, and ensure the efficient and stable operation of the application.
[0089] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.
[0090] In addition, in combination with the method for batch calculating the application health in the above embodiments, embodiments of the present application can provide a storage medium to implement. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the methods for batch calculating the application health in the above embodiments is implemented.
[0091] In an embodiment of the present application, an electronic device is also provided, and the electronic device can be a terminal. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for batch calculating the application health is implemented. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0092] In one embodiment, Figure 6 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. As Figure 6 shown, an electronic device is provided, and the electronic device can be a server. Its internal structure diagram can be as Figure 6As shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected through an internal bus. Among them, the non-volatile memory stores an operating system, computer programs, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with external terminals through a network connection. The internal memory is used to provide an environment for the operation of the operating system and computer programs. The computer program, when executed by the processor, implements a method for batch calculating the health of an application. The database is used to store data.
[0093] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0094] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in this application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0095] Those skilled in the art should understand that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0096] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for batch calculating the health degree of an application, characterized in that, it includes: Decompose the application health degree calculation task into sub-tasks in multiple stages, where the sub-tasks in multiple stages include basic monitoring and inspection tasks, resource utilization rate analysis tasks, and score calculation tasks; When executing the basic monitoring and inspection task, collect time series monitoring item data; When executing the resource utilization rate analysis task, group the time series monitoring item data and calculate the jitter identification and rising trend identification of each monitoring item in batches; When executing the score calculation task, score each monitoring item according to the time series monitoring item data, the jitter identification, and the rising trend identification, and calculate the score representing the application health degree; Among them, the batch calculation of the jitter identification of each monitoring item includes: based on the grouped time series monitoring item data, within a preset time period, determine whether there is jitter. If the number of jitters is greater than the preset number threshold, or the maximum value of the jitter duration is greater than the preset time threshold, the jitter identification is represented as abnormal jitter; The batch calculation of the rising trend identification of each monitoring item includes: perform continuity detection on the grouped time series monitoring item data. In the case of uninterrupted data, compare the monitoring item data at the previous moment with the monitoring item data at the current moment. If it is an upward trend, detect the length of the rising sequence. If the length of the rising sequence is greater than the length threshold, the rising trend identification of the monitoring item corresponding to the rising sequence is represented as abnormal rising trend until the following situations are encountered: in the case of a stable trend, record the number of times of stability. If it is higher than the number threshold, it is determined that the rise ends; in the case of a downward trend, it is determined that the rise ends.
2. The method according to claim 1, characterized in that, the grouping of the time series monitoring item data includes: Obtain a pre-constructed E-R relationship model, where the E-R relationship model is used to express the subordination relationship between applications, clusters, computer rooms, hosts, and monitoring items; According to the E-R relationship model, count the number of clusters where the application is located and the number of hosts included in each cluster; Group the clusters by a fixed number of hosts to obtain the time series monitoring item data of each group of clusters.
3. The method according to claim 1, characterized in that, the determination of whether there is jitter includes: Subtract the monitoring item data at adjacent moments and take the absolute value to obtain the amplitude; In the case where the amplitude is greater than the amplitude threshold, determine whether the change trend of the monitoring item data at the adjacent moments is opposite. If so, it is determined as jitter.
4. The method according to claim 3, characterized in that, the determination of whether the change trend of the monitoring item values at the adjacent moments is opposite includes: Subtract the monitoring item value at time t from the monitoring item value at time t-1 to obtain a first difference; Subtract the monitoring item value at time t+1 from the monitoring item value at time t to obtain a second difference; If the product of the first difference and the second difference is negative, it is determined that the change trend of the monitoring item values at time t and time t-1 is opposite.
5. The method according to any one of claims 1-4, characterized in that, The monitored items include CPU utilization rate, memory utilization rate, disk occupancy rate, number of garbage collections, application alarm volume, and alarm waiting response duration, and the scoring rules for each of the monitored items are different based on the time-series monitored item data, the jitter flag, and the upward trend flag.
6. An electronic device, comprising a memory and a processor, wherein, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the method according to any one of claims 1 to 4.
7. A storage medium, wherein, a computer program is stored in the storage medium, and the computer program is configured to execute the method according to any one of claims 1 to 4 when running.
Citation Information
Patent Citations
Task ranking system for electric power system operating state online / offline evaluation hybrid scheduling
CN103593724A
Method and system for evaluating cloud software health degree
CN103902442A