Macroscopic monitoring method and device for computer cluster and electronic equipment

By collecting and calculating the ratio of the variance of the index data of each server device to the 99-quantile value in the computer cluster, the problem of low monitoring efficiency in the prior art is solved, and accurate positioning and efficient detection of middleware abnormalities are achieved.

CN120196497APending Publication Date: 2025-06-24NETSUNION CLEARING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311773879.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing computer cluster monitoring methods are inefficient and cannot accurately locate abnormal middleware and causes.

Method used

By collecting multiple indicator data in each server device in the same computer room or in the same cluster, the ratio of the variance of each indicator data to the 99-quantile value is calculated as the standard difference value, and it is monitored as macro monitoring data.

Benefits of technology

It realizes efficient detection of abnormal situations of individual servers in the computer cluster, can accurately locate target middleware for abnormal operation, and improves monitoring efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196497A_ABST
    Figure CN120196497A_ABST
Patent Text Reader

Abstract

The invention provides a macroscopic monitoring method and device for a computer cluster and electronic equipment. The method comprises the following steps of: executing the following operations on to-be-monitored middleware: acquiring various index data of target middleware in the same machine room and / or server equipment in the same cluster according to an inspection cycle; statistical calculation is carried out on each index data of each server device through the following method: calculating the variance of the target index data of each server device; calculating a ratio of the variance of the target index data of each server device to a 99 quantile, and taking the ratio as a standard difference value of the target index data corresponding to the current inspection period; and monitoring by taking the standard difference value and the maximum value of various index data in the same machine room and / or the same cluster as macroscopic monitoring data. According to the scheme, the monitoring efficiency of the computer cluster can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology and can be used in the financial field. In particular, it relates to a method, device, and electronic device for macro monitoring of a computer cluster. Background Art

[0002] A computer cluster is a high-performance computing system that connects multiple computers in a local area network or the Internet and realizes resource sharing and task allocation through software and hardware to jointly complete a task. Computer clusters are mainly used in fields such as large-scale data processing, scientific computing, and digital media processing, and can provide higher computing efficiency and larger data storage capacity. Monitoring various indicators of a computer cluster can timely detect problems during the operation of the cluster system, so as to timely solve problems and avoid losses.

[0003] Existing computer cluster monitoring methods usually monitor indicators such as the memory usage, CPU occupancy, and disk occupancy of each server device in the computer cluster. For example, the min-max normalization method is used to normalize indicator data such as CPU occupancy, memory usage, and disk occupancy, and then the normalized indicator data is input into a pre-trained network model, and whether the computer cluster is operating abnormally is determined according to the output result of the network model.

[0004] However, the above monitoring method has low efficiency. Summary of the Invention

[0005] This specification provides a method, device, and electronic device for macro monitoring of a computer cluster to solve problems such as low efficiency of existing monitoring methods.

[0006] To solve the above technical problems, the first aspect of this specification provides a method for macro monitoring of a computer cluster, including performing the following operations on the middleware to be monitored: obtaining various indicator data of each server device in the same computer room and / or the same cluster of the target middleware according to the inspection cycle; performing statistical calculations on the indicator data of each server device through the following methods: determining the maximum value and the 99th percentile value in the target indicator data of each server device; calculating the variance of the target indicator data of each server device; calculating the ratio of the variance of the target indicator data of each server device to the 99th percentile value, and using the ratio as the standard deviation value of the target indicator data corresponding to the current inspection cycle; monitoring the standard deviation value and the maximum value of various indicator data in the same computer room and / or the same cluster as macro monitoring data.

[0007] In some embodiments, the 99th percentile value is calculated by the following method: sorting the target metric data of each server device from smallest to largest; taking the 99%×nth target metric data in the sorted sequence from smallest to largest as the 99th percentile value, where n is a natural number.

[0008] In some embodiments, after monitoring the standard deviation value and the maximum value of various metric data within the same computer room and / or the same cluster as the macro monitoring data, it further includes: screening out the target standard deviation values greater than or equal to a predetermined threshold; generating an alarm message, where the alarm message is used to indicate that the values of the metric data corresponding to the cluster or computer room of the target standard deviation value are abnormal.

[0009] In some embodiments, the value range of the predetermined threshold is 4% to 6%.

[0010] In some embodiments, the method further includes: analyzing the cause of the anomaly based on the alarm message; the cause of the anomaly includes at least one of the following: capacity restriction of server devices, continuous increase in service traffic, sporadic service traffic, sporadic failures of the network or devices, software vulnerabilities.

[0011] In some embodiments, the inspection cycle is daily; correspondingly, after monitoring the standard deviation value and the maximum value of various metric data within the same computer room and / or the same cluster as the macro monitoring data, it further includes: summarizing the average values of the maximum values and the 99th percentile values of the metrics of each middleware monitored daily; summarizing the daily data weekly and recording the average values of the maximum values and the 99th percentile values per week; summarizing the weekly data monthly and recording the average values of the maximum values and the 99th percentile values per month; summarizing the monthly data annually and recording the average values of the maximum values and the 99th percentile values per year; the daily inspection report shows the synchronous and month-on-month growth data calculated daily, weekly, and annually.

[0012] In some embodiments, the middleware is a middleware for processing financial business requests.

[0013] The second aspect of this specification provides a macro monitoring device for a computer cluster, including: an acquisition unit for acquiring various indicator data of each server device of a target middleware in the same computer room and / or the same cluster according to an inspection cycle; a calculation unit for statistically calculating the indicator data of each server device through the following means: determining the maximum value and the 99th percentile value in the target indicator data of each server device; calculating the variance of the target indicator data of each server device; calculating the ratio of the variance of the target indicator data of each server device to the 99th percentile value, and using the ratio as the standard deviation value of the target indicator data corresponding to the current inspection cycle; a determination unit for monitoring the standard deviation value and the maximum value of various indicator data in the same computer room and / or the same cluster as macro monitoring data.

[0014] In some embodiments, the device further includes: a screening unit for screening out target standard deviation values whose standard deviation values are greater than or equal to a predetermined threshold; an alarm unit for generating an alarm message, where the alarm message is used to indicate that the value of the indicator data of the cluster or computer room corresponding to the target standard deviation value is abnormal.

[0015] In some embodiments, the inspection cycle is daily; correspondingly, the device further includes: a first summarization unit for summarizing the average values of the maximum value and the 99th percentile value of the indicator data of each middleware monitored daily; a second summarization unit for summarizing the daily data weekly and recording the average values of the maximum value and the 99th percentile value of each week; a third summarization unit for summarizing the weekly data monthly and recording the average values of the maximum value and the 99th percentile value of each month; a fourth summarization unit for summarizing the monthly data annually and recording the average values of the maximum value and the 99th percentile value of each year; a fifth summarization unit for displaying the synchronous and year-on-year growth data calculated daily, weekly, and annually in the daily inspection report.

[0016] The third aspect of this specification provides an electronic device, including: a memory and a processor, which are communicatively connected to each other, and the memory stores computer instructions, and the processor realizes the steps of the method described in any one of the first aspects by executing the computer instructions.

[0017] The fourth aspect of this specification provides a computer storage medium, which stores computer program instructions, and when the computer program instructions are executed, the steps of the method described in any one of the first aspects are realized.

[0018] The fifth aspect of this specification provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of the first aspects are realized.

[0019] The macroscopic monitoring method of the computer cluster provided in this specification can achieve multi-dimensional monitoring of various clusters, computer rooms, and various index data, and realize horizontal summarization; by using the standard deviation value obtained by dividing the variance by the 99th percentile value and the maximum value, the abnormal conditions of individual servers in the computer cluster can be discovered, and the target middleware with abnormal operation can be efficiently detected.

[0020] The macroscopic monitoring method of the computer cluster provided in this specification has the following beneficial effects by using the standard deviation value obtained by dividing the variance by the 99th percentile value as the macroscopic monitoring data of the computer cluster: 1. By directly processing various index data of the middleware in the computer cluster, it can more accurately locate whether the computer is abnormal and accurately locate which middleware is abnormal, improving the monitoring efficiency; 2. Compared with the min-max normalization method that needs to mix all the data in all clusters and computer rooms for processing, and when new clusters or computer rooms to be monitored need to be added, all the collected data of all clusters and computer rooms need to be recalculated, and the processing method is very inflexible. In this solution, by using the method of dividing the variance by the 99th percentile value, only the 99th percentile value and the variance need to be determined within the collected data of the clusters or computer rooms to be monitored. When new clusters or computer rooms to be monitored are added, only the 99th percentile value and the variance need to be determined within the newly added clusters or computer rooms, without reprocessing the monitoring data of other clusters or computer rooms. The processing method is very flexible and improves the monitoring efficiency. Brief Description of the Drawings

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 Shows a schematic diagram of the macroscopic monitoring method of the computer cluster provided in this specification;

[0023] Figure 2 Shows another schematic diagram of the macroscopic monitoring method of the computer cluster provided in this specification;

[0024] Figure 3 Shows yet another schematic diagram of the macroscopic monitoring method of the computer cluster provided in this specification;

[0025] Figure 4 Shows a schematic diagram of the macroscopic monitoring device of the computer cluster provided in this specification;

[0026] Figure 5 The structural schematic diagram of the electronic device provided in this specification is shown. Specific embodiments

[0027] In order to enable those skilled in the art of this technology to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0028] In the technical solutions of this application, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.

[0029] In the existing computer cluster monitoring methods, usually, indicators such as the memory usage, CPU occupancy, and disk occupancy of each server device in the computer cluster are monitored. However, multiple middleware may be deployed in a server device. There may be a situation where the indicator data of one middleware is relatively high, but the indicator data of other middleware is relatively low, and finally the indicator data of the server device does not belong to an abnormal situation. That is to say, only monitoring the overall indicators of the server device may not be able to detect abnormal middleware, thus unable to accurately locate the abnormality and the cause of the abnormality (i.e., which middleware has an abnormality), resulting in a low efficiency of the monitoring method.

[0030] In addition, in the existing computer cluster monitoring methods, the min-max normalization method is used to normalize indicator data such as CPU occupancy, memory usage, and disk occupancy, and then the normalized indicator data is input into a pre-trained network model, and it is determined whether the operation of the computer cluster is abnormal according to the output result of the network model. Since the volume of business requests processed by the computer cluster is changing, and the resource volume of the server device is also changing, and different volumes of business requests and resource volumes of the server device will affect the monitoring results of the network model, therefore, the above monitoring method needs to spend a lot of time to frequently update the network model, and each time the network model is updated, a lot of time and effort are required to collect appropriate training data, which results in a low efficiency of the above monitoring method.

[0031] Middleware is a type of software between application systems and system software. It uses the basic services (functions) provided by system software to connect various parts of application systems or different applications on the network, and can achieve the purpose of resource sharing and function sharing. It does not have a very strict definition. The general definition is: middleware is an independent system software service program. Distributed application software uses this software to share resources between different technologies. Middleware is located on the operating system of the client server and manages computing resources and network communications. In this sense, middleware can be expressed by an equation: middleware = platform + communication, which limits it to being called middleware only when it is used in distributed systems, and also distinguishes it from support software and utility software.

[0032] The middleware in the computer cluster includes: cache middleware, storage middleware, message queue middleware, etc. Specifically, for example, Zookeeper, Redis, Kafka. Each middleware uses different resources (CPU, memory, disk), and the configuration of server equipment in each computer room and cluster is not exactly the same. Therefore, it is impossible to simply judge whether a middleware is in a normal state based on the value of the indicator (such as CPU usage, memory usage, disk usage). For example, the CPU usage of the same middleware in computer room A is x, the memory usage is y, and the disk usage is z. Then, based on the values ​​x, y, and z, it is impossible to know whether the middleware is running normally in computer room A.

[0033] In a computer cluster, a balancing strategy is usually adopted, that is, the business traffic (i.e., the number of business requests) is evenly distributed to each server device in the cluster. In addition, for the convenience of operation and maintenance, the configuration of server devices in the same cluster or computer room is usually consistent. Therefore, the same indicators of each server device in the same computer room and the same cluster of the same middleware should theoretically be consistent. Based on this theoretical basis, the macro-monitoring method of the computer cluster provided in this specification proposes to collect indicator data of each server device of the target middleware "in the same computer room and / or the same cluster", and use the ratio of the variance to the 99th percentile as monitoring data.

[0034] This specification provides a macro-monitoring method for a computer cluster, which can perform simple processing on various indicator data with less computational effort, and the user can clearly see the abnormal situation of the middleware based on the simple processing results.

[0035] like Figure 1 As shown, the macro monitoring method of the computer cluster includes the following steps for the middleware to be monitored:

[0036] S10: Acquire various indicator data of each server device in the same computer room and / or the same cluster of the target middleware according to the inspection cycle.

[0037] The target middleware is a middleware to be monitored.

[0038] When monitoring a computer cluster, metric data is usually obtained at regular intervals, and the interval period is the inspection cycle. The inspection cycle can be daily, every N hours (N is a natural number), or set every M minutes (M is a natural number).

[0039] "Multiple metric data of the target middleware on each server device in the same computer room and / or the same cluster" can be multiple metric data of the target middleware on each server in the same computer room, or multiple metric data of the target middleware on each server device in the same cluster, or multiple metric data of the target middleware on each server device in the same computer room and the same cluster.

[0040] Table 1 below shows an example of the collected data of the target middleware Zookeeper in a certain cluster and each computer room.

[0041] Table 1

[0042]

[0043] S20: Statistically calculate the metric data of each server device through the following method: determine the maximum value and the 99th percentile value in the target metric data of each server device; calculate the variance of the target metric data of each server device; calculate the ratio of the variance of the target metric data of each server device to the 99th percentile value, and use the ratio as the standard deviation value of the target metric data corresponding to the current inspection cycle.

[0044] S30: Monitor the standard deviation values and maximum values of various metric data in the same computer room and / or the same cluster as macro monitoring data.

[0045] For example, the metrics X of each server device located in the same computer room (or in the same cluster, or in the same cluster and the same computer room) are respectively: x1, x2, x3... xn (n is a natural number). Then, determine the maximum value and the 99th percentile value among x1, x2, x3... xn, and calculate the variance of x1, x2, x3... xn; then calculate var, and then calculate var / 99th percentile value, and use var / 99th percentile value as the standard deviation value, and use the standard deviation value as the monitoring data of metric X to reflect the macro health report of the computer cluster.

[0046] Similarly, the metrics Y of each server device located in the same computer room (or in the same cluster, or in the same cluster and the same computer room) are respectively: y1, y2, y3... yn (n is a natural number). Then, determine the maximum value and the 99th percentile value among y1, y2, y3... yn, and calculate the variance of y1, y2, y3... yn; then calculate var, and then calculate var / 99th percentile value. Take the var / 99th percentile value as the standard deviation value, and use the standard deviation value as the monitoring data of the metric Y to reflect the macroscopic health report of the computer cluster.

[0047] The 99th percentile value is calculated by the following method: Sort the target metric data of each server device from smallest to largest, and take the 99% × nth target metric data in the sorted sequence from smallest to largest as the 99th percentile value, where n is a natural number.

[0048] This solution reflects the deviation of the metric data of each server device in the cluster and / or computer room by calculating the variance of the metric data of each server device. Since the computer cluster adopts an equilibrium strategy and the configurations of server devices in the same cluster or computer room are usually the same, therefore, in theory, the variance of each server device in a cluster and / or computer room with better overall operating conditions should be very small and close to 0. If the variance is large, it means that the operating state of at least one server device in the cluster and / or computer room is abnormal. For example, a large difference in the value of a certain metric of each server device in the cluster or computer room can reflect problems such as uneven disk usage or uneven task allocation of the middleware.

[0049] However, it is difficult to determine exactly how much the variance should be greater than to be considered abnormal, and this is the key to data monitoring. In response to this, this solution proposes to divide the variance by the 99th percentile value to obtain the standard deviation value, and use the standard deviation value as the macroscopic data of the computer cluster for monitoring.

[0050] The macroscopic monitoring method of the computer cluster provided in this specification can achieve multi-dimensional monitoring of each cluster, each computer room, various metric data, etc., and achieve horizontal aggregation; through the two metrics of the standard deviation value and the maximum value obtained by dividing the variance by the 99th percentile value, it is possible to discover abnormal situations of individual servers in the computer cluster and efficiently detect the target middleware with abnormal operation.

[0051] The macro monitoring method of the computer cluster provided in this specification has the following beneficial effects by using the standard deviation value obtained by dividing the variance by the 99th percentile as the macro monitoring data of the computer cluster: 1. It directly processes various index data of the middleware in the computer cluster, can relatively accurately locate whether the computer has an abnormality, and accurately locate which middleware has an abnormality, improving the monitoring efficiency; 2. Compared with the min-max normalization method that needs to mix all the data in all clusters and all computer rooms for processing, and needs to recalculate all the collected data of all clusters and all computer rooms when adding clusters or computer rooms to be monitored, the processing method is very inflexible. In this solution, the method of dividing the variance by the 99th percentile only needs to determine the 99th percentile and variance within the collected data of the cluster or computer room to be monitored. When adding a cluster or computer room to be monitored, it only needs to determine the 99th percentile and variance in the newly added cluster or computer room, without reprocessing the monitoring data of other clusters or computer rooms. The processing method is very flexible and improves the monitoring efficiency.

[0052] This solution can unify the monitoring data in each cluster and each computer room into the same data range, so that it is convenient to obtain which cluster and which computer room have relatively poor overall operation status of the server devices by comparing the monitoring data in each cluster and each computer room.

[0053] This solution divides the variance by the 99th percentile instead of the maximum value, which can avoid extremely special situations in the collected data of the cluster or computer room from affecting the monitoring data and causing the monitoring data to not reflect the true operation status of most server devices in the cluster or computer room.

[0054] Since the overall operation status of the server devices in the cluster or computer room needs to be determined by the combination relationship of the values of multiple indicators such as CPU occupancy, memory usage, and network traffic, and the value ranges of the variances of each indicator data vary greatly, it is difficult to intuitively determine the combination relationship based on the actual values of the indicator data. This solution uses the standard deviation value obtained by dividing the variance of each indicator data by the 99th percentile as the monitoring data, so that the overall operation status of the server devices in the cluster or computer room can be intuitively judged through the combination relationship of the standard deviation values of each indicator data.

[0055] The definition of the standard deviation value in this solution not only unifies the data of each cluster or computer room into the same value range, but also unifies different indicator data used as monitoring data into the same value range, simplifying the effective monitoring of the computer cluster through simple calculation operations.

[0056] The following uses a specific example in another field to elaborate on the fourth point above: Assume that the metric data is height and income. The height of a cluster is 0.8 meters, and this difference is already very large; while the variance of income is 50, and actually the difference is not large. If only looking at the values 0.8 and 50, it is impossible to evaluate the degree of difference. In the case where the types of metric data are few, perhaps it is possible to manually remember what values of each metric data are considered large. However, the number of types of monitoring metric data of computer clusters is very large, and the monitoring metric data is highly technical, and it is simply impossible to require manual memory of what values of each metric are considered large.

[0057] If further assume that the 99th percentile value of height is 1.9 meters, then the standard deviation value is 0.8÷1.9 = 0.42; assume that the 99th percentile value of income is 50,000 yuan, then the standard deviation value is 50÷50,000 = 0.001. Based on 0.42 and 0.001, it can be concluded that the height difference is relatively large. If 0.42 and 0.001 are the standard deviation values of two metric data of a computer cluster, then it can be judged whether the overall operating state of the server devices in the cluster or computer room is normal according to whether the standard deviation values of these two metrics are within the allowable range of the corresponding metrics.

[0058] In some embodiments, as Figure 2 shown, after step S30, the following steps S40 and S50 are included.

[0059] S40: Screen out the target standard deviation values that are greater than or equal to a predetermined threshold.

[0060] S50: Generate an alarm message, and the alarm message is used to indicate that the value of the metric data of the cluster or computer room corresponding to the target standard deviation value is abnormal.

[0061] According to the foregoing analysis, theoretically, the various metric parameters of each server device in the same cluster or computer room should be the same. Therefore, the standard deviation value should also be very small. Correspondingly, the predetermined threshold in step S40 should also be small. The value range of the predetermined threshold can be 4% to 6%. For example, the predetermined threshold can take one of the following values: 4%, 4.5%, 5%, 5.5%, 6%. Preferably, the predetermined threshold can be 5%. Through practice, it is found that in the case of a predetermined threshold of 5%, using Figure 2 the described method can not only detect all abnormal problems, but also generate fewer false alarm messages. When the predetermined threshold is 6%, some abnormal problems may be difficult to detect, and when the predetermined threshold is 4%, there are relatively more false alarm messages.

[0062] The following Table II is an example of the monitoring results of the memory idle rate of multiple computer rooms of a cluster.

[0063] Table II

[0064] Redis - A Certain Business Cluster Computer Room 1 Computer Room 2 Computer Room 3 Memory Free Rate (%) Average Value: 11 12 13 Memory Free Rate (%) Standard Deviation Value: 3.85 5.48 3.39

[0065] Based on the detection results in Table 2, when the predetermined threshold value corresponding to the memory free rate (%) is 5%, it can be seen that Machine Room 2 is abnormal. Further, the memory free rates of each server in Machine Room 2 are processed, and it is determined that there are 4 server IPs with abnormal memory free rates as shown in Table 3 below, thus locating the servers that may have faults. These servers that may have faults can be monitored key points. When the memory free rate index remains abnormal for a predetermined duration, these servers can be cut off or replaced in time, so as to predict and eliminate in advance before the fault occurs, and avoid data processing errors in the computer cluster.

[0066] To determine the servers with abnormal memory free rates, servers with a memory free rate greater than the 99th percentile value can be used as abnormal servers, or the memory free rates of each server can be sorted from small to large first, calculate the average value of the memory free rates of the first 99%, then calculate the difference between each memory free rate and this average value divided by this average value respectively, and use the server corresponding to the quotient greater than the predetermined threshold as the abnormal server. This predetermined threshold can be 50% - 70%.

[0067] Table 3

[0068] Computer Room IP Memory Free Rate (%) Computer Room 2 10.0.0.1 19 Computer Room 2 10.0.0.2 19 Computer Room 2 10.0.0.3 20 Computer Room 2 10.0.0.4 22

[0069] In some embodiments, the inspection cycle is daily. For example, financial business requests usually show periodic changes on a daily basis, then the inspection cycle of the middleware index data for processing financial business requests can be set to daily.

[0070] Correspondingly, as Figure 3 shown, after step S30, the following steps S61 to S65 may further be included.

[0071] S61: Summarize the maximum value and the average value of the 99th percentile of each index data of each middleware monitored daily.

[0072] S62: Summarize the daily data by week, and record the maximum value and the average value of the 99th percentile of each week.

[0073] The maximum value of a week can be directly determined according to the maximum value determined daily, and the 99th percentile values determined daily are averaged to obtain the average value of the 99th percentile of a week.

[0074] S63: Summarize the weekly data by month, and record the maximum value and the average value of the 99th percentile of each month.

[0075] The maximum value for a month can be directly determined based on the maximum value determined weekly. The 99th percentile values determined weekly are averaged to obtain the average of the 99th percentile values for a month.

[0076] S64: Aggregate the monthly data by year, and record the annual maximum value and the average of the 99th percentile values.

[0077] The maximum value for a year can be directly determined based on the maximum value determined monthly. The 99th percentile values determined monthly are averaged to obtain the average of the 99th percentile values for a year.

[0078] S65: The daily inspection report shows the synchronous and month-on-month growth data calculated by day, week, and year.

[0079] Steps S61 to S65 do not need to directly process the original data collected in each inspection cycle, so there is no need to store a large amount of collected data, reducing the memory occupation.

[0080] Furthermore, in some embodiments, the method further includes the following step S70.

[0081] S70: Analyze the reasons for the anomaly based on the alarm information; the reasons for the anomaly include at least one of the following: capacity constraints of server devices, continuous increase in business traffic, occasional business traffic, occasional network or device failures, software vulnerabilities.

[0082] By aggregating the monitoring data according to the time dimension, it is possible to predict in a timely manner the risks that may occur in the future from the development trend of the monitoring data, so as to avoid risks in a timely manner.

[0083] For example, in the case where it is determined that an anomaly occurs based on the alarm information, it can be manually judged by a professional post to analyze the reasons for the anomaly. The results of the judgment include the following situations:

[0084] (1) First, analyze whether the abnormal indicator triggers the capacity management threshold. If it triggers, in response to the alarm information, work such as capacity expansion can be initiated. The evaluation before the capacity expansion work will analyze the change trend of business traffic in the past year, the change trend of capacity indicators, major business changes that may affect capacity, etc. to determine the target value of the capacity expansion.

[0085] (2) At the same time, analyze the reasons for the anomaly:

[0086] a) If the reason for the abnormal indicator is business traffic, continue to evaluate through the business department whether the business traffic is occasional or needs to be considered normally; if it is not occasional, relevant capacity expansion work needs to be done.

[0087] b) If the cause of the abnormal indicator is an occasional failure such as network or server, the SA (security audit) team can be contacted to analyze the logs, enhance monitoring, and find the cause; after finding the cause, it can be resolved by methods such as system software version upgrade or hardware replacement; if the cause cannot be found in a short time, a fault tolerance mechanism can be added to ensure that the software design can handle the situation, and enhance monitoring to analyze the pattern of abnormal occurrences.

[0088] c) If the cause of the abnormal indicator is a software BUG, the software version will be updated after repair.

[0089] This specification provides a macroscopic monitoring device for a computer cluster, which can be used to execute Figure 1 the macroscopic monitoring method of the computer cluster shown. As Figure 4 shown, the device includes an acquisition unit 10, a calculation unit 20, and a determination unit 30.

[0090] The acquisition unit 10 is used to acquire various index data of each server device of the target middleware in the same computer room and / or the same cluster according to the inspection cycle.

[0091] The calculation unit 20 is used to perform statistical calculations on the index data of each server device through the following means: determine the maximum value and the 99th percentile value in the target index data of each server device; calculate the variance of the target index data of each server device; calculate the ratio of the variance of the target index data of each server device to the 99th percentile value, and use the ratio as the standard difference value of the target index data corresponding to the current inspection cycle.

[0092] The determination unit 30 is used to monitor the standard difference value and the maximum value of various index data in the same computer room and / or the same cluster as macroscopic monitoring data.

[0093] In some embodiments, the device further includes a screening unit and an alarm unit.

[0094] The screening unit is used to screen out the target standard difference values whose standard difference values are greater than or equal to a predetermined threshold.

[0095] The alarm unit is used to generate an alarm message, and the alarm message is used to indicate that the value of the index data of the cluster or computer room corresponding to the target standard difference value is abnormal.

[0096] In some embodiments, the inspection cycle is daily; correspondingly, the device further includes a first summary unit, a second summary unit, a third summary unit, a fourth summary unit, and a fifth summary unit.

[0097] The first summary unit is used to summarize the average values of the maximum value and the 99th percentile value of the index data of each middleware monitored every day.

[0098] The second summarization unit is used to summarize the daily data on a weekly basis and record the weekly maximum value and the average value of the 99th percentile.

[0099] The third summarization unit is used to summarize the weekly data on a monthly basis and record the monthly maximum value and the average value of the 99th percentile.

[0100] The fourth summarization unit is used to summarize the monthly data on an annual basis and record the annual maximum value and the average value of the 99th percentile.

[0101] The fifth summarization unit is used to display the daily inspection report on a daily, weekly, and annual basis to calculate synchronous and year-on-year growth data.

[0102] The descriptions and functions of the above devices can be understood by referring to the content of the section on the macro monitoring method of the computer cluster, and will not be elaborated here.

[0103] An embodiment of the present invention also provides an electronic device, as Figure 5 shown. The electronic device may include a processor 501 and a memory 502, where the processor 501 and the memory 502 may be connected through a bus or other means, Figure 5 taking the connection through the bus as an example.

[0104] The processor 501 may be a central processing unit (CPU). The processor 501 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.

[0105] The memory 502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the macro monitoring method of the computer cluster in the embodiment of the present invention (for example, Figure 4 the obtaining unit 10, the calculating unit 20, and the determining unit 30 shown). The processor 501 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 502, that is, implements the macro monitoring method of the computer cluster in the above method embodiments.

[0106] The memory 502 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor 501 and the like. In addition, the memory 502 may include a high-speed random access memory and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 502 may optionally include a memory remotely disposed relative to the processor 501, and these remote memories may be connected to the processor 501 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0107] The one or more modules are stored in the memory 502 and, when executed by the processor 501, perform the foregoing macroscopic monitoring method of the computer cluster.

[0108] Specific details of the above electronic device can be understood by referring to the corresponding relevant descriptions and effects in the method embodiments, and will not be elaborated here.

[0109] This specification also provides a computer storage medium storing computer program instructions, and when the computer program instructions are executed, the steps of the foregoing macroscopic monitoring method of the computer cluster are implemented.

[0110] This specification also provides a computer program product including a computer program, and when the computer program is executed by a processor, the steps of the foregoing macroscopic monitoring method of the computer cluster are implemented.

[0111] Those skilled in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when the program is executed, it may include the processes of the embodiments of the above methods. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.

[0112] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized.

[0113] The systems, devices, modules or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.

[0114] For the convenience of description, when describing the above devices, they are divided into various units according to their functions and described separately. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0115] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of certain parts of each embodiment of the present application.

[0116] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0117] The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0118] Although the present application is depicted through embodiments, those of ordinary skill in the art know that the present application has many variations and changes without departing from the spirit of the present application. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of the present application.

Claims

1. A macroscopic monitoring method for a computer cluster, characterized in that Including performing the following operations on the middleware to be monitored: Obtaining various metric data of each server device of the target middleware in the same computer room and / or the same cluster according to the inspection cycle; Statistically calculating the metric data of each server device through the following method: determining the maximum value and the 99th percentile value in the target metric data of each server device; Calculating the variance of the target metric data of each server device; Calculating the ratio of the variance of the target metric data of each server device to the 99th percentile value, and using the ratio as the standard deviation value of the target metric data corresponding to the current inspection cycle; Monitoring the standard deviation values and the maximum values of various metric data in the same computer room and / or the same cluster as the macro monitoring data.

2. The method according to claim 1, characterized in that, The 99th percentile value is calculated through the following method: Sorting the target metric data of each server device from small to large; Taking the 99%×nth target metric data in the sorted sequence from small to large as the 99th percentile value, where n is a natural number.

3. The method according to claim 1, wherein After monitoring the standard deviation values and the maximum values of various metric data in the same computer room and / or the same cluster as the macro monitoring data, it further includes: Screening out the target standard deviation values whose standard deviation values are greater than or equal to the predetermined threshold; Generating an alarm message, where the alarm message is used to indicate that the value of the metric data of the cluster or computer room corresponding to the target standard deviation value is abnormal.

4. The method according to claim 3, wherein The value range of the predetermined threshold is 4% to 6%.

5. The method according to claim 3, characterized in that, The method further includes: Analyzing the cause of the anomaly according to the alarm message; the cause of the anomaly includes at least one of the following: capacity restriction of the server device, continuous increase in business traffic, occasional business traffic, occasional failure of the network or device, software vulnerability.

6. The method according to claim 1, wherein The inspection cycle is daily; correspondingly, after monitoring the standard deviation values and the maximum values of various metric data in the same computer room and / or the same cluster as the macro monitoring data, it further includes: Summarizing the average values of the maximum values and the 99th percentile values of the metric data of each middleware monitored every day; Summarizing the daily data weekly and recording the average values of the maximum values and the 99th percentile values per week; Summarizing the weekly data monthly and recording the average values of the maximum values and the 99th percentile values per month; Summarizing the monthly data annually and recording the average values of the maximum values and the 99th percentile values per year; The daily inspection report shows the synchronous and month-on-month growth data calculated daily, weekly, and annually.

7. The method according to claim 5, characterized in that The middleware is the middleware for processing financial business requests.

8. A macro monitoring device for a computer cluster, characterized in that, Including: An acquisition unit for obtaining various metric data of each server device of the target middleware in the same computer room and / or the same cluster according to the inspection cycle; A calculation unit for statistically calculating the metric data of each server device through the following means: determining the maximum value and the 99th percentile value in the target metric data of each server device; Calculating the variance of the target metric data of each server device; Calculating the ratio of the variance of the target metric data of each server device to the 99th percentile value, and using the ratio as the standard deviation value of the target metric data corresponding to the current inspection cycle; A determination unit is used to monitor the standard difference value and the maximum value of various metric data within the same computer room and / or the same cluster as macro monitoring data.

9. An electronic device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. Computer instructions are stored in the memory, and the processor realizes the steps of the method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer storage medium, characterized in that, The computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the method according to any one of claims 1 to 7 are realized.

11. A computer program product, characterized in that, It contains a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are realized.