Adaptive adjustment method and device for monitoring index collection and storage medium

By evaluating the status of the message queue and the message change rate in real time, and dynamically adjusting the acquisition configuration of monitoring indicators, the problem of mismatch between monitoring indicator collection and processing capabilities in large-scale computing systems is solved, and the stability and reliability of the system are achieved.

CN120223592AActive Publication Date: 2025-06-27国家超级计算天津中心 +1
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510662147.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-27
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

In large-scale computing systems, the number of monitoring indicators collected and generated rates change dynamically, resulting in the traditional static configuration being unable to match the processing capabilities of the monitoring system, resulting in system instability.

Method used

By evaluating the total length of the message queue and the message change rate in real time, dynamically adjusting the acquisition configuration of monitoring indicators to ensure that the upstream inflow rate matches the downstream processing rate.

Benefits of technology

The dynamic balance between the collection and configuration of monitoring indicators and the processing capabilities of monitoring systems is achieved, avoiding the abnormal impact of system fluctuations on monitoring indicators, and improving the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223592A_ABST
    Figure CN120223592A_ABST
Patent Text Reader

Abstract

The invention discloses a self-adaptive adjustment method and device for monitoring index collection and a storage medium. The method comprises the steps that the total length of a message queue at the t moment is determined; if the total length is smaller than or equal to zero, monitoring the total length of the message queues at the preset number of time points after the t moment, and adjusting the acquisition configuration of the monitoring indexes according to the total length of the message queues at the preset number of time points after the t moment; and if the total length is greater than zero, adjusting the acquisition configuration of the monitoring index according to the change rates of the messages in the message queues at the moment t and the preset number of time points after the moment t, and the change rates of the messages in the message queues at a plurality of time points before the moment t. According to the technical scheme, the collection number of the monitoring indexes can be dynamically adjusted according to the real-time collection condition of the monitoring indexes so as to adapt to the digestion processing capacity of the monitoring system on the monitoring indexes, and dynamic balance between the monitoring coverage degree and the monitoring processing capacity is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of monitoring index collection, and particularly to an adaptive adjustment method, device and storage medium for monitoring index collection. Background Art

[0002] Industrial control systems, data centers, supercomputing systems, etc. often involve a large number of computing nodes. Each node cooperates through communication to complete computing services. When an abnormality or error occurs in the hardware unit where a certain node is located, the system will enter an unstable operating state. And due to the large scale of the system, the probability of having problem nodes will also increase accordingly. The monitoring system provides in-band (monitoring indexes related to the operating system, such as CPU occupancy rate, memory occupancy rate, network IO) and out-of-band (monitoring indexes related to hardware, such as temperature, current, voltage, etc.) operating state monitoring of the system by collecting, observing and analyzing the operation data of each node (also called monitoring indexes) in real time, and realizes operation guarantees such as health analysis and anomaly prediction.

[0003] The number of monitoring indexes grows proportionally with the number of monitored object nodes and the number of services. Taking a supercomputing system as an example, when the number of nodes reaches 5000, even if the number of in-band monitoring indexes per single node is 56, the daily in-band monitoring index data of the whole system can reach 18GB. The number of out-of-band monitoring indexes is generally 300 per month, and the daily out-of-band monitoring index data of the whole system can reach 200GB, and a large amount of alarms are not considered here. Moreover, the generation rate of monitoring indexes fluctuates dynamically under the influence of various system events and anomalies.

[0004] Therefore, it is necessary to control and manage the collection of monitoring indexes to match the processing capacity and storage capacity of the monitoring system.

[0005] In view of this, the present invention is specifically proposed. Summary of the Invention

[0006] In order to solve the above technical problems, the present invention provides an adaptive adjustment method, device and storage medium for monitoring index collection, which can dynamically adjust the collection quantity of monitoring indexes according to the real-time collection situation of monitoring indexes, so as to adapt to the digestion and processing capacity of the monitoring system for monitoring indexes, and realize the dynamic balance between the monitoring coverage and the monitoring processing capacity.

[0007] In a first aspect, an embodiment of the present invention provides an adaptive adjustment method for monitoring index collection, and the method includes:

[0008] Determine the total length of the message queue at time t;

[0009] If the total length is less than or equal to zero, monitor the total length of the message queue at a preset number of time points after time t, and adjust the acquisition configuration of the monitoring metrics according to the total length of the message queue at the preset number of time points after time t;

[0010] If the total length is greater than zero, adjust the acquisition configuration of the monitoring metrics according to the change rate of the messages in the message queue at time t and at a preset number of time points after time t, and the change rate of the messages in the message queue at multiple time points before time t;

[0011] Among them, the monitoring metrics are temporarily stored in the message queue after being acquired, and dequeue from the message queue when being analyzed and processed.

[0012] In a second aspect, an embodiment of the present invention provides an electronic device, the electronic device includes:

[0013] a processor and a memory;

[0014] The processor is configured to execute the steps of the adaptive adjustment method for monitoring metric acquisition according to any one of the embodiments by calling a program or an instruction stored in the memory.

[0015] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, the computer-readable storage medium stores a program or an instruction, and the program or the instruction causes a computer to execute the steps of the adaptive adjustment method for monitoring metric acquisition according to any one of the embodiments.

[0016] The embodiments of the present invention have the following technical effects:

[0017] By determining the total length of the message queue at a certain moment, it is possible to determine whether the acquisition configuration of the monitoring metrics matches the processing capacity of the monitoring system for the monitoring metrics. When the acquisition configuration of the monitoring metrics does not match the processing capacity of the monitoring system for the monitoring metrics, the acquisition configuration of the monitoring metrics is dynamically adjusted, so as to achieve the purpose of matching the acquisition configuration of the monitoring metrics with the processing capacity of the monitoring system for the monitoring metrics, and realize the dynamic balance between the monitoring coverage and the monitoring processing capacity. Specifically, if the total length of the message queue at time t is zero or less than zero, it indicates that the processing capacity of the monitoring system is sufficient to adapt to the current acquisition configuration. To avoid the abnormal impact of system fluctuations on the monitoring metrics, on this premise, continue to monitor the total length of the message queue at a preset number of time points after time t, and adjust the acquisition configuration of the monitoring metrics according to the total length of the message queue at a preset number of time points after time t, so as to improve the reliability and accuracy of the adjustment method. If the total length of the message queue at time t is greater than zero, it means that there are remaining unprocessed monitoring metrics, indicating that the processing capacity of the monitoring system is not sufficient to adapt to the current acquisition configuration. On this premise, to avoid the abnormal impact of system fluctuations on the monitoring metrics, the acquisition configuration of the monitoring metrics is adjusted according to the change rate of the messages in the message queue at time t and at a preset number of time points after time t, as well as the change rate of the messages in the message queue at multiple time points before time t. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is the flowchart of an adaptive adjustment method for monitoring metric acquisition provided by an embodiment of the present invention Figure 1 ;

[0020] Figure 2 is the flowchart of an adaptive adjustment method for monitoring metric acquisition provided by an embodiment of the present invention Figure 2 ;

[0021] Figure 3 is the structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope protected by the present invention.

[0023] Generally, the monitoring system of a large-scale computing system generally operates at level 2 (or above), respectively using the collection agents and BMU (Unit Management Unit) deployed in the nodes to collect the in-band status (monitoring metrics related to the operating system, such as CPU occupancy, memory occupancy, network IO) and out-of-band status (monitoring metrics related to hardware, such as temperature, current, voltage, etc.) of the large-scale computing system. The collected monitoring metrics are confirmed by the metric whitelist in the configuration file. The collected metrics are forwarded by the intermediate monitoring server (for monitoring systems with three or more levels) or directly sent to the system-level monitoring server. Given the large number of nodes in a large-scale system, the monitoring server often cannot complete the storage and processing of all node monitoring metrics in real time. Therefore, the system-level monitoring server usually uses a message queue (such as Kafka) as a buffer middleware to send all the monitoring metrics reported by the nodes into multiple message queues, and the backend storage component and the upper-layer monitoring and analysis component carry out monitoring services by consuming the message data in the queues. The occurrence of abnormal collection of monitoring metrics is dynamic, and the traditional method of statically configuring the monitoring metric collection list and collection frequency does not have the ability to flexibly operate and solve abnormalities. In response to this, the technical solution of the present invention proposes a dynamic adjustment strategy for monitoring metrics based on backpressure, aiming to dynamically adjust the amount and frequency of the collection metrics of the underlying nodes (upstream) through the real-time evaluation of the system-level (downstream) queue length and node rate, so that the upstream inflow rate and the downstream processing rate match, thereby stabilizing the data production and consumption process.

[0024] The technical solution of the present invention mainly includes an anomaly assessment component and a dynamic adjustment component based on backpressure, which are respectively used to detect whether the reporting of monitoring metrics is too fast, difficult to process, etc., and adaptively adjust the collection configuration when abnormal reporting of monitoring metrics is found, so as to achieve the purpose of matching the collection configuration of monitoring metrics with the processing ability of the monitoring system for monitoring metrics, and realizing the dynamic balance between the monitoring coverage and the monitoring processing ability.

[0025] Among them, the functional implementation of the anomaly evaluation component includes: the collection agent continuously sends node monitoring metrics to the message queue of the system monitoring server. The production operation of the message (i.e., the monitoring metric) entering the queue will increase the length of the message queue. The system monitoring service consumes the messages from the message queue for the storage and processing of the monitoring metrics. The consumption operation of the message leaving the queue will shorten the length of the message queue. In the system server, by observing the production and consumption conditions of m message queues, the production rate of each message queue at time t can be calculated , the consumption rate , where i = 1, 2, 3....m, representing the identifier of the message queue. When there are abnormal situations such as too fast or too slow collection of monitoring metrics, there will be a trend of message accumulation or sudden drop in the message queue. This trend can be examined through rate difference. Therefore, calculate the difference between the production rate and the consumption rate of the th message queue . This difference can also be called the change rate of the message. Accumulate the differences between the production rates and consumption rates of all message queues to obtain the rate difference between the enqueueing and dequeueing of the global monitoring metric .

[0026] Considering that in large-scale systems, distributed nodes are prone to network transmission anomalies, monitoring metric anomalies, etc., resulting in frequent changes in the change rate of messages in the message queue. However, not all changes in the production rate and consumption rate represent abnormal collection of monitoring metrics. The abnormal collection of monitoring metrics that affects the system should show a certain persistence. Therefore, maintain a dynamic sliding window to record the statistical time series data of the difference between the production rate and the consumption rate within a period of time , thereby smoothing the interference of instantaneous or short-term fluctuations on the detection of abnormal collections and improving the stability and accuracy of detection. Specifically, at time, calculate the maximum value and the minimum value of the rate differences recorded in the sliding window respectively as the upper limit threshold and the lower limit threshold of the rate difference at the current time to determine whether exceeds the expected range. If it exceeds, it is marked as a significant fluctuation. To avoid misjudgment caused by short-term data fluctuations, when is determined to be a significant fluctuation, the system will continue to detect the rate differences at the subsequent n times (where n is the specified number of consecutive monitoring times). If in the consecutive n times of monitoring, the signs of the rate differences are the same as the sign of the rate difference at time and all exceed the limit, it is determined that there is an abnormal collection of monitoring metrics, and the corresponding subsequent processing strategy is triggered; otherwise, it is considered as short-term data fluctuation and no processing is required. This method effectively improves the ability to identify abnormal data and reduces the possibility of misjudgment

[0027] The functional implementation of the dynamic adjustment component based on backpressure includes:

[0028] At the system management server, obtain the current moment The lengths of each message queue (where i = 1, 2, 3....m), sum up the lengths of each message queue to obtain the total message queue length . First, by checking Whether it is 0 to determine whether the queue is in a piled-up state. If is equal to 0, it means that the consumption rate is greater than or equal to the production rate. Detect n times in a loop. If there is no change, control the script, reverse modify the acquisition configuration, and increase the acquisition frequency in an additive increase manner (such as the acquisition interval -a) or increase the number of metrics ( ), so as to increase the data production volume.

[0029] When it is determined that the consumption rate is less than the production rate through the total message queue length , further combine the change rate for determination. Specifically, when and , it is determined that the system is in a high-load state (HIGH_LOAD), control the script, reverse modify the acquisition configuration, and reduce the acquisition frequency in a multiplicative decrease manner (the acquisition interval ) or decrease the number of metrics ( ).

[0030] When and , it is determined that the system is in a low-load state (LOW_LOAD), control the script, reverse modify the acquisition configuration, and increase the acquisition frequency in an additive increase manner (the acquisition interval -a) or increase the number of metrics ( ), so as to increase the data production volume.

[0031] Generally speaking, the adaptive adjustment method for monitoring metric acquisition provided by the embodiments of the present invention aims to dynamically adjust the acquisition metric quantity and acquisition frequency of the underlying nodes (upstream) through real-time evaluation of the system-level (downstream) queue length and message change rate, so that the upstream inflow rate and downstream processing rate match, thereby stabilizing the data production and consumption processes. The adaptive adjustment method for monitoring metric acquisition provided by the embodiments of the present invention can be executed by an electronic device.

[0032] Figure 1 is the flowchart of an adaptive adjustment method for monitoring metric acquisition provided by the embodiments of the present invention. See Figure 1, the adaptive adjustment method for collecting the monitoring metrics specifically includes the following steps:

[0033] S110. Determine the total length of the message queue at time t.

[0034] Among them, time t can refer to any specific moment. For example, the current moment is marked as time t.

[0035] The number of message queues can be multiple. In this scenario, the total length refers to the sum of the lengths of each message queue. For example represents the length of the i-th message queue at time t, , and m represents the total number of message queues.

[0036] S120. If the total length is less than or equal to zero, monitor the total length of the message queue at a preset number of time points after time t, and adjust the acquisition configuration of the monitoring metrics according to the total length of the message queue at the preset number of time points after time t.

[0037] Among them, if the total length is less than or equal to zero, it means that the consumption rate of the message queue is greater than or equal to the production rate, indicating that the processing capacity of the monitoring system is sufficient to adapt to the current acquisition configuration.

[0038] Furthermore, considering that the acquisition of monitoring metrics is easily affected by system fluctuations, therefore, in order to avoid abnormal effects of system fluctuations on monitoring metrics and improve detection reliability, continue to monitor the total length of the message queue at a continuous number of time points after time t, and adjust the acquisition configuration of the monitoring metrics according to the total length of the message queue at the preset number of time points after time t.

[0039] Exemplarily, if the sum of the total lengths of the message queue at the preset number of time points after time t is zero, control the acquisition frequency of the monitoring metrics to increase, and / or control the number of monitoring metrics in one acquisition to increase.

[0040] That is, if , control the acquisition frequency of the monitoring metrics to increase, and / or control the number of monitoring metrics in one acquisition to increase. Among them, represents the total length of the message queue at time t + i.

[0041] S130. If the total length is greater than zero, adjust the acquisition configuration of the monitoring metrics according to the change rate of the messages in the message queue at time t and at a preset number of time points after time t, and the change rate of the messages in the message queue at multiple time points before time t.

[0042] Among them, the monitoring metrics are temporarily stored in the message queue after being acquired, and dequeue from the message queue when the monitoring metrics are analyzed and processed.

[0043] If the total length of the message queue at time t is greater than zero, it indicates that there are remaining unprocessed monitoring metrics, which means that the processing capacity of the monitoring system is insufficient to adapt to the current collection configuration. On this premise, in order to avoid the abnormal impact of system fluctuations on monitoring metrics, the collection configuration of monitoring metrics is adjusted according to the change rate of messages in the message queue at time t and at a preset number of time points after time t, as well as the change rate of messages in the message queue at multiple time points before time t.

[0044] Among them, the change rate of messages in the message queue is the difference between the generation rate and the consumption rate of messages. When there are multiple message queues, the change rate of messages in the message queue is the sum of the respective change rates of each message queue, that is, the difference between the production rate and the consumption rate of all message queues is accumulated to obtain the difference between the enqueue and dequeue rates of global monitoring metrics. 。

[0045] Considering that network transmission anomalies, monitoring metric anomalies, etc. are likely to occur in distributed nodes of large-scale systems, resulting in frequent changes in the change rate of messages in the message queue, but not all changes in production rate and consumption rate represent abnormal monitoring metric collection. Abnormal monitoring metric collection that affects the system should exhibit a certain persistence. Therefore, a dynamic sliding window is maintained to record the statistical time series data of the difference between the production rate and the consumption rate within a period of time to smooth the interference of instantaneous or short-term fluctuations on abnormal collection detection and improve the stability and accuracy of detection. Specifically, at time, calculate the maximum value and the minimum value in the rate difference recorded in the sliding window respectively as the upper threshold and the lower threshold of the rate difference at the current moment to determine whether it exceeds the expected range. If it exceeds, it is marked as a significant fluctuation. To avoid misjudgment caused by short-term data fluctuations, when is determined to be a significant fluctuation, the system will continue to detect the rate difference at the subsequent n time points (where n is the specified number of consecutive monitoring times). If the signs of the rate differences in the consecutive n monitoring times are the same as the sign of the rate difference at time and all exceed the limit, it is determined that the monitoring metric collection is abnormal, and the corresponding subsequent processing strategy is triggered; otherwise, it is considered as short-term data fluctuation and no processing is required. This method effectively improves the ability to identify abnormal data and reduces the possibility of misjudgment.

[0046] Exemplarily, adjusting the acquisition configuration of the monitoring metrics according to the change rate of the messages in the message queue at the time point t and the preset number of time points after the time point t, and the change rate of the messages in the message queue at multiple time points before the time point t includes:

[0047] Determine the upper limit threshold and the lower limit threshold according to the change rate of the messages in the message queue at multiple time points before the time point t; if the change rate of the messages in the message queue at the time point t is greater than the upper limit threshold and the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t is greater than the upper limit threshold (this indicates that the message production rate is greater than the consumption rate and exceeds the processing capacity range of the monitoring system), then control the acquisition frequency of the monitoring metrics to decrease, and / or control the number of monitoring metrics in one acquisition to decrease, so that the acquisition configuration of the monitoring metrics and the processing capacity of the monitoring system reach dynamic balance.

[0048] If the change rate of the messages in the message queue at the time point t is not greater than the upper limit threshold or the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t is not greater than the upper limit threshold, then compare the change rate of the messages in the message queue at the time point t and the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t with the lower limit threshold, and adjust the acquisition configuration of the monitoring metrics according to the comparison result.

[0049] The comparing the change rate of the messages in the message queue at the time point t and the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t with the lower limit threshold, and adjusting the acquisition configuration of the monitoring metrics according to the comparison result includes:

[0050] If the change rate of the messages in the message queue at the time point t is less than the lower limit threshold, and the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t is less than the lower limit threshold (this indicates that the message production rate is less than the consumption rate and the processing capacity of the monitoring system is not fully utilized. To balance the monitoring coverage and comprehensiveness, the acquisition frequency of the monitoring metrics can be increased at this time, or the acquisition quantity can be increased), then control the acquisition frequency of the monitoring metrics to increase, and / or control the number of monitoring metrics in one acquisition to increase.

[0051] If the change rate of the messages in the message queue at the time point t is not less than the lower limit threshold, or the average value of the change rate of the messages in the message queue at the preset number of time points after the time point t is not less than the lower limit threshold, then keep the acquisition frequency of the monitoring metrics and the number of monitoring metrics in one acquisition unchanged.

[0052] In some embodiments, reducing the number of monitored metrics in one collection by the control includes: determining a first target number in a multiplicative subtraction manner, and reducing the number of monitored metrics in one collection by the first target number. Among them, multiplicative subtraction is a mathematical strategy used when adjusting parameters such as the number or frequency of metric collection, that is, reducing according to a certain ratio. For example, assuming the initial number of metric collection is x, when multiplicative subtraction operation is required, multiply by a coefficient r less than 1 to get the new collection number y = x×r. If r = 0.5, then after each multiplicative subtraction operation, the number of metric collection will become half of the original. The purpose of such setting is to avoid excessive adjustment, and at the same time can quickly reduce the resource consumption of the monitoring system and relieve the processing pressure of the monitoring system.

[0053] Alternatively, determine the number of times of continuously controlling the reduction of the number of monitored metrics in one collection; determine a matching first target metric according to the number of times; control the first target metric to be deleted from the collection list. This way can selectively determine which monitored metrics to give up for collection. Preferably, give up less important monitored metrics first to ensure the monitoring effect as much as possible when the number of monitored metrics is reduced.

[0054] Increasing the number of monitored metrics in one collection by the control includes:

[0055] Determining a second target number in an additive increase manner, and increasing the number of monitored metrics in one collection by the second target number; Additive increase is a way of increasing by a fixed value. For example, assuming the initial number of metric collection is x, when additive increase operation is required, add a fixed increment a to get the new collection number y = x + a. The purpose of such setting is to stably increase the consumption of the processing resources of the monitoring system, accurately control the growth amplitude, and ensure the stability of the system.

[0056] Alternatively, determine the number of times of continuously controlling the increase of the number of monitored metrics in one collection; determine a matching second target metric according to the number of times; control the second target metric to be added to the collection list. This way can selectively determine which monitored metrics to increase for collection. Preferably, increase more important monitored metrics first to ensure the monitoring effect.

[0057] Reducing the collection frequency of the monitored metrics by the control includes:

[0058] Controlling the reduction of the collection frequency of the monitored metrics in a multiplicative subtraction manner; or, identifying the configured value of the collection frequency in the collection configuration of the monitored metric; calculating the difference between the configured value and the set value; modifying the configured value of the collection frequency in the collection configuration of the monitored metric to the difference.

[0059] Increasing the collection frequency of the monitored metrics by the control includes:

[0060] Control the acquisition frequency of the monitoring index to increase in an additive increase manner; or, identify the configured value of the acquisition frequency in the acquisition configuration of the monitoring index; calculate the sum of the configured value and the set value; modify the configured value of the acquisition frequency in the acquisition configuration of the monitoring index to the sum.

[0061] Correspondingly, reference can be made to the Figure 2 flow schematic diagram of an adaptive adjustment method for monitoring index acquisition as shown, specifically including the following steps:

[0062] S1. Obtain the total length len(t) of all message queues, and the difference △s(t) between the production rate and the consumption rate at time t.

[0063] S2. Determine whether len(t) is greater than zero. If so, execute S3; otherwise, execute S4.

[0064] S3. Determine the upper limit threshold and the lower limit threshold for anomaly inspection.

[0065] S4. Determine whether holds. If so, execute S5.

[0066] Among them, represents the sum of the total lengths of all message queues at n time points after time t, represents the total length of all message queues at time t + i.

[0067] S5. Increase the acquisition frequency of the upstream or increase the number of indicators.

[0068] S6. Determine whether △s(t) > upd(t) and holds. If so, execute S7; otherwise, execute S8.

[0069] Among them, upd(t) represents the upper limit threshold.

[0070] represents the average value of the difference between the production rate and the consumption rate at n time points after time t. represents the difference between the production rate and the consumption rate of all message queues at time t + i.

[0071] S7. Decrease the acquisition frequency of the upstream or reduce the number of indicators.

[0072] S8. Determine whether △s(t) < lwd(t) and holds. If so, execute S9; otherwise, execute S10.

[0073] Among them, lwd(t) represents the lower limit threshold.

[0074] S9. Increase the upstream collection frequency or the number of metrics.

[0075] S10. Keep the collection configuration unchanged.

[0076] Based on the continuous and significant change in the production-consumption rate difference of the monitoring metric message queue, the embodiments of the present invention detect anomalies in the metric collection process of large-scale computer systems. When metric collection anomalies occur in large-scale computer systems, the number and frequency of source-end metric collections are automatically adjusted in reverse based on the backpressure conduction method, and dynamic adaptation of metric collection is performed through additive increase and multiplicative decrease, so that the upstream inflow rate and the downstream processing rate match, thereby stabilizing the data production and consumption processes.

[0077] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, the electronic device 200 includes one or more processors 201 and a memory 202.

[0078] The processor 201 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 200 to perform desired functions.

[0079] The memory 202 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 201 may run the program instructions to implement the adaptive adjustment method for monitoring metric collection of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters and thresholds may also be stored in the computer-readable storage medium.

[0080] In one example, the electronic device 200 may further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 203 may include, for example, a keyboard, a mouse, etc. The output device 204 may output various information to the outside, including warning prompt information, braking force, etc. The output device 204 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0081] Of course, for simplicity, Figure 3Only some of the components related to the present invention in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 may further include any other appropriate components.

[0082] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps of the adaptive adjustment method for monitoring metric collection provided by any embodiment of the present invention.

[0083] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0084] In addition, an embodiment of the present invention may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the adaptive adjustment method for monitoring metric collection provided by any embodiment of the present invention.

[0085] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but not be limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0086] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one" and / or "the" do not specifically refer to the singular and may also include the plural. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such a process, method or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method or device comprising the said element.

[0087] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0088] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive adjustment method for monitoring index collection, characterized in that Including: Determine the total length of the message queue at time t; If the total length is less than or equal to zero, monitor the total length of the message queue at a preset number of time points after time t, and adjust the acquisition configuration of the monitoring metrics according to the total length of the message queue at the preset number of time points after time t; If the total length is greater than zero, adjust the acquisition configuration of the monitoring metrics according to the change rate of the messages in the message queue at time t and at a preset number of time points after time t, and the change rate of the messages in the message queue at multiple time points before time t; Among them, the monitoring metrics are temporarily stored in the message queue after being acquired, and dequeue from the message queue when the monitoring metrics are analyzed and processed.

2. The method according to claim 1, wherein The adjustment of the acquisition configuration of the monitoring metrics according to the total length of the message queue at a preset number of time points after time t includes: If the sum of the total lengths of the message queue at a preset number of time points after time t is zero, control the acquisition frequency of the monitoring metrics to increase, and / or control the number of monitoring metrics in one acquisition to increase.

3. The method according to claim 1, characterized in that The adjustment of the acquisition configuration of the monitoring metrics according to the change rate of the messages in the message queue at time t and at a preset number of time points after time t, and the change rate of the messages in the message queue at multiple time points before time t includes: Determine the upper threshold and the lower threshold according to the change rate of the messages in the message queue at multiple time points before time t; If the change rate of the messages in the message queue at time t is greater than the upper threshold and the average value of the change rate of the messages in the message queue at a preset number of time points after time t is greater than the upper threshold, control the acquisition frequency of the monitoring metrics to decrease, and / or control the number of monitoring metrics in one acquisition to decrease; If the change rate of the messages in the message queue at time t is not greater than the upper threshold or the average value of the change rate of the messages in the message queue at a preset number of time points after time t is not greater than the upper threshold, compare the change rate of the messages in the message queue at time t and the average value of the change rate of the messages in the message queue at a preset number of time points after time t with the lower threshold, and adjust the acquisition configuration of the monitoring metrics according to the comparison result.

4. The method according to claim 3, characterized in that The comparison of the change rate of the messages in the message queue at time t and the average value of the change rate of the messages in the message queue at a preset number of time points after time t with the lower threshold, and the adjustment of the acquisition configuration of the monitoring metrics according to the comparison result includes: If the change rate of the messages in the message queue at time t is less than the lower threshold, and the average value of the change rate of the messages in the message queue at a preset number of time points after time t is less than the lower threshold, control the acquisition frequency of the monitoring metrics to increase, and / or control the number of monitoring metrics in one acquisition to increase; If the change rate of messages in the message queue at time t is not less than the lower threshold, or the average value of the change rates of messages in the message queue at a preset number of time points after time t is not less than the lower threshold, then the acquisition frequency of the monitoring metrics and the number of monitoring metrics in one acquisition are kept unchanged.

5. The method according to claim 4, wherein The control to reduce the number of monitoring metrics in one acquisition includes: Determining a first target quantity in a multiplicative decrease manner, and controlling the number of monitoring metrics in one acquisition to decrease by the first target quantity; Alternatively, determining the number of consecutive times of controlling the number of monitoring metrics in one acquisition to decrease; determining a matching first target metric according to the number of times; and controlling the first target metric to be deleted from the acquisition list.

6. The method according to claim 4, wherein The control to increase the number of monitoring metrics in one acquisition includes: Determining a second target quantity in an additive increase manner, and controlling the number of monitoring metrics in one acquisition to increase by the second target quantity; Alternatively, determining the number of consecutive times of controlling the number of monitoring metrics in one acquisition to increase; determining a matching second target metric according to the number of times; and controlling the second target metric to be added to the acquisition list.

7. The method according to claim 4, wherein The control to decrease the acquisition frequency of the monitoring metrics includes: Controlling the acquisition frequency of the monitoring metrics to decrease in a multiplicative decrease manner; Alternatively, identifying the configured value of the acquisition frequency in the acquisition configuration of the monitoring metric; calculating the difference between the configured value and the set value; and modifying the configured value of the acquisition frequency in the acquisition configuration of the monitoring metric to the difference.

8. The method according to claim 4, characterized in that The control to increase the acquisition frequency of the monitoring metrics includes: Controlling the acquisition frequency of the monitoring metrics to increase in an additive increase manner; Alternatively, identifying the configured value of the acquisition frequency in the acquisition configuration of the monitoring metric; calculating the sum of the configured value and the set value; and modifying the configured value of the acquisition frequency in the acquisition configuration of the monitoring metric to the sum.

9. An electronic device, characterized in that, The electronic device includes: A processor and a memory; The processor, by invoking the program or instruction stored in the memory, is configured to execute the steps of the adaptive adjustment method for monitoring metric acquisition according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instruction, and the program or instruction causes a computer to execute the steps of the adaptive adjustment method for monitoring metric acquisition according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and memory manager for managing memory

    CN101685409A

  • A method and apparatus for configuring messages of a message queue

    CN109240836A

  • Data transmission method and device and storage medium

    CN110113782A

  • Message queue management method and device

    CN113138860A

  • Dynamic regulation and control log collection and processing method and device and storage medium

    CN115168030A