Method, device, and computer program for storing metric values of a monitored object
By dynamically filtering and reducing client metric values, the problem of excessive data volume in large-scale storage systems is solved, and storage space is saved and query accuracy is guaranteed.
Patent Information
- Application Number
- CN202011625408.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-31
AI Technical Summary
When monitoring and storing client metric values, large-scale storage systems face the problem of excessive data volume, and existing deletion and downsampling methods cannot meet the rapidly growing cluster system needs.
By receiving multiple sets of indicator values collected at multiple time points and selecting fewer indicator values from each set of indicator values, a reduction indicator value is generated for storage. Dynamically update the number of reduced metric values to ensure balance of query accuracy and storage space.
It significantly reduces the amount of monitoring data stored in the database, and also satisfies the precise query of the top monitoring objects in any time interval, ensuring the accuracy of query and optimization of storage space.
Smart Images

Figure CN114691021B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to a method, an electronic device, a computer-readable storage medium, and a computer program product for storing metric values of monitored objects. Background Art
[0002] A typical use case for monitoring a large-scale storage system is to list a number of monitored objects at the top according to a certain metric or measurement. For example, in a monitoring scenario, it is necessary to list, for example, 30 clients with the highest latency, 50 clients with the highest number of disk accesses per second, and so on. Since the distribution of various metric values of the clients is unknown, it is necessary to store all data of all clients in the monitoring database to accurately give the corresponding ranking. This poses a challenge to the capacity of the database. Since the distribution of various metric values of the clients is unknown, it is necessary to store all data of all clients in the monitoring database to accurately give the corresponding accurate ranking. This poses a challenge to the capacity of the database.
[0003] To reduce the amount of data in the database, some existing methods adopt deletion and downsampling. For example, data stored for more than a certain period of time (e.g., more than 30 days) is deleted, or the data is downsampled to a lower resolution (e.g., from data per minute to data per hour). However, since the scale of the cluster system may grow rapidly, the methods of deletion and downsampling may not meet the user's requirements. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for efficiently storing monitoring data, which can significantly reduce the storage space requirement for storing monitoring data while ensuring the query accuracy rate.
[0005] According to a first aspect of the present disclosure, there is provided a method for storing metric values of monitored objects, including: receiving multiple sets of metric values collected at multiple time points within a period of time, each set of metric values including a first number of metric values corresponding to a respective monitored object. The method further includes, for each of the multiple sets of metric values, selecting a second number of metric values from the first number of metric values, the second number being less than the first number, to generate a reduced multiple sets of metric values for storage. The method further includes updating the second number by performing at least one of the following for subsequent reduction of metric values: generating a first list of monitored objects based on at least a portion of the multiple sets of metric values; generating a second list of monitored objects based on at least a portion of the reduced multiple sets of metric values; and updating the second number according to a comparison between the first list and the second list.
[0006] According to a second aspect of the present disclosure, there is provided an electronic device, including: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform actions. The actions include: receiving multiple sets of metric values collected at multiple time points within a period of time, each set of metric values including a first number of metric values corresponding to a respective monitored object. The actions further include, for each of the multiple sets of metric values, selecting a second number of metric values from the first number of metric values, the second number being less than the first number, to generate a reduced multiple sets of metric values for storage. The actions further include updating the second number by performing the following at least once for subsequent reduction of metric values: generating a first list of monitored objects based on at least a part of the multiple sets of metric values; generating a second list of monitored objects based on at least a part of the reduced multiple sets of metric values; and updating the second number according to a comparison between the first list and the second list.
[0007] According to a third aspect of the present disclosure, there is provided a non-transitory computer storage medium including machine-executable instructions that, when executed by a device, cause the device to perform the method as described in the first aspect of the present disclosure.
[0008] According to a fourth aspect of the present disclosure, there is also provided a computer program product including machine-executable instructions that, when executed by a device, cause the device to perform the method as described in the first aspect of the present disclosure.
[0009] In this way, it is possible to significantly reduce the amount of monitored data stored in the database while satisfying the precise query of the top-ranked monitored objects within any time interval.
[0010] It should be understood that the summary section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it intended to be used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Through the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the embodiments of the present disclosure will become more readily understood. In the drawings, multiple embodiments of the present disclosure will be illustrated by way of example and not limitation, where:
[0012] Figure 1A The raw data of exemplary metric values collected within a period of time according to an embodiment of the present disclosure is shown.
[0013] Figure 1BShows the average value of exemplary metric values obtained from raw data over a period of time according to an embodiment of the present disclosure.
[0014] Figure 2 Shows a method for storing metric values of a monitored object according to an embodiment of the present disclosure.
[0015] Figure 3 Shows an exemplary system in which a method for storing metric values of a monitored object according to an embodiment of the present disclosure can be implemented.
[0016] Figure 4 Shows a schematic diagram of the process of dynamically correcting screening values according to an embodiment of the present disclosure. Calibration.
[0017] Figure 5 Shows a schematic diagram of the process 500 of dynamically correcting screening values for different time intervals..
[0018] Figure 6 Shows a curve graph of disk savings rate and accuracy under different data distributions according to an embodiment of the present disclosure.
[0019] Figure 7 Shows a schematic block diagram of a device that can be used to implement an embodiment of the present disclosure. Detailed implementation manners
[0020] Now, the concept of the present disclosure will be described with reference to various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is only for enabling those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that, where feasible, similar or identical reference numerals may be used in the figures, and similar or identical reference numerals may represent similar or identical elements. Those skilled in the art will understand that, from the following description, alternative embodiments of the structures and / or methods described herein may be employed without departing from the principles and concepts of the present disclosure described.
[0021] In the context of the present disclosure, the term "comprising" and its various variants can be understood as open-ended terms, meaning "including but not limited to"; the term "based on" can be understood as "at least partially based on"; the term "one embodiment" can be understood as "at least one embodiment"; the term "another embodiment" can be understood as "at least one other embodiment". Other terms that may occur but are not mentioned here should not be interpreted or defined in a manner contrary to the concepts on which the embodiments of the present disclosure are based, unless explicitly stated.
[0022] As described above, a typical use case for monitoring a large-scale storage system is to list a number of monitored objects at the top according to a certain metric or measurement. For example, a user may use a query such as "set a time period and then list the 30 clients with the largest latency values among all clients". Such a time period can be freely selected, such as "the last 5 minutes", "the last 30 minutes", "the last hour", and so on. Although the user only hopes to view the top-ranked objects, all object data needs to be maintained to support such a query. The reason is that different time period settings may lead to different query results for the top-ranked objects.
[0023] As an example, refer to Figure 1A , which shows the raw data of exemplary metric values collected within a time period. As Figure 1A shown, this data has a matrix form. Each row represents an individual client (c-1, c-2…, c-E), and each column represents the latency value collected every minute from 9:00 to 10:00. Assume that the user wants to query the top 3 clients with the highest latency values over different time periods, sorted by the average latency value. After setting the time period, for example, the user sets three time periods, 9:00 - 9:05, 9:00 - 9:10, 9:00 - 10:00, based on Figure 1A the shown raw data, calculate the average latency value of each client over these three time periods, as Figure 1B shown.
[0024] Figure 1B shows the average value of the exemplary metric values within a time interval based on the raw data. Figure 1B The left part of shows the average latency value of each client within the time interval 9:00 - 9:05. Among them, the clients c-2, c-3, and c-5 shown in shadow are the top three clients. Figure 1B The middle part of shows the average latency value of each client within the time interval 9:00 - 9:10. Among them, the clients c-3, c-5, and c-6 shown in shadow are the top three clients. Figure 1B The right part of shows the average latency value of each client within the time interval 9:00 - 10:00. Among them, the clients c-1, c-5, and c-7 shown in shadow are the top three clients. It can be seen that with different time intervals, the top three clients are different.
[0025] Embodiments of the present disclosure propose a solution that can be referred to as dynamic screening, which can significantly reduce the amount of data in the storage database while satisfying precise queries for the top-ranked monitored objects within any time interval. Different from existing deletion and downsampling solutions that reduce the amount of data after the data has been stored in the database, the solution proposed in the present disclosure screens the data to be stored in the database before writing to the database to reduce the amount of data written to the database.
[0026] In this document, the metric values in the monitoring storage cluster are not limited to latency values. For example, they can also include the number of read and write operations per second, average response time, number of transactions per second, traffic, etc. Additionally, in this document, the time point is not necessarily an instantaneous metric value collected at a specific specific time, but can also include metric values obtained within a time range.
[0027] Specific examples of the present disclosure will be described below. For example purposes, some specific environments, systems, configurations, numerical values, etc. may be involved in the description of the embodiments. It should be understood that these are merely exemplary, and the scope of the present disclosure is not limited to the described example implementations.
[0028] The following refers to Figure 2 and Figure 3 to describe a method for storing metric values of monitored objects according to an embodiment of the present disclosure. Figure 2 FIG. 200 shows a method for storing metric values of monitored objects according to an embodiment of the present disclosure. Figure 3 FIG. shows an exemplary system in which a method for storing metric values of monitored objects according to an embodiment of the present disclosure can be implemented.
[0029] As Figure 2 shown, at block 210, multiple sets of metric values collected at multiple time points within a time period are received. Each set of metric values includes a first number of metric values corresponding to a respective monitored object. In some embodiments, the monitored objects can be multiple clients 311 connected to Figure 3 the storage cluster 310 shown. The storage cluster 310 can be a parallel or distributed storage cluster composed of multiple nodes interconnected via a network. Each node constitutes a unit of the storage system 310, and each node can support thousands or even more clients 311. The clients 311 can back up their data to the storage cluster 310 via the network.
[0030] When the storage cluster 310 has a first number of online clients, their respective metric values, such as protocol latency values, number of I / O operations per second, average response time, number of transactions per second, traffic, etc., can be collected from these clients 311 in real time, not limited thereto. In some embodiments, it can be by Figure 3The illustrated real-time data collector 320 receives multiple sets of metric values collected at multiple time points within a time period from the storage cluster 310. For example, the real-time data collector 320 can, according to settings, regularly (such as every minute, every 5 minutes, every half hour, etc.) receive corresponding metric values regarding the client 311 from the storage cluster 310 at multiple time points within a time period.
[0031] In some embodiments, the real-time data collector 320 can send requests for associated metric values to the storage cluster 310 at these time points and receive multiple sets of metric values as responses. Alternatively or additionally, the storage cluster 310 can push these metric values to the real-time data collector 320 at these time points.
[0032] In some embodiments, the multiple time points for collecting metric values can be evenly distributed (such as but not limited to every minute, every 5 minutes, every half hour) within a time period. Metric values evenly distributed over time can more accurately reflect the working state of the client 311, enabling users to obtain useful analysis results.
[0033] The metric values from the storage cluster 310 can be organized as a set of ordered metric values, where each metric value corresponds to the metric value of a single client. By way of example only, if there are 1000 clients online in the current storage cluster 310, the real-time data collector 320 can collect multiple sets of metric values with a dimension of 1000 at multiple time points. The metric values received from the storage cluster 310 are also referred to as raw data. The raw data received by the real-time data collector 320 can be transmitted to a corrector 330 and a data filter 340 as Figure 3 illustrated. In some embodiments, multiple sets of metric values can be saved and transmitted in the form of a matrix, where the rows correspond to the clients 311 and the columns correspond to specific time points at which the collection actions occur.
[0034] In block 220, for each of the multiple sets of metric values, a second number of metric values is selected from a first number of metric values, the second number being less than the first number, to generate a reduced multiple sets of metric values for storage. In other words, the number of metric values in each set of metric values is reduced from the first number to the second number, such that the storage capacity required to store them is reduced.
[0035] In some embodiments, it can be by Figure 3The data filter 340 shown is used to select a second number of metric values from a first number of metric values. Specifically, the data filter 340 can be configured with a filtering value and use this filtering value to select a second number of metric values from the first number of metric values. According to an embodiment of the present disclosure, the filtering value is dynamically configured by the corrector 330. The corrector 330 can generate a corrected filtering value regularly or irregularly and send it to the data filter 340 for the data filter 340 to filter the received metric values in the next time period.
[0036] In some embodiments, the filtering value can be an integer and be directly used as the second number, in which case the filtering value and the second number can be used interchangeably. Alternatively, the filtering value can be a number in the range of 0 to 1, as the ratio of the second number to the first number. In this case, the data filter 340 can multiply the filtering value by the first number and round up or down as needed to obtain the second number. It should be noted that, compared with the original data, from a matrix perspective, some elements in the matrix of the reduced metric values are empty, that is, sparser, thus saving the storage space required to store the metric values in the monitoring database 360.
[0037] In some embodiments, the data filter 340 can sort the first number of metric values and select the second number of metric values with the highest rankings among the sorted metric values. In the scenario where a user queries the top several monitored objects within a time interval, in this way, the reduced metric values stored in the monitoring database 360 can better match such a query. In some embodiments, the sorting of the metric values can be from largest to smallest or from smallest to largest, depending on the type of the metric values. For example, for latency values, users usually pay more attention to clients 311 with higher latency values to detect potential or already occurred faults.
[0038] Merely by way of example, for instance, when there are 1000 clients online in the storage cluster 310, a set of metric values can include up to 1000 latency values. At this time, the data filter 340 can sort these 1000 latency values from largest to smallest and select the second number (e.g., 100) of latency values from the sorted latency values for storage in the monitoring database 360. Referring to Figure 3 , in the simplified example 380 shown for illustrative purposes only, the data filter 340 receives a set of latency values for clients c-1 to c-E at time point 9:00, then sorts them, and selects and retains the data of clients c-1, c-5, and c-E with the largest latency values when the second number is 3.
[0039] As described above, the data filter 340 generates multiple sets of reduced metric values. Referring to Figure 3The reduced multiple sets of metric values can be transmitted to the database writer 350 in the form of a sparser matrix. The database writer 350 can store the reduced multiple sets of metric values received from the data filter 340 into the monitoring database 360 in real time, periodically, or based on events, thereby reducing the amount of data stored in the monitoring database 360.
[0040] After the reduced multiple sets of metric values are stored in the monitoring database 360, a user can access the monitoring database to query the monitored objects of interest within a certain time interval. As Figure 3 shown, the system 300 further includes a monitoring control 370. The user can use the monitoring control 370 to access the monitoring database 360. For example, the user can query the top N monitored objects within a certain time interval. By way of example only, N = 30, which is typically much smaller than the total number of monitored objects (e.g., compared to 1000 online clients). The monitoring database 360 stores the reduced multiple sets of metric values, and the resulting query results may be different from the query results generated from the original data. In this document, the accuracy rate is considered to evaluate the data filter 340. In some embodiments, the user can also transmit the number N of the monitored objects of interest to the corrector 330 via the monitoring control 370 for correcting the filtering values, which will be described below.
[0041] According to an embodiment of the present disclosure, when the data distribution of the metric values changes or the user's query requirements change, the second number may not meet the accuracy rate requirement, that is, there is a significant difference between the query results generated based on the reduced metric values and the query results generated based on the original data. In this case, the second number can be dynamically updated by the corrector 330 to control the accuracy rate of the query results within an acceptable range, as described in detail below.
[0042] Return Figure 2 , block 230 is executed at least once to update the second number for subsequent reduction of metric values. Specifically, in block 231, a first list of monitored objects is generated based on at least a part of the multiple sets of metric values. In block 232, a second list of monitored objects is generated based on at least a part of the reduced multiple sets of metric values. In block 433, the second number is updated according to the comparison between the first list and the second list. In other words, the degree of reduction of the original data can be dynamically adjusted to simultaneously meet the requirements of the accuracy rate of the query results and the reduction of the storage amount. Here, the first list is generated based on the original data and serves as a true benchmark; the second list is generated based on the reduced data and will be compared with the first list. The comparison result can be used to update the second number.
[0043] The actions described above with reference to block 230 can be performed by Figure 3The corrector 330 shown is used to perform dynamic correction or update of the filtering value. The corrector 330 can dynamically update the filtering value and transmit the updated filtering value to the data filter 340, whereby the data filter 340 can obtain a second number for filtering multiple groups of metric values received in a subsequent time period.
[0044] Figure 4 FIG. shows a schematic diagram of dynamically correcting a filtering value according to some embodiments of the present disclosure. Refer to Figure 4 The process described can be regarded as a specific implementation of block 230 and can be executed by the corrector 330. A buffer (not shown in the figure) can be included in the corrector 330 for storing the raw data 410 within a time period received from the real-time data collector 320. The raw data 410 can be stored in the buffer in the form of a matrix, where the rows of the matrix represent the monitored objects and the columns represent each time point of the collected data within that time period. For example, merely by way of example, if the monitored objects are 1000 clients, the time period is 1 hour, and the latency values of the clients are collected every minute, then the matrix of the raw data stored in the buffer includes 1000 rows and 60 columns of latency values. It should be understood that the length of the time period related to the raw data stored in the buffer is configurable, for example, half an hour, 10 minutes, 5 minutes, etc.; the number or interval of time points within each time period is also configurable, for example, every 30 seconds, 1 minute, 5 minutes, 10 minutes, etc., and is not limited to the above examples.
[0045] In block 420, the corrector 330 can generate a list of the top N monitored objects based on the raw data. In some embodiments, for a time interval, the raw data within that time interval, i.e., the metric values of the corresponding columns of the raw data matrix, is used to generate a first list of monitored objects. This first list of monitored objects can be regarded as the ground truth. Merely by way of example, for the entire time period (e.g., one hour) in the buffer, the raw data is used to calculate the average latency value of each client and sort them, so as to obtain a list of the top (e.g., the top 30) clients within the entire time period. Alternatively or additionally, for a part of the time interval in the buffer, the corresponding part of the raw data can be used to generate a first list of monitored objects within that time interval.
[0046] In block 430, the corrector 330 can set the initial filtering value K = N. Here, N can be, for example, from Figure 3The number of monitored objects of interest to the user received by the monitoring control 370 shown. Then, at block 440, filtered data is generated using the filtering values. Here, similar to the data filter 340 described above, the corrector 330 can select and retain the top K metric values from a set of metric values of the raw data, thereby generating the filtered data. In other words, the metric values in a column of the raw data matrix are sorted, and the top K metric values are selected and retained.
[0047] At block 450, the corrector 330 can generate a second list of the top N monitored objects based on the filtered data. It should be understood that the filtered data used to generate the second list of monitored objects should correspond to the raw data used to generate the first list of monitored objects that serves as the true baseline. Similar to block 420, by way of example only, for the entire time period (e.g., one hour) in the buffer, the filtered data can be used to calculate the average latency value for each client and sort them, so as to obtain a list of the top (e.g., top 30) clients over the entire time period. Alternatively or additionally, for a part of the time interval in the buffer, a corresponding part of the filtered data can be used to generate a list of monitored objects within that time interval. It should be noted that compared to the raw data, from a matrix perspective, some elements in the filtered data are empty, i.e., sparser. In this case, non-empty elements should be considered when generating the second list of monitored objects. By way of example only, a certain client has 60 latency values (e.g., for each minute within one hour) in the raw data, but after filtering, there may only be 10. In this case, when calculating the average latency value of this client, only these 10 latency values are used.
[0048] At block 460, the corrector 330 can calculate the hit rate 360. According to an embodiment of the present disclosure, the hit rate refers to the accuracy rate of the list of monitored objects generated from the filtered data compared to the true baseline. Therefore, the number of monitored objects generated from the filtered data that hit the true baseline can be counted, and the hit rate can be calculated.
[0049] In some embodiments, the corrector 330 can determine the proportion of the same monitored objects in the first list and the second list among the monitored objects in the second list; and determine a reduced second number of subsequent metric values based on the proportion. In other words, when the monitored objects in the second list also appear in the first list, the hit monitored objects are correct, and thus this proportion can be used as the accuracy rate of the list generated from the reduced metric values. For example, by way of example only, in the case where the user's query of interest is to list the top 30 clients, if 18 out of the 30 monitored objects in the list of monitored objects generated from the filtered data appear in the corresponding true baseline, the hit rate can be calculated as 60%.
[0050] At block 470, the corrector 330 can determine whether the hit rate is higher than a threshold. By way of example only, the threshold can be an acceptable accuracy rate, such as, for example, 85%, 90%, 95%, etc., without limitation. If the hit rate exceeds the threshold, indicating that the current screening value or the corresponding second number is acceptable, then at block 490, the current screening value is determined as the screening value for subsequent use. For example, the screening value that meets the threshold can be sent to the data filter 340 to filter data in the next time period.
[0051] If the hit rate is lower than the threshold, proceed to block 480 to update the screening value. For example, the screening value can be increased by a preset amount (K = K + ΔK). By trying a larger screening value and repeating the above blocks 440, 450, 460, and 470 iteratively, the hit rate can be increased until the hit rate. According to an embodiment of the present disclosure, the screening value can be increased iteratively until the hit rate exceeds the threshold. By way of example only, the preset amount can be any integer, such as 1, 2, 3, etc., without limitation.
[0052] In other words, if the proportion of monitored objects in the second list that hit the first list is lower than the threshold, the corrector 330 needs to perform the following steps at least once until the proportion exceeds the threshold: increase the screening value or the second number by a preset amount for updating; regenerate the reduced sets of metric values using the updated screening value or second number; and re-determine the proportion based on the regenerated reduced sets of metric values.
[0053] Additionally, at block 490, the corrector 330 can determine the current screening value as the screening value for subsequent use. For example, the screening value that meets the threshold can be sent to the data filter 340 to filter data in the next time period.
[0054] It can be seen that through the process described above with reference to Figure 4 different degrees of filtered data can be continuously tried to dynamically generate a screening value that meets the accuracy requirement, thereby ensuring the accuracy of the list of monitored objects generated from the subsequent reduced metric values.
[0055] The process of the corrector 330 correcting the filtered value for a single time interval is described above. The corrector 330 can also correct the filtered value for multiple different time intervals. In some embodiments, the columns in the original data used in the correction process can be configurable, that is, the corrector 330 does not have to use all the data in the buffer to correct the filtered value and can use only the original data for a part of the time intervals. For example, by way of example only, the latency values of columns 1 to 5 (the first 5 minutes), columns 1 - 10 (the first 10 minutes), columns 56 - 60 (the last 5 minutes), and columns 51 - 60 (the last 10 minutes) are used, but not limited thereto. Further, the corrector 330 can also correct multiple times for different time intervals to obtain multiple corrected filtered values, and then determine the filtered value for subsequent filtering transmitted to the data filter 340 based on these multiple corrected filtered values. In some embodiments, the maximum value among the multiple corrected filtered values can be transmitted to the data filter 340. Alternatively, the average value among the multiple corrected filtered values can be transmitted to the data filter 340.
[0056] In some embodiments, for different time intervals within the time period of the buffer of the corrector 330, the second number can be updated respectively based on the corresponding parts of the multiple sets of metric values and the corresponding parts of the reduced multiple sets of metric values to obtain multiple alternative second numbers; and the average value or the maximum value of the multiple alternative second numbers is used to update the second number. Thus, an optimized value of the second number for different time intervals can be provided.
[0057] The following refers to Figure 5 Describe such an embodiment in detail. Figure 5 A schematic diagram of a method 500 for dynamically correcting filtered values for different time intervals is shown. The method 500 can be implemented by the corrector 330. Since the time interval of the user query can be arbitrary, for example, by way of example only, a shorter time interval may be 5 minutes or 10 minutes, and a longer time interval may be 1 day, one week or even longer. Therefore, it is necessary to take into account different time intervals to correct the filtered value.
[0058] The method 500 includes respectively determining multiple corrected filtered values for multiple different time intervals (time intervals 1, 2,... m). In some embodiments, time intervals 1 to m can be time intervals of different lengths. For example, by way of example only, time interval 1 can be 5 minutes in length, time interval 2 can be 10 minutes in length, and time interval 3 can be 30 minutes, etc.
[0059] In block 510-1, for time interval 1, the corrector 330 can use the raw data 310 within time interval 1 in its buffer (for example, for a 1-hour buffer, it includes 12 time intervals 1 each with a length of 5 minutes) and the corresponding filtered data to correct the filtered value, obtaining K1. At this time, K1 can be the average value, the maximum value, or any combination thereof of the corrected filtered values for each of the 12 time intervals. Similarly, in block 510-2, for time interval 2, the corrector 330 can use the raw data 310 within time interval 2 in its buffer (for example, for a 1-hour buffer, it includes 6 time intervals 2 each with a length of 10 minutes) and the corresponding filtered data to correct the filtered value, obtaining K2. At this time, K2 can be the average value, the maximum value, or any combination thereof of the corrected filtered values for each of the 6 time intervals. By analogy, corrected filtered values K, K1, K2... Km for time intervals of different lengths are obtained.
[0060] Then, in block 520, the average value or the maximum value of the filtered values (K1, K2... Km) for time intervals of different lengths is determined as the reduced filtered value K for subsequent time periods and sent to a data filter 340 such as Figure 3 shown. Thereby, optimized values for any different time intervals can be provided.
[0061] As can be seen from the above, the method for storing the metric values of monitored objects according to the embodiments of the present disclosure can significantly reduce the amount of data stored in the database while satisfying the accurate query of the top-ranked monitored objects within any time interval.
[0062] Figure 6 A graph showing the disk saving rate and the accuracy rate under different data distributions according to the embodiments of the present disclosure is shown. Figure 6 In the horizontal axis represents the range of the dynamic filtered value, and the vertical axes on both sides respectively represent the accuracy rate and the disk saving rate, where it is assumed that the number of clients is 1000.
[0063] Figure 6The test results showing three different types of data distributions are presented. Among them, curve 610 represents a normal distribution with a large standard deviation, curve 620 represents a normal distribution with a small standard deviation, and curve 630 represents a Poisson distribution (λ = 10). It can be seen that for curve 610, when the accuracy requirement is 90%, the screening value is at least 84, that is, only 84 of the 1000 index values need to be stored. Correspondingly, the disk savings rate is 91.6%. For curve 620, when the accuracy requirement is 90%, the screening value is at least 306, that is, only 306 of the 1000 index values need to be stored. Correspondingly, the disk savings rate is 69.4%. For curve 630, when the accuracy requirement is 90%, the screening value is at least 450, that is, only 450 of the 1000 index values need to be stored. Correspondingly, the disk savings rate is 55%. Thus, the technology provided by the present disclosure is particularly applicable to index values with a large data distribution deviation. Since the embodiments of the present disclosure can dynamically correct the screening value over time, the screening scale can be adjusted in a timely manner when the data distribution of the index value changes, thereby obtaining the best balance between the disk savings rate and the data accuracy rate.
[0064] Figure 7 A schematic block diagram of a device 700 that can be used to implement the embodiments of the present disclosure is shown. As shown, the device 700 includes a central processing unit (CPU) 701, which can execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The CPU 701, ROM 702, and RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0065] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disc, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0066] Each of the methods or processes described above may be executed by the processing unit 701. For example, in some embodiments, the method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU 701, one or more steps or actions of the methods or processes described above may be performed.
[0067] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0068] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium would include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0069] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to respective computing / processing devices, or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0070] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages and conventional procedural programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN)-or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0071] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The computer - readable program instructions can also be stored in a computer - readable storage medium, and these instructions cause a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0072] The computer - readable program instructions can also be loaded onto a computer, other programmable data - processing apparatus, or other device, such that a series of operation steps are executed on the computer, other programmable data - processing apparatus, or other device to produce a computer - implemented process, so that the instructions executed on the computer, other programmable data - processing apparatus, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0074] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.
Claims
1. A method for storing indicator values of a monitoring object, include: Receiving a plurality of groups of indicator values collected at a plurality of time points within a time period, each group of indicator values in the plurality of groups of indicator values including a first number of indicator values corresponding to a corresponding monitoring object; for each of the plurality of sets of indicator values, selecting a second number of indicator values from the first number of indicator values, the second number being less than the first number, to generate a reduced plurality of sets of indicator values for storage; as well as Based on a predetermined threshold ratio of matching items, updating the second number for reduction of subsequent index values, the updating being performed at least once and comprising: generating a first list of monitored objects based on at least a portion of the plurality of sets of indicator values; generating a second list of monitored objects based on at least a portion of the reduced plurality of sets of indicator values; and The second number is updated based on determining whether a ratio of the same monitored objects in the first list and the second list to the monitored objects in the second list meets the predetermined threshold ratio.
2. The method of claim 1, wherein the multiple time points are evenly distributed within the time period.
3. The method of claim 1 , wherein a second number of indicator values are selected from the first number of indicator values in each group of indicator values. include: sorting the first number of index values to obtain sorted index values; as well as The second highest index value among the sorted index values is selected.
4. The method according to claim 1, further comprising: include: If the ratio is lower than the predetermined threshold ratio, performing the following steps at least once until the ratio exceeds the threshold: increasing the second number by a preset amount to update the second number; regenerating the reduced sets of index values using the updated second number; as well as re-determining the ratio based on the regenerated reduced sets of indicator values; as well as The updated second number satisfying the ratio exceeding the predetermined threshold ratio is determined as a reduced second number for a subsequent indicator value.
5. The method according to any one of claims 1 to 4, further comprising: include: For different time intervals within the time period, based on corresponding parts of the multiple groups of indicator values and corresponding parts of the reduced multiple groups of indicator values, respectively update the second number to obtain multiple candidate second numbers; as well as The second number is updated using an average value or a maximum value of the plurality of candidate second numbers.
6. The method according to claim 1, further comprising: include: identifying, according to the criteria, a first subset of monitored objects in the first list that satisfies a predetermined number of objects ranked high; as well as identifying, according to the criterion, a second subset of monitored objects in the second list that satisfies a predetermined number of objects ranked high, Wherein, determining whether the ratio of the same monitoring objects in the first list and the second list to the monitoring objects in the second list meets the predetermined threshold ratio includes: Determine whether the ratio of the same monitored objects in the first subset of monitored objects and the second subset of monitored objects to the monitored objects in the second subset of monitored objects meets the predetermined threshold ratio.
7. An electronic device, comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit cause the electronic device to perform actions, the actions including: Receiving multiple sets of metric values collected at multiple time points within a time period, each set of metric values in the multiple sets of metric values including a first number of metric values corresponding to a respective monitored object; For each of the multiple sets of metric values, select a second number of metric values from the first number of metric values, the second number being less than the first number, to generate reduced multiple sets of metric values for storage; Based on a predetermined threshold ratio of matching items, update the second number for subsequent reduction of metric values, the update being performed at least once and including: Generating a first list of monitored objects based on at least a portion of the multiple sets of metric values; Generating a second list of monitored objects based on at least a portion of the reduced multiple sets of metric values; Updating the second number based on determining whether the ratio of the same monitored objects in the first list and the second list to the monitored objects in the second list meets the predetermined threshold ratio; and Regenerating the reduced multiple sets of metric values.
8. The electronic device according to claim 7, wherein the multiple time points are evenly distributed within the time period.
9. The electronic device according to claim 7, wherein selecting a second number of metric values from the first number of metric values in each set of metric values includes: Sorting the first number of metric values to obtain sorted metric values; and Selecting the top-ranked second number of sorted metric values.
10. The electronic device according to claim 7, the actions further including: If the ratio is lower than the predetermined threshold ratio, perform the following steps at least once until the ratio exceeds the threshold: Increase the second number by a preset amount; Regenerate the reduced multiple sets of metric values using the increased second number; and Based on the regenerated reduced multiple sets of metric values, re-determine the ratio.
11. The electronic device according to any one of claims 7-10, the actions further including: For different time intervals within the time period, update the second number respectively based on the corresponding portions of the multiple sets of metric values and the corresponding portions of the reduced multiple sets of metric values to obtain multiple alternative second numbers; and Update the second number using the average or maximum value of the multiple alternative second numbers.
12. The electronic device according to claim 7, wherein the actions further include: Identifying a first subset of monitored objects in the first list that meet a predetermined number of top-ranked objects according to a criterion; and Identify a second subset of monitored objects in the second list that meet a predetermined number of top-ranked objects according to the standard. Among them, determining whether the ratio of the same monitored objects in the first list and the second list to the monitored objects in the second list meets the predetermined threshold ratio includes: Determining whether the ratio of the same monitored objects in the first subset of monitored objects and the second subset of monitored objects to the monitored objects in the second subset of monitored objects meets the predetermined threshold ratio.
13. A non-transitory computer storage medium, including machine-executable instructions that, when executed by a device, cause the device to execute the method according to any one of claims 1 to 6.
14. A computer program product, including machine-executable instructions that, when executed by a device, cause the device to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and system for monitoring interested subjects
CN104252461A
Intelligent index scheduling
US20140006413A1