System monitoring method and device and storage medium
By extracting and classifying the data in the monitoring system, using different acquisition strategies and dynamic adjustments, the problem of single data acquisition methods in the existing monitoring system is solved, efficient data acquisition and storage management is achieved, and the stability and resource utilization of the system are improved.
Patent Information
- Application Number
- CN202510553269.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-29
AI Technical Summary
The data collection method in the existing monitoring system is relatively single and cannot be adjusted adaptively, which makes it difficult to balance the contradiction between real-time data and storage resources, affecting the stable operation of the system.
By determining the numerical change rate, business correlation and historical data volatility of the target data, the classified data is core dynamic data, business association data and auxiliary record data, and different data acquisition strategies are adopted for different types, including dynamic adjustments based on PID algorithms and data type conversion rules.
The refinement and intelligence of monitoring data are realized, the contradiction between data collection and storage is balanced, and the system operation stability and resource utilization efficiency are improved.
Smart Images

Figure CN120386687A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of system monitoring, and particularly to a system monitoring method, device, and storage medium. Background Art
[0002] With the development of information technology, the stable operation of various complex systems such as enterprise-level information management systems, large-scale industrial control systems, and financial trading systems has become the key to ensuring the normal operation of the social economy. These complex systems usually cover numerous subsystems and components, and their operating states are affected by various factors. Therefore, they highly rely on an efficient monitoring system. Only by monitoring various indicators of the system in real time and accurately can potential problems be discovered in a timely manner and effective countermeasures be taken to ensure the continuous and stable operation of the system. However, most current monitoring systems have the problem of being unable to balance the real-time acquisition of monitoring data and the large amount of data storage.
[0003] Current monitoring systems usually adopt a relatively simple architecture, collect various types of data in the system at a fixed collection frequency without considering data differences, and the collection method is relatively single. If the collection interval is too short, although the real-time nature of the data can be guaranteed, a huge amount of data will be generated, causing great pressure on storage resources and increasing storage costs and management difficulties. On the contrary, if the collection interval is too long, although the amount of data can be reduced, the real-time nature of the data will be severely weakened, making it impossible to detect and handle system problems in a timely manner. Summary of the Invention
[0004] This application provides a system monitoring method, device, and storage medium to at least solve the technical problem in the related art that the data collection method is relatively single and cannot be adaptively adjusted, and by classifying target data and adopting different collection strategies, the accuracy and flexibility of system monitoring are improved.
[0005] This application provides a system monitoring method, including:
[0006] Determine the data characteristics corresponding to the current target data in the system to be monitored; the data characteristics include the numerical change rate of the target data, the business relevance between the target data and the key business of the system to be monitored, and the data volatility of the historical data corresponding to the target data;
[0007] Determine the data type of the target data based on the numerical change rate, the business relevance, and the data volatility;
[0008] Determine the target data collection strategy corresponding to the data type of the target data from several preset data collection strategies; different preset data collection strategies are configured with different data collection intervals;
[0009] Collect the target data in the current system to be monitored based on the target data collection strategy, and use the collected target data to determine the system monitoring result corresponding to the system to be monitored.
[0010] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above system monitoring methods when executing the computer program.
[0011] This application also provides a computer-readable storage medium storing a computer program, where the computer program implements the steps of any of the above system monitoring methods when executed by a processor.
[0012] Through this application, since it is possible to first determine data characteristics such as the numerical change rate corresponding to the target data in the current system to be monitored, the business relevance with the key services of the system to be monitored, and the data volatility of historical data before data collection, and then determine the data type corresponding to the current target data based on the determined data characteristics, and determine the target data collection strategy corresponding to the target data from the preset data collection strategies corresponding to each data type, so as to collect the target data based on the data collection interval in the target data collection strategy and realize the monitoring of the system to be monitored. Therefore, it is possible to solve the technical problems of the current relatively single data collection method and poor system monitoring effect. By extracting and classifying data features, the monitoring data is divided into various data types, and corresponding data collection intervals are adopted for different types of data according to their respective strategies, so as to realize the refinement and intelligence of the monitoring system, balance the contradiction between data collection and storage, and further improve the stability of system operation. Description of the Drawings
[0013] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a flowchart of a system monitoring method provided by an embodiment of the present application;
[0015] Figure 2 It is a system monitoring flowchart provided by an embodiment of the present application;
[0016] Figure 3 It is a specific system monitoring method flowchart provided by an embodiment of the present application;
[0017] Figure 4Flowchart of a method for adjusting the acquisition interval of core dynamic data based on the PID algorithm provided by an embodiment of the present application;
[0018] Figure 5 Flowchart of a method for adjusting the acquisition interval of service-related data provided by an embodiment of the present application;
[0019] Figure 6 Schematic diagram of the acquisition and storage rules for auxiliary record data provided by an embodiment of the present application;
[0020] Figure 7 Schematic diagram of the structure of a system monitoring device provided by an embodiment of the present application. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0023] In order to enable those skilled in the art in the technical field to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0024] See Figure 1 As shown, an embodiment of the present application provides a system monitoring method, including:
[0025] Step S11, determining the data characteristics corresponding to the current target data in the system to be monitored; the data characteristics include the numerical change rate of the target data, the business relevance between the target data and the key services of the system to be monitored, and the data volatility of the historical data corresponding to the target data.
[0026] In this embodiment, as Figure 2As shown, it is first necessary to determine the data characteristics corresponding to the current target data in the system to be monitored. Among them, the above data characteristics include, but are not limited to, the numerical change rate of the target data, the business relevance between the target data and the key services of the system to be monitored, and the data volatility of the historical data corresponding to the target data. That is to say, in a specific embodiment, the multi-dimensional characteristics of the monitoring data can be deeply analyzed, such as: numerical change rate, relevance to the key services of the system, historical data volatility, etc., and based on this, the monitoring data can be accurately classified into corresponding data types. In this way, by conducting a comprehensive and multi-dimensional in-depth analysis of the monitoring data and making a fine classification according to the data characteristics, it can ensure that various types of data are accurately defined.
[0027] Step S12: Determine the data type of the target data based on the numerical change rate, the business relevance, and the data volatility.
[0028] In this embodiment, the data type of the target data can be determined based on the numerical change rate, the business relevance, and the data volatility among the above data characteristics. And it can be understood that corresponding importance identifiers can be pre-assigned to different data types, so as to more quickly adopt different data collection strategies for data of different data types. In a specific embodiment, the above data types may include: core dynamic data, business-related data, and auxiliary record data. Among them, core dynamic data is data related to the operation of the system core of the system to be monitored. It can be understood that core dynamic data is closely connected to the operation of the system core and is directly related to the key performance indicators and stability of the system, such as the response time of key business services and the utilization rate of core server resources. Therefore, even a small change in these data may have a significant impact on the overall operation of the system. Also, because the numerical value of core dynamic data changes rapidly and a large amount of data is generated, high-frequency data collection is required to capture these changes in a timely manner so that the system can quickly respond, and it has high requirements for storage and processing capabilities. Therefore, when collecting core dynamic data, the highest importance weight in the entire monitoring system can be assigned to ensure the reliability of its storage and the efficiency of its reading. For example, in a server monitoring system, data such as CPU (Central Processing Unit) utilization rate and memory utilization rate, which have high real-time requirements and a relatively large impact on the system, are core dynamic data. Among them, on a busy server, when a large number of users access the website simultaneously, the CPU utilization rate may rise sharply. Therefore, the CPU utilization rate reflects the usage of CPU resources in real time. If the CPU utilization rate exceeds 90% for a long time, it may cause the system response to slow down or even freeze; the memory utilization rate shows the usage of system memory. If the memory utilization rate is too high, it may cause a memory shortage error in the system, affecting the normal operation of the application program.
[0029] Business - related data is data that reflects the operation status and development trend of the preset key business processes of the system to be monitored. Among them, although the business - related data has a relatively indirect immediate impact on the system, it plays a crucial role in the long - term business planning, performance optimization, and fault prediction. And the amount of business - related data is relatively stable, not as volatile as the core dynamic data, but it will gradually increase with the development of the business. Therefore, the business - related data is of high importance to the business and needs to be effectively stored and analyzed to provide support for business decisions. Compared with the core dynamic data, the change frequency of business - related data is relatively moderate. It does not need to be collected at an extremely high frequency like the core dynamic data, but also needs to be reasonably adjusted according to the changes in the business. For example, in a server monitoring system, business - related data can include data such as the number of network connections and the distribution of process time consumption, which are closely related to the business. The number of network connections reflects the connection situation between the operating system and the external network; the distribution of process time consumption can reflect the average time from when the user initiates a request to obtaining a response, and the proportion of processing time consumption of different business modules (such as data retrieval, file transfer).
[0030] Auxiliary record data is auxiliary information including the system log of the system to be monitored, the operation status record of non - critical modules, and temporary cache data. The auxiliary record data has little effect during the normal operation of the system, but in specific situations, such as in - depth fault troubleshooting or system behavior auditing, it can provide valuable references. And the amount of auxiliary record data varies, depending on the operation of the system and the level of detail of the records. At the same time, the change frequency of the auxiliary record data is low and does not need to be collected and stored frequently. Therefore, the importance weight of the auxiliary record data is relatively low, and a lower collection frequency and a more economical storage method can be adopted. It should be noted that when classifying the target data, it is not limited to the above classification results. The target data can be divided into several different data types according to the data characteristics corresponding to different data, and different data collection strategies can be adopted for the corresponding data types.
[0031] Step S13: Determine the target data collection strategy corresponding to the data type of the target data from several preset data collection strategies; different preset data collection strategies are configured with different data collection intervals.
[0032] In this embodiment, as Figure 2As shown, after determining the data type corresponding to the target data, the target data collection strategy corresponding to the data type of the target data can be determined from several preset data collection strategies. It can be understood that based on the above steps, different data collection intervals are configured in different preset data collection strategies, and different data collection intervals can correspond to data of different importance levels. Therefore, before determining the target data collection strategy corresponding to the data type of the target data, it is necessary to construct the data collection strategies corresponding to each data type.
[0033] Specifically, a first data collection strategy for collecting core dynamic data based on the system load of the system to be monitored, disk read / write utilization rate, and the first target data collection interval can be configured; a second data collection strategy for collecting business-related data based on the data change gradient of the target data and the second target data collection interval can be configured; a third data collection strategy for collecting auxiliary record data based on the data change information of the target data and the third target data collection interval can be configured. Among them, the above-mentioned first target data collection interval, second target data collection interval, and third target data collection interval increase in sequence. That is to say, in this embodiment, the data collection interval and data importance are in an inverse relationship. The data collection interval for data with high importance is small, and the data collection interval for data with low importance is large. In this way, it helps to achieve efficient monitoring and optimized utilization of storage resources, so as to balance the real-time nature of monitoring data and the data storage pressure.
[0034] Step S14: Collect the target data in the currently monitored system based on the target data collection strategy, and use the collected target data to determine the system monitoring result corresponding to the monitored system.
[0035] In this embodiment, the target data in the currently monitored system can be collected based on the target data collection strategy, and the system monitoring result corresponding to the monitored system can be determined using the collected target data. It can be understood that after the target data collection is completed, the collected target data can be stored in a preset cache unit. It can be understood that based on the timeliness characteristics of monitoring data, generally, data in the short term has high value for real-time monitoring and fault troubleshooting, while data beyond a certain time range can be regarded as having a low priority. Therefore, in order to ensure the efficient operation of the storage system and avoid data redundancy, a regular cleaning rule can be set to clean the target data in the cache unit. For example, for monitoring data that exceeds 15 days, the system will automatically perform a cleaning operation. In this way, by regularly cleaning old data, not only can the storage space be released, but also the database performance can be optimized, the query latency can be reduced, and the system can always be in the best operating state.
[0036] Correspondingly, in the process of determining the system monitoring results corresponding to the system to be monitored, visualization components such as Grafana can be used. Through its flexible configuration and visualization functions, the collected data can be presented in the form of intuitive charts, dashboards, etc., so as to realize the intuitive display and efficient analysis of monitoring data, and facilitate users to obtain the system monitoring results more intuitively.
[0037] And in this embodiment, considering that the importance of the target data may change due to system changes at different time points, during the data collection process, the data type of the target data can be judged in real time, so as to adjust the corresponding data collection strategy, providing a strong guarantee for the stable operation and efficient operation and maintenance of the system.
[0038] Specifically, if the data type of the current target data is determined to be core dynamic data, it is judged whether the current corresponding first target data collection interval of the core dynamic data is greater than the first preset interval threshold. If so, the data type of the current target data is adjusted to business-related data. For example, if the collection interval of the core dynamic data is extended to 10 seconds (i.e., reaching the default collection interval of business-related data), data classification degradation is triggered, and the data flow is transferred to the business-related data queue, and the corresponding data collection rules are enabled.
[0039] If the data type of the current target data is determined to be business-related data, it is judged whether the current corresponding second target data collection interval of the business-related data is less than the second preset interval threshold. If so, and within the fourth preset time period, the data change gradient of the business-related data is greater than the second preset gradient threshold, then the data type of the current target data is adjusted to core dynamic data. For example, when the collection interval of the business-related data is shortened to 1 second (i.e., reaching the minimum collection interval threshold of core dynamic data) due to a sharp increase in the gradient, and it still maintains a high-gradient fluctuation (such as gradient > 8%) for 3 consecutive collection cycles (3 seconds), data classification upgrade is triggered. At this time, the data flow will switch to the dynamic core data queue, and the corresponding data collection strategy is enabled.
[0040] If the data type of the current target data is determined to be business-related data, it is judged whether the current corresponding second target data collection interval of the business-related data is greater than the third preset interval threshold. If so, and within the fifth preset time period, the data change gradient of the business-related data is less than the third preset gradient threshold, then the data type of the current target data is adjusted to auxiliary record data. For example, if the collection interval of the business-related data is extended to 30 seconds (i.e., reaching the default collection interval of auxiliary record data) due to a low-gradient state, and there is no effective data change (gradient < 2%) for 5 consecutive minutes (10 collection cycles), data classification degradation is triggered, and the data flow is transferred to the auxiliary record data queue, and the corresponding data collection rules are enabled.
[0041] It should be noted that the auxiliary recorded data includes system logs, non-critical module running status records, temporary cache data, etc. These data have little effect during the normal operation of the system. Therefore, in the conversion strategy of this embodiment, these data usually adopt a lower collection frequency. Only when an anomaly is detected or a specific event is triggered, the collection frequency is increased and the stored data is updated.
[0042] Through the above technical solution, before data collection in this embodiment, the numerical change rate of the target data corresponding to the current system to be monitored, the business relevance with the key business of the system to be monitored, and the data volatility of historical data and other data characteristics are first determined. Then, based on the determined data characteristics, the data type corresponding to the current target data is determined, and the target data collection strategy corresponding to the target data is determined from the preset data collection strategies corresponding to each data type, so as to collect the target data based on the data collection interval in the target data collection strategy. And during this process, the data type of the current target data is adjusted in real time according to the monitoring situation of the target data, and then the obtained detection result is displayed through a chart to realize the monitoring of the system to be monitored. It solves the technical problem that the current data collection method is relatively single and the system monitoring effect is poor. By extracting and classifying the data features, the monitoring data is divided into core dynamic data, business-related data, and auxiliary recorded data, and a dynamic adjustment strategy for the collection interval adapted to the characteristics and change rules of different types of data is customized to ensure the timeliness and accuracy of data collection, and data type conversion rules for different types of data are formulated, so as to realize efficient monitoring and optimal utilization of storage resources. In this way, the purpose of balancing the real-time nature of monitoring data and the data storage pressure is achieved, and the system operation stability is further improved.
[0043] Based on the previous embodiment, it can be known that this application can classify the data in the system to be monitored according to characteristics, and adopt different data collection strategies for data of different data types. Next, the real-time adjustment process of the data collection strategy will be elaborated in detail in this embodiment. See Figure 3 As shown, an embodiment of this application provides a specific system monitoring method, including:
[0044] Step S21, determine the target data collection strategy corresponding to the data type of the target data from several preset data collection strategies; different data collection intervals are configured in different preset data collection strategies.
[0045] Step S22, collect the target data in the current system to be monitored based on the target data collection strategy, and use the collected target data to determine the system monitoring result corresponding to the system to be monitored.
[0046] In this embodiment, based on the previous embodiment, different data acquisition strategies are formulated for different types of data. Specifically, if the current target data is core dynamic data, the system load and disk read / write utilization rate of the current system to be monitored are monitored, the system load error is determined based on the preset system load threshold and the system load, the disk read / write utilization rate error is determined based on the preset disk read / write utilization rate threshold and the disk read / write utilization rate, and then using the preset control strategy, the current data acquisition interval adjustment amount is determined according to the system load error and the disk read / write utilization rate error, and the first target data acquisition interval is determined according to the data acquisition interval adjustment amount and the first data acquisition strategy, so as to collect the target data in the current system to be monitored according to the first target data acquisition interval. Among them, the preset control strategy is a strategy constructed based on proportional-integral-derivative control (PID, Proportional, Integral, Differential), and the first target data acquisition interval is greater than the preset minimum acquisition interval and less than the preset maximum acquisition interval.
[0047] That is to say, in this embodiment, when monitoring core dynamic data, a PID data acquisition interval dynamic adjustment strategy based on multi-source fusion is designed, and the acquisition interval of dynamic core data is adjusted in real time by combining the system load and the disk read / write utilization rate, so as to achieve efficient data acquisition and storage management. Specifically, as Figure 4 shown, first obtain the system load of the current system. Specifically, tools such as top and htop provided by the operating system can be used to capture the CPU usage rate ( ) and the memory usage rate ( ) in real time, and average the two, which represents the current load status of the system. The specific formula is expressed as:
[0048] ;
[0049] For example, at a certain moment, it is monitored that the CPU usage rate is (40%) and the memory usage rate is (30%). At this time, the system load is ( ).
[0050] In this embodiment, a threshold can be set for the ratio of system load and disk I / O utilization. For example, if the system load is desired to be maintained at 40%, 40% is the preset system load threshold, also known as the system load setpoint. Correspondingly, if the disk I / O utilization is desired to be 50%, 50% is the preset disk read / write utilization threshold, also known as the disk read / write utilization setpoint. Based on the thresholds, a corresponding error calculation can be performed. The system load error is the actual system load minus the setpoint, and the disk I / O utilization error is the actual disk I / O utilization minus the setpoint. For example, if the actual system load is 50% and the setpoint is 40%, the system load error is 50% - 40% = 10%. The preset control strategy can then be used to determine the current data collection interval adjustment based on the system load error and the disk read / write utilization error. It is understood that in a data collection scenario, a high system load indicates a shortage of system resources. In this case, collecting data at a fixed high frequency may further increase the system burden and affect overall system performance. After data is collected, it needs to be stored on disk. If the disk I / O load is too high, indicating that the disk is busy with read and write operations, continuing to collect data at a high frequency may cause data storage delays or even data loss. This embodiment dynamically adjusts the collection interval through PID control based on the system load error. This can appropriately extend the collection interval when the system load is high, avoiding excessive use of system resources. At the same time, by adjusting the collection interval based on the disk I / O load error, the data collection frequency can be reasonably reduced when storage resources are limited, avoiding problems such as storage overflow.
[0051] Specifically, a first system load adjustment amount and a first disk read / write utilization adjustment amount can be determined based on a proportional coefficient of a preset control strategy and based on the system load error and the disk read / write utilization error, respectively. A total system load error and a total disk read / write utilization error within a first preset time period can be determined based on the system load error and the disk read / write utilization error, and a second system load adjustment amount and a second disk read / write utilization adjustment amount can be determined based on the total system load error and the total disk read / write utilization error, respectively, based on an integral coefficient of a preset control strategy. The first preset time period is the time period between the current moment and a target moment, and the target moment is the moment when the system load error or the disk read / write utilization error exceeds a first preset error threshold. A third system load adjustment amount and a third disk read / write utilization adjustment amount can be determined based on the differential coefficient of a preset control strategy and based on the system load error and the disk read / write utilization error, respectively. A data collection interval adjustment amount can then be determined based on the first system load adjustment amount, the first disk read / write utilization adjustment amount, the second system load adjustment amount, the second disk read / write utilization adjustment amount, the third system load adjustment amount, and the third disk read / write utilization adjustment amount.
[0052] That is to say, in the multi-source fusion PID data acquisition strategy corresponding to the core dynamic data in this embodiment, a control strategy combining three links of proportional (P), integral (I), and derivative (D) is adopted. Specifically, in the proportional (P) link, it includes a system load ratio term and a disk I / O utilization ratio term. In the system load ratio term, the adjustment amount is calculated according to the error of the system load (the difference between the current load and the set point), and this adjustment amount is obtained by multiplying the proportional coefficient by the error. The utilization of disk I / O is similar to the system load, and the adjustment amount is obtained by multiplying the utilization coefficient of disk I / O by the utilization error of disk I / O. The specific formula for the proportional P adjustment amount is as follows:
[0053] ;
[0054] where E is the error of the system load or the utilization of disk I / O, and P is the corresponding proportional coefficient.
[0055] It should be noted that when determining the current data acquisition interval adjustment amount according to the system load error and the disk read / write utilization error, if the system load error or the disk read / write utilization error is greater than the second preset error threshold, or the system load error or the disk read / write utilization error is less than the third preset error threshold, then the exponentially weighted moving average of the system load error or the disk read / write utilization error within the second preset time period is determined, and the proportional coefficient is adjusted based on the exponentially weighted moving average. It can be understood that the role of the proportional coefficient (P) is to directly amplify or reduce the current error. If the error is large, increasing the value of P can make the system respond faster. For example, if the absolute value of the error is greater than a certain threshold, the value of P can be increased by a certain proportion; conversely, if the error is small, the value of P can be reduced by a certain proportion. Specifically, the proportional coefficient (P) is dynamically adjusted using the exponentially weighted moving average (EMA, Exponential Moving Average) to adapt to the change trend of the system load error. That is to say, if the error of the recent system load or the utilization of disk I / O has been increasing, the EMA will increase the proportional coefficient according to this trend to respond faster to the increasing load. The specific formula for adjusting the proportional coefficient (P) is:
[0056] ;
[0057] where is the smoothing coefficient of the EMA.
[0058] The integral (I) link specifically includes the system load integral term, which accumulates the errors of the system load over a period of time and then multiplies by the integral coefficient. It can be understood that if the system load is always higher than the set point, the integral term will become larger and larger, that is, the acquisition interval needs to be adjusted; the integral term of the disk I / O utilization rate is similar to the system load integral term, which integrates the errors of the disk I / O utilization rate over a period of time and multiplies by the integral coefficient. The specific formula is as follows:
[0059] ;
[0060] where I is the integral coefficient and E(t) is the error of the system load or the disk I / O utilization rate within the time period . And it can be understood that in the integral (I) link, when the error exceeds a certain threshold, the calculation of the integral term will pause to prevent integral saturation. For example: when , the integral term pauses to update, where is the error threshold. If the system load continues to be above the set point and the error exceeds the threshold, the integral term will pause to update.
[0061] The derivative (D) link includes the system load derivative term, which can specifically determine the change rate of the system load error, multiplies the derivative coefficient by the error change rate, and if the system load suddenly changes rapidly, the derivative term can make an early response to adjust the acquisition interval. Correspondingly, it also includes the derivative term of the disk I / O utilization rate, which can be obtained by multiplying the derivative coefficient of the disk I / O utilization rate by the change rate of the disk I / O utilization rate error. When the change rate of the disk I / O utilization rate changes, the derivative term will take effect. The specific formula is as follows:
[0062] ;
[0063] where is the derivative coefficient, is the change rate of the system load or the disk I / O utilization rate error.
[0064] And in this process, the change rate of the current system load error and / or the disk read / write utilization rate error can be determined. If the absolute value of the change rate is not equal to the preset change rate threshold, the derivative coefficient is adjusted according to the preset step size. In this way, when the system is in a stable state and the error change rate is small, the value of is appropriately reduced to avoid over-adjustment; when there is a large error change in the system, such as a sudden load change or a change in the disk I / O utilization rate, the value of is increased so that the system can respond quickly. For example, a change rate threshold can be set. When , it is increased according to a certain ratio value; when happens, decrease the value in accordance with a certain ratio .
[0065] Correspondingly, considering the influence of system load and disk I / O utilization rate on the acquisition interval, after calculating their adjustment amounts in the proportional (P), integral (I), and derivative (D) links respectively, add the adjustment amounts of the system load and disk I / O utilization rate in the same link to obtain the comprehensive proportional, integral, and derivative adjustment amounts, and then add the comprehensive proportional, integral, and derivative adjustment amounts to obtain the total adjustment amount .
[0066] ;
[0067] Among them, are the adjustment amounts of the proportional, integral, and derivative links of the system load respectively, are the adjustment amounts of the proportional, integral, and derivative links of the disk I / O utilization rate respectively.
[0068] And when adjusting the acquisition interval, an initial acquisition interval can be preset, and multiply the initial interval by (1 + total adjustment amount ) to obtain the new acquisition interval . And in this embodiment, in order to prevent the acquisition interval from becoming too long, it is stipulated that the maximum acquisition interval is 10 seconds and the minimum acquisition interval is 1 second, that is:
[0069] ;
[0070] For example: the initial interval is 3 seconds, the calculated adjustment amount is 0.5, and the new interval is 3×(1 + 0.5) = 4.5 seconds; if the adjustment amount is 3, the new interval should be 3×(1 + 3) = 12 seconds, but since the maximum acquisition interval is 10 seconds, the new interval is 10 seconds; if the adjustment amount is -0.7, the new acquisition interval should be 3×(1 - 0.7) = 0.9 seconds, but because the minimum acquisition interval is 1 second, the new interval is 1 second. In summary, in this embodiment, a dynamic adjustment and optimization strategy for the PID data acquisition interval based on multi-source fusion is constructed. By combining the proportional, integral, and derivative links, the acquisition interval of data can be dynamically adjusted according to the actual operation situation of the system, and parameter adaptive adjustment is introduced, which can more effectively respond to the changes in system load and disk I / O utilization rate, thereby optimizing the data acquisition interval and improving the efficiency and response speed of the system. The aim is to improve the response speed and efficiency of the system.
[0071] In another specific embodiment, if the current target data is business-related data, determine the data change gradient of the current target data, and determine whether the absolute value of the data change gradient is greater than a first preset gradient threshold; if the absolute value of the data change gradient is greater than the first preset gradient threshold, then reduce the second target data collection interval in the second data collection strategy; if the absolute value of the data change gradient is not greater than the first preset gradient threshold, and within a third preset time period, the absolute value of the data change gradient is not greater than the first preset gradient threshold, then increase the second target data collection interval in the second data collection strategy; so as to collect business-related data according to the current second target data collection interval. Specifically, as Figure 5 shown, the initial collection interval can be set to 10 seconds. The gradient threshold is . Then every 1 minute, calculate the data change gradient . When , shorten the collection interval to , where . When and the data is in a low-gradient state for more than 5 minutes, extend the collection interval to , where . In this way, when the data change gradient is large (exceeding the threshold), the system will shorten the collection interval to capture the rapid changes in data more quickly, which is suitable for scenarios where the data changes dynamically and requires real-time monitoring; when the data change gradient is small (below the threshold) and lasts for a period of time, the system will appropriately extend the collection interval to save resources and reduce unnecessary data collection operations, which is suitable for situations where the data changes slowly.
[0072] In another specific embodiment, as Figure 6 shown, if the current target data is auxiliary record data, then based on the third data collection strategy, collect the auxiliary record data in the system to be monitored, and compare the currently collected auxiliary record data with the previously saved auxiliary record data; if the currently collected auxiliary record data is the same as the previously saved auxiliary record data, then prohibit the save operation for the currently collected auxiliary record data; if the currently collected auxiliary record data is different from the previously saved auxiliary record data, then update the previously saved auxiliary record data based on the currently collected auxiliary record data. For example, set the initial collection interval to = 30 seconds. At the same time, compare whether the currently collected data is equal to the data saved last time. If the data has changed, save the current data; if the data has not changed, do not save the current data and directly use the data saved last time. This strategy is mainly applied to scenarios where the data change frequency is very low. For example, for the status parameters of some industrial equipment (such as the on / off status and working mode of the equipment), their status may remain unchanged for a long time and only change when the equipment is operated or a fault occurs. In this way, using this strategy can reduce the storage of stable state data, and at the same time ensure that the data is updated and stored in a timely manner when the equipment status changes, facilitating subsequent analysis of the equipment operation status and fault troubleshooting.
[0073] Among them, for a more specific processing process of the above step S21, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0074] In this embodiment, an adaptive acquisition interval dynamic adjustment mechanism is established. For different types of data, the acquisition interval is adjusted according to their respective strategies, realizing refined and intelligent monitoring of the monitoring system, balancing data acquisition and storage, and providing a strong guarantee for the stable operation and efficient operation and maintenance of the system. Specifically, a PID data acquisition interval dynamic adjustment strategy based on multi-source fusion is formulated for core dynamic data, and the acquisition interval can be dynamically optimized according to the system load and disk I / O; for business-related data, a strategy for adjusting the acquisition interval according to the data change gradient is set, and the acquisition frequency is dynamically changed according to the data change gradient threshold. In this way, the acquisition interval of business-related data is adjusted according to the data change trend perception, which not only meets the sensitive response of the monitoring system to emergencies but also realizes the optimal utilization of resources under normal operation; for auxiliary record data, the data is acquired according to the initial acquisition interval, and the current data is compared with the data saved last time, and only the changed data is saved to cope with the low-frequency change data scenario and reduce unnecessary storage. Through the above technical solutions, the technical problem that it is difficult to balance the data acquisition interval and the amount of stored data in the existing monitoring system is solved. By introducing a dynamic adjustment mechanism, the data acquisition interval is flexibly adjusted according to the type of monitoring data, so as to achieve efficient data acquisition and storage management, and further optimize the acquisition strategies of various types of data to improve the accuracy and flexibility of adaptive adjustment, ensuring that the monitoring system reaches the best balance in terms of data acquisition efficiency and storage resource utilization.
[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0076] See Figure 7 As shown, the embodiment of the present application further provides a system monitoring device, including:
[0077] A feature determination module 11, configured to determine data features corresponding to current target data in a system to be monitored; the data features include the numerical change rate of the target data, the business relevance between the target data and the critical services of the system to be monitored, and the data volatility of the historical data corresponding to the target data.
[0078] A data classification module 12, configured to determine the data type of the target data based on the numerical change rate, the business relevance, and the data volatility.
[0079] A policy determination module 13, configured to determine a target data acquisition policy corresponding to the data type of the target data from a plurality of preset data acquisition policies; different preset data acquisition policies are configured with different data acquisition intervals.
[0080] A data acquisition module 14, configured to acquire the target data in the current system to be monitored based on the target data acquisition policy, and determine a system monitoring result corresponding to the system to be monitored by using the acquired target data.
[0081] In this embodiment, first, data features such as the numerical change rate corresponding to the target data in the current system to be monitored, the business relevance between the target data and the critical services of the system to be monitored, and the data volatility of the historical data can be determined. Then, based on the determined data features, the data type corresponding to the current target data is determined, and a target data acquisition policy corresponding to the target data is determined from the preset data acquisition policies corresponding to each data type, so as to acquire the target data based on the data acquisition interval in the target data acquisition policy, realizing the monitoring of the system to be monitored. In this way, by extracting and classifying the features of the data, the monitoring data is divided into data of multiple data types, and for different types of data, corresponding data acquisition intervals are adopted according to their respective policies, thereby realizing the refinement and intelligence of the monitoring system, balancing the contradiction between data acquisition and storage, and further improving the stability of the system operation.
[0082] In some specific embodiments, the system monitoring device further includes:
[0083] A first policy configuration module, configured to configure a first data acquisition policy for acquiring the core dynamic data based on the system load of the system to be monitored, the disk read / write utilization rate, and a first target data acquisition interval.
[0084] A second policy configuration module, configured to configure a second data acquisition policy for acquiring the service-related data based on the data change gradient of the target data and a second target data acquisition interval.
[0085] A third policy configuration module, configured to configure a third data collection policy for collecting the auxiliary record data based on the data change information of the target data and a third target data collection interval;
[0086] Wherein, the first target data collection interval, the second target data collection interval, and the third target data collection interval increase in sequence.
[0087] In some specific embodiments, the data collection module 14 specifically includes:
[0088] A load monitoring sub-module, configured to monitor the system load and the disk read / write utilization rate of the to-be-monitored system if the current target data is the core dynamic data;
[0089] An error determination sub-module, configured to determine a system load error based on a preset system load threshold and the system load, and determine a disk read / write utilization rate error based on a preset disk read / write utilization rate threshold and the disk read / write utilization rate;
[0090] An interval determination sub-module, configured to use a preset control policy to determine a current data collection interval adjustment amount according to the system load error and the disk read / write utilization rate error, and determine a first target data collection interval according to the data collection interval adjustment amount and the first data collection policy; the preset control policy is a policy constructed based on proportional-integral-derivative control, and the first target data collection interval is greater than a preset minimum collection interval and less than a preset maximum collection interval;
[0091] A data collection sub-module, configured to collect the target data in the to-be-monitored system according to the first target data collection interval.
[0092] In some specific embodiments, the interval determination sub-module specifically includes:
[0093] A first adjustment amount determination unit, configured to determine a first system load adjustment amount and a first disk read / write utilization rate adjustment amount respectively according to the proportional coefficient of the preset control policy and based on the system load error and the disk read / write utilization rate error;
[0094] An error determination unit, configured to determine a total system load error and a total disk read / write utilization rate error within a first preset time period according to the system load error and the disk read / write utilization rate error; the first preset time period is a time period before the current moment and a target moment, and the target moment is a moment when the system load error or the disk read / write utilization rate error is greater than a first preset error threshold;
[0095] A second adjustment amount determination unit, configured to determine a second system load adjustment amount and a second disk read / write utilization adjustment amount respectively based on the integral coefficient of the preset control strategy and based on the total system load error and the total disk read / write utilization error;
[0096] A third adjustment amount determination unit, configured to determine a third system load adjustment amount and a third disk read / write utilization adjustment amount respectively based on the differential coefficient of the preset control strategy and based on the system load error and the disk read / write utilization error;
[0097] A fourth adjustment amount determination unit, configured to determine the data acquisition interval adjustment amount based on the first system load adjustment amount, the first disk read / write utilization adjustment amount, the second system load adjustment amount, the second disk read / write utilization adjustment amount, the third system load adjustment amount, and the third disk read / write utilization adjustment amount.
[0098] In some specific embodiments, the data acquisition module 14 further includes:
[0099] A first coefficient adjustment sub-module, configured to determine the exponentially moving average of the system load error or the disk read / write utilization error within a second preset time period if the system load error or the disk read / write utilization error is greater than a second preset error threshold, or if the system load error or the disk read / write utilization error is less than a third preset error threshold, and adjust the proportional coefficient based on the exponentially moving average;
[0100] A second coefficient adjustment sub-module, configured to determine the change rate of the current system load error and / or the disk read / write utilization error, and adjust the differential coefficient according to a preset step length if the absolute value of the change rate is not equal to a preset change rate threshold.
[0101] In some specific embodiments, the data acquisition module 14 specifically includes:
[0102] A gradient determination unit, configured to determine the data change gradient of the current target data if the current target data is the service-related data;
[0103] A gradient judgment unit, configured to judge whether the absolute value of the data change gradient is greater than a first preset gradient threshold;
[0104] A first interval adjustment unit, configured to reduce the second target data acquisition interval in the second data acquisition strategy if the absolute value of the data change gradient is greater than the first preset gradient threshold;
[0105] A second interval adjustment unit, configured to increase the second target data acquisition interval in the second data acquisition strategy if the absolute value of the data change gradient is not greater than a first preset gradient threshold and the absolute value of the data change gradient is not greater than the first preset gradient threshold within a third preset time period;
[0106] A first data acquisition unit, configured to acquire the service-related data according to the current second target data acquisition interval.
[0107] In some specific embodiments, the data acquisition module 14 specifically includes:
[0108] A second data acquisition unit, configured to acquire the auxiliary record data in the to-be-monitored system based on the third data acquisition strategy if the current target data is the auxiliary record data;
[0109] A data storage unit, configured to compare the currently acquired auxiliary record data with the previously stored auxiliary record data; if the currently acquired auxiliary record data is the same as the previously stored auxiliary record data, prohibit performing a storage operation on the currently acquired auxiliary record data; if the currently acquired auxiliary record data is different from the previously stored auxiliary record data, update the previously stored auxiliary record data based on the currently acquired auxiliary record data.
[0110] In some specific embodiments, the system monitoring device further includes:
[0111] A first type adjustment module, configured to determine whether the first target data acquisition interval currently corresponding to the core dynamic data is greater than a first preset interval threshold if the data type of the current target data is determined to be the core dynamic data, and if so, adjust the data type of the current target data to the service-related data;
[0112] A second type adjustment module, configured to determine whether the second target data acquisition interval currently corresponding to the service-related data is less than a second preset interval threshold if the data type of the current target data is determined to be the service-related data, and if so, and the data change gradient of the service-related data is greater than a second preset gradient threshold within a fourth preset time period, adjust the data type of the current target data to the core dynamic data;
[0113] A third type of adjustment module is configured to, if the data type of the current target data is determined to be the service-related data, determine whether the second target data collection interval currently corresponding to the service-related data is greater than a third preset interval threshold. If so, and within a fifth preset time period, the data change gradient of the service-related data is less than a third preset gradient threshold, then adjust the data type of the current target data to the auxiliary record data.
[0114] For the description of the features in the embodiments corresponding to the system monitoring device, reference may be made to the relevant descriptions in the embodiments corresponding to the system monitoring method, which will not be elaborated here one by one.
[0115] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the system monitoring method.
[0116] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the system monitoring method when running.
[0117] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, etc., which can store computer programs.
[0118] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0119] The above has introduced in detail a system monitoring method, device, and storage medium provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A system monitoring method, characterized in that, Including: Determine the data characteristics corresponding to the current target data in the system to be monitored; the data characteristics include the numerical change rate of the target data, the business relevance between the target data and the key services of the system to be monitored, and the data volatility of the historical data corresponding to the target data; Determine the data type of the target data based on the numerical change rate, the business relevance, and the data volatility; Determine the target data acquisition strategy corresponding to the data type of the target data from a number of preset data acquisition strategies; Different data acquisition intervals are configured in different preset data acquisition strategies; Collect the target data in the current system to be monitored based on the target data acquisition strategy, and use the collected target data to determine the system monitoring result corresponding to the system to be monitored.
2. The system monitoring method according to claim 1, wherein The data types include: core dynamic data, business-related data, and auxiliary record data; the core dynamic data is data related to the operation of the system core of the system to be monitored, the business-related data is data reflecting the operation status and development trend of the preset system key business processes of the system to be monitored, and the auxiliary record data is auxiliary information including the system log, non-critical module operation status record, and temporary cache data of the system to be monitored; Correspondingly, before determining the target data acquisition strategy corresponding to the data type of the target data from a number of preset data acquisition strategies, it further includes: Configure a first data acquisition strategy for collecting the core dynamic data based on the system load, disk read / write utilization rate, and the first target data acquisition interval of the system to be monitored; Configure a second data acquisition strategy for collecting the business-related data based on the data change gradient of the target data and the second target data acquisition interval; Configure a third data acquisition strategy for collecting the auxiliary record data based on the data change information of the target data and the third target data acquisition interval; Wherein, the first target data acquisition interval, the second target data acquisition interval, and the third target data acquisition interval increase in sequence.
3. The system monitoring method according to claim 2, wherein The collecting the target data in the current system to be monitored based on the target data acquisition strategy includes: If the current target data is the core dynamic data, then monitor the system load and the disk read / write utilization rate of the current system to be monitored; Determine the system load error based on the preset system load threshold and the system load, and determine the disk read / write utilization rate error based on the preset disk read / write utilization rate threshold and the disk read / write utilization rate; Use a preset control strategy to determine the current data acquisition interval adjustment amount according to the system load error and the disk read / write utilization rate error, and determine the first target data acquisition interval according to the data acquisition interval adjustment amount and the first data acquisition strategy; the preset control strategy is a strategy constructed based on proportional-integral-derivative control, and the first target data acquisition interval is greater than the preset minimum acquisition interval and less than the preset maximum acquisition interval; Collect the target data in the current system to be monitored according to the first target data acquisition interval.
4. The system monitoring method according to claim 3, wherein, Using the preset control strategy to determine the current adjustment amount of the data acquisition interval according to the system load error and the disk read / write utilization rate error, includes: According to the proportional coefficient of the preset control strategy, and respectively based on the system load error and the disk read / write utilization rate error, determine the first system load adjustment amount and the first disk read / write utilization rate adjustment amount accordingly; According to the system load error and the disk read / write utilization rate error, determine the total system load error and the total disk read / write utilization rate error within the first preset time period; the first preset time period is the time period before the current moment and the target moment, and the target moment is the moment when the system load error or the disk read / write utilization rate error is greater than the first preset error threshold; According to the integral coefficient of the preset control strategy, and respectively based on the total system load error and the total disk read / write utilization rate error, determine the second system load adjustment amount and the second disk read / write utilization rate adjustment amount accordingly; According to the differential coefficient of the preset control strategy, and respectively based on the system load error and the disk read / write utilization rate error, determine the third system load adjustment amount and the third disk read / write utilization rate adjustment amount accordingly; Based on the first system load adjustment amount, the first disk read / write utilization rate adjustment amount, the second system load adjustment amount, the second disk read / write utilization rate adjustment amount, the third system load adjustment amount and the third disk read / write utilization rate adjustment amount, determine the adjustment amount of the data acquisition interval.
5. The system monitoring method according to claim 4, wherein, During the process of using the preset control strategy to determine the current adjustment amount of the data acquisition interval according to the system load error and the disk read / write utilization rate error, it further includes: If the system load error or the disk read / write utilization rate error is greater than the second preset error threshold, or the system load error or the disk read / write utilization rate error is less than the third preset error threshold, then determine the exponentially weighted moving average of the system load error or the disk read / write utilization rate error within the second preset time period, and adjust the proportional coefficient based on the exponentially weighted moving average; Determine the change rate of the current system load error and / or the disk read / write utilization rate error. If the absolute value of the change rate is not equal to the preset change rate threshold, then adjust the differential coefficient according to the preset step size.
6. The system monitoring method according to claim 2, wherein The collecting the target data in the current to-be-monitored system based on the target data acquisition strategy includes: If the current target data is the service-related data, then determine the data change gradient of the current target data; Judge whether the absolute value of the data change gradient is greater than the first preset gradient threshold; If the absolute value of the data change gradient is greater than the first preset gradient threshold, then reduce the second target data acquisition interval in the second data acquisition strategy; If the absolute value of the data change gradient is not greater than the first preset gradient threshold, and within the third preset time period, the absolute value of the data change gradient is not greater than the first preset gradient threshold, then increase the second target data acquisition interval in the second data acquisition strategy; Collect the service-related data according to the current second target data acquisition interval.
7. The system monitoring method according to claim 2, characterized in that, Collecting the target data in the to-be-monitored system currently based on the target data collection strategy includes: If the current target data is the auxiliary record data, collecting the auxiliary record data in the to-be-monitored system based on the third data collection strategy; Comparing the currently collected auxiliary record data with the previously saved auxiliary record data; If the currently collected auxiliary record data is consistent with the previously saved auxiliary record data, prohibiting the execution of the save operation on the currently collected auxiliary record data; If the currently collected auxiliary record data is inconsistent with the previously saved auxiliary record data, updating the previously saved auxiliary record data based on the currently collected auxiliary record data.
8. The system monitoring method according to any one of claims 2 to 7, characterized in that, It further includes: If the data type of the current target data is determined to be the core dynamic data, determining whether the current corresponding first target data collection interval of the core dynamic data is greater than the first preset interval threshold. If so, adjusting the data type of the current target data to the service-related data; If the data type of the current target data is determined to be the service-related data, determining whether the current corresponding second target data collection interval of the service-related data is less than the second preset interval threshold. If so, and within the fourth preset time period, the data change gradient of the service-related data is greater than the second preset gradient threshold, adjusting the data type of the current target data to the core dynamic data; If the data type of the current target data is determined to be the service-related data, determining whether the current corresponding second target data collection interval of the service-related data is greater than the third preset interval threshold. If so, and within the fifth preset time period, the data change gradient of the service-related data is less than the third preset gradient threshold, adjusting the data type of the current target data to the auxiliary record data.
9. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the system monitoring method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the system monitoring method according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Cited By
Self-adaptive data monitoring method and device, electronic equipment and storage medium
CN121233439A