Intelligent Monitoring Method and System for Server Performance

By analyzing the server load historical data, building a standard relationship of load-performance, the problem of poor monitoring accuracy and adaptability caused by not taking into account load conditions in the existing technology is solved, and efficient monitoring and safe operation of server performance is achieved.

CN119883852BActive Publication Date: 2025-07-08SHENZHEN YUNHAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510370279.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-08
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

In the prior art, server performance monitoring does not take into account different load conditions, resulting in poor monitoring accuracy and adaptability, making it difficult to ensure the safe operation of the server.

Method used

By collecting load history data of the target server, analyzing the load interval range and degree of change, splitting the load interval, building a standard relationship between load-performance, using the regression model to analyze the deviation situation, and outputting current performance information.

Benefits of technology

Improve the accuracy and adaptability of server performance monitoring to ensure efficient and safe operation of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883852B_ABST
    Figure CN119883852B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent server performance monitoring method and system, which relates to the technical field of server performance. It includes analyzing historical load data to determine the range and degree of change of the load, providing a reliable basis for subsequent splitting of the load range, taking into account the previous changes in the server load, and thus accurately describing the correlation between the detailed load range and performance data. Confirm the correlation between the load and the historical performance data in different partial ranges of the load to obtain the standard load-performance correlation. Among the corresponding relationships between the load and the historical performance data, select the normal values of the historical performance data. Input the load data and performance data into the standard load-performance correlation to determine the deviation situation, analyze the deviation situation and output the current performance information of the target server, improve the accuracy and adaptability of server performance monitoring, and effectively ensure the efficient and safe operation of the server.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server performance, and particularly to an intelligent monitoring method and system for server performance. Background Art

[0002] The background art of the intelligent monitoring solution for server performance mainly stems from the increasing requirements of the information technology field for the stability and reliability of servers. With the expansion of enterprise business and the acceleration of digital transformation, servers, as the core infrastructure supporting various application systems, their performance directly affects business continuity and user experience. Traditional server management methods often struggle to detect potential problems in real time and accurately, resulting in lag in fault warning and handling.

[0003] Therefore, an intelligent monitoring solution has emerged. It combines data collection, data analysis, and an early warning mechanism. By collecting key performance indicators such as the load, response time, and throughput of the server in real time, it uses machine learning algorithms for anomaly detection and performance trend prediction. These technologies not only improve the accuracy and efficiency of monitoring but also achieve a comprehensive evaluation and optimization of server performance. The intelligent monitoring solution can timely detect performance bottlenecks and potential fault risks, providing strong data support for the operation and maintenance team to ensure the stable operation of the server and business continuity.

[0004] In the prior art, the monitoring of server performance often monitors according to the parameter sizes of performance data, without considering the impact of different loads of the server on performance, resulting in poor accuracy and adaptability of server performance monitoring and unable to effectively ensure the safe operation of the server.

[0005] Therefore, how to improve the accuracy and adaptability of server performance monitoring is a technical problem to be solved at present. Summary of the Invention

[0006] The purpose of the present invention is to solve the problem of poor accuracy and adaptability of server performance monitoring in the prior art due to the lack of consideration of the influence of different load conditions, and to propose an intelligent monitoring method for server performance, which includes:

[0007] Collect the historical load data of the target server, analyze the historical load data to determine the interval range and change degree of the load, and split the interval range of the load according to the change degree to obtain multiple partial intervals of the load;

[0008] Collect the historical performance data of the target server according to multiple partial intervals of the load, confirm the correlation between the load and the historical performance data under different partial intervals of the load, and obtain the standard correlation between load and performance;

[0009] Obtain the load data and performance data of the target server within a certain period of time, and input the load data and performance data into the standard correlation relationship between load and performance to determine the deviation situation;

[0010] Analyze the deviation situation and output the current performance information of the target server, so as to realize the monitoring of the server performance.

[0011] In some embodiments of the present application, analyze the load historical data to determine the range and degree of change of the load, including,

[0012] Draw a histogram of the load based on the load historical data, mark the minimum and maximum values of the load on the histogram, and take the range between the minimum and maximum values of the load as the range of the load;

[0013] Identify the load gradual change segment and the load jump segment from the load historical data, and integrate the load gradual change segment and the load jump segment respectively to obtain the load gradual change segment record and the load jump segment record;

[0014] Based on the load gradual change segment record and the load jump segment record, determine the occurrence frequency of different ranges within the range of the load, determine the frequent weights of the load gradual change segment record and the load jump segment record under different ranges respectively through the occurrence frequency, perform difference on the different ranges of the load gradual change segment record and the load jump segment record respectively, and combine the difference weights under the corresponding ranges to obtain the change degree of the gradual change and the change degree of the jump;

[0015] Comprehensively determine the change degree of the range of the load based on the change degree of the gradual change and the change degree of the jump.

[0016] In some embodiments of the present application, split the range of the load according to the degree of change to obtain multiple partial ranges of the load, including,

[0017] Calculate the degree of dispersion of the load historical data, determine a splitting quantity according to the degree of dispersion and the change degree of the range of the load, and correspondingly split the range of the load through the splitting quantity to obtain multiple partial ranges of the load.

[0018] In some embodiments of the present application, confirm the correlation relationship between the load and the performance historical data under different partial ranges of the load to obtain the standard correlation relationship between load and performance, including,

[0019] Classify the performance data, determine the corresponding performance data categories and the distribution ranges of each performance data under different partial ranges of the load, calculate the sensitivity between the load and each performance data, and count the distribution parameters of each performance data;

[0020] The distribution parameters of the performance data include mean, standard deviation, quartiles, IQR, skewness and kurtosis;

[0021] Set the first Z - score threshold according to the mean and standard deviation, set the second Z - score threshold according to the quartiles and IQR, and set the third Z - score threshold according to skewness and kurtosis;

[0022] Combine the first Z - score threshold, the second Z - score threshold, and the third Z - score threshold to set the fourth Z - score threshold;

[0023] Distinguish normal values and outliers in the distribution range of each type of performance data through the fourth Z - score threshold, eliminate the outliers, and construct a standard correlation relationship between load and performance based on the normal values in the distribution range of performance data and the sensitivity between the load and each type of performance data.

[0024] In some embodiments of the present application, constructing a standard correlation relationship between load and performance based on the normal values in the distribution range of performance data and the sensitivity between the load and each type of performance data includes,

[0025] On the basis of the normal values in the distribution range of performance data, fit the relationship between the load and the distribution range of performance data in different partial intervals to obtain a regression model, and adjust the regression coefficients in the regression model through the sensitivity between the load and each type of performance data to describe the standard correlation relationship between load and performance.

[0026] In some embodiments of the present application, inputting the load data and performance data into the standard correlation relationship between load and performance to determine the deviation situation includes,

[0027] Input the load data and performance data into the corresponding regression model according to different partial intervals of the load data, and count the deviation under each load partial interval, and draw the deviation curve of each load partial interval;

[0028] Integrate the deviation curves of all load partial intervals, extract the curve features of all deviation curves, and generate the deviation situation.

[0029] In some embodiments of the present application, analyzing the deviation situation and outputting the current performance information of the target server includes,

[0030] Determine the curve features under each type of performance data, generate a matching coefficient under each type of performance data by synthesizing the curve features, determine the basic performance index of each type of performance data according to the parameters of each type of performance data, and determine the performance index by combining the basic performance index and the matching coefficient of each type of performance data;

[0031] Integrate the performance indexes under all performance data categories to output the current performance information of the target server.

[0032] In some embodiments of the present application, the performance metrics under all performance data categories are integrated to output the current performance information of the target server, including,

[0033] The performance metrics under all performance data categories are integrated to obtain the overall performance metric of the target server;

[0034] The performance metrics under all performance data categories are analyzed to obtain the performance status of each type of performance data, and the performance status, performance metrics, and overall performance metric of each type of performance data are used as the current performance information of the target server.

[0035] Correspondingly, the present application also provides a server performance intelligent monitoring system, including,

[0036] A first module for collecting the load historical data of the target server, analyzing the load historical data to determine the range and degree of change of the load, and splitting the range of the load according to the degree of change to obtain multiple partial ranges of the load;

[0037] A second module for collecting the performance historical data of the target server according to multiple partial ranges of the load, confirming the correlation between the load and the performance historical data under different partial ranges of the load, and obtaining the standard load-performance correlation;

[0038] A third module for obtaining the load data and performance data of the target server within a current period of time, and inputting the load data and performance data into the standard load-performance correlation to determine the deviation situation;

[0039] A fourth module for analyzing the deviation situation and outputting the current performance information of the target server, thereby realizing the monitoring of the server performance.

[0040] Compared with the prior art, the beneficial effects of the present invention are:

[0041] 1. Analyzing the load historical data to determine the range and degree of change of the load provides a reliable basis for the subsequent splitting of the load range, taking into account the past changes in the server load, so as to accurately describe the correlation between the detailed load range and the performance data. Confirming the correlation between the load and the performance historical data under different partial ranges of the load, obtaining the standard load-performance correlation, and selecting the normal values of the performance historical data in the corresponding relationship between the load and the performance historical data to construct the standard load-performance correlation ensures the accuracy of the description of the correlation between the load and each type of performance data under different load ranges.

[0042] 2. Input the load data and performance data into the standard correlation relationship between load and performance to determine the deviation situation, analyze the deviation situation, output the current performance information of the target server, improve the accuracy and adaptability of server performance monitoring, and effectively ensure the efficient and secure operation of the server. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flow chart of the server performance intelligent monitoring method proposed by the present invention;

[0044] Figure 2 It is a schematic structural diagram of the server performance intelligent monitoring system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0046] Refer to Figure 1 , the server performance intelligent monitoring method includes the following steps.

[0047] Step S101, collect the load historical data of the target server, analyze the load historical data to determine the interval range and change degree of the load, and split the interval range of the load according to the change degree to obtain multiple partial intervals of the load.

[0048] In this embodiment, system commands (such as uptime, top, sar, etc.) or professional monitoring tools (such as Zabbix, Nagios, Prometheus, etc.) are used to collect the load historical data of the target server. Ensure that the selected tool can provide accurate and comprehensive load data, including but not limited to key indicators such as CPU usage, memory usage, network bandwidth, etc. Store the collected load historical data in a safe and reliable database or file system. Perform necessary data sorting work on the data, such as duplicate removal, formatting, etc., to ensure the integrity and consistency of the data. Preprocess the collected load historical data, including steps such as data cleaning (such as removing outliers, missing values, etc.) and data normalization. Ensure that the preprocessed data can accurately reflect the load situation of the server. Use statistical methods (such as histograms, box plots, etc.) to analyze the preprocessed load data to determine the interval range of the load. In the existing means, according to the distribution of the load data, it is divided into three intervals: low load, medium load, and high load. However, such a vague and general load interval is not suitable for analyzing the relationship between specific load and performance. Therefore, it is necessary to split the load interval according to the previous change situation of the server load to analyze the relationship between load and performance under specific load intervals.

[0049] In some embodiments of the present application, analyzing the load historical data to determine the range and degree of change of the load includes:

[0050] Drawing a histogram of the load based on the load historical data, marking the minimum and maximum values of the load on the histogram, and taking the range between the minimum and maximum values of the load as the range of the load;

[0051] Identifying the load gradual change segments and load jump segments from the load historical data, and integrating the load gradual change segments and load jump segments respectively to obtain the load gradual change segment record and the load jump segment record;

[0052] Based on the load gradual change segment record and the load jump segment record, determining the occurrence frequencies of different ranges within the range of the load, respectively determining the frequent weights of the load gradual change segment record and the load jump segment record in different ranges, taking the difference between different ranges of the load gradual change segment record and the load jump segment record, and combining the differential weighting with the frequent weights in the corresponding ranges to obtain the degree of change of the gradual change and the degree of change of the jump;

[0053] Comprehensively determining the degree of change of the range of the load by combining the degree of change of the gradual change and the degree of change of the jump.

[0054] In this embodiment, the range of the load is the range between the minimum and maximum values of the load on the histogram, which is all possible ranges of the load. There are two modes of load change on the server. One mode is gradual change, and the other mode is direct jump. Gradual change: The load continuously increases or decreases over a period of time. This change is usually relatively stable and is suitable for statistical frequency through interval division to reflect the overall trend of the load. Direct jump: The load changes significantly in a short period of time. This change is usually relatively drastic and is suitable for capturing the load fluctuation through the statistics of specific values. Integrating all the load gradual change segments and all the load jump segments respectively to obtain the load gradual change segment record and the load jump segment record. Under the premise of the two modes, calculating the occurrence frequencies of different ranges within the range of the load, assigning frequent weights, and when calculating the difference, the difference can be weighted according to the occurrence frequency of the load interval. The frequently occurring load intervals should be assigned higher weights in the differential calculation to more accurately reflect their impact on the overall load fluctuation, and combining the two modes to reflect the degree of change of the load.

[0055] In some embodiments of the present application, splitting the range of the load according to the degree of change to obtain multiple partial ranges of the load includes:

[0056] Calculating the degree of dispersion of the load historical data, determining a splitting quantity according to the degree of dispersion and the degree of change of the range of the load, and correspondingly splitting the range of the load through the splitting quantity to obtain multiple partial ranges of the load.

[0057] In this embodiment, the degree of dispersion of the load historical data also indicates the change of the load. By combining the degree of dispersion and the degree of change of the interval range of the load (such as weighted summation), a splitting number is determined. The higher the degree of dispersion and the degree of change, the larger the splitting number, and the more accurately the relationship between the load and the performance data can be analyzed. Too many interval numbers may increase the complexity of the analysis, while too few interval numbers may not accurately reflect the relationship between the load and the performance. Therefore, a reasonable splitting number needs to be determined.

[0058] Step S102: Collect the performance historical data of the target server according to multiple partial intervals of the load, confirm the correlation between the load and the performance historical data under different partial intervals of the load, and obtain the standard correlation between the load and the performance.

[0059] In this embodiment, according to the application scenario and business requirements of the server, appropriate performance metrics are selected. The performance metrics should be able to comprehensively reflect the performance status of the server, such as response time, throughput, number of concurrent users, etc. Use monitoring tools or system commands to collect performance historical data according to multiple partial intervals of the load. Ensure that the collected performance data is consistent with the load data in time for subsequent correlation analysis.

[0060] The performance of the server is reflected in multiple aspects or dimensions, and each aspect or dimension directly reflects different performance situations of the server. The following is a detailed summary and description of these aspects or dimensions:

[0061] 1. CPU Performance

[0062] Reflected aspects: The processing speed and concurrent processing ability of the server.

[0063] Performance situation:

[0064] Processing speed: The CPU is the core component of the server and is responsible for executing various computing tasks. A high-performance CPU can execute computing tasks faster, improving the response speed and performance of the server. For example, in application scenarios that require a large amount of computing, such as data analysis and scientific computing, the performance of the CPU is crucial.

[0065] Concurrent processing ability: Modern CPUs usually have a multi-core design and can handle multiple tasks simultaneously. A multi-core CPU can enhance the concurrent processing ability of the server, enabling it to handle more user requests or tasks simultaneously. This is very important for running large databases, virtualization environments, or high-load applications.

[0066] 2. Memory Performance

[0067] Reflected aspects: The task processing ability, system stability, and database performance of the server.

[0068] Performance:

[0069] Task processing ability: Memory is used to temporarily store data generated by the CPU. Sufficient memory enables the server to process more data and requests, improving the response speed. At the same time, memory also supports more programs to run simultaneously, reducing the switching time between programs and enhancing the system's response speed.

[0070] System stability: A larger memory capacity can alleviate the problem of tight system resources, providing more ample running space, thereby enhancing the system's stability and reliability.

[0071] Database performance: Memory, as the cache area of the database, has an important impact on the database's performance. A larger memory capacity can provide more cache space, reducing the number of disk accesses and significantly improving the database's query and write speeds.

[0072] 3. Disk Performance

[0073] Aspects reflected: The data storage and reading speed of the server, and the overall system efficiency.

[0074] Performance:

[0075] Data storage and reading speed: High-performance hard drives (such as SSDs or NVMe SSDs) offer data read and write speeds far higher than those of traditional mechanical hard drives (HDDs). This directly affects the response time of application programs and the speed of data processing. For example, operations such as database queries, logging, and file transfers can be significantly accelerated.

[0076] Overall system efficiency: Faster disk read and write speeds mean shorter startup times for the operating system and application programs, and higher overall system efficiency. At the same time, high-performance hard drives can also support higher IOPS (Input / Output Operations Per Second), ensuring stable operation under high loads.

[0077] 4. Network Performance

[0078] Aspects reflected: The data transfer speed of the server, and the concurrent request processing ability.

[0079] Performance:

[0080] Data transfer speed: Network bandwidth determines the data transfer speed in the network. High network bandwidth can reduce data transfer latency and improve the efficiency of distributed computing and remote access.

[0081] Concurrent request processing ability: In high-concurrency scenarios, the server needs to handle a large number of requests simultaneously. Sufficient network bandwidth can support more concurrent connections, ensuring that the server can process these requests in a timely manner.

[0082] 5. Reliability and Stability

[0083] Embodiment aspects: The fault tolerance ability and long-term stable operation ability of the server.

[0084] Performance situation:

[0085] Fault tolerance ability: High-quality server hardware and software designs can improve the fault tolerance ability of the server. For example, redundant power supplies and fans, predictable hard drive and fan failures, and RAID systems are all common technologies used to improve server reliability.

[0086] Long-term stable operation ability: The stability and reliability of the server are also reflected in its ability to operate stably for a long time without failures or performance degradation. This is crucial for application scenarios that require continuous service provision (such as websites, databases, etc.).

[0087] In this embodiment, performance historical data of the target server is collected according to multiple partial intervals of the load. Taking the partial intervals of the load as the standard, performance historical data of the target server corresponding to each interval is collected.

[0088] In some embodiments of the present application, the association relationship between the load and the performance historical data under different partial intervals of the load is confirmed to obtain the standard association relationship of load-performance, including,

[0089] Classify the performance data, determine the performance data categories corresponding to different partial intervals of the load and the distribution range of each type of performance data, calculate the sensitivity between the load and each type of performance data, and statistically analyze the distribution parameters of each type of performance data;

[0090] The distribution parameters of the performance data include mean, standard deviation, quartiles, IQR, skewness, and kurtosis;

[0091] Set the first Z-score threshold according to the mean and standard deviation, set the second Z-score threshold according to the quartiles and IQR, and set the third Z-score threshold according to the skewness and kurtosis;

[0092] Combine the first Z-score threshold, the second Z-score threshold, and the third Z-score threshold to set the fourth Z-score threshold;

[0093] Distinguish the normal values and abnormal values in the distribution range of each type of performance data through the fourth Z-score threshold, eliminate the abnormal values, and construct the standard association relationship of load-performance based on the normal values in the distribution range of the performance data and the sensitivity between the load and each type of performance data.

[0094] In this embodiment, for the distribution range of each performance data (the distribution range of parameter values), the sensitivity between the computing load and each performance data is calculated. Here, methods such as the Pearson correlation coefficient can be used to calculate the sensitivity. In order to accurately describe the relationship between the load and performance, the abnormal data (errors, abnormal performance, etc.) in the performance data are removed. Here, the Z-score method is used to identify normal data and abnormal data. The Z-score of each data is calculated, and the Z-score threshold is compared to determine whether it is abnormal data.

[0095] In this embodiment, the distribution parameters of the performance data include the mean, standard deviation, quartiles, IQR, skewness, and kurtosis, which are specifically as follows:

[0096] Mean: The mean is the most common method, which can help us understand the central tendency of the data. In the Z-score method, the mean is used to calculate the difference between each data point and the mean.

[0097] Median: The median is another indicator of the central tendency of the data set. It is the value in the middle after the data is arranged in ascending or descending order. The median is not affected by extreme values. When there are extreme values in the data set, the median can more accurately reflect the central position of the data than the mean.

[0098] Mode: The mode is the value that appears most frequently in the data set. When analyzing the data distribution, the mode can help us understand the most common situation.

[0099] Quartiles: Quartiles can divide the data into four parts, each part containing 25% of the data points. The distance between the first quartile (Q1) and the third quartile (Q3) is called the interquartile range (IQR), which is an indicator to measure the dispersion degree of the data. Quartiles can help us quickly identify outliers. Generally, data points falling outside Q1 - 1.5IQR and Q3 + 1.5IQR are considered outliers.

[0100] Skewness: Skewness describes the symmetry of the data distribution. Positive skewness indicates that the data distribution is skewed to the right, that is, there are more larger values; negative skewness indicates that the data distribution is skewed to the left, that is, there are more smaller values. Skewness can affect the setting of the Z-score threshold because a data distribution with a larger skewness may be more likely to produce extreme values.

[0101] Kurtosis: Kurtosis describes the peakedness of the data distribution. High kurtosis indicates that the data is concentrated near the mean, and low kurtosis indicates that the data is more dispersed. Kurtosis can also affect the setting of the Z-score threshold because a data distribution with a higher kurtosis may be more likely to produce data points close to the mean.

[0102] Based on the mean and standard deviation: In the case where the data is approximately normally distributed, the mean and standard deviation are usually used to calculate the Z-score, and a fixed threshold (such as 3) is set to identify outliers. However, if the data distribution deviates from the normal distribution or has large fluctuations, the threshold may need to be adjusted.

[0103] Consider quartiles and IQR: For data with large fluctuations or data that deviates from the normal distribution, quartiles and IQR can be used to assist in setting the Z-score threshold. For example, the IQR of the data can be calculated, and Q1 - 1.5IQR and Q3 + 1.5IQR can be used as the boundaries for outliers. Then, the Z-score threshold is adjusted based on these boundaries.

[0104] Combined with skewness and kurtosis: When considering setting the Z-score threshold, the skewness and kurtosis of the data can also be combined. If the data has a large skewness or high kurtosis, a more lenient threshold may need to be set to avoid misjudging normal fluctuations as outliers.

[0105] In this embodiment, the fourth Z-score threshold is set by combining the first Z-score threshold, the second Z-score threshold, and the third Z-score threshold. The calculation formula is as follows:

[0106] ;

[0107] where, is the fourth Z-score threshold corresponding to the th performance data in the th partial interval of the load, , , are the contribution weights of the first Z-score threshold, the second Z-score threshold, and the third Z-score threshold respectively, , , are the first Z-score threshold, the second Z-score threshold, and the third Z-score threshold corresponding to the th performance data in the th partial interval of the load, and are respectively , , the maximum and minimum values among the three, is the first constant corresponding to the th performance data in the th partial interval of the load;

[0108] In this embodiment, Represents the correction of the average of the maximum and minimum influence amounts of the first Z-score threshold, the second Z-score threshold, and the third Z-score threshold on the average of their sum, To balance the magnitude of the correction function.

[0109] In some embodiments of the present application, a standard load-performance correlation relationship is constructed based on the normal values in the distribution range of performance data and the sensitivity between the load and each type of performance data, including,

[0110] Based on the normal values in the distribution range of performance data, the relationship between the load and the distribution range of performance data in different partial intervals is fitted to obtain a regression model, and the regression coefficients in the regression model are adjusted through the sensitivity between the load and each type of performance data, so as to describe the standard load-performance correlation relationship.

[0111] In this embodiment, based on the normal values in the distribution range of performance data, a suitable regression model (such as linear regression, polynomial regression, etc.) is selected to fit the relationship between the load and the distribution range of performance data in different partial intervals. The model is trained using the collected data to obtain preliminary regression coefficients. The sensitivity coefficient reflects the degree of influence of load changes on performance data changes. If the sensitivity coefficient is large, it indicates that the load has a significant impact on performance data; otherwise, the impact is small. According to the results of sensitivity analysis, the regression coefficients in the regression model are adjusted. If a certain performance data is highly sensitive to the load, the regression coefficient corresponding to this performance data can be appropriately increased to enhance the influence of the load on this performance data. Conversely, if a certain performance data is less sensitive to the load, the regression coefficient corresponding to this performance data can be appropriately decreased to weaken the influence of the load on this performance data. The adjusted regression model is verified using the validation dataset to evaluate the prediction accuracy and stability of the model. If the model performs poorly, the above steps can be repeated for model optimization until a satisfactory model is obtained, and the standard load-performance correlation relationship in different load intervals is described through different regression models.

[0112] Step S103, obtain the load data and performance data of the target server in a current period of time, and input the load data and performance data into the standard load-performance correlation relationship to determine the deviation situation.

[0113] In this embodiment, real-time data collection: Use monitoring tools to obtain the load data and performance data of the target server in a current period of time in real-time. Data recording: Record the obtained data and store it in a secure and reliable database or file system. The data recording should include two parts: load data and performance data, for subsequent correlation analysis and monitoring. Input the load data and performance data into the corresponding regression model, and count the deviation situation of the performance data.

[0114] In some embodiments of the present application, the load data and performance data are input into the standard correlation relationship between load and performance to determine the deviation situation, including,

[0115] According to different partial intervals of the load data, the load data and performance data are input into the corresponding regression model, and the deviation under each load partial interval is counted, and the deviation curve of each load partial interval is drawn;

[0116] Integrate the deviation curves of all load partial intervals, extract the curve characteristics of all deviation curves, and generate the deviation situation.

[0117] In this embodiment, the performance deviation is the difference between the predicted value and the actual value corresponding to different loads under the regression model. The smaller the deviation, the better the matching between the actual load and the performance, which also describes the matching degree between the relationship of the actual load - performance and the standard correlation relationship of load - performance. The curve characteristics of the deviation curve include slope, extreme value, smoothness, etc.

[0118] Step S104, analyze the deviation situation and output the current performance information of the target server, so as to realize the monitoring of the server performance.

[0119] In this embodiment, the current performance situation of the target server is analyzed by combining the parameter size of the performance data and the deviation situation, and the current performance is comprehensively analyzed from the parameter size of the performance data and the matching situation of the load - performance relationship.

[0120] In some embodiments of the present application, analyzing the deviation situation and outputting the current performance information of the target server includes,

[0121] Determine the curve characteristics under each type of performance data, generate the matching coefficient under each type of performance data by synthesizing the curve characteristics, determine the basic performance index of each type of performance data according to the parameters of each type of performance data, and determine the performance index by combining the basic performance index and the matching coefficient of each type of performance data;

[0122] Integrate the performance indexes under all performance data categories and output the current performance information of the target server.

[0123] In some embodiments of the present application, integrating the performance indexes under all performance data categories and outputting the current performance information of the target server includes,

[0124] Integrate the performance indexes under all performance data categories to obtain the overall performance index of the target server;

[0125] Analyze the performance indexes under all performance data categories to obtain the performance status of each type of performance data, and use the performance status, performance index and overall performance index of each type of performance data as the current performance information of the target server.

[0126] In this embodiment, all curve features are integrated (weighted sum after normalization and mapped to obtain a matching coefficient), a matching coefficient between this type of performance data and the load is obtained, the parameter magnitudes of the index data for each type of performance are evaluated to obtain the basic performance index (the quality represented by the performance parameter), and the current performance information of the target server is output by integrating the performance indexes under all performance data categories.

[0127] ;

[0128] Among them, is the current performance information (total server performance) of the target server, is the number of performance data categories, is the matching coefficient of the th performance data, is the basic performance index of the th performance data, is the number of performance data categories where is the performance difference conversion coefficient, is the matching coefficient of the th performance data less than the corresponding threshold, is the basic performance index of the th performance data less than the corresponding threshold, is a preset constant, represents the correction of the sum of performances of all performance data for the sum of thresholds less than the corresponding ones, which is for balancing the correction function.

[0129] Correspondingly, the present application also provides a server performance intelligent monitoring system, as Figure 2 shown, including,

[0130] The first module is used to collect the load historical data of the target server, analyze the load historical data to determine the interval range and change degree of the load, and split the interval range of the load according to the change degree to obtain multiple partial intervals of the load;

[0131] The second module is used to collect the performance historical data of the target server according to the multiple partial intervals of the load, confirm the correlation between the load and the performance historical data under different partial intervals of the load, and obtain the standard correlation of load - performance;

[0132] The third module is used to obtain the load data and performance data of the target server within a certain period of time, and input the load data and performance data into the standard correlation relationship between load and performance to determine the deviation situation;

[0133] The fourth module is used to analyze the deviation situation and output the current performance information of the target server, so as to realize the monitoring of the server performance.

[0134] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0135] 1. Analyze the load historical data to determine the interval range and change degree of the load, provide a reliable basis for the subsequent splitting of the interval range of the load, take into account the previous change situation of the server load, and thus accurately describe the correlation relationship between the detailed interval range of the load and the performance data. Confirm the correlation relationship between the load and the performance historical data in different partial intervals of the load, obtain the standard correlation relationship between load and performance, and select the normal values of the performance historical data in the corresponding relationship between the load and the performance historical data to construct the standard correlation relationship between load and performance, ensuring the accuracy of the description of the correlation relationship between the load and each performance data in different load intervals by the standard correlation relationship between load and performance.

[0136] 2. Input the load data and performance data into the standard correlation relationship between load and performance to determine the deviation situation, analyze the deviation situation and output the current performance information of the target server, improve the accuracy and adaptability of the server performance monitoring, and effectively ensure the efficient and safe operation of the server.

[0137] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present invention.

[0138] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0139] Those skilled in the art can understand that the modules in the devices in the implementation scenarios can be distributed in the devices in the implementation scenarios according to the description of the implementation scenarios, or can be correspondingly changed and located in one or more devices different from the present implementation scenario. The modules in the above implementation scenarios can be combined into one module, or can be further split into multiple sub-modules.

[0140] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. Server performance intelligent monitoring method, characterized in that, including collecting the load historical data of the target server, analyzing the load historical data to determine the interval range and the degree of change of the load, and splitting the interval range of the load according to the degree of change to obtain multiple partial intervals of the load; collecting the performance historical data of the target server according to the multiple partial intervals of the load, confirming the correlation between the load and the performance historical data under different partial intervals of the load, and obtaining the standard load-performance correlation; obtaining the load data and performance data of the target server within a current period of time, and inputting the load data and performance data into the standard load-performance correlation to determine the deviation situation; analyzing the deviation situation and outputting the current performance information of the target server, thereby realizing the monitoring of the server performance; wherein analyzing the load historical data to determine the interval range and the degree of change of the load, including drawing a histogram of the load according to the load historical data, marking the minimum value and the maximum value of the load on the histogram, and taking the range between the minimum value and the maximum value of the load as the interval range of the load; identifying the load gradual change segment and the load jump segment from the load historical data, and respectively integrating the load gradual change segment and the load jump segment to obtain the load gradual change segment record and the load jump segment record; determining the occurrence frequency of different ranges under the interval range of the load based on the load gradual change segment record and the load jump segment record, respectively determining the frequent weights of different ranges of the load gradual change segment record and the load jump segment record through the occurrence frequency, taking the difference between different ranges of the load gradual change segment record and the load jump segment record, and combining the differential weighting with the frequent weights under the corresponding ranges to obtain the change degree of the gradual change and the change degree of the jump; determining the change degree of the interval range of the load by synthesizing the change degree of the gradual change and the change degree of the jump; 2. The server performance intelligent monitoring method according to claim 1, characterized in that, splitting the interval range of the load according to the degree of change to obtain multiple partial intervals of the load including calculating the degree of dispersion of the load historical data, determining a splitting number according to the degree of dispersion and the change degree of the interval range of the load, and correspondingly splitting the interval range of the load through the splitting number to obtain multiple partial intervals of the load; 3. The intelligent server performance monitoring method according to claim 1, characterized in that confirming the correlation between the load and the performance historical data under different partial intervals of the load to obtain the standard load-performance correlation, including classifying the performance data, determining the corresponding performance data categories and the distribution ranges of each performance data under different partial intervals of the load, calculating the sensitivity between the load and each performance data, and statistically analyzing the distribution parameters of each performance data; the distribution parameters of the performance data include mean, standard deviation, quartile, IQR, skewness and kurtosis; setting a first Z-score threshold according to the mean and the standard deviation, setting a second Z-score threshold according to the quartile and the IQR, and setting a third Z-score threshold according to the skewness and the kurtosis; setting a fourth Z-score threshold by combining the first Z-score threshold, the second Z-score threshold and the third Z-score threshold; Normal values and outliers in the distribution range of each type of performance data are distinguished by the fourth Z-score threshold, and the outliers are removed. Based on the normal values in the distribution range of the performance data and the sensitivity between the load and each type of performance data, a standard association relationship between load and performance is constructed.

4. The server performance intelligent monitoring method according to claim 3, wherein Constructing a standard association relationship between load and performance based on the normal values in the distribution range of the performance data and the sensitivity between the load and each type of performance data includes Based on the normal values in the distribution range of the performance data, fitting the relationship between the load and the distribution range of the performance data in different partial intervals to obtain a regression model, and adjusting the regression coefficients in the regression model through the sensitivity between the load and each type of performance data to describe the standard association relationship between load and performance.

5. The server performance intelligent monitoring method according to claim 4, wherein And inputting the load data and performance data into the standard association relationship between load and performance to determine the deviation situation, including Inputting the load data and performance data into the corresponding regression model according to different partial intervals of the load data, and counting the deviation in each load partial interval, and plotting the deviation curve of each load partial interval; Integrating the deviation curves of all load partial intervals, extracting the curve characteristics of all deviation curves, and generating the deviation situation.

6. The intelligent server performance monitoring method according to claim 5, characterized in that Analyzing the deviation situation and outputting the current performance information of the target server, including Determining the curve characteristics under each type of performance data, comprehensively generating the matching coefficient under each type of performance data based on the curve characteristics, determining the basic performance index of each type of performance data according to the parameters of each type of performance data, and determining the performance index by combining the basic performance index and the matching coefficient of each type of performance data; Integrating the performance indicators under all performance data categories to output the current performance information of the target server.

7. The intelligent server performance monitoring method according to claim 6, characterized in that, Integrating the performance indicators under all performance data categories to output the current performance information of the target server, including Integrating the performance indicators under all performance data categories to obtain the overall performance indicator of the target server; Analyzing the performance indicators under all performance data categories to obtain the performance status of each type of performance data, and taking the performance status, performance indicators and overall performance indicator of each type of performance data as the current performance information of the target server.

8. Server performance intelligent monitoring system, characterized in that, For implementing the server performance intelligent monitoring method according to any one of claims 1-7, the system includes The first module is used to collect the load historical data of the target server, analyze the load historical data to determine the interval range and change degree of the load, and split the interval range of the load according to the change degree to obtain multiple partial intervals of the load; The second module is used to collect the performance historical data of the target server according to multiple partial intervals of the load, confirm the association relationship between the load and the performance historical data in different partial intervals of the load, and obtain the standard association relationship between load and performance; The third module is used to obtain the load data and performance data of the target server within a current period of time, and input the load data and performance data into the standard association relationship between load and performance to determine the deviation situation; The fourth module is used to analyze the deviation situation and output the current performance information of the target server, so as to realize the monitoring of the server performance.

Citation Information

Patent Citations

  • Method, device and system for predicting application performance risk

    CN105354092A

  • Server load balancing method and system based on comparison service

    CN117311984A