Server operation and maintenance data monitoring method and system

By monitoring server operation and maintenance data in each period and setting targeted thresholds, the problems of incomplete data collection and inapplicable thresholds in the prior art are solved, and the accuracy and business stability of operation and maintenance data monitoring are improved.

CN120086085APending Publication Date: 2025-06-03JIANGSU QIANYUANTONG INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411430252.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-14
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

In the monitoring of server operation and maintenance data, the comprehensiveness of data collection and storage security cannot be guaranteed, and targeted thresholds cannot be set according to the time period, which affects the accuracy of operation and maintenance data identification and business stability during peak business periods.

Method used

By arranging monitoring moments in each period, monitoring the operation and maintenance data of the target server, setting thresholds for various operation and maintenance parameters based on historical operation and maintenance information, analyzing the status of operation and maintenance parameters, and performing corresponding operation and maintenance operations.

Benefits of technology

It improves the effectiveness and accuracy of operation and maintenance data monitoring, ensures the normal operation of the business during peak periods, promptly warns and improves the operation and maintenance supervision effect of the server.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086085A_ABST
    Figure CN120086085A_ABST
Patent Text Reader

Abstract

The invention discloses a server operation and maintenance data monitoring method, and relates to the technical field of server operation and maintenance, and the method comprises the steps: selecting a server operation and maintenance data monitoring tool for analysis, selecting a suitable monitoring tool for monitoring operation and maintenance data, and guaranteeing the comprehensiveness of data collection and the safety of storage. Therefore, reference data is provided for subsequent operation and maintenance data recognition, the operation and maintenance data monitoring effect is improved, targeted threshold values are set for various operation and maintenance data of each target server in each time period according to historical operation and maintenance data, then the states of various operation and maintenance data in each target server are recognized, the operation and maintenance data recognition accuracy is improved, and the operation and maintenance data recognition efficiency is improved. And meanwhile, normal operation of the service is guaranteed in the service peak period, the operation effect of the server is improved, early warning is carried out in time when the operation and maintenance data are abnormal, stable operation of the server is guaranteed, and the operation and maintenance supervision effect of the server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of server operation and maintenance, and particularly relates to a method and system for monitoring server operation and maintenance data. Background Art

[0002] A server is a high-performance computer that can provide various services to multiple users in a network environment, such as data storage, data communication, application program running, etc. By continuously monitoring the server, abnormal states of the server can be detected in a timely manner, and thus measures can be taken promptly to prevent system crashes or performance degradation, ensuring system stability.

[0003] The prior art, such as a method and system for monitoring server operation and maintenance data disclosed in the application with the publication number CN117931577A, belongs to the technical field of data processing and monitoring. By obtaining the multi-dimensional monitoring data to be identified corresponding to the server to be monitored, the operation state and operation and maintenance state of the server can be understood more comprehensively, thereby improving the transparency of server operation and maintenance and ensuring the compliance of operation and maintenance activities; by obtaining the target characteristic parameters under the initial constraint conditions, it is helpful to identify potential abnormal operation states or abnormal operation and maintenance states of the server, and by correcting the target characteristic parameters through a pre-deployed parameter correction model, the noise in the data can be effectively removed, thereby making the monitoring more accurate; by identifying the corrected target characteristic parameters through a pre-deployed characteristic recognition model to determine the operation and maintenance data monitoring results, operation and maintenance monitoring and a certain prediction function can be realized, and finally, by issuing an enhanced monitoring strategy, the operation and maintenance supervision of the server can be effectively strengthened.

[0004] Regarding the above solution, it has at least the following deficiencies: 1. The monitoring of operation and maintenance data in the server is related to the monitoring capabilities of the monitoring tools. Different monitoring tools use different storage methods and monitoring scopes, which affect the subsequent data storage efficiency and the comprehensiveness of data collection. However, in the above solution, the selection of the monitoring tools for server operation and maintenance data is not analyzed, and the comprehensiveness of data collection and the security of data storage cannot be guaranteed, so reference data for the subsequent identification of operation and maintenance data cannot be provided, reducing the monitoring effect of operation and maintenance data.

[0005] 2. The business volumes of different servers are different at different times. Different thresholds can be set for various operation and maintenance data at different times to ensure the normal operation of the business. However, in the above solution, when identifying server operation and maintenance data, targeted thresholds are not set for various operation and maintenance data according to the collection time, so the accuracy of operation and maintenance data identification cannot be improved, and the normal operation of the business cannot be guaranteed during the business peak period, reducing the operation effect of the server. Summary of the Invention

[0006] Aiming at the above-mentioned technical deficiencies, the purpose of the present invention is to provide a server operation and maintenance data monitoring method and system.

[0007] To solve the above technical problems, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a server operation and maintenance data monitoring method, including the following steps: S1. Arrange each monitoring moment in each time period, and monitor the operation and maintenance data corresponding to each monitoring moment of each target server in each time period, where the operation and maintenance data includes various operation and maintenance parameters.

[0008] S2. Extract the historical operation and maintenance information of each target server in each historical time period, and set the thresholds of various operation and maintenance parameters of each target server in each time period.

[0009] S3. Use the operation and maintenance data corresponding to each monitoring moment of each target server in each time period and the thresholds of various operation and maintenance parameters in each time period to analyze the status of various operation and maintenance parameters in each target server, and confirm the operation and maintenance status of each server.

[0010] S4. Perform corresponding operation and maintenance operations according to the operation and maintenance status of each target server.

[0011] In the second aspect, the present invention provides a server operation and maintenance data monitoring system, including: a monitoring module, which is used to arrange each monitoring moment in each time period and monitor the operation and maintenance data corresponding to each monitoring moment of each target server in each time period, where the operation and maintenance data includes various operation and maintenance parameters.

[0012] A threshold setting module, which is used to extract the historical operation and maintenance information of each target server in each historical time period and set the thresholds of various operation and maintenance parameters of each target server in each time period.

[0013] A status analysis module, which is used to use the operation and maintenance data corresponding to each monitoring moment of each target server in each time period and the thresholds of various operation and maintenance parameters in each time period to analyze the status of various operation and maintenance parameters in each target server, and confirm the operation and maintenance status of each server.

[0014] An execution module, which is used to perform corresponding operation and maintenance operations according to the operation and maintenance status of each target server.

[0015] The beneficial effects of the present invention are as follows: The present invention provides a method and system for monitoring server operation and maintenance data. By analyzing the selection of server operation and maintenance data monitoring tools, appropriate monitoring tools are selected to monitor operation and maintenance data, ensuring the comprehensiveness of data collection and the security of storage, thereby providing reference data for the subsequent identification of operation and maintenance data, improving the monitoring effect of operation and maintenance data. According to historical operation and maintenance data, targeted thresholds are set for various types of operation and maintenance data of each target server at each time period, and then the status of various types of operation and maintenance data in each target server is identified, improving the accuracy of operation and maintenance data identification. At the same time, during the peak business period, the normal operation of the business is ensured, the operation effect of the server is improved, and timely warnings are issued when the operation and maintenance data is abnormal, ensuring the stable operation of the server and improving the effect of server operation and maintenance supervision. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0017] Figure 1 It is a schematic flowchart of the implementation steps of the method of the present invention.

[0018] Figure 2 It is a schematic connection diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0020] Please refer to Figure 1 As shown, a method for monitoring server operation and maintenance data includes the following steps: S1. Set monitoring times in each time period, and monitor the operation and maintenance data corresponding to each monitoring time in each time period for each target server, where the operation and maintenance data includes various operation and maintenance parameters.

[0021] Among them, various operation and maintenance parameters include CPU usage rate, disk capacity, network bandwidth usage rate, etc.

[0022] In a specific embodiment, the process of monitoring the operation and maintenance data corresponding to each monitoring moment of each target server in each time period is as follows: Obtain the monitoring data and storage data corresponding to each historical monitoring of each monitoring tool, and at the same time obtain the storage method of each monitoring tool.

[0023] It should be noted that the monitoring data and storage data corresponding to each historical monitoring of each monitoring tool are obtained from the historical records; the monitoring data includes the number of monitored servers, the number of types of operation and maintenance parameters in each monitored server, and the recording moments of various operation and maintenance parameters in each monitored server; the storage data includes the storage capacity of the operation and maintenance data of each monitored server and the storage method.

[0024] Each monitoring tool includes Zabbix, Nagios, and SolarWinds Server & Application Monitor, etc.

[0025] Count the number of target servers, and then use the operation and maintenance data and collection time corresponding to each historical monitoring of each monitoring tool to analyze the monitoring ability matching coefficient of each monitoring tool. At the same time, use the storage data corresponding to each historical monitoring of each monitoring tool and the storage method of each monitoring tool to analyze the storage ability matching coefficient corresponding to each monitoring tool.

[0026] Among them, the storage method includes local storage and cloud storage.

[0027] Preferably, the process of analyzing the monitoring ability matching coefficient of each monitoring tool is as follows: Obtain the number of monitored servers corresponding to each historical monitoring of each monitoring tool and the number of types of operation and maintenance parameters in each monitored server from the monitoring data corresponding to each historical monitoring of each monitoring tool, and at the same time extract the recording moments of various operation and maintenance parameters in each monitored server; Accumulate the number of types of operation and maintenance parameters in each monitored server corresponding to each historical monitoring of each monitoring tool to obtain the total number of types of operation and maintenance parameters monitored corresponding to each historical monitoring of each monitoring tool.

[0028] Obtain various operation and maintenance parameters corresponding to each target server and accumulate them to obtain the total number of types of operation and maintenance parameters to be monitored; Compare the recording moments of various operation and maintenance parameters in each monitored server corresponding to each historical monitoring of each monitoring tool, and calculate the time deviation coefficient corresponding to each monitoring tool, denoted as , where x represents the number corresponding to each monitoring tool, and x is a positive integer.

[0029] It should be added that by comparing the recording times of various historical monitoring operations of each monitoring tool corresponding to various operation and maintenance parameters in each monitoring server, the recording time differences between various operation and maintenance parameters in each monitoring server corresponding to various historical monitoring operations of each monitoring tool are obtained. The maximum time difference is selected as the recording time difference for each historical monitoring operation of each monitoring tool. The recording time differences for each historical monitoring operation of each monitoring tool are averaged to obtain the average recording time difference. The recording time difference for each historical monitoring operation of each monitoring tool is divided by the average recording time difference, and then averaged to obtain the time deviation coefficient corresponding to each monitoring tool.

[0030] Compare the number of monitoring servers corresponding to each historical monitoring operation of each monitoring tool with the number of target servers. Consider each historical monitoring operation in each monitoring tool where the difference between the number of monitoring servers and the number of target servers is less than the preset server number difference threshold as each type of valid monitoring. Thus, count the number of times of each type of valid monitoring in each monitoring tool, denoted as .

[0031] It should be noted that the preset server number difference threshold is a reference value for judging whether the number of monitoring servers is close to the number of target servers, which is set by technical personnel. Suppose the preset server number difference threshold is 2, and the difference between the number of monitoring servers and the number of target servers is 1, and 1 < 2, then it indicates that the number of monitoring servers is close to the number of target servers.

[0032] Compare the total number of operation and maintenance parameter types corresponding to each historical monitoring operation of each monitoring tool with the total number of operation and maintenance parameter types to be monitored. Consider each historical monitoring operation in each monitoring tool where the difference between the total number of operation and maintenance parameter types and the total number of operation and maintenance parameter types to be monitored is less than the preset total number of operation and maintenance parameter type difference threshold as each type of secondary valid monitoring. Thus, count the number of times of each type of secondary valid monitoring in each monitoring tool, denoted as .

[0033] Based on the monitoring ability matching coefficient, analyze and obtain the monitoring ability matching coefficient of each monitoring tool.

[0034] In the above, the expression of the monitoring ability matching coefficient is: , obtain the monitoring ability matching coefficient of the x-th monitoring tool , where z represents the number of monitoring tools, , are respectively the proportionality coefficients of the set number of times of the first type of valid monitoring and the proportionality coefficients of the number of times of the second type of valid monitoring.

[0035] It should be noted that the proportionality coefficients of the number of effective monitoring times of the first type and the proportionality coefficients of the number of effective monitoring times of the second type are jointly discussed and formulated by multiple technical personnel based on the historical operation and maintenance data of the server and the influence degree of the number of effective monitoring times of the first type and the number of effective monitoring times of the second type on the ability matching coefficient of the monitoring tool. For example, when the influence degree of the number of effective monitoring times of the first type on the ability matching coefficient of the monitoring tool is greater, it is 0.6, and it is 0.4.

[0036] Preferably, the storage ability matching coefficients corresponding to each monitoring tool are analyzed as follows: Obtain the operation and maintenance data storage amounts corresponding to each historical monitoring of each target server from the historical storage records, select the mode as the reference operation and maintenance data storage amount corresponding to each target server, and then perform accumulation to obtain the total reference operation and maintenance data storage amount, denoted as WC.

[0037] Accumulate the operation and maintenance data storage amounts corresponding to each historical monitoring of each monitoring server by each monitoring tool to obtain the total operation and maintenance data storage amount corresponding to each historical monitoring of each monitoring tool. If there are more than 1 storage methods for the operation and maintenance data of each monitoring server in a certain historical monitoring of a certain monitoring tool, then remove this historical monitoring. In this way, the remaining historical monitorings are denoted as each marked monitoring, and thus extract the total operation and maintenance data storage amount corresponding to each marked monitoring of each monitoring tool, and denote it as .

[0038] According to the analysis formula , obtain the storage ability matching coefficient of the x-th monitoring tool. In the formula, e represents the natural constant, b represents the total number of marked monitorings, represents the total number of marked monitorings of the x-th monitoring tool.

[0039] When storing data, store it in the same way, and the positions of the data are close. When searching for subsequent data, there is no need to search in multiple data sources, which saves time and effort and improves data management efficiency.

[0040] Accumulate the monitoring ability matching coefficients of each monitoring tool and the storage ability matching coefficients corresponding to each monitoring tool to obtain the ability matching coefficient of each monitoring tool, and select the monitoring tool corresponding to the maximum ability matching coefficient as the target monitoring tool.

[0041] Use the target monitoring tool to record the operation and maintenance data of each target server at each monitoring moment in each time period to obtain the operation and maintenance data corresponding to each target server at each monitoring moment in each time period.

[0042] S2. Extract the historical operation and maintenance information of each target server in each historical time period, and set the thresholds of various operation and maintenance parameters of each target server in each time period.

[0043] Herein, historical operation and maintenance information of each target server in each historical period is extracted from the operation and maintenance records.

[0044] In a specific embodiment, the thresholds of various operation and maintenance parameters of each target server in each time period are set, and the specific setting process is as follows: the thresholds of various operation and maintenance parameters, fault status and performance data of each historical moment are obtained from the historical operation and maintenance information of each target server in each historical time period; the performance data of each target server at each historical moment in each historical time period is used to analyze the performance status of each target server in each historical time period, where the performance status includes good and degraded.

[0045] Preferably, the performance data includes CPU temperature and packet loss rate, etc. The performance data of each target server at each historical moment in each historical period is compared with a preset performance data threshold. If the performance data of a target server at a historical moment in a historical period is greater than the preset performance data threshold, it indicates that the performance status of the target server in the historical period has declined, otherwise it is good, so as to analyze the performance status of each target server in each historical period.

[0046] The historical time periods in each target server that are the same as each time period are taken as the reference time periods corresponding to each time period, thereby extracting the thresholds and fault states of various operation and maintenance parameters of each target server in each time period corresponding to each reference time period, and then integrating the thresholds of various operation and maintenance parameters of each target server in each time period corresponding to each reference time period to obtain the threshold set of various operation and maintenance parameters of each target server in each time period.

[0047] If the fault status of a target server in a certain time period corresponding to a reference time period is faulty or the performance status is degraded, then various abnormal operation and maintenance parameters of the target server in the time period corresponding to the reference time period are obtained from the historical fault records, and the thresholds of various abnormal operation and maintenance parameters are eliminated from the threshold set of various operation and maintenance parameters corresponding to the target server in the time period, so as to obtain the reference thresholds of various operation and maintenance parameters of each target server in each time period and the performance data corresponding to each reference threshold.

[0048] By using the reference thresholds of various operation and maintenance parameters of each target server in each time period and the performance data corresponding to each reference threshold, the impact coefficient of the performance of each target server corresponding to various operation and maintenance parameter thresholds in each time period is analyzed.

[0049] Preferably, the performance impact coefficient of each target server corresponding to each operation and maintenance parameter threshold in each time period is analyzed. The specific analysis process is as follows: using each reference threshold of each operation and maintenance parameter corresponding to each target server in each time period and the performance data corresponding to each reference threshold, the reference threshold of each operation and maintenance parameter of each target server in each time period under each performance data is counted, which is recorded as where \(i\) represents the number of each target server, \(t\) represents the number of each time period, \(j\) represents the number of each type of operation and maintenance parameter, \(n\) represents the number of each performance data, \(m\) represents the number of each reference threshold, and \(i\), \(t\), \(j\), \(n\), and \(m\) are all positive integers. Cluster the reference thresholds of each type of operation and maintenance parameter of each target server in each time period under each performance data, and use the cluster center as the target threshold, denoted as .

[0050] Substitute and into the performance impact coefficient analysis formula to obtain the performance impact coefficients of each target server corresponding to the thresholds of each type of operation and maintenance parameter in each time period.

[0051] Among them, the performance impact coefficient analysis formula is:[[]]END]] , in the formula,[[]]END]] represents the performance impact coefficient of the \(i\)-th target server corresponding to the threshold of the \(j\)-th type of operation and maintenance parameter in the \(t\)-th time period,[[]]END]] represents the target threshold of the \(i\)-th target server corresponding to the \(j\)-th type of operation and maintenance parameter in the \(t\)-th time period under the \((n + 1)\)-th performance data, \(w\) represents the number of performance data, \(r\) represents the number of reference thresholds,[[]]END]] represents the unit value of the \(j\)-th type of operation and maintenance parameter, \(XN\) represents the unit value of the performance data,[[]]END]] , respectively represent the \((n + 1)\)-th performance data of the \(i\)-th target server corresponding to the \(j\)-th type of operation and maintenance parameter in the \(t\)-th time period, \(e\) is the natural constant,[[]]END]] , are respectively the proportionality coefficients of the set internal parameter threshold difference and the external parameter threshold difference.[[]]END]]

[0052] It should be noted that , and , are set in the same way and will not be elaborated here.[[]]END]]

[0053] Take the types of operation and maintenance parameters of each target server in each time period whose impact coefficients are less than or equal to the impact coefficient threshold as each type of operation and maintenance parameter; then cluster the reference thresholds of each target server in each time period corresponding to each type of operation and maintenance parameter to obtain the value corresponding to the cluster center as the threshold of each target server corresponding to each type of operation and maintenance parameter in each time period.[[]]END]]

[0054] Take various operation and maintenance parameters of each target server in each time period with an influence coefficient greater than the influence coefficient threshold as each second-class operation and maintenance parameter, obtain the initial performance data corresponding to each target server, calculate the value coefficient of each second-class operation and maintenance parameter corresponding to each reference threshold for each target server in each time period, and select the reference threshold with the largest value coefficient as the threshold of each second-class operation and maintenance parameter corresponding to each target server in each time period.

[0055] It should be noted that the setting method of the influence coefficient threshold is the same as that of the preset server quantity difference threshold, which will not be elaborated here.

[0056] In the above, the specific process of calculating the value coefficient of each second-class operation and maintenance parameter corresponding to each reference threshold for each target server in each time period is as follows: Cluster each second-class operation and maintenance parameter corresponding to each reference threshold for each target server in each time period, and use the value of the cluster center as the target threshold, denoted as , k represents the number of each second-class operation and maintenance parameter, k is a positive integer, and denote the performance data of each second-class operation and maintenance parameter corresponding to each reference threshold for each target server in each time period as , denote the initial performance data corresponding to each target server as , and denote each second-class operation and maintenance parameter corresponding to each reference threshold for each target server in each time period as .

[0057] The value coefficient analysis formula is: , where represents the value coefficient of the i-th target server corresponding to the k-th second-class operation and maintenance parameter corresponding to the m-th reference threshold in the t-th time period, and e represents the natural constant.

[0058] S3. Use the operation and maintenance data corresponding to each monitoring moment of each target server in each time period and the thresholds of various operation and maintenance parameters in each time period to analyze the status of various operation and maintenance parameters in each target server, and confirm the operation and maintenance status of each server.

[0059] In a specific embodiment, the specific process of analyzing the status of various operation and maintenance parameters in each target server is as follows: Compare the operation and maintenance data corresponding to each monitoring moment of each target server in each time period with the thresholds of various operation and maintenance parameters in each time period. If the operation and maintenance data corresponding to a certain monitoring moment of a certain target server in a certain time period is greater than the threshold of this type of operation and maintenance parameter in this time period, it indicates that the status of this type of operation and maintenance parameter in this target server is in an abnormal state, otherwise it is in a normal state. Analyze the status of various operation and maintenance parameters in each target server in this way.

[0060] S4. Perform corresponding operation and maintenance operations according to the operation and maintenance status of each target server. The specific process is as follows: When the status of a certain type of operation and maintenance parameter in a certain target server is in an abnormal state, display the number of the target server on the monitor, and at the same time display the type of operation and maintenance parameter in the abnormal state, and prompt the staff to view it.

[0061] Please refer to Figure 2 As shown in the figure, a server operation and maintenance data monitoring system includes: a monitoring module, a threshold setting module, a status analysis module, and an execution module.

[0062] The monitoring module is used to arrange monitoring moments in each time period and monitor the operation and maintenance data corresponding to each monitoring moment of each target server in each time period, where the operation and maintenance data includes various types of operation and maintenance parameters.

[0063] The threshold setting module is used to extract the historical operation and maintenance information of each target server in each historical time period and set the thresholds of various operation and maintenance parameters of each target server in each time period.

[0064] The status analysis module is used to analyze the status of various operation and maintenance parameters in each target server by using the operation and maintenance data corresponding to each monitoring moment of each target server in each time period and the thresholds of various operation and maintenance parameters in each time period, and confirm the operation and maintenance status of each server.

[0065] The execution module is used to perform corresponding operation and maintenance operations according to the operation and maintenance status of each target server.

[0066] In the embodiment of the present invention, by analyzing the selection of the server operation and maintenance data monitoring tool, an appropriate monitoring tool is selected to monitor the operation and maintenance data, ensuring the comprehensiveness of data collection and the security of storage, so as to provide reference data for the subsequent identification of operation and maintenance data, improving the monitoring effect of operation and maintenance data. And according to the historical operation and maintenance data, targeted thresholds are set for various operation and maintenance data of each target server in each time period, and then the status of various operation and maintenance data in each target server is identified, improving the accuracy of operation and maintenance data identification. At the same time, during the peak business period, the normal operation of the business is ensured, improving the operation effect of the server, and timely warning is given when the operation and maintenance data is abnormal, ensuring the stable operation of the server and improving the effect of server operation and maintenance supervision.

[0067] The above content is only an example and explanation of the concept of the present invention. Those skilled in the art of the present technology can make various modifications or supplements or use similar methods to replace the specific embodiments described, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all belong to the protection scope of the present invention.

Claims

1. A server operation and maintenance data monitoring method, characterized in that: The steps include: S1. Arrange each monitoring moment in each time period, monitor the operation and maintenance data corresponding to each monitoring moment in each time period of each target server, wherein the operation and maintenance data includes various operation and maintenance parameters; S2. Extract the historical operation and maintenance information of each target server in each historical period, and set the thresholds of various operation and maintenance parameters of each target server in each period; S3. Analyze the status of various operation and maintenance parameters in each target server by using the operation and maintenance data corresponding to each monitoring time in each time period and the threshold values ​​of various operation and maintenance parameters in each time period, and confirm the operation and maintenance status of each server; S4. Execute corresponding operation and maintenance operations according to the operation and maintenance status of each target server.

2. A server operation and maintenance data monitoring method according to claim 1, characterized in that: The specific process of monitoring the operation and maintenance data of each target server corresponding to each monitoring time in each time period is as follows: Obtain the monitoring data and storage data corresponding to each historical monitoring of each monitoring tool, and obtain the storage method of each monitoring tool; Count the number of target servers, and then use the operation and maintenance data and collection time corresponding to each historical monitoring of each monitoring tool to analyze the monitoring capacity matching coefficient of each monitoring tool. At the same time, use the storage data corresponding to each historical monitoring of each monitoring tool and the storage method of each monitoring tool to analyze the storage capacity matching coefficient of each monitoring tool. Accumulate the monitoring capability matching coefficients of each monitoring tool and the storage capability matching coefficients corresponding to each monitoring tool to obtain the capability matching coefficients of each monitoring tool, and select the monitoring tool corresponding to the maximum capability matching coefficient as the target monitoring tool; At each monitoring moment in each time period, the target monitoring tool is used to record the operation and maintenance data of each target server, so as to obtain the operation and maintenance data corresponding to each target server at each monitoring moment in each time period.

3. A server operation and maintenance data monitoring method according to claim 2, characterized in that: The specific process of analyzing the monitoring capability matching coefficient of each monitoring tool is as follows: Obtain the number of monitoring servers corresponding to each historical monitoring of each monitoring tool and the number of operation and maintenance parameter types in each monitoring server from the monitoring data corresponding to each historical monitoring of each monitoring tool, and extract the recording time of each type of operation and maintenance parameter in each monitoring server; accumulate the number of operation and maintenance parameter types in each monitoring server corresponding to each historical monitoring of each monitoring tool to obtain the total number of operation and maintenance parameter types monitored by each historical monitoring of each monitoring tool; Obtain various operation and maintenance parameters corresponding to each target server and add them up to get the total number of operation and maintenance parameter types that need to be monitored; compare the recording time of various operation and maintenance parameters in each monitoring server corresponding to each historical monitoring of each monitoring tool, and calculate the time deviation coefficient corresponding to each monitoring tool, which is recorded as , x represents the number corresponding to each monitoring tool, and x is a positive integer; The number of monitored servers corresponding to each historical monitoring of each monitoring tool is compared with the number of target servers. Each historical monitoring in which the difference between the number of monitored servers and the number of target servers in each monitoring tool is less than the preset server number difference threshold is regarded as a valid monitoring of the first category. The number of valid monitoring of the first category in each monitoring tool is counted and recorded as ; The total number of operation and maintenance parameter types corresponding to each historical monitoring of each monitoring tool is compared with the total number of operation and maintenance parameter types to be monitored. The historical monitoring in which the difference between the total number of operation and maintenance parameter types in each monitoring tool and the total number of operation and maintenance parameter types to be monitored is less than the preset total number difference threshold of operation and maintenance parameter types is regarded as each second-class effective monitoring. The number of second-class effective monitoring in each monitoring tool is counted and recorded as ; According to the monitoring capability matching coefficient, the monitoring capability matching coefficient of each monitoring tool is analyzed and obtained.

4. A server operation and maintenance data monitoring method according to claim 3, characterized in that: The expression of the monitoring capability matching coefficient is: , get the monitoring capability matching coefficient of the xth monitoring tool , where z represents the number of monitoring tools, , They are respectively the proportional coefficient of the first category effective monitoring times and the proportional coefficient of the second category effective monitoring times.

5. A server operation and maintenance data monitoring method according to claim 1, characterized in that: The specific setting process of setting the thresholds of various operation and maintenance parameters of each target server in each time period is as follows: Obtaining thresholds, fault states, and performance data of various operation and maintenance parameters at various historical moments from the historical operation and maintenance information of each target server at various historical time periods; analyzing the performance states of each target server at various historical moments in various historical time periods, wherein the performance states include good and degraded; The historical periods of each target server that are the same as each period are taken as the reference periods corresponding to each period, thereby extracting the thresholds and fault states of various operation and maintenance parameters of each target server in each period corresponding to each reference period, and then integrating the thresholds of various operation and maintenance parameters of each target server in each period corresponding to each reference period to obtain the threshold sets of various operation and maintenance parameters of each target server in each period; If the fault state of a target server in a certain period corresponding to a certain reference period is faulty or the performance state is degraded, then various abnormal operation and maintenance parameters of the target server in the period corresponding to the reference period are obtained from the historical fault records, and the thresholds of various abnormal operation and maintenance parameters are removed from the threshold set of various operation and maintenance parameters corresponding to the target server in the period, so as to obtain various reference thresholds of various operation and maintenance parameters corresponding to each target server in each period and performance data corresponding to each reference threshold; By using the reference thresholds of various operation and maintenance parameters of each target server in each time period and the performance data corresponding to each reference threshold, the influence coefficient of the performance of each target server corresponding to various operation and maintenance parameter thresholds in each time period is analyzed; The various operation and maintenance parameters of each target server in each time period whose influence coefficient is less than or equal to the influence coefficient threshold are taken as the first-class operation and maintenance parameters; then the reference thresholds of the first-class operation and maintenance parameters corresponding to each target server in each time period are clustered to obtain the values ​​corresponding to the cluster centers as the thresholds of the various operation and maintenance parameters corresponding to each target server in each time period; Take various types of operation and maintenance parameters of each target server in each time period whose influence coefficient is greater than the influence coefficient threshold as each type-two operation and maintenance parameter, obtain the initial performance data corresponding to each target server, calculate the value coefficient of each type-two operation and maintenance parameter of each target server in each time period corresponding to each reference threshold, and select the reference threshold with the largest value coefficient as the threshold of each type-two operation and maintenance parameter corresponding to each target server in each time period.

6. A server operation and maintenance data monitoring method according to claim 1, characterized in that: The performance impact coefficient of each target server corresponding to each operation and maintenance parameter threshold in each time period is analyzed. The specific analysis process is as follows: Using the reference thresholds of various operation and maintenance parameters of each target server in each time period and the performance data corresponding to each reference threshold, the reference thresholds of various operation and maintenance parameters of each target server in each time period under each performance data are counted and recorded as , i represents the number of each target server, t represents the number of each time period, j represents the number of each operation and maintenance parameter, n represents the number of each performance data, and m represents the number of each reference threshold. i, t, j, n and m are all positive integers. The reference thresholds of each operation and maintenance parameter of each target server in each time period under each performance data are clustered, and the cluster center is used as the target threshold, which is recorded as ; Will and Substituting into the performance impact coefficient analysis formula, the performance impact coefficient of each target server corresponding to each type of operation and maintenance parameter threshold in each time period is obtained.

7. A server operation and maintenance data monitoring method according to claim 6, characterized in that: The performance impact coefficient analysis formula is: , where It represents the performance impact coefficient of the j-th class operation and maintenance parameter threshold of the i-th target server in the t-th period, represents the target threshold of the j-th class operation and maintenance parameter of the i-th target server in the t-th period under the n+1-th performance data, w represents the number of performance data, r represents the number of reference thresholds, represents the unit value of the j-th class operation and maintenance parameter, XN represents the unit value of the performance data, , They represent the n+1th performance data of the jth class operation and maintenance parameter of the i-th target server in the t-th period, respectively. e is a natural constant. , They are respectively the proportional coefficient of the set internal parameter threshold difference and the proportional coefficient of the external parameter threshold difference.

8. A server operation and maintenance data monitoring method according to claim 6, characterized in that: The specific process of calculating the value coefficient of each reference threshold value corresponding to each second type of operation and maintenance parameter of each target server in each time period is as follows: Cluster the second-class operation and maintenance parameters of each target server in each time period corresponding to each reference threshold, and use the value of the cluster center as the target threshold, denoted as , k represents the number of each type II operation and maintenance parameter, k is a positive integer, and the performance data of each type II operation and maintenance parameter of each target server corresponding to each reference threshold in each time period is recorded as , the initial performance data corresponding to each target server is recorded as And the reference thresholds corresponding to the second type of operation and maintenance parameters of each target server in each period are recorded as ; The value coefficient analysis formula is: , where It represents the value coefficient of the mth reference threshold value corresponding to the kth type II operation and maintenance parameter of the ith target server in the tth time period, and e represents a natural constant.

9. A server operation and maintenance data monitoring method according to claim 1, characterized in that: The specific process of analyzing the status of various operation and maintenance parameters in each target server is as follows: The operation and maintenance data corresponding to each monitoring moment in each time period of each target server are compared with the threshold value of each type of operation and maintenance parameters in each time period. If the operation and maintenance data corresponding to a certain monitoring moment in a certain time period of a target server is greater than the threshold value of this type of operation and maintenance parameter in the time period, it indicates that the state of this type of operation and maintenance parameter in the target server is in an abnormal state, otherwise it is in a normal state, so as to analyze the state of various operation and maintenance parameters in each target server.

10. A server operation and maintenance data monitoring system for executing the server operation and maintenance data monitoring method according to any one of claims 1 to 9, characterized in that: include: The monitoring module is used to arrange each monitoring moment in each time period, monitor the operation and maintenance data corresponding to each monitoring moment in each time period of each target server, wherein the operation and maintenance data includes various operation and maintenance parameters; The threshold setting module is used to extract the historical operation and maintenance information of each target server in each historical period and set the thresholds of various operation and maintenance parameters of each target server in each period; The status analysis module is used to analyze the status of various operation and maintenance parameters in each target server by using the operation and maintenance data corresponding to each monitoring time in each time period and the threshold values ​​of various operation and maintenance parameters in each time period, and confirm the operation and maintenance status of each server; The execution module is used to perform corresponding operation and maintenance operations according to the operation and maintenance status of each target server.

Citation Information

Patent Citations

  • Server operation and maintenance data monitoring method and system

    CN117931577A