A server failure remote monitoring system and method

By monitoring the key performance indicators of the server in real time, analyzing potential risk parameters, and adjusting the monitoring frequency based on the analysis results, the problem of the inability to detect potential server failures in the existing technology is solved, and the stability and security of the server are improved.

CN118885356BActive Publication Date: 2025-08-26RUIHE ENERGY (SHENZHEN) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411022866.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-08-26
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

The existing technology cannot conduct in-depth potential risk analysis on the server, resulting in the inability to detect and handle potential failures in a timely manner, affecting the stability and security of the server.

Method used

By monitoring the server's key performance indicators in real time, counting fault information, building monitoring periods, analyzing potential risk parameters, and adjusting the monitoring frequency based on the risk analysis results to trigger an alarm mechanism.

Benefits of technology

It realizes timely warning and rapid handling of server failures, improves the stability and security of the server, and improves the accuracy and efficiency of fault monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118885356B_ABST
    Figure CN118885356B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of server operation monitoring, and specifically relates to a server fault remote monitoring system and method. This invention collects the real-time operating parameters of key performance indicators of the server, performs risk analysis in combination with potential risk parameters, and updates the monitoring frequency in real time based on the results of the risk analysis, thereby achieving efficient remote monitoring of server faults. During the real-time monitoring process, the ratio of the number and duration of positive fluctuation parameters to negative fluctuation parameters is calculated, and the distribution state of the indicator parameters is evaluated based on the ratio of the number and duration, providing strong data support for risk analysis. This allows accurate judgment of the real-time operating state of the server, improving the accuracy and efficiency of fault monitoring. In addition, the present invention also achieves the associated storage of risk indicators, risk parameters, and risk nodes by constructing a risk database, providing a convenient way for subsequent fault detection and repair.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of server operation monitoring, and in particular relates to a server fault remote monitoring system and method. Background Art

[0002] With the rapid development of information technology, servers, as core devices for data storage, processing, and transmission, are of vital importance to both enterprises and individuals in terms of stability and security. However, during long-term operation, servers may fail due to various factors, such as hardware aging, software vulnerabilities, and network attacks, thus affecting their normal operation. Therefore, how to effectively remotely monitor servers and promptly detect and handle faults has become an urgent issue that needs to be addressed.

[0003] Although there are some methods for remote monitoring of servers in the prior art, they can often only perform simple monitoring of the server's operating status and cannot conduct in-depth analysis and prediction of the server's potential risks, thereby failing to timely discover and handle potential failures. To address this problem, the present invention proposes a server failure remote monitoring system and method, which aims to achieve timely warning and rapid handling of server failures through real-time monitoring of key performance indicators of the server and potential risk analysis. Summary of the Invention

[0004] The purpose of the present invention is to provide a server fault remote monitoring system and method, which can realize real-time monitoring of key performance indicators of the server and analyze the potential risks of the server, so as to timely discover potential faults and improve the stability and security of the server.

[0005] The technical solutions adopted by the present invention are as follows:

[0006] A method for remotely monitoring server failures, comprising:

[0007] Obtaining the server's operating status information, extracting key performance indicators from the operating status information, and collecting statistics on fault information associated with the key performance indicators, recording the information as historical fault information, and then matching the initial monitoring frequency of each key performance indicator based on the historical fault information;

[0008] Counting the occurrence nodes of each of the historical fault information and recording them as monitoring nodes, establishing a monitoring period based on the monitoring nodes, and outputting potential risk parameters when the historical fault information occurs based on the key performance indicators within the monitoring period;

[0009] According to the initial monitoring frequency, real-time operating parameters of each key performance indicator of the server are collected, and risk analysis is performed on the real-time operating parameters in combination with the potential risk parameters, and the real-time operating status of the server is output according to the result of the risk analysis, wherein the real-time operating status includes a normal state and a risk state;

[0010] Under the normal state, statistical analysis is performed on the real-time operating parameters of the key performance indicators, and the monitoring frequency of each key performance indicator is reallocated according to the statistical analysis results;

[0011] Under the risk state, the alarm mechanism is triggered and an alarm signal is issued simultaneously.

[0012] In a preferred solution, the step of obtaining the operating status information of the server and extracting key performance indicators from the operating status information includes:

[0013] Obtain historical operating status information of the server and perform data cleaning to remove redundant and abnormal data in the historical status information;

[0014] Standardize the historical operating status information after data cleaning to obtain a standardized operating status data set;

[0015] Perform feature selection on the standardized operating status dataset to screen out key performance indicators that reflect server performance.

[0016] In a preferred solution, the step of matching the initial monitoring frequency of each key performance indicator according to the historical fault information includes:

[0017] Obtaining the historical fault frequency and duration of each of the key performance indicators;

[0018] quantifying the historical fault occurrence frequency and duration, and recording the quantified historical fault occurrence frequency and duration as a first evaluation parameter and a second evaluation parameter;

[0019] Obtaining an evaluation function, inputting the first evaluation parameter and the second evaluation parameter into the evaluation function, and recording an output result of the evaluation function as an allocation weight;

[0020] The allowed allocation resources are obtained and allocated according to the allocation weights to obtain the initial monitoring frequency of each key performance indicator.

[0021] In a preferred solution, the step of establishing a monitoring period based on the monitoring node and outputting potential risk parameters when historical fault information occurs based on key performance indicators within the monitoring period includes:

[0022] Acquire the monitoring node, perform reverse offset on the monitoring node, and output a monitoring period according to the offset result;

[0023] Collecting key performance indicators within the monitoring period and indicator parameters corresponding to the key performance indicators;

[0024] Arrange the indicator parameters within the monitoring period according to the chronological order of occurrence, perform differential processing on the indicator parameters under adjacent positions, and record the differential results as the fluctuation parameters;

[0025] Outputting the distribution state of the indicator parameters within the monitoring period according to the fluctuation parameter, wherein the distribution state includes an ordered state and a disordered state;

[0026] In the ordered state, a first calculation function is obtained, the indicator parameter is input into the first calculation function, and an output result of the first calculation function is calibrated as a potential risk parameter;

[0027] In the disordered state, a second measurement function is obtained, the indicator parameters are input into the second measurement function, and the output result of the second measurement function is calibrated as a potential risk parameter.

[0028] In a preferred embodiment, the step of outputting the distribution state of the indicator parameter within the monitoring period according to the fluctuation parameter includes:

[0029] Classify the fluctuation parameters during the monitoring period according to their numerical types, recording the fluctuation parameters with positive values ​​as positive fluctuation parameters and recording the fluctuation parameters with negative values ​​as negative fluctuation parameters;

[0030] Counting the number of the positive fluctuation parameters and the negative fluctuation parameters, and calculating the ratio of the number of the positive fluctuation parameters to the number of the negative fluctuation parameters;

[0031] Counting the duration of the positive fluctuation parameter and the negative fluctuation parameter, and calculating the ratio of the duration of the positive fluctuation parameter to the duration of the negative fluctuation parameter;

[0032] Performing a weighted summation process on the quantity ratio and the duration ratio, and recording the weighted summation result as a comprehensive ratio;

[0033] Obtaining a distribution state evaluation interval, and comparing the comprehensive ratio with the distribution state evaluation interval;

[0034] If the comprehensive ratio is within the distribution state evaluation interval, it indicates that the distribution of the indicator parameters during the monitoring period is stable, and the distribution state of the indicator parameters is calibrated as an ordered state;

[0035] If the comprehensive ratio does not belong to the distribution state evaluation interval, it indicates that the distribution of the indicator parameters in the monitoring period is scattered, and the distribution state of the indicator parameters is calibrated as a disordered state.

[0036] In a preferred embodiment, the step of collecting the real-time operating parameters of each key performance indicator of the server, performing risk analysis on the real-time operating parameters in combination with the potential risk parameters, and outputting the real-time operating status of the server based on the results of the risk analysis includes:

[0037] Taking the current node as the reference node, backtracking offset is performed to obtain the backtracking node, and then the evaluation period is constructed based on the backtracking node and the reference node;

[0038] Collecting real-time operating parameters of each key performance indicator during the evaluation period and arranging them according to the collection order;

[0039] According to the arrangement order of the real-time operating parameters, performing difference processing on the real-time operating parameters of adjacent positions to obtain the real-time fluctuation parameter and the distribution state of the real-time operating parameters;

[0040] When the distribution state of the real-time operating parameters is an ordered state, a first verification function is obtained, and the real-time operating parameters within the evaluation period and the potential risk parameters output by the first measurement function are output to the first verification function, and the output result of the first verification function is calibrated as a first state evaluation parameter;

[0041] When the distribution state of the real-time operating parameters is in a disordered state, a second verification function is obtained, and the real-time operating parameters within the evaluation period and the potential risk parameters output by the second measurement function are input into the second verification function, and the output result of the second verification function is calibrated as a second state evaluation parameter;

[0042] Obtain an allowable offset threshold and compare the allowable offset threshold with a first state evaluation parameter or a second state evaluation parameter. When the first state evaluation parameter or the second state evaluation parameter is greater than the allowable offset threshold, determine that the real-time operating state of the server is a normal state; otherwise, determine that the real-time operating state of the server is an abnormal state.

[0043] In a preferred embodiment, in the normal state, the steps of performing statistical analysis on the real-time operating parameters of the key performance indicators and reallocating the monitoring frequencies of the key performance indicators according to the statistical analysis results include:

[0044] Obtaining a first state evaluation parameter or a second state evaluation parameter of each of the key performance indicators under a normal state;

[0045] sorting all the first state evaluation parameters or the second state evaluation parameters according to their values;

[0046] A weight allocation function is obtained, and the first state evaluation parameter or the second state evaluation parameter is input into the weight allocation function, and an output result of the weight allocation function is calibrated as an updated weight, and the allowed allocation resources are reallocated according to the updated weight.

[0047] In a preferred solution, under the risk state, the corresponding key performance indicator is output and calibrated as the risk indicator, and then the real-time operating parameters under the risk indicator are calibrated as risk parameters, and the occurrence time of the risk parameter is recorded as the risk node;

[0048] The risk indicators, risk parameters and risk nodes are associated and stored to form a risk database.

[0049] The present invention also provides a server fault remote monitoring system, which uses the above-mentioned server fault remote monitoring method, including:

[0050] A data acquisition module, which is used to obtain the operating status information of the server, extract key performance indicators from the operating status information, collect statistics on fault information under the key performance indicators, and record them as historical fault information, and then match the initial monitoring frequency of each key performance indicator based on the historical fault information;

[0051] a risk analysis module, the risk analysis module being used to count the occurrence nodes of each of the historical fault information and record them as monitoring nodes, establish a monitoring period based on the monitoring nodes, and output potential risk parameters when the historical fault information occurs based on the key performance indicators within the monitoring period;

[0052] a status assessment module, the status assessment module being configured to collect real-time operating parameters of various key performance indicators of the server based on the initial monitoring frequency, perform risk analysis on the real-time operating parameters in combination with the potential risk parameters, and output the real-time operating status of the server based on the results of the risk analysis, wherein the real-time operating status includes a normal state and a risk state;

[0053] An updating module, configured to perform statistical analysis on the real-time operating parameters of the key performance indicators in the normal state, and reallocate the monitoring frequency of each key performance indicator based on the statistical analysis results;

[0054] An alarm module is used to trigger an alarm mechanism and simultaneously send out an alarm signal in the risk state.

[0055] And, an electronic device, comprising:

[0056] at least one processor;

[0057] and a memory communicatively coupled to the at least one processor;

[0058] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned server fault remote monitoring method.

[0059] The technical effects achieved by the present invention are:

[0060] The present invention collects the real-time operating parameters of the key performance indicators of the server, combines them with potential risk parameters for risk analysis, and updates the monitoring frequency in real time according to the results of the risk analysis, thereby realizing efficient remote monitoring of server failures. In the real-time monitoring process, the distribution state of the indicator parameters is evaluated based on the number ratio and duration ratio of the positive fluctuation parameters and the negative fluctuation parameters, providing strong data support for risk analysis, thereby being able to

[0061] Accurately judging the real-time operating status of the server improves the accuracy and efficiency of fault monitoring. In addition, the present invention also realizes the associated storage of risk indicators, risk parameters and risk nodes by constructing a risk database, providing a convenient way for subsequent fault detection and repair. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 It is a schematic flow chart of the method of the present invention;

[0063] Figure 2 It is a schematic diagram of the system modules of the present invention;

[0064] Figure 3 It is a schematic structural diagram of an electronic device of the present invention. DETAILED DESCRIPTION

[0065] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0066] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0067] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in a preferred embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive of other embodiments.

[0068] See also Figure 1 As shown, the present invention provides a method for remote monitoring of server failures, comprising:

[0069] S1. Obtain the server's operating status information, extract key performance indicators from the operating status information, collect statistics on fault information under the key performance indicators, and record them as historical fault information. Then, match the initial monitoring frequency of each key performance indicator based on the historical fault information.

[0070] S2. Count the occurrence nodes of each historical fault information and record them as monitoring nodes. Based on the monitoring nodes, a monitoring period is established. Based on the key performance indicators within the monitoring period, the potential risk parameters when the historical fault information occurs are output;

[0071] S3. Based on the initial monitoring frequency, collect real-time operating parameters of each key performance indicator of the server, perform risk analysis on the real-time operating parameters in combination with potential risk parameters, and output the real-time operating status of the server based on the results of the risk analysis, where the real-time operating status includes a normal state and a risk state;

[0072] S4. Under normal conditions, statistical analysis is performed on the real-time operating parameters of the key performance indicators, and the monitoring frequency of each key performance indicator is reallocated based on the statistical analysis results;

[0073] S5. Under risk conditions, the alarm mechanism is triggered and an alarm signal is issued simultaneously.

[0074] As shown in the above steps, as the server runs, its operating status may change with various factors such as time, load, environment, etc. Therefore, remote monitoring of server faults is very necessary. In this embodiment, the server's operating status information is first obtained, and key performance indicators are selected from the operating status information. Then, the fault information under the key performance indicators is counted and recorded as historical fault information. At the same time, an initial monitoring frequency is set for each key performance indicator based on the historical fault information. Then, the time nodes of the historical fault information are counted and recorded as monitoring nodes. At the same time, a monitoring period is constructed according to the monitoring node, and the historical fault information is output according to the key performance indicators within the monitoring period. The potential risk parameters when the information occurs, and then according to the set initial monitoring frequency, the real-time operating parameters of each key performance indicator of the server are collected, and the real-time operating parameters are analyzed for risk in combination with the potential risk parameters. According to the results of the risk analysis, the real-time operating status of the server is output. The real-time operating status in this embodiment includes a normal state and a risk state. In the normal state, the real-time operating parameters of the key performance indicators are statistically analyzed, and according to the statistical analysis results, the monitoring frequency of each key performance indicator is reallocated. In the risk state, the alarm mechanism will be triggered and an alarm signal will be issued simultaneously, so that timely measures can be taken to prevent the occurrence or expansion of the fault, thereby realizing effective monitoring and prevention of server faults.

[0075] In a preferred embodiment, the steps of obtaining the operating status information of the server and extracting key performance indicators from the operating status information include:

[0076] S101. Obtain historical operating status information of the server and perform data cleaning to remove redundant and abnormal data in the historical status information;

[0077] S102, standardizing the historical operating status information after data cleaning to obtain a standardized operating status data set;

[0078] S103: Perform feature selection on the standardized operating status data set to screen out key performance indicators reflecting server performance.

[0079] As described in steps S101-S103 above, after the server's operating status information is collected, the collected operating status information must first be cleaned to eliminate any possible errors or irrelevant information. This includes removing redundant data points (e.g., repeated readings or useless information) and identifying and eliminating abnormal data (e.g., data caused by temporary failures or abnormal operations). After data cleaning, to ensure data consistency and comparability, the historical operating status information must be standardized, including scaling the data to a uniform scale or converting it to numerical values ​​with the same reference point to facilitate effective comparison between different time points or different server states. Through this standardization process, a unified operating status dataset can be created. In order to accurately identify and extract key indicators that reflect server performance, feature selection must be performed on the standardized dataset. For example, based on decision trees, random forests, or support vector machines, the most representative key performance indicators can be screened out. The key performance indicators are used for subsequent server fault monitoring and early warning.

[0080] In a preferred embodiment, the step of matching the initial monitoring frequency of each key performance indicator according to historical fault information includes:

[0081] S104. Obtain the historical fault frequency and duration of each key performance indicator;

[0082] S105, quantifying the historical fault occurrence frequency and duration, and recording the quantified historical fault occurrence frequency and duration as a first evaluation parameter and a second evaluation parameter;

[0083] S106, obtaining an evaluation function, inputting the first evaluation parameter and the second evaluation parameter into the evaluation function, and recording the output result of the evaluation function as the distribution weight;

[0084] S107: Acquire allowed allocation resources, and allocate the allowed allocation resources according to allocation weights to obtain initial monitoring frequencies of various key performance indicators.

[0085] As described in steps S104-S107 above, in order to ensure the stability and operating efficiency of the server monitoring system, it is necessary to determine the initial monitoring frequency of each key performance indicator. Specifically, first, historical fault information related to each key performance indicator is collected and sorted, including the number of faults and the duration of each fault. Then, this historical fault information is quantified. Specifically, the frequency of faults is converted into a specific numerical value. For example, the frequency value can be obtained by dividing the number of faults by the total operating time. Then, the quantified historical fault frequency and duration are recorded as the first evaluation parameter and the second evaluation parameter, respectively. Then, a preset evaluation function is introduced. The evaluation function can calculate the distribution weight of each key performance indicator based on the first evaluation parameter and the second evaluation parameter. The expression of the evaluation function is: , where represents the assigned weight of the key performance indicator, represents the frequency of failures of key performance indicators, Indicates the duration of historical failures of key performance indicators. represents the total occurrence frequency of all key performance indicators, It represents the total historical fault duration of all key performance indicators. After the weight allocation is output, the allowable allocation resources are obtained. The allowable allocation resources are the monitoring resources of each key performance indicator. Finally, the allowable allocation resources are allocated according to the obtained allocation weights. In this way, the initial monitoring frequency of each key performance indicator can be output, ensuring the optimal utilization of system resources, realizing effective monitoring of key performance indicators, and ensuring the stable operation of the system.

[0086] In a preferred embodiment, the steps of establishing a monitoring period based on monitoring nodes and outputting potential risk parameters when historical fault information occurs based on key performance indicators within the monitoring period include:

[0087] S201, obtaining a monitoring node, performing a reverse offset on the monitoring node, and outputting a monitoring period based on the offset result;

[0088] S202. Collect key performance indicators (KPIs) within the monitoring period, and the corresponding indicator parameters of the KPIs;

[0089] S203: Arrange the indicator parameters within the monitoring period according to the time sequence of occurrence, perform subtraction processing on the indicator parameters in adjacent positions, and record the subtraction result as the fluctuation parameter;

[0090] S204: Outputting the distribution state of the indicator parameters within the monitoring period according to the fluctuation parameters, wherein the distribution state includes an ordered state and a disordered state;

[0091] S205: In the ordered state, obtain a first measurement function, input the indicator parameter into the first measurement function, and calibrate the output result of the first measurement function as a potential risk parameter;

[0092] S206. In the disordered state, obtain a second measurement function, input the indicator parameters into the second measurement function, and calibrate the output result of the second measurement function as a potential risk parameter.

[0093] As described in the above steps S201-S206, after the monitoring node of the historical fault information is output, the monitoring node will be reversely offset to output the time period before the server failure occurs. This embodiment records it as the monitoring period. The specific offset length needs to be set according to the actual situation. After the monitoring period is constructed, the indicator parameters of the key performance indicators in the monitoring period will be collected and arranged synchronously according to the occurrence time sequence of the indicator parameters. Then, the indicator parameters of adjacent positions are subjected to difference processing, so that a series of fluctuation parameters can be obtained. The fluctuation parameters can reflect the changes in the indicator parameters over time. After the fluctuation parameters are output, the distribution state of the indicator parameters in the monitoring period will be output based on them. Here, the distribution state of the indicator parameters includes an ordered state and a disordered state. In the ordered state, a first measurement function will be introduced to analyze the indicator parameters, wherein the expression of the first measurement function is: , where represents the potential risk parameter, Indicates the number of indicator parameters in each fluctuation period within the monitoring period. The fluctuation period is set according to the staggered nodes of positive and negative fluctuation parameters. The time period between adjacent staggered nodes will be recorded as the fluctuation period. represents the number of fluctuation periods, Indicates the length of the fluctuation period, and Represents adjacent volatility parameters. Through the first measurement function, the potential risk parameters in the orderly state can be obtained. In the disordered state, the second measurement function is used for analysis. The expression of the second measurement function is: , where represents the potential risk parameter, represents the number of indicator parameters, Represents the indicator parameters. Similarly, by inputting the indicator parameters into the second measurement function, the potential risk parameters under the disordered state can be obtained.

[0094] In a preferred embodiment, the step of outputting the distribution state of the indicator parameter within the monitoring period according to the fluctuation parameter includes:

[0095] Step 1: Classify the fluctuation parameters according to their numerical types during the monitoring period, recording the fluctuation parameters with positive values ​​as positive fluctuation parameters and the fluctuation parameters with negative values ​​as negative fluctuation parameters;

[0096] Step 2: Count the number of positive fluctuation parameters and negative fluctuation parameters, and calculate the ratio of the number of positive fluctuation parameters to the number of negative fluctuation parameters;

[0097] Step 3: Count the duration of positive and negative fluctuation parameters, and calculate the ratio of the duration of positive and negative fluctuation parameters.

[0098] Step 4: Perform weighted summation on the quantity ratio and the duration ratio, and record the result of the weighted summation as the comprehensive ratio;

[0099] Step 5. Obtain the distribution status evaluation interval and compare the comprehensive ratio with the distribution status evaluation interval;

[0100] If the comprehensive ratio is within the distribution state evaluation interval, it indicates that the distribution of the indicator parameters during the monitoring period is stable, and the distribution state of the indicator parameters is calibrated as an ordered state;

[0101] If the comprehensive ratio does not belong to the distribution status evaluation interval, it means that the distribution of the indicator parameters during the monitoring period is scattered, and the distribution status of the indicator parameters is calibrated as a disordered state.

[0102] As described in the above steps Step 1-Step 5, after the fluctuation parameters are output, they need to be classified accordingly according to the specific numerical types of the fluctuation parameters during the monitoring period, that is, all fluctuation parameters with positive values ​​are accurately recorded as positive fluctuation parameters, and similarly, all fluctuation parameters with negative values ​​are accurately recorded as negative fluctuation parameters. Then, the number of occurrences of positive and negative fluctuation parameters during the monitoring period, as well as the duration during the monitoring period, are counted and recorded, so that the ratio of the number of positive fluctuation parameters to negative fluctuation parameters, as well as the duration ratio of the positive fluctuation parameters to negative fluctuation parameters can be calculated. Finally, the quantity ratio and the duration ratio are further weighted and summed. By processing, an indicator that comprehensively reflects the distribution state of the fluctuation parameters can be obtained. In this embodiment, it is recorded as a comprehensive ratio. In order to evaluate the distribution state of the fluctuation parameters, a distribution state evaluation interval is preset, and the calculated comprehensive ratio is compared with the distribution state evaluation interval to judge whether the distribution state of the indicator parameters during the monitoring period is stable. If the comprehensive ratio falls within the preset evaluation interval, it indicates that the distribution of the indicator parameters during the monitoring period is stable, and it will be calibrated as an ordered state. On the contrary, if the comprehensive ratio fails to fall within this evaluation interval, it is considered that the distribution of the indicator parameters during the monitoring period is scattered, and it will be calibrated as a disordered state.

[0103] In a preferred embodiment, the steps of collecting real-time operating parameters of various key performance indicators of the server, performing risk analysis on the real-time operating parameters in combination with potential risk parameters, and outputting the real-time operating status of the server based on the results of the risk analysis include:

[0104] S301: Using the current node as the reference node, perform backtracking offset to obtain a backtracking node, and then construct an evaluation period based on the backtracking node and the reference node;

[0105] S302. Collect real-time operating parameters of each key performance indicator during the evaluation period and arrange them according to the collection order;

[0106] S303: performing difference processing on adjacent real-time operating parameters according to the arrangement order of the real-time operating parameters to obtain the real-time fluctuation parameter and the distribution state of the real-time operating parameters;

[0107] S304: When the distribution state of the real-time operating parameters is in an ordered state, obtain a first verification function, output the real-time operating parameters within the evaluation period and the potential risk parameters output by the first measurement function to the first verification function, and calibrate the output result of the first verification function as a first state evaluation parameter;

[0108] S305: When the distribution state of the real-time operating parameters is in a disordered state, obtain a second verification function, input the real-time operating parameters within the evaluation period and the potential risk parameters output by the second measurement function into the second verification function, and calibrate the output result of the second verification function as a second state evaluation parameter;

[0109] S306. Obtain an allowable offset threshold, and compare the allowable offset threshold with the first state evaluation parameter or the second state evaluation parameter. When the first state evaluation parameter or the second state evaluation parameter is greater than the allowable offset threshold, determine that the real-time operating state of the server is normal; otherwise, determine that the real-time operating state of the server is abnormal.

[0110] As described in the above steps S301-S306, during the actual operation of the server, the current node will first be selected as the reference node, and then the backtracking node will be found by backtracking offset. Then, based on the backtracking node and the reference node, an evaluation period will be constructed. The time length of the evaluation period is consistent with the time length of the monitoring period. During the evaluation period, it is necessary to collect the real-time operating parameters of each key performance indicator, and arrange these parameters in the order of collection. Then, according to the arrangement order of the real-time operating parameters, the parameters of adjacent positions are subjected to difference processing to obtain the real-time fluctuation parameters and the distribution status of the real-time operating parameters. The process of determining the distribution status of the real-time operating parameters is consistent with the process of determining the distribution status of the above-mentioned indicator parameters, and will not be repeated here. If the distribution status of the real-time operating parameters is ordered, the first verification function is introduced, and the real-time operating parameters in the evaluation period and the potential risk parameters output by the first measurement function are input into the first verification function. The expression of the first verification function is: , where represents the first state evaluation parameter, It indicates the number of real-time operating parameters in each real-time fluctuation period within the evaluation period. The determination method of the real-time fluctuation period is the same as that of the fluctuation period. Indicates the number of real-time fluctuation periods, Indicates the length of the real-time fluctuation period. and Represents adjacent real-time operating parameters. The output result of the first verification function is calibrated as the first state evaluation parameter. If the distribution state of the real-time operating parameters is disordered, the second verification function is introduced. The real-time operating parameters within the evaluation period and the potential risk parameters output by the second measurement function are input into the second verification function. The expression of the second verification function is: , where represents the second state evaluation parameter, Indicates the number of real-time running parameters during the evaluation period, Represents the real-time operating parameters within the evaluation period. Similarly, the output result of the second verification function will be calibrated as the second state evaluation parameter. Then, it is necessary to obtain the preset allowable offset threshold and compare the allowable offset threshold with the first state evaluation parameter or the second state evaluation parameter. If the first state evaluation parameter or the second state evaluation parameter is greater than the allowable offset threshold, it indicates that the real-time operating state of the server is within the normal range. At this time, it can be judged that the server is in a normal state and can continue to operate stably. Otherwise, the real-time operating state of the server is judged to be abnormal.

[0111] In a preferred embodiment, under normal conditions, the steps of performing statistical analysis on the real-time operating parameters of the key performance indicators and reallocating the monitoring frequency of each key performance indicator based on the statistical analysis results include:

[0112] Obtaining a first state evaluation parameter or a second state evaluation parameter of each key performance indicator under a normal state;

[0113] Sorting all first state evaluation parameters or second state evaluation parameters according to their values;

[0114] A weight allocation function is obtained, and the first state evaluation parameter or the second state evaluation parameter is input into the weight allocation function, and the output result of the weight allocation function is calibrated as the updated weight, and the allowed allocation resources are reallocated according to the updated weight.

[0115] In this embodiment, the server, under normal conditions, updates the monitoring frequency according to its corresponding first state evaluation parameter or second state evaluation parameter. Specifically, it is necessary to collect the first state evaluation parameter and the second state evaluation parameter of each key performance indicator under normal operating conditions, and sort the collected first state evaluation parameter or the second state evaluation parameter. The sorting is based on the value of the first state evaluation parameter or the second state evaluation parameter. Then, a weight distribution function is introduced. The weight distribution function maps the input first state evaluation parameter or the second state evaluation parameter to a weight value so that the monitoring resources can be reasonably distributed according to the updated weight. The expression of the weight distribution function is: , where represents the updated weight of the key performance indicator, It represents the sum of the first state evaluation parameter or the second state evaluation parameter under each key performance indicator. By inputting the sorted first state evaluation parameter or the second state evaluation parameter into the weight allocation function and obtaining the output result, the updated weight of the updated key performance indicator can be obtained. Finally, the allowed allocation resources can be reallocated according to the updated weight to ensure that each key performance indicator is monitored more reasonably, thereby improving the overall performance of the system.

[0116] In a preferred embodiment, under the risk state, the corresponding key performance indicators are output and calibrated as risk indicators, and then the real-time operating parameters under the risk indicators are calibrated as risk parameters, and the occurrence time of the risk parameters is recorded as the risk node;

[0117] Risk indicators, risk parameters and risk nodes are associated and stored to form a risk database.

[0118] In this implementation, when the server is in a risk state, it is first necessary to clarify the corresponding key performance indicators and record the key performance indicators as risk indicators. After determining these risk indicators, the real-time operating parameters under the indicators will be further classified as risk parameters. Similarly, the occurrence time of the risk parameters will also be recorded in detail and marked as risk nodes to ensure that a close association is established between the risk indicators, risk parameters and their corresponding risk nodes. The association is stored in the system's database to construct a corresponding risk database, which will help to promptly discover and respond to risk events and take effective countermeasures to ensure the stable operation of the server.

[0119] See also Figure 2 A server fault remote monitoring system, using the above-mentioned server fault remote monitoring method, comprises:

[0120] The data acquisition module is used to obtain the server's operating status information, extract key performance indicators from the operating status information, and collect statistics on fault information under the key performance indicators and record them as historical fault information. The initial monitoring frequency of each key performance indicator is then matched based on the historical fault information.

[0121] Risk analysis module: The risk analysis module is used to count the occurrence nodes of each historical fault information and record them as monitoring nodes. It also builds monitoring periods based on the monitoring nodes and outputs potential risk parameters when historical fault information occurs based on the key performance indicators within the monitoring period.

[0122] The status assessment module is used to collect the real-time operating parameters of each key performance indicator of the server based on the initial monitoring frequency, and conduct risk analysis on the real-time operating parameters in combination with potential risk parameters. Based on the results of the risk analysis, the module outputs the real-time operating status of the server, where the real-time operating status includes normal status and risk status;

[0123] An update module is used to perform statistical analysis on the real-time operating parameters of key performance indicators under normal conditions, and reallocate the monitoring frequency of each key performance indicator based on the statistical analysis results;

[0124] Alarm module, the alarm module is used to trigger the alarm mechanism in a risk state and simultaneously send out an alarm signal.

[0125] In the above, the system includes a data acquisition module, a risk analysis module, a status assessment module, an update module and an alarm module. The data acquisition module is responsible for collecting the operating status information of the server and extracting key performance indicators from all operating status information. After extracting the key performance indicators, the fault information of the key performance indicators will be sorted out and recorded as historical fault information, and the initial monitoring frequency of each key performance indicator will be determined based on the historical fault information. The risk analysis module will be responsible for locating the nodes where the fault information occurs for the historical fault information, and record the nodes where the historical fault information occurs as monitoring nodes. After collecting the monitoring nodes, a monitoring period will be synchronously constructed. During the monitoring period, the system will output the potential risk parameters when the historical fault information occurs according to the key performance indicators, providing relevant information for subsequent risk analysis. Based on the corresponding, the status assessment module will collect the operating parameters of each key performance indicator of the server in real time according to the initial monitoring frequency. After collecting the real-time operating parameters, it will conduct risk analysis in combination with potential risk parameters. After completing the risk analysis, it will output the real-time operating status of the server according to the analysis results. Here, the real-time operating status includes normal status and risk status. The update module is responsible for performing corresponding statistical analysis on the real-time operating parameters of the key performance indicators when the server is in a normal state, and will reallocate the monitoring frequency of each key performance indicator according to the statistical analysis results to ensure the effectiveness of the monitoring work. The alarm module is responsible for triggering the alarm mechanism when the server is in a risky state, and synchronously issuing an alarm signal so as to promptly notify relevant personnel to take corresponding measures to avoid or reduce possible losses.

[0126] See also Figure 3 , an electronic device, the electronic device comprising:

[0127] at least one processor;

[0128] and a memory communicatively coupled to the at least one processor;

[0129] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can execute the above-mentioned server fault remote monitoring method.

[0130] The processor of the electronic device may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor. The memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, a hard disk, a solid-state drive (SSD), an optical disk, etc. The computer program stored in the memory includes instructions. When the processor executes the instructions, it can execute the above-mentioned server fault remote monitoring method. In addition, the electronic device may be equipped with corresponding input / output devices, such as a keyboard, mouse, and display, to facilitate user interaction. Users can use the input / output devices to configure and query the electronic device, as well as view the real-time operating status and fault information of the server. The operator may also be a logic operator, an arithmetic operator, etc., which is used to perform various mathematical and logical operations to support the normal operation of the electronic device and the execution of the server fault remote monitoring method.

[0131] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0132] The foregoing is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art may make various improvements and modifications without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained herein shall, unless otherwise specified or limited, be implemented in accordance with conventional means in the art.

Claims

1. A method for remote monitoring of server failures, characterized by: include: Obtaining the server's operating status information, extracting key performance indicators from the operating status information, and collecting statistics on fault information associated with the key performance indicators, recording the information as historical fault information, and then matching the initial monitoring frequency of each key performance indicator based on the historical fault information; Counting the occurrence nodes of each of the historical fault information and recording them as monitoring nodes, establishing a monitoring period based on the monitoring nodes, and outputting potential risk parameters when the historical fault information occurs based on the key performance indicators within the monitoring period; According to the initial monitoring frequency, real-time operating parameters of each key performance indicator of the server are collected, and risk analysis is performed on the real-time operating parameters in combination with the potential risk parameters, and the real-time operating status of the server is output according to the result of the risk analysis, wherein the real-time operating status includes a normal state and a risk state; Under the normal state, statistical analysis is performed on the real-time operating parameters of the key performance indicators, and the monitoring frequency of each key performance indicator is reallocated according to the statistical analysis results; Under the risk state, the alarm mechanism is triggered and an alarm signal is issued simultaneously; The step of establishing a monitoring period based on the monitoring node and outputting potential risk parameters when historical fault information occurs based on key performance indicators within the monitoring period includes: Acquire the monitoring node, perform reverse offset on the monitoring node, and output a monitoring period according to the offset result; Collecting key performance indicators within the monitoring period and indicator parameters corresponding to the key performance indicators; Arrange the indicator parameters within the monitoring period according to the chronological order of occurrence, perform differential processing on the indicator parameters under adjacent positions, and record the differential results as the fluctuation parameters; Outputting the distribution state of the indicator parameters within the monitoring period according to the fluctuation parameter, wherein the distribution state includes an ordered state and a disordered state; In the ordered state, a first calculation function is obtained, and the indicator parameter is input into the first calculation function, and the output result of the first calculation function is calibrated as a potential risk parameter. The expression of the first calculation function is: , where represents the potential risk parameter, Indicates the number of indicator parameters in each fluctuation period within the monitoring period. The fluctuation period is set according to the staggered nodes of positive and negative fluctuation parameters. The time period between adjacent staggered nodes will be recorded as the fluctuation period. represents the number of fluctuation periods, Indicates the length of the fluctuation period, and represents the adjacent fluctuation parameters; In the disordered state, a second calculation function is obtained, and the indicator parameter is input into the second calculation function, and the output result of the second calculation function is calibrated as the potential risk parameter. The expression of the second calculation function is: , where represents the potential risk parameter, represents the number of indicator parameters, Indicates indicator parameters.

2. A server failure remote monitoring method according to claim 1, characterized in that: The step of obtaining the operating status information of the server and extracting key performance indicators from the operating status information includes: Obtain historical operating status information of the server and perform data cleaning to remove redundant and abnormal data in the historical status information; Standardize the historical operating status information after data cleaning to obtain a standardized operating status data set; Perform feature selection on the standardized operating status dataset to screen out key performance indicators that reflect server performance.

3. A server failure remote monitoring method according to claim 1, characterized in that: The step of matching the initial monitoring frequency of each key performance indicator according to the historical fault information includes: Obtaining the historical fault frequency and duration of each of the key performance indicators; quantifying the historical fault occurrence frequency and duration, and recording the quantified historical fault occurrence frequency and duration as a first evaluation parameter and a second evaluation parameter; Obtain an evaluation function, input the first evaluation parameter and the second evaluation parameter into the evaluation function, and record the output result of the evaluation function as the distribution weight. The expression of the evaluation function is: , where represents the assigned weight of the key performance indicator, represents the frequency of failures of key performance indicators, Indicates the duration of historical failures of key performance indicators. represents the total occurrence frequency of all key performance indicators, Indicates the total historical fault duration of all key performance indicators; The allowed allocation resources are obtained and allocated according to the allocation weights to obtain the initial monitoring frequency of each key performance indicator.

4. A server failure remote monitoring method according to claim 1, characterized in that: The step of outputting the distribution state of the indicator parameters within the monitoring period according to the fluctuation parameters includes: Classify the fluctuation parameters during the monitoring period according to their numerical types, recording the fluctuation parameters with positive values ​​as positive fluctuation parameters and recording the fluctuation parameters with negative values ​​as negative fluctuation parameters; Counting the number of the positive fluctuation parameters and the negative fluctuation parameters, and calculating the ratio of the number of the positive fluctuation parameters to the number of the negative fluctuation parameters; Counting the duration of the positive fluctuation parameter and the negative fluctuation parameter, and calculating the ratio of the duration of the positive fluctuation parameter to the duration of the negative fluctuation parameter; Performing a weighted summation process on the quantity ratio and the duration ratio, and recording the weighted summation result as a comprehensive ratio; Obtaining a distribution state evaluation interval, and comparing the comprehensive ratio with the distribution state evaluation interval; If the comprehensive ratio is within the distribution state evaluation interval, it indicates that the distribution of the indicator parameters during the monitoring period is stable, and the distribution state of the indicator parameters is calibrated as an ordered state; If the comprehensive ratio does not belong to the distribution state evaluation interval, it indicates that the distribution of the indicator parameters in the monitoring period is scattered, and the distribution state of the indicator parameters is calibrated as a disordered state.

5. A server failure remote monitoring method according to claim 1, characterized in that: The step of collecting the real-time operating parameters of each key performance indicator of the server, performing risk analysis on the real-time operating parameters in combination with the potential risk parameters, and outputting the real-time operating status of the server according to the result of the risk analysis includes: Taking the current node as the reference node, backtracking offset is performed to obtain the backtracking node, and then the evaluation period is constructed based on the backtracking node and the reference node; Collecting real-time operating parameters of each key performance indicator during the evaluation period and arranging them according to the collection order; According to the arrangement order of the real-time operating parameters, performing difference processing on the real-time operating parameters of adjacent positions to obtain the real-time fluctuation parameter and the distribution state of the real-time operating parameter; When the distribution state of the real-time operating parameters is an ordered state, a first verification function is obtained, and the real-time operating parameters within the evaluation period and the potential risk parameters output by the first measurement function are output to the first verification function, and the output result of the first verification function is calibrated as a first state evaluation parameter; When the distribution state of the real-time operating parameters is in a disordered state, a second verification function is obtained, and the real-time operating parameters within the evaluation period and the potential risk parameters output by the second measurement function are input into the second verification function, and the output result of the second verification function is calibrated as a second state evaluation parameter; Obtain an allowable offset threshold and compare the allowable offset threshold with a first state evaluation parameter or a second state evaluation parameter. When the first state evaluation parameter or the second state evaluation parameter is greater than the allowable offset threshold, determine that the real-time operating state of the server is a normal state; otherwise, determine that the real-time operating state of the server is an abnormal state.

6. A server failure remote monitoring method according to claim 1, characterized in that: Under the normal state, the steps of performing statistical analysis on the real-time operating parameters of the key performance indicators and reallocating the monitoring frequency of each key performance indicator according to the statistical analysis results include: Obtaining a first state evaluation parameter or a second state evaluation parameter of each of the key performance indicators under a normal state; sorting all the first state evaluation parameters or the second state evaluation parameters according to their values; A weight allocation function is obtained, and the first state evaluation parameter or the second state evaluation parameter is input into the weight allocation function, and an output result of the weight allocation function is calibrated as an updated weight, and the allowed allocation resources are reallocated according to the updated weight.

7. A server failure remote monitoring method according to claim 1, characterized in that: Under the risk state, the corresponding key performance indicators are output and calibrated as risk indicators, the real-time operating parameters under the risk indicators are then calibrated as risk parameters, and the occurrence time of the risk parameters is recorded as the risk node; The risk indicators, risk parameters and risk nodes are associated and stored to form a risk database.

8. A server fault remote monitoring system, using the server fault remote monitoring method according to any one of claims 1 to 7, characterized in that: include: A data acquisition module, which is used to obtain the operating status information of the server, extract key performance indicators from the operating status information, collect statistics on fault information under the key performance indicators, and record them as historical fault information, and then match the initial monitoring frequency of each key performance indicator based on the historical fault information; a risk analysis module, the risk analysis module being used to count the occurrence nodes of each of the historical fault information and record them as monitoring nodes, establish a monitoring period based on the monitoring nodes, and output potential risk parameters when the historical fault information occurs based on the key performance indicators within the monitoring period; a status assessment module, the status assessment module being configured to collect real-time operating parameters of various key performance indicators of the server based on the initial monitoring frequency, perform risk analysis on the real-time operating parameters in combination with the potential risk parameters, and output the real-time operating status of the server based on the results of the risk analysis, wherein the real-time operating status includes a normal state and a risk state; An updating module, configured to perform statistical analysis on the real-time operating parameters of the key performance indicators in the normal state, and reallocate the monitoring frequency of each key performance indicator based on the statistical analysis results; An alarm module is used to trigger an alarm mechanism and simultaneously send out an alarm signal in the risk state.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the server fault remote monitoring method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Remote monitoring system of server hardware and monitoring method thereof

    CN117271267A

  • Server hardware monitoring acquisition method based on time sequence analysis algorithm

    CN117312094A