An AI-based server operation status monitoring and management system
By trending the request volume and response time of the AI server, the CPU usage and response adaptation deviation parameters are calculated, and the server status is evaluated in combination with the weight coefficient, the problem of insufficient early warning in the existing technology is solved, and efficient and accurate fault identification and monitoring is achieved.
Patent Information
- Application Number
- CN202411718914.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-11-28
AI Technical Summary
In the prior art, AI servers cannot detect operating status failures in a timely manner when the number of requests is low, the warning accuracy is insufficient, and it is difficult to analyze fault problems under complex operating conditions, resulting in inaccurate fault identification and analysis.
By dividing the target time period into n time nodes, obtaining and sorting the server request data, calculating the changing feature amount of CPU usage and response time, calculating the CPU usage deviation and response adaptation deviation parameters based on these feature amounts, integrating the server's current operating status parameters with weight coefficients, and setting a threshold to generate an alarm signal.
It realizes accurate identification of server operating status failures under complex operating conditions, improves the accuracy and efficiency of fault analysis, reduces disordered data interference, and improves the accuracy of monitoring and evaluation.
Smart Images

Figure CN119668976B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a server operation status monitoring and management system based on AI intelligence. Background Art
[0002] A server is a specific IT device that provides computing power and runs software applications in a network environment. It provides computing or application services for other client devices such as personal computers and smartphones in the network. Generally, a server has the ability to undertake response service requests, undertake services, and guarantee services. An AI intelligent server, that is, an artificial intelligence server, is a high-performance server specially designed to run artificial intelligence (AI) applications and algorithms. The server operation status monitoring and management system based on AI intelligence is a system for monitoring and managing the operation status of an AI server.
[0003] For example, the application publication number is CN105591816A, the application publication date is May 18, 2016, and the title is "Method for Monitoring the Operation Status of an IT Operation and Maintenance Server". Its monitoring method includes obtaining parameters in three aspects of server performance, server capacity, and server status simultaneously, so that the local area can evaluate the overall status of the server based on these three parameters simultaneously. When a problem occurs in a certain aspect, an alarm is used to inform the user.
[0004] Similar to the above application, in the prior art, when monitoring the operation status of an AI server, key parameters such as CPU usage rate, memory occupancy, and disk space of the operation status of the AI server are mostly collected, and then reasonable thresholds are set for the key parameters of the server. When these parameters exceed the thresholds, an alarm is triggered to notify relevant personnel so as to take measures in time, thereby realizing the monitoring of the server operation status. However, when the server is running with a small request volume, even if there are problems with the server stability or performance, the key parameters collected from the server will not exceed the thresholds, and it is impossible to discover and give an early warning in time, resulting in insufficient early warning accuracy. Relying only on a single parameter or a simple threshold judgment, it is difficult to accurately identify the server operation status fault problems under complex working conditions. Moreover, when a server operation fault early warning occurs, a large amount of monitoring data and server operation logs are directly retrieved for analysis, resulting in troublesome analysis and processing of the fault, and at the same time affecting the accuracy of fault identification and analysis. Summary of the Invention
[0005] The purpose of the present invention is to provide a server operation status monitoring and management system based on AI intelligence to solve the above deficiencies in the prior art.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A server operation status monitoring and management system based on AI intelligence, comprising:
[0007] An operating status monitoring module, which is communicatively connected to the server and is used to access the server and retrieve the server's operating status data. The operating status data includes the CPU usage rate, request volume data, and response time during the server's operation in the target time period.
[0008] A data processing module that integrates and processes the collected server operating status data, and calculates the CPU usage adaptation deviation parameter K based on the relationship between the change in the request volume and the change in the CPU usage rate. (S,Q) ;
[0009] Calculates the response adaptation deviation parameter K based on the relationship between the change in the CPU usage rate and the change in the response time. (Q,T) ;
[0010] An integration calculation module that integrally calculates the current operating status parameters of the server based on the CPU usage adaptation deviation parameter and the calculated response adaptation deviation parameter.
[0011] A status evaluation module that sets the operating status parameter threshold, evaluates the server's operating status based on the server's current operating parameters, and displays the server's operating status failure.
[0012] As a further description of the above technical solution: The integration and processing of the collected server operating status data specifically includes:
[0013] Divides the target time period into n time nodes, and collects the request volume data of the server corresponding to the n time nodes.
[0014] Sequentially sorts the server request volume data from smallest to largest and records them as S1, S2... Sn, where Sn > Sn-1.
[0015] Calculates the change characteristic quantities between adjacent two request volume data and records them as K s1 、K s2 ...K sm , where where K sm represents the change characteristic quantity between the server request volume data Sn and Sn-1.
[0016] As a further description of the above technical solution: Calculating the CPU usage coordination parameter based on the relationship between the change in the request volume and the change in the CPU usage rate specifically includes:
[0017] Sequentially obtains the CPU usage rates corresponding to the server request volume data and records them as Q1, Q2... Qn, where the time node corresponding to the server CPU usage rate Qn is the same as the time node corresponding to the server request volume data Sn.
[0018] Calculates the change characteristic quantities between adjacent two CPU usage rates and records them as Kq1 , K q2 ...K qm , where where K qm represents the change characteristic quantity between the server request volume data Qn and Qn-1;
[0019] The calculation logic for calculating the cpu usage adaptation deviation parameter is as follows:
[0020] where i is an integer from 1 to m, K si represents the change characteristic quantity between the server request volume data at the i-th time node and the (i + 1)-th time node, K qi represents the change characteristic quantity between the server cpu usage rate at the i-th time node and the (i + 1)-th time node, K (S,Q) represents the cpu usage adaptation deviation parameter, where
[0021] As a further description of the above technical solution: calculating the response adaptation deviation parameter based on the relationship between the change in cpu usage rate and the change in response time is specifically as follows:
[0022] Obtain the response time corresponding to the server request volume data and record it as T1, T2,... Tn, where the time node corresponding to the server response time Tn is the same as the time node corresponding to the server request volume data Sn;
[0023] Calculate the change characteristic quantities between adjacent response times and record them as K t1 , K t2 ...K tm , where where K tm represents the change characteristic quantity between the server request volume data Tn and Tn-1;
[0024] The calculation logic for calculating the response adaptation deviation parameter is as follows:
[0025] where i is an integer from 1 to m, K ti represents the change characteristic quantity between the server response times at the i-th time node and the (i + 1)-th time node, K (Q,T) represents the response adaptation deviation parameter, where
[0026] As a further description of the above technical solution: integrating and calculating the current operating state parameter of the server based on the cpu usage adaptation deviation parameter and the calculated response adaptation deviation parameter is specifically as follows:
[0027] Preset the weight coefficients of the CPU usage adaptation deviation parameter and the response adaptation deviation parameter;
[0028] Perform weighted summation calculation on the CPU usage adaptation deviation parameter and the response adaptation deviation parameter to obtain the current operating state parameter of the server.
[0029] As a further description of the above technical solution: Set a parameter threshold, evaluate the operating state of the server based on the current operating parameter of the server and display it specifically as follows:
[0030] Set the threshold of the operating state parameter;
[0031] Retrieve the current operating state parameter of the server and compare it with the threshold of the operating state parameter. When the operating state parameter of the server is greater than the threshold of the operating state parameter, generate an alarm signal and display it.
[0032] As a further description of the above technical solution: It further includes an abnormal node recognition module;
[0033] The abnormal node recognition module is communicatively connected to the data processing module;
[0034] The abnormal node recognition module is used to retrieve the change characteristic quantity data K between the CPU usage rates of two adjacent ones q1 , K q2 ...K qm and the change characteristic quantity data K between the response times of two adjacent ones t1 , K t2 ...K tm , and verify and identify the abnormal time node.
[0035] As a further description of the above technical solution: Verifying and identifying the abnormal time node specifically is as follows:
[0036] Calculate the corresponding node change difference between the change characteristic quantity data of the CPU usage rate and the change characteristic quantity data of the response time, and obtain the node change difference set: {C1, C2, C3...C m}), where C m =K qm -K tm ;
[0037] Calculate the average value of all node change differences, compare each node change difference with the average value of all node change differences, and mark the node difference greater than the average value of all node differences as an abnormal node.
[0038] As a further description of the above technical solution: It further includes an information collection module, which is used to retrieve abnormal node information, collect the operation logs of the server at the corresponding time nodes based on the abnormal node information, and integrate the collected operation logs to obtain the operation log of the server failure characteristics.
[0039] As a further description of the above technical solution: Collecting the operation logs of the server at the corresponding time nodes based on the abnormal node information specifically means that when the node change difference C x is an abnormal node, where x is an integer from 1 to m. At this time, collect the operation log data of the Xth time node and the (X + 1)th time node of the server.
[0040] In the above technical solution, a server operation status monitoring and management system based on AI intelligence provided by the present invention divides the target time period into n time nodes, then distributes and obtains the server request volume data corresponding to the n time nodes, and sorts the data in an increasing manner, so as to realize the processing of the collected data in a trend, and then associates and integrates the cpu usage rate and response time during the server operation with the corresponding nodes of the request volume data, avoiding the interference of the disorder of the server operation data changes on the subsequent processing, and improving the accuracy and efficiency of data analysis;
[0041] By calculating the cpu usage adaptation deviation parameter based on the relationship between the change in the request volume and the change in the cpu usage rate, and calculating the response adaptation deviation parameter based on the relationship between the change in the cpu usage rate and the change in the response time, it realizes the integrated evaluation of the server operation status from two dimensions of the associated adaptation situation of the cpu usage rate and the associated adaptation situation of the changes in the cpu usage rate and the response time when the server request volume data changes in an increasing trend, avoiding the interference of the amount of server operation requests, accurately identifying the server operation status failure problems under complex working conditions, and improving the accuracy of the monitoring and evaluation of the server operation assembly. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0043] Figure 1 It is a schematic diagram of a server operation status monitoring and management system based on AI intelligence provided by an embodiment of the present invention. Detailed Embodiments
[0044] In order to enable those skilled in the art to better understand the technical solutions of the present invention, the following will further introduce the present invention in detail in conjunction with the drawings.
[0045] Please refer to Figure 1 , an embodiment of the present invention provides a technical solution: a server operation status monitoring and management system based on AI intelligence, including:
[0046] An operation status monitoring module, which is communicatively connected to the server and is used to access the server and retrieve the operation status data of the server. The operation status data includes the CPU usage rate, request volume data, and response time of the server during the target time period; where the CPU usage rate represents the busy degree of the CPU during the target time period and reflects the utilization of the server's CPU resources. The request volume data represents the number of requests received and processed by the server during the target time period, and the response time refers to the time consumed from when the client sends a request to when the server returns a response result.
[0047] A data processing module that integrates and processes the collected server operation status data,
[0048] divides the target time period into n time nodes, and collects the request volume data of the server corresponding to the n time nodes;
[0049] sorts the server request volume data from smallest to largest and records them as S1, S2... Sn in sequence, where Sn > Sn-1;
[0050] calculates the change characteristic quantities between adjacent two request volume data and records them as K s1 、K s2 ...K sm , where where K sm represents the change characteristic quantity between the server request volume data Sn and Sn-1. It should be noted that due to the disorder of the server operation data changes, by dividing the target time period into n time nodes, then obtaining the server request volume data corresponding to the n time nodes respectively, and sorting the data in an increasing manner, the collected data is processed in a trend-based manner. Then, the CPU usage rate and response time during the server operation are associated and integrated based on the corresponding nodes of the request volume data, avoiding the interference of the disorder of the server operation data changes on the subsequent processing, and improving the accuracy and efficiency of data analysis.
[0051] Calculates the CPU usage adaptation deviation parameter K (S,Q) based on the change situation of the request volume and the change situation of the CPU usage rate;
[0052] Specifically: sequentially obtain the CPU usage rates corresponding to the server request volume data and record them as Q1, Q2... Qn, where the time node corresponding to the server CPU usage rate Qn is the same as the time node corresponding to the server request volume data Sn;
[0053] Calculate the change characteristic quantities between the CPU utilization rates of adjacent two and record them as K q1 、K q2 ...K qm , where where K qm represents the change characteristic quantity between the server request volume data Qn and Qn-1;
[0054] The calculation logic for calculating the CPU usage adaptation deviation parameter is as follows:
[0055] where i is an integer from 1 to m, K si represents the change characteristic quantity between the server request volume data at the i-th time node and the i+1-th time node, K qi represents the change characteristic quantity between the server CPU utilization rate at the i-th time node and the i+1-th time node, K (S,Q) represents the CPU usage adaptation deviation parameter, where where the CPU usage adaptation deviation parameter represents the deviation situation when the server CPU utilization rate adapts and adjusts the utilization rate of CPU resources based on the growth of server request volume data. Obviously, the larger the CPU usage adaptation deviation parameter, the worse the CPU usage situation adapts to the server operation.
[0056] Calculate the response adaptation deviation parameter K by associating the change situation of the response time with the change situation of the CPU utilization rate (Q,T) ;
[0057] Specifically: Obtain the response times corresponding to the server request volume data and record them as T1, T2,... Tn, where the time node corresponding to the server response time Tn is the same as the time node corresponding to the server request volume data Sn;
[0058] Calculate the change characteristic quantities between adjacent two response times and record them as K t1 、K t2 ...K tm , where where K tm represents the change characteristic quantity between the server request volume data Tn and Tn-1;
[0059] The calculation logic for calculating the response adaptation deviation parameter is as follows:
[0060] where i is an integer from 1 to m, K ti represents the change characteristic quantity between the server response times at the i-th time node and the i+1-th time node, K (Q,T) represents the response adaptation deviation parameter, where It should be noted that when the server request volume data increases, the server CPU usage rate increases synchronously so that the response time of the server remains stable. The response adaptation deviation parameter represents the deviation of adapting the server response time when the CPU usage rate changes. The larger the response adaptation deviation parameter is, the worse the adaptation of the CPU usage rate change to the response time operation is.
[0061] An integration calculation module that integrally calculates the current operating state parameters of the server based on the CPU usage adaptation deviation parameter and the calculated response adaptation deviation parameter;
[0062] A status evaluation module that sets a threshold for the operating state parameters and evaluates and displays the operating state of the server based on the current operating parameters of the server.
[0063] This embodiment provides a server operating state monitoring and management system based on AI intelligence. It calculates the CPU usage adaptation deviation parameter by associating the change in the CPU usage rate with the change in the request volume, and calculates the response adaptation deviation parameter by associating the change in the CPU usage rate with the change in the response time. It realizes the integrated evaluation of the server operating state from two dimensions: the associated adaptation of the CPU usage rate when the server request volume data shows an increasing trend and the associated adaptation of the change in the CPU usage rate and the change in the response time, avoiding the interference of the amount of server operation requests, accurately identifying the server operating state failure problems under complex working conditions, and improving the accuracy of monitoring and evaluating the server operation and assembly.
[0064] The specific integration calculation of the current operating state parameters of the server based on the CPU usage adaptation deviation parameter and the calculated response adaptation deviation parameter is as follows:
[0065] Preset the weight coefficient of the CPU usage adaptation deviation parameter and the weight coefficient of the response adaptation deviation parameter;
[0066] Perform weighted summation calculation on the CPU usage adaptation deviation parameter and the response adaptation deviation parameter to obtain the current operating state parameters of the server. The calculation logic of the current operating state parameters of the server is: Fz = K (S,Q) *β1 + K (Q,T) *β2, where β1 represents the weight coefficient of the CPU usage adaptation deviation parameter, β2 represents the weight coefficient of the response adaptation deviation parameter, and optionally β1 = 0.64 and β2 = 0.36.
[0067] The specific setting of the parameter threshold, evaluating the operating state of the server based on the current operating parameters of the server and displaying it is as follows:
[0068] Set the threshold of the operating state parameters;
[0069] Retrieve the current operating status parameters of the server and compare them with the threshold values of the operating status parameters. When the server operating status parameters are greater than the threshold values of the operating status parameters, an alarm signal is generated, and the server operating status failure is displayed.
[0070] In another embodiment provided by the present invention, preferably, it further includes an abnormal node identification module;
[0071] The abnormal node identification module is communicatively connected to the data processing module. When the status evaluation module evaluates that the server operating status fails based on the current operating parameters of the father, the abnormal node identification module is used to retrieve the change characteristic quantity data K between the CPU usage rates of two adjacent ones q1 、K q2 ...K qm and the change characteristic quantity data K between the response times of two adjacent ones t1 、K t2 ...K tm , and verify and identify the abnormal time node.
[0072] Verifying and identifying the abnormal time node specifically is:
[0073] Calculate the change difference of the corresponding nodes between the change characteristic quantity data of the CPU usage rate and the change characteristic quantity data of the response time, and obtain the node change difference set: {C1, C2, C3... C m}, where C m =K qm -K tm ;
[0074] Calculate the average value of all node change differences, compare each node change difference with the average value of all node change differences, and mark the node difference greater than the average value of all node differences as an abnormal node.
[0075] It further includes an information collection module, which is used to retrieve the abnormal node information, and collect the operation logs of the server at the corresponding time nodes based on the abnormal node information, and integrate the collected operation logs to obtain the server fault characteristic operation logs.
[0076] Collecting the operation logs of the server at the corresponding time nodes based on the abnormal node information specifically means that when the node change difference C xWhen it is an abnormal node, where x is an integer from 1 to m, at this time, the operation log data of the Xth time node and the (X + 1)th time node of the acquisition server are collected. Specifically, the abnormal time node is obtained by verifying and identifying the change characteristic quantity data between the cpu usage rate and the change characteristic quantity data between the response times, and based on the abnormal node information, the operation logs of the corresponding time node servers are collected and integrated to obtain the operation log of the server failure characteristics, so that when the server has an operation status failure, it is possible to specifically collect the operation logs of the servers with significant fault characteristics for targeted analysis, even if the collected data has a stronger correlation with the server operation status failure, there is no need to analyze all the operation logs of the server, with strong pertinence and at the same time improving the accuracy of fault identification and analysis.
[0077] Only some exemplary embodiments of the present invention have been described above by way of illustration. Without doubt, for those of ordinary skill in the art, various different ways can be used to modify the described embodiments without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A server operation status monitoring and management system based on AI intelligence, characterized in that, Including: An operating status monitoring module, which is communicatively connected to the server and is used to access the server and retrieve the server's operating status data. The operating status data includes the CPU usage rate, request volume data, and response time during the server's operation in the target time period; The data processing module integrates and processes the collected server operation status data, and calculates the CPU usage adaptation deviation parameter by associating the change in CPU usage with the change in the request volume. ; The calculation logic for calculating the CPU usage adaptation deviation parameter is: , where i is an integer from 1 to m, represents the change characteristic quantity between the server request volume data of the i-th time node and the (i + 1)-th time node, represents the change characteristic quantity between the server CPU usage rate of the i-th time node and the (i + 1)-th time node, represents the CPU usage adaptation deviation parameter, where , ; Calculate the response adaptation deviation parameter by correlating the change in response time with the change in CPU usage ; The calculation logic for calculating the response adaptation deviation parameter is: , where i is an integer from 1 to m, represents the change characteristic quantity between the server response times of the i-th time node and the (i + 1)-th time node, represents the response adaptation deviation parameter, where ; An integration calculation module, which integrally calculates the current operating status parameter of the server based on the CPU usage adaptation deviation parameter and the calculated response adaptation deviation parameter; A status evaluation module, which sets an operating status parameter threshold, evaluates the server's operating status based on the server's current operating parameter, and displays the server's operating status failure.
2. The server operation status monitoring and management system based on AI intelligence according to claim 1, characterized in that, The specific integration processing of the collected server operating status data is: Dividing the target time period into n time nodes, and collecting the request volume data of the server corresponding to the n time nodes; Sequentially sorting the server request volume data from smallest to largest and recording them as S1, S2... Sn, where Sn > Sn-1; Calculate the change characteristic quantities between adjacent two request volume data and record them respectively as 、 ... , where , where represents the change characteristic quantity between the server request volume data Sn and Sn-1.
3. The server operation status monitoring and management system based on AI intelligence according to claim 2, characterized in that, The specific calculation of the CPU usage adaptation deviation parameter by associating the CPU usage rate change situation based on the request volume change situation is: Sequentially obtaining the CPU usage rates corresponding to the server request volume data and recording them as Q1, Q2... Qn, where the time node corresponding to the server CPU usage rate Qn is the same as the time node corresponding to the server request volume data Sn; Calculate the change feature amounts between the CPU usages of adjacent two and record them respectively as , ... , where , where represents the change feature amount between the server request volume data Qn and Qn-1.
4. An AI intelligence-based server operation status monitoring and management system according to claim 2, characterized in that, The specific calculation of the response adaptation deviation parameter by associating the response time change situation based on the CPU usage rate change situation is: Obtaining the response times corresponding to the server request volume data and recording them as T1, T2,... Tn, where the time node corresponding to the server response time Tn is the same as the time node corresponding to the server request volume data Sn; Calculate the change characteristic quantity between the response times of two adjacent ones and record them respectively as 、 ... , where , where represents the change characteristic quantity between the server request volume data Tn and Tn-1.
5. The server operation status monitoring and management system based on AI intelligence according to claim 4, characterized in that The specific integration calculation of the current operating status parameter of the server based on the CPU usage adaptation deviation parameter and the calculated response adaptation deviation parameter is: Presetting the weight coefficient of the CPU usage adaptation deviation parameter and the weight coefficient of the response adaptation deviation parameter; Performing a weighted sum calculation on the CPU usage adaptation deviation parameter and the response adaptation deviation parameter to obtain the current operating status parameter of the server.
6. The server operation status monitoring and management system based on AI intelligence according to claim 5, characterized in that, Setting the parameter threshold, evaluating the server's operating status based on the server's current operating parameter, and displaying it specifically as: Setting the operating status parameter threshold; Retrieving the current operating status parameter of the server and comparing it with the operating status parameter threshold. When the server operating status parameter is greater than the operating status parameter threshold, an alarm signal is generated and displayed.
7. An AI intelligence-based server operation status monitoring and management system according to claim 4, characterized in that, It further includes an abnormal node identification module; The abnormal node identification module is communicatively connected to the data processing module; The abnormal node recognition module is used to retrieve the change characteristic quantity data between the CPU usage rates of two adjacent ones 、 ... and the change characteristic quantity data between the response times of two adjacent ones 、 ... and verify and identify the abnormal time node.
8. An AI intelligence-based server operation status monitoring and management system according to claim 7, characterized in that, The specific verification and identification to obtain the abnormal time node is: Calculate the difference in the corresponding nodes between the characteristic quantity data of the change in the calculated CPU usage rate and the characteristic quantity data of the change in the response time, and obtain a set of node change differences: { , , ... }, where ; Calculating the average value of the change differences of all nodes, comparing the change differences of each node with the average value of the change differences of all nodes, and marking the node differences greater than the average value of all node differences as abnormal nodes.
9. An AI intelligence-based server operation status monitoring and management system according to claim 8, characterized in that, It further includes an information collection module, which is used to retrieve the abnormal node information, collect the operation logs of the server corresponding to the corresponding time nodes based on the abnormal node information, and integrate the collected operation logs to obtain the operation logs with server failure characteristics.
10. An AI intelligence-based server operation status monitoring and management system according to claim 9, characterized in that, Based on the abnormal node information, the server at the corresponding time node is used to collect the running logs. Specifically, when the difference in node changes is an abnormal node, where x is an integer from 1 to m. At this time, the running log data of the Xth time node and the (X + 1)th time node of the server are collected.
Citation Information
Patent Citations
Detection method for detecting running state of IT operation server
CN105591816A
Temperature administration system
CN104603699A
Application overlay for power estimation mechanisms
CN116635813A