A Server Cluster Operation and Maintenance Management and Control System and Method
By collecting and analyzing the power data of the server cluster in real time, combining system resources and network node data, evaluating risk values and implementing differentiated control, the problems of real-time monitoring and precise prevention in server cluster operation and maintenance are solved, and operation and maintenance efficiency and overall stability are improved.
Patent Information
- Application Number
- CN202411002544.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-07-25
AI Technical Summary
The existing technology is difficult to achieve real-time power data acquisition and analysis of server clusters, resulting in the inability to detect power abnormalities in time, and it is difficult to accurately locate performance bottlenecks and potential risks. The scientificity and accuracy of operation and maintenance decisions are insufficient, and the operation and maintenance efficiency needs to be improved.
By collecting power data from the server cluster in real time, building an evaluation cycle to obtain the operating status, comprehensively analyzing system resources and network node data, evaluating risk values and implementing differentiated operation and maintenance management, building feedback cycle monitoring of the operating status, and dynamically adjusting the operation and maintenance strategy.
It realizes efficient operation and maintenance management of the server cluster, quickly identify abnormal servers, avoid potential failures, improve operation and maintenance efficiency and effectiveness, reduce operation costs, and enhance the reliability and stability of the server cluster.
Smart Images

Figure CN118838781B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of server operation and maintenance, and particularly relates to a server cluster operation and maintenance management and control system and method. Background Art
[0002] With the rapid development of information technology, server clusters have become an indispensable infrastructure for modern enterprises and organizations. By combining multiple servers into an integrated whole to jointly provide services, server clusters not only improve the performance and reliability of the system but also enhance the scalability and maintainability of the system. However, with the continuous expansion of the scale and the increase in complexity of server clusters, their operation and maintenance management and control also face many challenges.
[0003] In recent years, with the continuous development of technologies such as cloud computing, big data, and artificial intelligence, new opportunities and challenges have been brought to the operation and maintenance management and control of server clusters. Cloud computing technology provides the ability of elastic scaling and on-demand resource allocation, enabling server clusters to more flexibly respond to changes in business requirements. Big data technology can monitor and analyze the running status of server clusters in real time and promptly discover potential problems and hidden dangers. Artificial intelligence technology can achieve intelligent operation and maintenance management and control of server clusters, improving the automation level and intelligence degree of operation and maintenance.
[0004] However, despite the support of these new technologies, the operation and maintenance management and control of server clusters still face some technical problems. It is difficult to achieve real-time collection and analysis of server power data, resulting in the inability to promptly discover and handle power anomalies. It is difficult to accurately locate the performance bottlenecks and potential risks of servers, leading to insufficient scientificity and accuracy of operation and maintenance decisions. It is difficult to quickly respond to and handle sudden problems in server clusters, and the operation and maintenance efficiency needs to be improved. Summary of the Invention
[0005] The purpose of the present invention is to provide a server cluster operation and maintenance management and control method, which can utilize data collection, analysis, evaluation, and feedback mechanisms to achieve efficient operation and maintenance management and control of server clusters, and has significant practical application value.
[0006] The technical solutions adopted by the present invention are specifically as follows:
[0007] A server cluster operation and maintenance management and control method includes:
[0008] Real-time collect the power data of each server in the server cluster;
[0009] Judge whether the power data of each server meets the preset requirements. If not, mark it as a target server;
[0010] Construct an evaluation period, obtain the running status of the target server within the evaluation period, and mark the target running data;
[0011] Obtain multiple system resource usage data from the target operation data, and obtain the first status value according to the multiple system resource usage data;
[0012] Obtain multiple network node data from the target operation data, and obtain the second status value according to the multiple network node data;
[0013] Evaluate the risk value of the target server according to the first status value and the second status value;
[0014] Obtain the corresponding control level according to the risk value of the target server;
[0015] Perform operation and maintenance control on the server cluster according to the control level;
[0016] Construct a feedback cycle, obtain the running status of each server within the feedback cycle after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback running status;
[0017] Return the feedback running status as the current running status to the step of judging whether the power data of each server meets the preset requirements. If not, mark it as the target server.
[0018] In a preferred solution, the step of judging whether the power data of each server meets the preset requirements. If not, mark it as the target server includes:
[0019] Obtain the standard power threshold;
[0020] Judge whether the power data of each server exceeds the standard power threshold;
[0021] If the power data of the server exceeds the standard power threshold, determine that the power data of the server is abnormal, and mark the server with the power data exceeding the standard power threshold as the target server
[0022] If the power data of the server exceeds the standard power threshold, determine that the power data of the server is normal.
[0023] In a preferred solution, the step of constructing an evaluation cycle, obtaining the running status of the target server within the evaluation cycle, and marking the target operation data includes:
[0024] Obtain the time when the power data of the target server is determined to exceed the standard power threshold, and mark it as the start time of the evaluation cycle;
[0025] Obtain the power data when the target server is determined to exceed the standard power threshold, and mark it as the target power value;
[0026] Obtain the standard evaluation duration and the power weight coefficient;
[0027] Obtain the evaluation duration function, input the target power value, the standard evaluation duration, and the power weight coefficient into the evaluation duration function, and mark the output result as the target evaluation duration;
[0028] Obtain the end time of the evaluation period based on the target evaluation duration and the start time of the evaluation period;
[0029] Obtain the running status of the target server within the start time and the end time of the evaluation period, and mark the target running data.
[0030] In a preferred solution, the step of obtaining multiple system resource usage data from the target running data and obtaining the first status value according to the multiple system resource usage data includes:
[0031] Obtain multiple system resource usage data from the target running data;
[0032] Obtain multiple system resource utilization rates according to the multiple system resource usage data;
[0033] Obtain multiple system resource idle rates according to the multiple system resource usage data;
[0034] Obtain the first status function;
[0035] Input the multiple system resource utilization rates and the multiple system resource idle rates into the first status function, and mark the output result as the first status value.
[0036] In a preferred solution, the step of obtaining multiple network node data from the target running data and obtaining the second status value according to the multiple network node data includes:
[0037] Obtain multiple network node data from the target running data;
[0038] Obtain the corresponding multiple network node values according to the multiple network node data;
[0039] Obtain the second status function;
[0040] Input the multiple network node values into the second status function, and mark the output result as the second status value.
[0041] In a preferred solution, the step of evaluating the risk value of the target server according to the first status value and the second status value includes:
[0042] Obtain the risk function;
[0043] Obtain the corresponding first weight coefficient according to the first status value;
[0044] Obtain the corresponding second weight coefficient according to the second status value;
[0045] Input the first status value, the first weight coefficient, the second status value, and the second weight coefficient into the risk function, and mark the output result as the risk value.
[0046] In a preferred solution, the step of obtaining the corresponding control level according to the risk value of the target server includes:
[0047] Obtain an operation and maintenance level table, where the operation and maintenance level table includes multiple risk assessment interval values and the corresponding control levels for each risk assessment interval value;
[0048] Obtain the corresponding target risk assessment interval value according to the risk value of the target server;
[0049] Obtain the corresponding control level from the operation and maintenance level table according to the target risk assessment interval value.
[0050] In a preferred solution, the step of constructing a feedback period, obtaining the operating status of each server within the feedback period after performing operation and maintenance control on the server cluster according to the control level, and marking it as the feedback operating status includes:
[0051] Obtain the time for performing operation and maintenance control on the server cluster according to the control level, and mark it as the start time of the feedback period;
[0052] Obtain the acquisition frequency of the power data, and obtain the acquisition time interval according to the acquisition frequency;
[0053] Obtain the number of times the power data of the target server exceeds the standard power threshold within the evaluation period according to the acquisition time interval, and mark it as the excess number;
[0054] Obtain a feedback duration function;
[0055] Input the excess number into the feedback duration function, and mark the output result as the feedback period duration;
[0056] Obtain the end time of the feedback period according to the feedback period duration and the start time of the feedback period;
[0057] Obtain the operating status of each server within the start time and end time of the feedback period, and mark it as the feedback operating status.
[0058] The present invention also provides a server cluster operation and maintenance control system for the above server cluster operation and maintenance control method, including:
[0059] A real-time acquisition module for real-time acquiring the power data of each server in the server cluster;
[0060] A judgment module for judging whether the power data of each server meets the preset requirements, and if not, marking it as the target server;
[0061] An operating status module, configured to construct an evaluation period, obtain the operating status of a target server within the evaluation period, and mark target operating data;
[0062] A first status module, configured to obtain multiple system resource usage data from the target operating data, and obtain a first status value according to the multiple system resource usage data;
[0063] A second status module, configured to obtain multiple network node data from the target operating data, and obtain a second status value according to the multiple network node data;
[0064] A risk module, configured to evaluate the risk value of the target server according to the first status value and the second status value;
[0065] A control level module, configured to obtain a corresponding control level according to the risk value of the target server;
[0066] An operation and maintenance control module, configured to perform operation and maintenance control on a server cluster according to the control level;
[0067] A feedback period module, configured to construct a feedback period, obtain the operating status of each server within the feedback period after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback operating status;
[0068] A feedback module, configured to return the feedback operating status as the current operating status to the step of determining whether the power data of each server meets a preset requirement, and if not, mark it as the target server.
[0069] And, an operation and maintenance control terminal for a server cluster, including:
[0070] One or more processors;
[0071] A storage device, on which one or more programs are stored;
[0072] When the one or more programs are executed by the one or more processors, the one or more processors implement the operation and maintenance control method for the server cluster.
[0073] The technical effects achieved by the present invention are:
[0074] The present invention can quickly identify abnormal servers by collecting and analyzing power data in real time, and take corresponding measures to avoid potential failures. By comprehensively analyzing system resources and network node data, it comprehensively evaluates the operating status of the servers, accurately calculates the risk value, thereby achieving precise prevention. According to different risk levels, it implements differentiated operation and maintenance control measures, improves the efficiency and effectiveness of operation and maintenance, avoids excessive or insufficient operation and maintenance operations, dynamically adjusts and optimizes operation and maintenance strategies, can continuously self-improve, improves the overall operating stability and performance, monitors and adjusts the power consumption and resource usage of the servers, improves the energy utilization efficiency of the server cluster, reduces operating costs, discovers and processes potential problems in a timely manner, reduces the server failure rate, and improves the reliability and stability of the entire server cluster. Brief Description of the Drawings
[0075] Figure 1 is the flowchart of the method provided by the present invention;
[0076] Figure 2 is the system module diagram provided by the present invention. Detailed Embodiments
[0077] To make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given in conjunction with the accompanying drawings of the specification.
[0078] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0079] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that can be included in at least one implementation manner of the present invention. The "in a preferred embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.
[0080] Furthermore, the present invention is described in detail in conjunction with the schematic diagrams. When describing the embodiments of the present invention in detail, for the sake of explanation, the schematic diagrams are only examples and should not limit the scope of protection of the present invention here.
[0081] Please refer to the attached Figure 1 As shown, a method for operation and maintenance control of a server cluster is provided, including:
[0082] S1. Collect the power data of each server in the server cluster in real time;
[0083] S2. Determine whether the power data of each server meets the preset requirements. If not, mark it as the target server;
[0084] S3. Construct an evaluation period, obtain the running status of the target server during the evaluation period, and mark the target running data;
[0085] S4. Obtain multiple system resource usage data from the target running data, and obtain the first status value according to the multiple system resource usage data;
[0086] S5. Obtain multiple network node data from the target running data, and obtain the second status value according to the multiple network node data;
[0087] S6. Evaluate the risk value of the target server according to the first status value and the second status value;
[0088] S7. Obtain the corresponding control level according to the risk value of the target server;
[0089] S8. Perform operation and maintenance control on the server cluster according to the control level;
[0090] S9. Construct a feedback period, obtain the running status of each server during the feedback period after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback running status;
[0091] S10. Return the feedback running status as the current running status to the step of determining whether the power data of each server meets the preset requirements. If not, mark it as the target server.
[0092] In the above steps S1 to S10, power consumption data of each server in the server cluster is collected in real time through sensors or other monitoring devices, and the collected power data is analyzed to determine whether it meets the preset power standard. If the power data of a certain server does not meet the requirements, it is marked as a "target server". A specific time period is defined as the evaluation period, and the running status of the target server during this period is continuously monitored, and relevant data is recorded. The usage of multiple system resources, such as CPU, memory, storage, etc., is extracted from the running data of the target server, and a comprehensive "first status value" is calculated. Data of multiple network nodes, such as network traffic, connection count, etc., is extracted from the running data of the target server, and a comprehensive "second status value" is calculated. Combining the first status value and the second status value, the risk value of the target server is evaluated. The risk value can be comprehensively calculated according to the weights of different indicators, reflecting the current running risk of the server. According to the risk value of the target server, the corresponding control level is determined. The control level can be divided into different levels such as low, medium, and high, to reflect the severity of the operation and maintenance measures that need to be taken. According to the determined control level, corresponding operation and maintenance control measures are taken for the server cluster. These measures may include adjusting the server load, optimizing resource allocation, performing preventive maintenance, etc. A feedback period is defined to monitor the running status of the server cluster after the control measures are implemented, and it is recorded and marked as the "feedback running status". The running status data during the feedback period is used as the current running status, and the compliance of the power data is re-evaluated. If new anomalies are found, the target server is re-marked and the above steps are looped. Collecting and analyzing power data in real time can quickly identify abnormal servers and take corresponding measures to avoid the occurrence of potential failures. By comprehensively analyzing system resource and network node data, the running status of the server is comprehensively evaluated, and the risk value is accurately calculated, so as to achieve precise prevention. According to different risk levels, differential operation and maintenance control measures are implemented to improve the operation and maintenance efficiency and effect, and avoid excessive or insufficient operation and maintenance operations. Dynamically adjusting and optimizing the operation and maintenance strategy can continuously self-improve, improve the overall running stability and performance. By monitoring and adjusting the power consumption and resource usage of the server, the energy utilization efficiency of the server cluster is improved, the operation cost is reduced, potential problems are timely discovered and processed, the server failure rate is reduced, and the reliability and stability of the entire server cluster are improved.
[0093] In a specific embodiment, the step of determining whether the power data of each server meets the preset requirements, and if not, marking it as a target server includes:
[0094] S201. Obtain the standard power threshold;
[0095] S202. Determine whether the power data of each server exceeds the standard power threshold;
[0096] If the power data of the server exceeds the standard power threshold, it is determined that the power data of the server is abnormal, and the server with power data exceeding the standard power threshold is marked as the target server
[0097] If the power data of the server exceeds the standard power threshold, it is determined that the power data of the server is normal.
[0098] In the above steps S201 to S202, first, a standard power threshold is set, which is a reasonable power range obtained based on the historical data of the normal operation of the server, the manufacturer's suggestions, or the operation and maintenance experience. This threshold should be able to reflect the power consumption of the server under normal workload. After collecting the real-time power data of the server, it is compared with the preset standard power threshold. If the power data of a certain server exceeds the standard power threshold, it is determined that the power data of this server is abnormal, and it is marked as the target server. If the power data of a certain server does not exceed the standard power threshold, it is determined that the power data of this server is normal. By setting a reasonable standard power threshold, servers with abnormal power consumption can be accurately detected, avoiding misjudgment due to small fluctuations, thereby improving the accuracy of anomaly detection. Comparing the power data with the standard threshold in real time can quickly identify and respond to servers with power anomalies, reducing the time window for faults to occur and improving the operation and maintenance efficiency. Using the standard power threshold as the judgment criterion can effectively reduce false alarms caused by normal fluctuations or short-term high loads, thereby reducing the workload of operation and maintenance personnel and the system alarm frequency. By accurately marking the target servers, these servers can be analyzed and processed in a targeted manner, avoiding large-scale and unnecessary interventions on the entire server cluster, saving resources and time. Adopting a data-driven approach, through historical data and threshold setting, the power threshold can be continuously optimized and adjusted to improve the flexibility and accuracy of judgment. Timely discovery and marking of servers with power anomalies can enable preventive maintenance, handle potential problems in advance, and avoid the expansion of faults or more serious impacts.
[0099] In a specific embodiment, the steps of constructing an evaluation period, obtaining the operating status of the target server during the evaluation period, and marking the target operating data include:
[0100] S301. Obtain the time when the power data of the target server is determined to exceed the standard power threshold, and mark it as the start time of the evaluation period;
[0101] S302. Obtain the power data when the target server is determined to exceed the standard power threshold, and mark it as the target power value;
[0102] S303. Obtain the standard evaluation duration and the power weight coefficient;
[0103] S304. Obtain the evaluation duration function, input the target power value, the standard evaluation duration, and the power weight coefficient into the evaluation duration function, and mark the output result as the target evaluation duration;
[0104] S305. Obtain the end time of the evaluation period based on the target evaluation duration and the start time of the evaluation period;
[0105] S306. Obtain the running status of the target server during the start time and the end time of the evaluation period, and mark the target running data.
[0106] In the above steps S301 to S306, when the power data of a certain server is determined to exceed the standard power threshold, record this time point and mark it as the start time of the evaluation period, record the specific power data of the target server when it exceeds the standard power threshold, and mark this power value as the target power value. The standard evaluation duration is a preset time length, representing the length of the evaluation period under default circumstances. The power weight coefficient is a parameter used to adjust the evaluation duration, dynamically adjusting the evaluation duration according to the abnormality degree of the power data. The evaluation duration function is a mathematical formula or algorithm used to calculate the target evaluation duration based on the target power value, the standard evaluation duration, and the power weight coefficient. The evaluation duration function is , where represents the target evaluation duration, represents the target power value, represents the standard power threshold, q represents the power weight coefficient, represents the standard evaluation duration. Add the target evaluation duration to the start time of the evaluation period to obtain the end time of the evaluation period. During the evaluation period (from the start time to the end time), continuously monitor the running status of the target server and collect relevant data. These running status data include CPU usage rate, memory usage rate, network traffic, etc. Mark these data as the target running data. Through the power weight coefficient and the evaluation duration function, the length of the evaluation period can be dynamically adjusted according to the power abnormality degree of the target server, making the evaluation more flexible and accurate. Record the start and end times of the evaluation period to ensure that the abnormal period of the target server can be fully covered, which helps to deeply analyze the root cause of the problem. During the evaluation period, comprehensively collect various running status data of the target server to provide detailed basic data for subsequent analysis and risk assessment. Through the precise evaluation period and detailed running data, targeted operation and maintenance and optimization can be better carried out, improving the overall performance and stability of the server cluster. The dynamically adjusted evaluation period and comprehensive running data collection help to improve the system's prediction ability for abnormal situations and take preventive measures in advance.
[0107] In a specific embodiment, the steps of obtaining a plurality of system resource usage data from the target operation data and obtaining a first status value according to the plurality of system resource usage data include:
[0108] S401. Obtain a plurality of system resource usage data from the target operation data;
[0109] S402. Obtain a plurality of system resource utilization rates according to the plurality of system resource usage data;
[0110] S403. Obtain a plurality of system resource idle rates according to the plurality of system resource usage data;
[0111] S404. Obtain a first status function;
[0112] S405. Input the plurality of system resource utilization rates and the plurality of system resource idle rates into the first status function, and mark the output result as the first status value.
[0113] In the above steps S401 to S405, the usage conditions of a plurality of system resources are extracted from the target operation data, including but not limited to CPU usage data, memory usage data, storage usage data, network bandwidth usage data, etc. According to the obtained system resource usage data, the utilization rate of each system resource is calculated. According to the obtained system resource usage data, the idle rate of each system resource is calculated. The first status function is , where Z represents the first status value, v represents the number of the system resource utilization rate and the number of the system resource idle rate, u represents the total number of the system resource utilization rate and the total number of the system resource idle rate, represents the vth system resource utilization rate, represents the vth system resource idle rate. Taking the utilization rates and idle rates of a plurality of system resources as inputs and substituting them into the first status function, the first status value is calculated. This status value can comprehensively reflect the operation status of the server during the evaluation period. By obtaining the usage data, utilization rates and idle rates of a plurality of system resources, the usage conditions of each resource of the server can be comprehensively understood, providing a detailed data basis for subsequent status evaluation. The first status function calculates a comprehensive status value by comprehensively calculating the utilization rate and idle rate of system resources, which can more accurately reflect the overall operation situation of the server rather than the performance of a single indicator. Through detailed calculation of resource utilization rates and idle rates, the bottleneck resources in the server operation can be accurately identified, helping the operation and maintenance personnel to optimize and adjust targeted. Dynamically obtaining and calculating the usage conditions of system resources can timely reflect the load change situation, improving the operation stability and efficiency.
[0114] In a specific embodiment, the step of obtaining a plurality of network node data from the target operation data and obtaining a second state value according to the plurality of network node data includes:
[0115] S501. Obtain a plurality of network node data from the target operation data;
[0116] S502. Obtain corresponding multiple network node values according to the plurality of network node data;
[0117] S503. Obtain a second state function;
[0118] S504. Input the multiple network node values into the second state function, and mark the output result as the second state value.
[0119] In the above steps S501 to S504, relevant data of multiple network nodes are extracted from the operation data of the target server. These data may include but are not limited to: network traffic, packet loss rate, latency, number of connections, etc. According to the obtained network node data, corresponding multiple network node values are calculated. For example, calculate the traffic value, packet loss rate value, latency value, connection number value, etc. of each node. The second state function is , where Y represents the second state value, i represents the number of the network node value, n represents the total number of network node values, represents the i-th network node value. Taking the multiple network node values as inputs and substituting them into the second state function, the second state value is calculated. This state value can comprehensively reflect the network state of the server during the evaluation period. By obtaining the data of multiple network nodes, the network performance and condition of the server can be comprehensively understood, providing a detailed data basis for subsequent state evaluation. The second state function calculates a comprehensive state value by comprehensively calculating the values of multiple network nodes, which can more accurately reflect the overall network state of the server rather than the performance of a single network metric. Through detailed calculation of network node data, network problems in the server operation can be accurately identified, such as traffic bottlenecks, severe packet loss, high latency, etc., helping the operation and maintenance personnel to optimize and adjust targeted. Dynamically obtaining and calculating the values of network nodes can timely reflect the changes in the network state, improving the stability and efficiency of network operation.
[0120] In a specific embodiment, the step of evaluating the risk value of the target server according to the first state value and the second state value includes:
[0121] S601. Obtain a risk function;
[0122] S602. Obtain a corresponding first weight coefficient according to the first state value;
[0123] S603. Obtain a corresponding second weight coefficient according to the second state value;
[0124] S604. Input the first status value, the first weight coefficient, the second status value, and the second weight coefficient into the risk function, and mark the output result as the risk value.
[0125] In the above steps S601 to S604, the risk function is a mathematical model or algorithm for comprehensively evaluating the risk value of the server. This function performs operations on the input multiple status values and weight coefficients and outputs a comprehensive value reflecting the running risk of the server. The risk function is F = Z×α + Y×β, where F represents the risk value, Z represents the first status value, α represents the first weight coefficient, Y represents the second status value, and β represents the second weight coefficient. According to the first status value (a comprehensive status value reflecting the system resource usage), the corresponding first weight coefficient is determined. This weight coefficient can be determined based on experience, historical data, or preset rules, reflecting the importance of system resources in the overall risk assessment. According to the second status value (a comprehensive status value reflecting the network status), the corresponding second weight coefficient is determined. This weight coefficient can also be determined based on experience, historical data, or preset rules, reflecting the importance of the network status in the overall risk assessment. Taking the first status value, the first weight coefficient, the second status value, and the second weight coefficient as inputs and substituting them into the risk function for calculation, the output result is the risk value. This risk value comprehensively reflects the running risk of the server during the evaluation period. Through the risk function, the system resource status and the network status are comprehensively considered to obtain a comprehensive risk assessment result. This comprehensive assessment method avoids the one-sidedness that may be brought by a single indicator and improves the accuracy of risk assessment. By setting the first weight coefficient and the second weight coefficient, the weights of system resources and network status in the overall risk assessment can be flexibly adjusted. This flexibility allows the risk assessment to be optimized according to the actual situation and experience. The comprehensive risk value provides a scientific and intuitive measurement standard for operation and maintenance personnel, helping them make more reasonable operation and maintenance decisions, take timely measures to handle high-risk servers, reduce the failure rate, and can timely reflect the changes in the running state of the server, quickly identify and respond to potential risks, and improve the running stability.
[0126] In a specific embodiment, the step of obtaining the corresponding control level according to the risk value of the target server includes:
[0127] S701. Obtain the operation and maintenance level table, where the operation and maintenance level table includes multiple risk assessment interval values and the corresponding control levels for each risk assessment interval value;
[0128] S702. Obtain the corresponding target risk assessment interval value according to the risk value of the target server;
[0129] S703. Obtain the corresponding control level from the operation and maintenance level table according to the target risk assessment interval value.
[0130] In the above steps S701 to S703, the operation and maintenance level table is a predefined table or list that includes multiple risk assessment interval values and the corresponding control levels for each interval value. This table can be formulated based on the company's policies, experience, or industry standards. According to the previously calculated target server risk value, determine the specific risk assessment interval value to which it belongs. This interval value can be divided according to the settings of the operation and maintenance level table. According to the determined target risk assessment interval value, search for the corresponding control level in the operation and maintenance level table. Usually, the operation and maintenance level table will define the specific control measures and priorities corresponding to each interval value. The operation and maintenance level table defines the mapping relationship between different risk assessment interval values and control levels, which can help the operation and maintenance team quickly determine appropriate control strategies and measures according to the risk value. According to the specific risk assessment interval value of the target server, the operation and maintenance personnel can accurately formulate targeted control plans, optimize resource allocation and operation and maintenance activities, improve efficiency, quickly determine the control level of the target server by quickly searching the operation and maintenance level table, so as to quickly respond to potential risks, take appropriate preventive and emergency measures, reduce system risks, the operation and maintenance level table provides a unified reference standard, which can help the management and operation and maintenance team better understand and manage the risk status of the server cluster, improve management efficiency and decision-making quality, and through refined risk assessment and control level hierarchical management, the operation and maintenance costs caused by system problems can be effectively reduced, and the overall service level and user satisfaction can be improved.
[0131] In a specific embodiment, the step of constructing the feedback cycle, obtaining the running status of each server within the feedback cycle after performing operation and maintenance control on the server cluster according to the control level, and marking it as the feedback running status includes:
[0132] S901. Obtain the time for performing operation and maintenance control on the server cluster according to the control level, and mark it as the start time of the feedback cycle;
[0133] S902. Obtain the acquisition frequency of the power data, and obtain the acquisition time interval according to the acquisition frequency;
[0134] S903. Obtain the number of times the power data of the target server exceeds the standard power threshold within the evaluation cycle according to the acquisition time interval, and mark it as the excess number of times;
[0135] S904. Obtain the feedback duration function;
[0136] S905. Input the excess number of times into the feedback duration function, and mark the output result as the feedback cycle duration;
[0137] S906. Obtain the end time of the feedback period based on the duration of the feedback period and the start time of the feedback period;
[0138] S907. Obtain the running status of each server within the start time and end time of the feedback period, and mark it as the feedback running status.
[0139] In the above steps S901 to S907, perform corresponding operation and maintenance control on the server cluster according to the previously determined control level, determine the time point when the control operation starts as the start time of the feedback period, determine the collection frequency of the server power data, for example, collect data once per minute or per hour, and then calculate the corresponding collection time interval. Within the feedback period, according to the set collection time interval, count the number of times the target server power data exceeds the preset standard power threshold. These times reflect the power abnormality of the server during the control period. The feedback duration function is , where represents the duration of the feedback period, represents the excess times, represents the standard times, represents the standard feedback duration. Take the excess times counted within the evaluation period as the input, and calculate the duration of the feedback period through the feedback duration function. This duration reflects the specific abnormality and its duration of the server during the control period. According to the determined duration and start time of the feedback period, calculate the end time of the feedback period. This end time marks the end point of the feedback period. Within the determined feedback period, obtain the running status of each server between the start time and the end time. These statuses can include various indicators such as system load, network status, service availability, etc. Constructing the feedback period can ensure continuous monitoring of the running status of the server, timely discover and handle potential problems. By counting specific indicators such as the excess times of power data, the abnormality of the server during the control period can be accurately located, providing a basis for the rapid location and solution of problems. The feedback period can help the operation and maintenance team understand the change of the running status of the server in real time, contribute to the rapid response and adjustment of management strategies. By analyzing the running status within the feedback period, the control strategy can be optimized and adjusted, improving the control efficiency and effectiveness, reducing unnecessary operation and maintenance costs and risks. The detailed data collected within the feedback period can be used as the basis for data-driven decision-making, helping the management and operation and maintenance teams make scientific and reasonable decisions based on the actual situation. Through timely feedback and monitoring, system failures and performance degradation can be effectively prevented and reduced, enhancing the overall stability and availability of the system.
[0140] Please refer to the appendix Figure 2 As shown in the figure, the present invention also provides a server cluster operation and maintenance control system for the above server cluster operation and maintenance control method, including:
[0141] A real-time acquisition module, which is used to acquire the power data of each server in the server cluster in real time;
[0142] A judgment module, which is used to judge whether the power data of each server meets the preset requirements. If not, it is marked as a target server;
[0143] An operating status module, which is used to construct an evaluation period, obtain the operating status of the target server during the evaluation period, and mark the target operating data;
[0144] A first status module, which is used to obtain multiple system resource usage data from the target operating data and obtain a first status value according to the multiple system resource usage data;
[0145] A second status module, which is used to obtain multiple network node data from the target operating data and obtain a second status value according to the multiple network node data;
[0146] A risk module, which is used to evaluate the risk value of the target server according to the first status value and the second status value;
[0147] A control level module, which is used to obtain the corresponding control level according to the risk value of the target server;
[0148] An operation and maintenance control module, which is used to perform operation and maintenance control on the server cluster according to the control level;
[0149] A feedback period module, which is used to construct a feedback period, obtain the operating status of each server during the feedback period after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback operating status;
[0150] A feedback module, which is used to return the feedback operating status as the current operating status to the step of judging whether the power data of each server meets the preset requirements. If not, it is marked as a target server.
[0151] As described above, the real-time data acquisition module regularly obtains the power data of each server at a set acquisition frequency to ensure the timeliness and accuracy of the data. The judgment module compares the power data of each server with a preset standard power threshold, and if it exceeds, marks it as a target server that requires further management and monitoring. The operating status module collects various operating data of the target server during the set evaluation period, including but not limited to system resource usage, network node status, etc., for subsequent risk assessment and determination of the control level. The first status module analyzes the system resource usage of the target server during the evaluation period, calculates data such as the system resource utilization rate and idle rate, and synthesizes them into a first status value, which is one of the important bases for subsequent risk assessment. The second status module collects the network node status data of the target server during the evaluation period, such as latency, transmission rate, etc., calculates and synthesizes them into a second status value for subsequent risk assessment and calculation of the control level. The risk module uses a preset risk assessment algorithm or model, combines the first status value and the second status value, calculates the risk level of the target server during the evaluation period, and determines its risk degree. The control level module determines the corresponding control level according to the calculated risk value based on a pre-set operation and maintenance level table, and identifies the control measures and priorities to be taken. The operation and maintenance control module performs corresponding management and maintenance operations according to the control levels of each server, including but not limited to resource allocation, performance optimization, fault handling, etc., to ensure the stable operation of the server cluster. The feedback cycle module collects the detailed operating status data of each server during the feedback cycle according to the start time of the control operation and the set feedback cycle duration for subsequent operation and maintenance decision-making and continuous optimization. The feedback module feeds back the latest operating status data collected during the feedback cycle to the judgment module to update the status of the target server and continue to monitor and control the servers that do not meet the requirements. It can collect data in real time and automate processing, improving the operation and maintenance efficiency and response speed. Through the collaborative action of multiple modules, it can more accurately identify and locate problems in the server cluster, prevent and solve potential risks in advance. The feedback module can help the system continuously learn and optimize, improve the management level and service quality, reduce the operation and maintenance cost, support data-driven operation and maintenance decision-making based on the large amount of data collected, improve the scientificity and accuracy of decision-making, and effectively improve the stability and reliability of the server cluster through a comprehensive control and feedback mechanism to ensure the continuous operation of the business.
[0152] And, a server cluster operation and maintenance control terminal, comprising:
[0153] One or more processors;
[0154] A storage device on which one or more programs are stored;
[0155] When the one or more programs are executed by the one or more processors, the one or more processors implement the server cluster operation and maintenance management method.
[0156] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. The structures, devices, and operation methods not specifically described and explained in the present invention are implemented according to the conventional means in the art without special explanation and limitation.
Claims
1. A method for operation and maintenance management and control of a server cluster, characterized in that, Including: Real-time collect the power data of each server in the server cluster; Judge whether the power data of each server meets the preset requirements. If not, mark it as the target server; Construct an evaluation period, obtain the running status of the target server during the evaluation period, and mark the target running data; Obtain multiple system resource usage data from the target running data, and obtain the first status value according to the multiple system resource usage data; Obtain multiple network node data from the target running data, and obtain the second status value according to the multiple network node data; Evaluate the risk value of the target server according to the first status value and the second status value; Obtain the corresponding control level according to the risk value of the target server; Perform operation and maintenance control on the server cluster according to the control level; Construct a feedback period, obtain the running status of each server during the feedback period after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback running status; Take the feedback running status as the current running status and return it to the step of judging whether the power data of each server meets the preset requirements. If not, mark it as the target server; The step of evaluating the risk value of the target server according to the first status value and the second status value includes: Obtain the risk function. The risk function is F = Z×α + Y×β, where F represents the risk value, Z represents the first status value, α represents the first weight coefficient, Y represents the second status value, and β represents the second weight coefficient; Obtain the corresponding first weight coefficient according to the first status value; Obtain the corresponding second weight coefficient according to the second status value; Input the first status value, the first weight coefficient, the second status value, and the second weight coefficient into the risk function, and mark the output result as the risk value.
2. The server cluster operation and maintenance management and control method according to claim 1, wherein, The step of judging whether the power data of each server meets the preset requirements. If not, mark it as the target server includes: Obtain the standard power threshold; Judge whether the power data of each server exceeds the standard power threshold; If the power data of the server exceeds the standard power threshold, it is determined that the power data of the server is abnormal, and the server with the power data exceeding the standard power threshold is marked as the target server If the power data of the server exceeds the standard power threshold, it is determined that the power data of the server is normal.
3. The server cluster operation and maintenance management and control method according to claim 1, wherein The step of constructing an evaluation period, obtaining the running status of the target server during the evaluation period, and marking the target running data includes: Obtain the time when the power data of the target server is determined to exceed the standard power threshold, and mark it as the start time of the evaluation period; Obtain the power data when the target server is determined to exceed the standard power threshold, and mark it as the target power value; Obtain the standard evaluation duration and the power weight coefficient; Obtain the evaluation duration function, input the target power value, the standard evaluation duration, and the power weight coefficient into the evaluation duration function, and mark the output result as the target evaluation duration; Obtain the end time of the evaluation period according to the target evaluation duration and the start time of the evaluation period; Obtain the running status of the target server during the start time and the end time of the evaluation period, and mark the target running data.
4. The server cluster operation and maintenance management and control method according to claim 1, wherein The step of obtaining multiple system resource usage data from the target operation data and obtaining a first status value according to the multiple system resource usage data includes: Obtain multiple system resource usage data from the target operation data; Obtain multiple system resource utilization rates according to the multiple system resource usage data; Obtain multiple system resource idle rates according to the multiple system resource usage data; Obtain a first status function; Input the multiple system resource utilization rates and the multiple system resource idle rates into the first status function, and mark the output result as the first status value.
5. The server cluster operation and maintenance management and control method according to claim 1, characterized in that The step of obtaining multiple network node data from the target operation data and obtaining a second status value according to the multiple network node data includes: Obtain multiple network node data from the target operation data; Obtain corresponding multiple network node values according to the multiple network node data; Obtain a second status function; Input the multiple network node values into the second status function, and mark the output result as the second status value.
6. The server cluster operation and maintenance management and control method according to claim 1, wherein, The step of obtaining a corresponding control level according to the risk value of the target server includes: Obtain an operation and maintenance level table, where the operation and maintenance level table includes multiple risk assessment interval values and the corresponding control levels for each risk assessment interval value; Obtain a corresponding target risk assessment interval value according to the risk value of the target server; Obtain the corresponding control level from the operation and maintenance level table according to the target risk assessment interval value.
7. The server cluster operation and maintenance management and control method according to claim 1, characterized in that The step of constructing a feedback period, obtaining the running status of each server within the feedback period after performing operation and maintenance control on the server cluster according to the control level, and marking it as the feedback running status includes: Obtain the time for performing operation and maintenance control on the server cluster according to the control level, and mark it as the start time of the feedback period; Obtain the acquisition frequency of the power data, and obtain the acquisition time interval according to the acquisition frequency; Obtain the number of times that the power data of the target server exceeds the standard power threshold within the evaluation period according to the acquisition time interval, and mark it as the excess number; Obtain a feedback duration function; Input the excess number into the feedback duration function, and mark the output result as the feedback period duration; Obtain the end time of the feedback period according to the feedback period duration and the start time of the feedback period; Obtain the running status of each server within the start time and end time of the feedback period, and mark it as the feedback running status.
8. A server cluster operation and maintenance management and control system, which is applied to the server cluster operation and maintenance management and control method described in any one of claims 1 to 7, and is characterized in that, Includes: A real-time acquisition module for real-time acquisition of the power data of each server in the server cluster; A judgment module for judging whether the power data of each server meets the preset requirements, and if not, marking it as the target server; A running status module for constructing an evaluation period, obtaining the running status of the target server within the evaluation period, and marking the target operation data; A first status module for obtaining multiple system resource usage data from the target operation data and obtaining a first status value according to the multiple system resource usage data; A second status module for obtaining multiple network node data from the target operation data and obtaining a second status value according to the multiple network node data; A risk module for evaluating the risk value of the target server according to the first status value and the second status value; A control level module for obtaining a corresponding control level according to the risk value of the target server; The operation and maintenance control module is used to perform operation and maintenance control on the server cluster according to the control level; The feedback cycle module is used to construct a feedback cycle, obtain the running status of each server within the feedback cycle after performing operation and maintenance control on the server cluster according to the control level, and mark it as the feedback running status; The feedback module is used to return the feedback running status as the current running status to the step of judging whether the power data of each server meets the preset requirements. If not, it is marked as the target server.
9. An operation and maintenance control terminal for a server cluster, characterized in that, It includes: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the server cluster operation and maintenance control method according to any one of claims 1-7.
Citation Information
Patent Citations
Intelligent monitoring operation and maintenance method and system based on data twinning
CN116070802A