Server repair methods and electronic devices
By acquiring hardware and software monitoring data, utilizing fault prediction models and knowledge graphs to locate the root cause of faults, and combining a self-healing strategy library and sandbox verification to execute repair actions, the problems of single server monitoring dimensions and insufficient communication coverage are solved. This enables automatic prediction and rapid repair of server faults, improving operational efficiency and reliability.
Patent Information
- Application Number
- CN202511261799.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Server monitoring methods suffer from problems such as limited monitoring dimensions, delayed fault response, and insufficient communication coverage, resulting in low operational efficiency, poor reliability, and impact on business continuity and stability.
By acquiring hardware and software monitoring data, fault prediction models and fault knowledge graphs are used to locate the root cause of the fault. The self-healing strategy library and sandbox verification are combined to execute repair actions, and dual-mode redundant communication is used to ensure the effectiveness and reliability of the repair.
It enables automatic prediction, precise location, and rapid repair of server failures, improving operational efficiency and reliability, reducing operational costs, and ensuring business continuity and stability.
Smart Images

Figure CN120743613B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more particularly to server repair methods and electronic devices. Background Technology
[0002] To remotely query and control the working status of the servers, all servers must first be connected to a local area network (LAN) using network cables, and then connected to the Internet. Maintenance personnel need to access the server ports on the Internet to achieve remote monitoring and management.
[0003] The relevant technology acquires server temperature and voltage, and sends an alarm to the maintenance terminal when the detected values are outside the safe range. Maintenance personnel can then remotely log in to the management and control device to understand the cause or location of the anomaly.
[0004] However, the server monitoring methods in related technologies have the following problems: (1) The monitoring dimensions are relatively simple, focusing only on a certain part of the hardware or software indicators, making it difficult to comprehensively and accurately reflect the actual operating status of the server; (2) The monitoring system is slow to respond, and after a fault occurs, manual intervention is often required for troubleshooting and repair, making it impossible to take effective self-healing measures in a timely manner, resulting in the server being in a fault state for a long time, affecting the normal operation of the business. In addition, in the scenario of remote server monitoring, the server is also constrained by a single communication method, among which the problem of insufficient communication coverage is more prominent, especially in some areas with unstable signals, where the reliability of data transmission is difficult to guarantee, which greatly reduces the effectiveness of remote monitoring and management, and maintenance personnel cannot obtain accurate server status information in a timely manner, thus making it impossible to make corresponding decisions and operations in a timely manner. Summary of the Invention
[0005] This invention provides a server repair method and electronic device to at least solve the problems of single monitoring dimensions, delayed fault response and insufficient communication coverage in server monitoring methods, significantly improve the efficiency and reliability of server operation and maintenance, reduce operation and maintenance costs, and ensure the continuity and stability of business.
[0006] This invention provides a method for repairing a server, comprising the following steps:
[0007] Acquire hardware monitoring data and software monitoring data;
[0008] The hardware monitoring data and the software monitoring data are input into a preset fault prediction model to obtain the fault time point. If the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset association conditions, the root cause event of the fault is located based on the preset fault knowledge graph and the hardware event and the software event. The preset fault prediction model is obtained by training a target neural network from historical monitoring data.
[0009] Based on the fault time point and the fault root cause event, a repair action is determined from a preset self-healing strategy library. The repair action is then sandbox-tested. If the repair action passes the sandbox test, an execution signal is generated based on the repair action, and the repair action is executed based on the execution signal. The feedback result after executing the repair action is sent to a preset terminal. This invention also provides a server repair device, comprising:
[0010] The acquisition module is used to acquire hardware monitoring data and software monitoring data.
[0011] The prediction module is used to input the hardware monitoring data and the software monitoring data into a preset fault prediction model to obtain the fault time point, and, if the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data satisfy the preset association conditions, locate the root cause event of the fault based on the hardware event and the software event according to the preset fault knowledge graph. The preset fault prediction model is obtained by training a target neural network from historical monitoring data.
[0012] The repair module is used to determine a repair action from a preset self-healing strategy library based on the fault time point and the fault root cause event, perform sandbox verification on the repair action, and generate an execution signal based on the repair action if the repair action passes the sandbox verification, and execute the repair action based on the execution signal, and send the feedback result after executing the repair action to a preset terminal.
[0013] The present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the server repair method as described in the above embodiments.
[0014] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described server repair methods.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described server repair methods.
[0016] Therefore, by acquiring hardware and software monitoring data, and inputting these data into a pre-defined fault prediction model to obtain the fault time point, and determining that the hardware events identified by the hardware monitoring data and the software events identified by the software monitoring data meet pre-defined correlation conditions, the root cause event of the fault is located based on a pre-defined fault knowledge graph. The pre-defined fault prediction model is obtained by training a target neural network using historical monitoring data. Based on the fault time point and the root cause event, a repair action is determined from a pre-defined self-healing strategy library. This repair action is then sandboxed for verification. If the repair action passes the sandbox verification, an execution signal is generated based on the action, and the repair action is executed. The feedback result after the repair action is executed is then sent to a pre-defined terminal. This solves the problems of single monitoring dimensions, delayed fault response, and insufficient communication coverage in traditional server monitoring methods, significantly improving the efficiency and reliability of server operation and maintenance, reducing operation and maintenance costs, and ensuring business continuity and stability. Attached Figure Description
[0017] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram illustrating the working principle of a server monitoring system provided in one embodiment of the present invention;
[0019] Figure 2 A schematic flowchart illustrating the server repair method provided in an embodiment of the present invention;
[0020] Figure 3 A schematic diagram illustrating the working principle of a fault knowledge graph provided in one embodiment of the present invention;
[0021] Figure 4 A schematic diagram illustrating the working principle of a sandbox verification unit provided in one embodiment of the present invention;
[0022] Figure 5 A schematic diagram illustrating the working principle of a dual-mode redundant communication unit provided in an embodiment of the present invention;
[0023] Figure 6 A block diagram of a server repair device provided in an embodiment of the present invention;
[0024] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0026] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Before describing the server repair method of this embodiment of the invention, let's first introduce a server monitoring system to which the server repair method of this embodiment of the invention is applied.
[0029] like Figure 1 As shown, Figure 1 This is a schematic diagram of the server monitoring system according to an embodiment of the present invention. The server monitoring system includes: a perception module, an analysis module, and an execution module.
[0030] The system comprises the following modules: a perception module for acquiring hardware and software monitoring data; an analysis module for inputting the hardware and software monitoring data into a preset fault prediction model to obtain the fault time point, and, based on a preset fault knowledge graph, locating the root cause event of the fault when the hardware events determined by the hardware monitoring data and the software events determined by the software monitoring data meet preset correlation conditions; a decision module for determining repair actions from a preset self-healing strategy library based on the fault time point and the root cause event, and generating an execution signal based on the repair actions when the repair actions pass sandbox verification; and an execution module for executing the repair actions based on the execution signal and sending the feedback results after executing the repair actions to a preset terminal.
[0031] Specifically, the perception module includes a hardware probe cluster deployed on the server motherboard and a software probe cluster injected into the operating system kernel; the analysis module is connected to the perception module and includes a fault knowledge graph unit and a lightweight AI prediction unit; the decision module is connected to the analysis module and includes a self-healing strategy library and a sandbox verification module; the execution module is connected to the decision module and includes a dual-mode redundant communication module and an automatic control interface. In this embodiment of the invention, the perception module synchronously collects hardware and software indicators; the analysis module matches the fault knowledge graph and runs the AI prediction model; the decision module calls the self-healing strategy library and verifies the repair plan in the sandbox; the execution module drives the control interface to execute repair actions and feeds back the results to the maintenance terminal via the dual-mode communication channel.
[0032] Therefore, this invention enables automatic fault prediction, precise fault location, and rapid fault repair, significantly improving system reliability and stability. By predicting the fault's timing in advance, preventative measures can be taken to minimize its impact on system operation. Furthermore, the use of fault knowledge graphs and sandbox verification ensures the accuracy and safety of repair actions, preventing secondary faults caused by erroneous repairs. Finally, the execution module feeds back the repair results to the terminal, facilitating monitoring and analysis by maintenance personnel to further optimize system operation.
[0033] Specifically, Figure 2 This is a flowchart illustrating a server repair method provided in an embodiment of the present invention.
[0034] like Figure 2 As shown, the data transmission method includes the following steps:
[0035] In step S101, hardware monitoring data and software monitoring data are acquired.
[0036] Hardware monitoring data refers to the indicators collected from physical components that reflect the operation of those components, while software monitoring data refers to the indicators collected from the software layer that reflect the software's performance.
[0037] Hardware monitoring data includes at least one of temperature, voltage, and fan speed, while software monitoring data includes at least one of software process status, application log keywords, and network traffic metrics.
[0038] Specifically, embodiments of the present invention can employ a server monitoring system. This system uses a perception module to acquire hardware and software monitoring data. Specifically, the perception module includes a first probe unit and a second probe unit. The first probe unit includes a first probe array deployed on multiple target hardware devices and configured to acquire hardware monitoring data from these devices. The second probe unit includes a second probe array and a policy engine. The policy engine is communicatively coupled to the second probe array and configured to load the second probe array to acquire software monitoring data based on the software process state. Furthermore, the second probe unit is deployed in a containerized form on the server node.
[0039] In actual execution, the first probe unit, through the first probe array deployed on multiple target hardware devices, can acquire real-time monitoring data of the hardware devices, providing a foundation for monitoring the hardware status. The second probe unit combines the second probe array and the policy engine. The policy engine dynamically loads the second probe array according to the status of the software process, thereby acquiring software monitoring data. This dynamic loading method can flexibly adjust the monitoring policy according to the actual running situation, improving monitoring efficiency. In addition, the second probe unit is deployed on the server node in a containerized form. Taking advantage of the advantages of containerization technology, it is easy to manage and expand, and can quickly adapt to different server environments and software requirements.
[0040] For example, the first probe unit in this embodiment of the invention includes: a temperature sensing unit: distributedly mounted at key locations in the CPU, GPU, and power supply circuits; a voltage detection unit: connected to the VCC pin of the motherboard chip, with a sampling frequency of 1kHz; and a speed feedback unit: capturing fan speed via a PWM signal. The first probe unit can collect hardware monitoring data such as temperature, voltage, and fan speed in real time.
[0041] The second probe unit is injected into the operating system kernel in a containerized form. It can dynamically collect key software monitoring data such as process status, application log keywords, and network traffic indicators. Furthermore, it can dynamically load the second probe array through the policy engine. For example, when a database service is detected to be running, the SQL slow query analysis probe is automatically loaded to conduct in-depth analysis of the database's operating efficiency and potential problems. When a GPU-intensive task is detected, the GPU memory leak monitoring probe is automatically loaded to promptly detect and prevent failures caused by GPU memory leaks.
[0042] Therefore, the embodiments of the present invention can acquire hardware and software monitoring data comprehensively and efficiently, providing more comprehensive and accurate data support for subsequent fault prediction and diagnosis, and improving the overall monitoring capability and reliability of the system.
[0043] In step S102, hardware monitoring data and software monitoring data are input into a preset fault prediction model to obtain the fault time point. If the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset correlation conditions, the root cause event of the fault is located based on the preset fault knowledge graph.
[0044] The preset fault prediction model is obtained by training the target neural network with historical monitoring data. The preset correlation can be preset by relevant personnel, and the preset fault knowledge graph can be preset by relevant personnel.
[0045] Specifically, this application inputs hardware monitoring data and software monitoring data into a preset fault prediction model. The preset fault prediction model extracts time-series features from the two types of data and fuses them to output the predicted fault time point. It also matches the hardware events determined by the hardware monitoring data and the software events determined by the software monitoring data according to preset association conditions. When the hardware events determined by the hardware monitoring data and the software events determined by the software monitoring data meet the preset association conditions, the root cause event of the fault is located based on the preset fault knowledge graph. For example, a graph traversal and root cause scoring algorithm is executed starting from the hardware events and software events, and the node with the highest score is finally located as the root cause event of the fault.
[0046] Therefore, the embodiments of this application can predict failures before they occur, and improve accuracy by using knowledge graphs to associate hardware events and software events.
[0047] According to one embodiment of the present invention, inputting hardware monitoring data and software monitoring data into a preset fault prediction model to obtain a fault time point includes: inputting hardware monitoring data and software monitoring data into a preset fault prediction model to obtain a fault time point and a prediction confidence level; and outputting the fault time point when the prediction confidence level is higher than a confidence level threshold.
[0048] The prediction confidence level is the probability or score given synchronously by the preset fault prediction model at the time of output fault, which is used to characterize the credibility of the prediction result. The confidence level threshold can be preset by the user, obtained through a limited number of experiments, or obtained through a limited number of computer simulations, and is not specifically limited here.
[0049] Specifically, in this embodiment of the invention, the model receives real-time collected hardware monitoring data and software monitoring data, first extracts and fuses time-series features in parallel, and then outputs the fault time point and the predicted confidence level of the fault time point. The system immediately compares the confidence level with a pre-calibrated confidence level threshold. If it is higher than the threshold, the fault time point is pushed to the subsequent root cause localization module or alarm platform. Otherwise, it is regarded as a low-confidence prediction and is silently discarded or downgraded, thereby avoiding false alarms.
[0050] Therefore, by introducing confidence threshold filtering, the system can significantly reduce false alarms and improve the reliability of prediction while ensuring the sensitivity of early warning.
[0051] According to one embodiment of the present invention, before determining that the hardware event determined by hardware monitoring data and the software event determined by software monitoring data satisfy a preset association condition, the method includes: assigning a first weight to the hardware event and assigning a second weight to the software event; determining whether the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than a first preset threshold; if the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than the first preset threshold, then determining that the hardware event determined by hardware monitoring data and the software event determined by software monitoring data satisfy the preset association condition.
[0052] Specifically, in this embodiment of the invention, weights can be assigned to hardware events and software events respectively, and weighted calculations can be performed. If the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than a first preset threshold, then the corresponding hardware event and software event are considered to satisfy a preset correlation relationship, and root cause analysis can be performed. Otherwise, it is regarded as a low-confidence prediction and is silently discarded or downgraded, thereby avoiding false alarms.
[0053] For example, such as Figure 3 As shown, embodiments of the present invention can establish association rules between hardware events and software events. By defining event association rules, such as triggering root cause analysis when "temperature anomaly event weight × 0.7 + process CPU utilization event weight × 0.3 > 0.8", events of different dimensions are associated to achieve comprehensive analysis of server failures, uncover potential root causes of failures, and store historical failure paths, such as hard disk failure prediction data - migration records - post-migration health indicator change curves, providing rich empirical data support for subsequent failure handling and prevention, and helping to optimize failure handling strategies and take preventive measures in advance.
[0054] Therefore, the embodiments of the present invention ensure that the importance of hardware and software events is reasonably assessed, avoid misjudgments caused by unclear event weights, and further narrow the scope of fault diagnosis by defining the relationship between events, thereby reducing unnecessary troubleshooting work.
[0055] According to one embodiment of the present invention, before inputting hardware monitoring data and software monitoring data into a preset fault prediction model, the method further includes: acquiring historical monitoring data and constructing a dataset based on the historical monitoring data; dividing the dataset into a training set, a validation set, and a test set based on a preset partitioning ratio; constructing a target neural network, inputting the training set into the target neural network for training to obtain initial model parameters; based on the initial model parameters, inputting the validation set into the target neural network for performance evaluation, and adjusting the initial model parameters according to the performance evaluation results until the joint loss function of the validation set converges to obtain optimal model parameters; based on the optimal model parameters, inputting the test set into the target neural network for model testing, and obtaining a preset fault prediction model when the test results meet preset requirements.
[0056] The target neural network can be an LSTM (Long Short-Term Memory) network. Historical monitoring data can include historical hardware monitoring data and historical software monitoring data, as well as corresponding indicators. The preset division ratio and preset requirements can be set by the user, obtained through a limited number of experiments, or obtained through a limited number of computer simulations. No specific limitations are made here.
[0057] Specifically, embodiments of the present invention can collect historical monitoring data, clean and label it to form a dataset, and split the dataset into a training set, a validation set, and a test set according to a preset ratio (e.g., 7:1.5:1.5). A target neural network (e.g., LSTM) suitable for the temporal characteristics is selected or designed. Initial training is performed using the training set to obtain initial model parameters. The joint loss function is calculated using the validation set until the validation set loss converges, locking in the optimal model parameters. Finally, the test set is input into the network. If the test results meet the test requirements, a preset fault prediction model is generated. Furthermore, after obtaining the preset fault prediction model, embodiments of the present invention can also perform lightweight processing on the model to obtain a lightweight AI prediction model.
[0058] Therefore, through the above-mentioned "training-verification-testing" closed loop, it can be ensured that the final fault prediction model has high accuracy, high robustness and good early warning capability.
[0059] In actual implementation, embodiments of the present invention can predict the time point of failure through an analysis module. The analysis module may include a prediction unit. The function of the prediction unit is to input the collected hardware monitoring data and software monitoring data into a pre-trained fault prediction model. This model is trained through complex algorithms and a large amount of historical data, and can predict the time point when a fault may occur and provide the confidence level of the prediction result. If the prediction confidence level is higher than the preset confidence level threshold, it indicates that the prediction result has high reliability. At this time, the prediction unit will output the time point of failure. This function enables the system not only to locate the root cause of the fault, but also to predict the time of occurrence of the fault in advance, providing an important basis for preventive maintenance.
[0060] In actual implementation, the LSTM algorithm is used to predict the time point of failure. The LSTM model is trained based on time series data. The input is a time series window of temperature, CPU utilization and network packet loss rate. After deep learning and analysis of the preset failure prediction model, the failure prediction result is output. When the prediction confidence exceeds 85%, an early warning is triggered 15 minutes in advance, so that the system can prepare in advance and reduce the impact of failure on server operation.
[0061] Therefore, by predicting the failure time, the system can take proactive measures, such as scheduling maintenance in advance, adjusting resource allocation, or activating backup systems, thereby reducing the impact of failures on system operation. The introduction of prediction confidence further improves the reliability of the prediction results; the failure time is only output when the confidence level exceeds a threshold, avoiding unnecessary maintenance or resource waste due to misjudgment. This combination of early warning and reliability assessment not only improves the system's reliability and stability but also reduces operational costs and enhances overall system performance and user experience.
[0062] In step S103, a repair action is determined from a preset self-healing strategy library based on the fault time point and the fault root cause event. The repair action is then sandbox verified. If the repair action passes the sandbox verification, an execution signal is generated based on the repair action. The repair action is then executed based on the execution signal, and the feedback result after executing the repair action is sent to a preset terminal.
[0063] Furthermore, according to one embodiment of the present invention, sandbox verification of the repair action includes: simulating the target server environment in a preset simulation container, performing the repair action on the target server to obtain the sandbox verification result.
[0064] The pre-defined self-healing strategy library includes detailed automated actions for various common fault events. For example: for memory leak events, restart the process and generate a heap dump file; for network congestion events, switch to a backup link and adjust QoS priority; for high-temperature alarm events, migrate computing tasks to idle nodes and increase fan speed to the maximum value. Sandbox verification refers to simulating the execution, evaluating the effect, and scanning the risks of the repair actions matched based on the root cause of the fault in an isolated virtual environment. If the simulation results meet the requirements, the repair action is deemed to have passed sandbox verification.
[0065] Specifically, based on the fault time point and root cause event, a repair action is determined from a preset self-healing strategy library. To prevent the repair action from adversely affecting the server, this embodiment of the invention requires sandbox verification of the repair action after it is determined. This embodiment of the invention can set up a sandbox verification unit in the decision module, wherein the sandbox verification unit includes: a receiving subunit, a simulation subunit, and a verification subunit. The receiving subunit is used to receive repair instructions generated based on the repair action; the simulation subunit is used to simulate the target server environment in a preset simulation container according to the repair instructions; the verification subunit is used to execute the repair action on the preset simulation container to obtain verification results.
[0066] After receiving the repair instruction, the receiving subunit simulates the target server's operating environment in a preset simulation container. This simulation environment is similar to the actual server environment but is completely isolated, thus allowing for safe testing of repair actions. The subunit verifies the execution of repair actions in the simulation environment, evaluates their effectiveness, and generates verification results. If the verification results indicate that the repair actions are safe and effective, the decision module will generate an execution signal to trigger the execution module to perform the repair actions. This process ensures that the repair actions have undergone rigorous testing and verification before actual application, reducing the risk of secondary failures caused by erroneous repairs.
[0067] In actual execution, the sandbox verification process is as follows: receive the repair instructions sent by the decision module; simulate the target server environment in a Docker container; execute the repair instructions and verify the validity of the results; if the verification is successful, send an execution signal to the automatic control interface.
[0068] Therefore, by testing repair actions in a simulated environment, the system can verify the effectiveness and safety of these actions without affecting the production environment. The receiving subunit ensures accurate reception of repair commands, the simulation subunit provides a highly similar testing environment, and the verification subunit rigorously evaluates the repair actions. This verification mechanism not only reduces the risk of secondary failures caused by erroneous repairs but also improves the overall stability and reliability of the system. Furthermore, the sandbox verification unit makes the testing of repair actions more efficient and flexible, enabling rapid adaptation to different failure scenarios and repair requirements, further enhancing the system's self-healing capabilities.
[0069] According to one embodiment of the present invention, sending the feedback result after performing the repair action to a preset terminal includes: acquiring the signal strength and the duration of the signal strength; sending the feedback result to the preset terminal using a first communication channel when the signal strength is higher than a first strength threshold; and sending the feedback result to the preset terminal using a second communication channel when the signal strength is lower than a second strength threshold and the duration of the signal strength is higher than a preset time threshold.
[0070] Signal strength refers to the strength of the communication signal, usually expressed in units such as decibels (dB); the duration of signal strength refers to the duration during which the signal strength is above or below a certain threshold; the first strength threshold, the second strength threshold, and the preset time threshold can be preset by the user, obtained through a limited number of experiments, or obtained through a limited number of computer simulations, and are not specifically limited here; the preset terminal can be a mobile phone, computer, etc.; the first communication channel can be 5G, and the second communication channel can be a LoRa self-organizing network.
[0071] Specifically, embodiments of the present invention can monitor the strength and duration of communication signals. When the signal strength is higher than a first strength threshold, the feedback result is sent to a preset terminal using a 5G channel. When the signal strength is lower than a second strength threshold and the duration of the signal strength is higher than a preset time threshold, the feedback result is sent to the preset terminal using a LoRa self-organizing network.
[0072] Furthermore, embodiments of the present invention can automatically switch communication channels based on signal strength and signal duration. Specifically, by monitoring the strength and duration of the communication signal, it is determined whether the current communication channel meets preset signal switching conditions. If the signal strength is below a certain threshold, or the signal duration is below a certain threshold, it indicates that the current communication channel may be unstable or unreliable. In this case, the dual-mode redundant communication unit will automatically switch to the backup communication channel to ensure the continuity and stability of communication. This mechanism is particularly important in complex network environments, effectively preventing communication interruptions caused by signal problems and ensuring the smooth execution of repair actions and the timely transmission of feedback results.
[0073] For example, a dual-mode redundant communication unit is set in the execution module. The dual-mode redundant communication unit integrates the 5G slice network and the LoRa self-organizing network, and automatically switches channels according to the signal strength threshold. The switching logic of the dual-mode redundant communication module is as follows: when the 5G signal strength is below 30dBm for 5 seconds, the LoRa channel is enabled to transmit data; when the 5G signal strength recovers to above 40dBm, the 5G channel is automatically switched back.
[0074] Therefore, the embodiments of the present invention can ensure that the transmission of the execution signal and feedback result of the repair action is not affected, further enhancing the system's self-healing capability and operation and maintenance efficiency.
[0075] According to one embodiment of the present invention, it is determined whether to receive an extended script corresponding to a new repair strategy; if an extended script corresponding to a new repair strategy is received, the extended script is sandbox verified, and if the extended script passes the sandbox verification, the preset self-healing strategy library is updated using the extended script.
[0076] Specifically, in this embodiment of the invention, a script extension interface can be set in the server monitoring system. The script extension interface is used to obtain the extension script corresponding to the repair strategy and to perform sandbox verification on the extension script. If the extension script passes the sandbox verification, the preset self-healing strategy library is updated using the extension script.
[0077] It should be noted that the main function of the script extension interface is to obtain the extension scripts corresponding to the repair strategies and perform sandbox verification on these scripts. If the extension scripts pass the sandbox verification, proving their safety and effectiveness, the decision module can use these scripts to update the preset self-healing strategy library. This mechanism allows the system to dynamically extend and update repair strategies based on new fault scenarios or user needs, improving the system's flexibility and adaptability. Through sandbox verification, the system can ensure that newly introduced repair strategies will not negatively impact the existing system, thereby guaranteeing the system's stability and reliability.
[0078] In actual implementation, this embodiment of the invention extends the interface through a script: it provides an open Python API for users to write custom repair strategies, which are then compiled and executed through a sandbox verification module.
[0079] Therefore, by allowing users or the system to provide extended scripts, the system can quickly adapt to new fault scenarios and requirements, and continuously update and optimize the self-healing strategy library.
[0080] According to one embodiment of the present invention, after locating the root cause event of a fault based on a preset fault knowledge graph and according to hardware events and software events, the method includes: generating a fault reminder instruction based on the fault time point and the root cause event of the fault, and performing acoustic fault reminder and / or optical fault reminder according to the fault reminder instruction.
[0081] Specifically, in this embodiment of the invention, a fault reminder module can be set up in the server monitoring system to generate fault reminder instructions based on the fault time point and the fault root cause event, and to provide fault reminders based on the fault reminder instructions. For example, a buzzer can be driven by a square wave with different frequencies / duty cycles according to the fault level, or the fault event can be encoded and mapped to the flashing mode of RGB LEDs (such as red fast flashing = severe, yellow slow flashing = normal), while text and icons are scrolling on the large screen.
[0082] Therefore, through acoustic and optical alerts, maintenance personnel can immediately learn about faulty equipment, fault levels, and handling prompts without constantly monitoring the screen, significantly reducing the time required for manual discovery.
[0083] In addition, embodiments of the present invention can also set up a linkage management unit in the server monitoring system. The linkage management module is communicatively connected to the perception module, analysis module, decision-making module and execution module, respectively, and provides a visual topology mapping and a policy customization interface. The visual topology mapping displays the real-time load distribution of the server cluster in the form of a heat map, supports node drill-down to view historical self-healing records, and the policy customization interface allows users to customize monitoring and self-healing strategies according to their own needs and scenarios.
[0084] Therefore, the introduction of the linkage management unit significantly enhances the management and monitoring capabilities of the server monitoring system. Through visual topology mapping, users can intuitively understand the real-time load distribution of the server cluster and promptly identify potential performance bottlenecks. The node drill-down function further provides detailed self-healing records, helping users to deeply analyze the historical issues and handling status of each node. The policy customization interface gives users greater autonomy, enabling them to flexibly adjust monitoring and self-healing strategies according to their own needs and scenarios.
[0085] In summary, the embodiments of the present invention have the following beneficial effects:
[0086] (1) Comprehensive and accurate monitoring: This embodiment of the invention can comprehensively and in real time collect the hardware and software indicators of the server, realizing all-round monitoring of the server's operating status. The first probe unit is precisely deployed and can obtain important data such as temperature, voltage, and rotation speed of key hardware components; the second probe unit is deployed in a containerized form and dynamically loads probe modules, which can flexibly adapt to different server service types and dynamically collect software-related information such as process status, application log keywords, and network traffic indicators. This multi-dimensional monitoring method combining hardware and software can more comprehensively and accurately reflect the actual operating status of the server compared with traditional single-dimensional monitoring, effectively avoiding missed and false alarms of faults.
[0087] (2) The system integrates the fault knowledge graph module and lightweight AI prediction module of the analysis unit, as well as the self-healing strategy library and sandbox verification module of the decision-making unit, enabling intelligent diagnosis, early prediction, and automatic repair of server faults. The fault knowledge graph module quickly establishes the association between hardware events and software events through preset event association rules, accurately locating the root cause of the fault; the lightweight AI prediction module, based on the LSTM algorithm, can predict the fault time point in advance, providing early warning for fault handling. The decision-making unit calls the corresponding automatic repair strategy from the self-healing strategy library according to the analysis results, and ensures the effectiveness of the repair strategy through the sandbox verification module. Then, the execution unit drives the control interface to execute the repair action, realizing the automatic repair of server faults. This intelligent linkage self-healing management process effectively reduces the downtime of server faults and the intervention cost of operation and maintenance personnel, significantly improving the efficiency and reliability of server operation and maintenance, and has a huge advantage compared with the traditional passive fault handling method.
[0088] To enable those skilled in the art to further understand the server repair method of the present invention, the following detailed description is provided in conjunction with specific embodiments.
[0089] Taking a server cluster in a medium-sized data center as an example: First, a probe unit is carefully deployed on the motherboard of each server. The temperature sensing unit is closely attached to the key heat-generating parts of the CPU, GPU, and power supply circuit. These parts are the core heat sources in the operation of the server hardware. The temperature changes are monitored in real time by a high-precision temperature sensor. The voltage detection unit is accurately connected to the VCC pin of the motherboard chip and continuously acquires voltage data at a high-frequency sampling frequency of 1kHz to ensure that subtle voltage fluctuations can be captured in time, providing a strong guarantee for the stable operation of the hardware. The speed feedback unit is connected to each fan of the server and accurately acquires the fan speed information through PWM signal to monitor the working status of the heat dissipation system in real time, so that the fan speed can be quickly adjusted for heat dissipation when abnormal temperature occurs.
[0090] The second probe unit is injected into the operating system kernel in a containerized manner. Leveraging its lightweight and highly isolated characteristics, it ensures that the probe's operation does not significantly impact the server's normal business operations. Simultaneously, a policy engine dynamically manages the software probe cluster, monitoring various services and task types running on the server in real time. For example, when a database service is detected to be starting and running, the policy engine automatically triggers the loading of an SQL slow query analysis probe. This probe delves into the database's operation process, monitoring the execution efficiency of SQL queries in real time, promptly identifying slow query issues, and providing data support for database performance optimization. Conversely, when the server performs GPU-intensive tasks, such as deep learning training or graphics rendering, the policy engine automatically loads a GPU memory leak monitoring probe, closely monitoring GPU memory usage, promptly identifying and warning of GPU memory leaks, and preventing GPU failures or task interruptions due to insufficient GPU memory.
[0091] The analysis module includes a fault knowledge graph unit, the principle of which is illustrated in the diagram below. Figure 3 As shown, during the system initialization phase, association rules between hardware and software events are constructed. Through learning and analysis of a large amount of historical fault data and normal operation data, association rules such as "temperature anomaly event weight × 0.7 + process CPU utilization event weight × 0.3 > 0.8" are determined. This associates abnormal hardware temperature conditions with excessively high CPU utilization in software processes. When this condition is met, a root cause analysis process is triggered to investigate whether a hardware failure is causing the software process abnormality or a software process problem is causing the hardware temperature rise, thereby quickly locating the root cause of the fault. Simultaneously, the fault knowledge graph module stores rich historical fault path information, such as hard drive fault prediction → data migration records → post-migration health indicator change curves. When a potential hard drive failure is predicted, the system records the data migration process and related data according to the pre-stored path and continuously tracks the changes in the hard drive's health indicators after migration, providing strong support for subsequent hard drive fault handling and storage strategy optimization.
[0092] The analysis module also includes a prediction unit. This unit trains and predicts based on the LSTM algorithm, collecting time-series data such as server temperature, CPU utilization, and network packet loss rate during long-term server operation. This data is divided into multiple time windows and used as input to the model for thorough training. The trained model can predict potential server failure times within a given period based on real-time input data. When the prediction confidence exceeds 85%, a warning signal is sent to the decision-making level 15 minutes in advance, allowing the system to prepare in advance, such as pre-allocating backup resources and adjusting task scheduling, thus reducing the impact of failures on server operation and business.
[0093] The decision-making module's self-healing strategy library has defined detailed automated actions for various common fault events. For example, when a memory leak event occurs, the system automatically performs a process restart and generates a heap dump file, facilitating in-depth analysis and remediation of the memory leak by operations and maintenance personnel. In the event of network congestion, the system quickly switches to a backup link and adjusts QoS priorities to ensure that the transmission of critical business data is not affected by congestion, guaranteeing smooth network communication. When faced with a high-temperature alarm event, the system immediately migrates computing tasks to idle nodes and increases fan speed to maximum to promptly reduce server temperature and protect hardware from overheating damage.
[0094] like Figure 4 As shown, after receiving the repair command from the decision module 300, the sandbox verification unit quickly simulates a running environment highly consistent with the target server in a Docker container. In this simulated environment, it strictly follows the repair command to execute corresponding operations and monitors changes in various indicators in the environment in real time to verify the effectiveness of the repair results. Only after successful verification will it send an execution signal to the automatic control interface, ensuring that the repair operation will not adversely affect the actual server and guaranteeing its stable operation.
[0095] like Figure 5 As shown, the dual-mode redundant communication unit of the execution module monitors the 5G signal strength in real time. When the 5G signal strength remains below 30dBm for 5 seconds, the LoRa channel is automatically activated for data transmission. Leveraging its advantages in low power consumption and long-distance transmission, the LoRa self-organizing network ensures that data can still be stably and reliably transmitted to the monitoring center even in areas with poor 5G signal. When the 5G signal strength recovers to above 40dBm, the system automatically switches back to the 5G channel, fully utilizing the high speed and low latency characteristics of the 5G network to ensure rapid data transmission and timely processing.
[0096] The dynamic topology mapping unit of the linkage management unit displays the load distribution of the entire server cluster in real time in the form of an intuitive heatmap. Operations personnel can use this heatmap to clearly understand the load of each server node, promptly identify nodes with excessive load, and proactively adjust tasks and optimize resource allocation. Simultaneously, it supports drill-down operations on nodes to view their historical self-healing records, including detailed information such as fault type, occurrence time, self-healing process, and results. This provides operations personnel with rich historical data for reference, helping to summarize lessons learned and further optimize operational strategies.
[0097] The script extension interface provides powerful customization capabilities for operations and maintenance personnel and developers. Through the open Python API, users can write personalized remediation strategies and management scripts according to their own business needs and specific scenarios. After writing, the scripts are compiled and verified using the sandbox verification module to ensure their correctness and security. Then, they are integrated into the system to further enrich the system's functionality, improve its adaptability and flexibility, and meet the diverse needs of different users in complex and ever-changing server operation and maintenance environments.
[0098] Therefore, this invention enables comprehensive and real-time monitoring of servers, provides early warnings of faults using fault knowledge graphs and lightweight AI prediction technology, achieves safe and reliable automatic repair based on a self-healing strategy library and sandbox verification module, and uses 5G+LoRa dual-mode communication to ensure the stability of remote monitoring and management. It effectively solves the problems of single monitoring dimensions, delayed response and insufficient communication coverage in existing technologies, significantly improves the efficiency and reliability of server operation and maintenance, reduces operation and maintenance costs, and ensures the continuity and stability of business.
[0099] According to the server repair method proposed in this embodiment of the invention, hardware monitoring data and software monitoring data are acquired, and the hardware monitoring data and software monitoring data are input into a preset fault prediction model to obtain the fault time point. If the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data satisfy preset association conditions, the root cause event of the fault is located based on a preset fault knowledge graph. The preset fault prediction model is obtained by training a target neural network from historical monitoring data. Repair actions are determined from a preset self-healing strategy library based on the fault time point and the root cause event. The repair actions are then sandboxed for verification. If the repair actions pass the sandbox verification, an execution signal is generated based on the repair actions, and the repair actions are executed based on the execution signal. The feedback result after executing the repair actions is sent to a preset terminal. This solves the problems of single monitoring dimensions, delayed fault response, and insufficient communication coverage in traditional server monitoring methods, significantly improving the efficiency and reliability of server operation and maintenance, reducing operation and maintenance costs, and ensuring the continuity and stability of business operations.
[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0101] Embodiments of the present invention also provide a server repair device.
[0102] Figure 6 This is a block diagram illustrating a server repair device according to an embodiment of the present invention.
[0103] like Figure 6 As shown, the server repair device 10 includes: an acquisition module 100, a prediction module 200, and a repair module 300.
[0104] The acquisition module 100 is used to acquire hardware monitoring data and software monitoring data.
[0105] The prediction module 200 is used to input hardware monitoring data and software monitoring data into a preset fault prediction model to obtain the fault time point, and, if the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset correlation conditions, locate the root cause event of the fault based on the preset fault knowledge graph, wherein the preset fault prediction model is obtained by training the target neural network from historical monitoring data.
[0106] The repair module 300 is used to determine the repair action from the preset self-healing strategy library based on the fault time point and the fault root cause event, perform sandbox verification on the repair action, and generate an execution signal based on the repair action if the repair action passes the sandbox verification, and execute the repair action based on the execution signal, and send the feedback result after the execution of the repair action to the preset terminal.
[0107] According to one embodiment of the present invention, before determining that the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset correlation conditions, the prediction module 200 includes: a judgment unit and a determination unit.
[0108] The judgment unit is used to assign a first weight to the hardware event and a second weight to the software event, and to determine whether the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than a first preset threshold.
[0109] The determination unit is configured to determine that the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data satisfy a preset association condition if the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than a first preset threshold.
[0110] According to one embodiment of the present invention, the prediction module 200 includes an output unit.
[0111] The output unit is used to input hardware monitoring data and software monitoring data into a preset fault prediction model to obtain the fault time point and prediction confidence. If the prediction confidence is higher than the confidence threshold, the fault time point is output.
[0112] According to one embodiment of the present invention, before inputting hardware monitoring data and software monitoring data into a preset fault prediction model, the prediction module 200 further includes: a first acquisition unit, a division unit, a construction unit, an adjustment unit, and a generation unit.
[0113] The first acquisition unit is used to acquire historical monitoring data and construct a dataset based on the historical monitoring data.
[0114] The partitioning unit is used to divide the dataset into training, validation, and test sets based on a preset partitioning ratio.
[0115] The building unit is used to construct the target neural network. It is obtained by inputting the training set into the target neural network for training and obtaining the initial model parameters.
[0116] The adjustment unit is used to input the validation set into the target neural network for performance evaluation based on the initial model parameters, and adjust the initial model parameters according to the performance evaluation results until the joint loss function of the validation set converges to obtain the optimal model parameters.
[0117] The generation unit is used to input the test set into the target neural network based on the optimal model parameters to test the model, and obtain the preset fault prediction model when the test results meet the preset requirements.
[0118] According to one embodiment of the present invention, the repair module 300 includes: a verification unit.
[0119] The verification unit is used to simulate the target server environment in a preset simulation container, perform repair actions on the target server, and obtain sandbox verification results.
[0120] According to one embodiment of the present invention, the repair module 300 includes: a second acquisition unit, a first transmission unit, and a second transmission unit.
[0121] The second acquisition unit is used to acquire the signal strength and the duration of the signal strength.
[0122] The first transmitting unit is used to send the feedback result to the preset terminal through the first communication channel when the signal strength is higher than the first strength threshold.
[0123] The second transmitting unit is used to send the feedback result to the preset terminal through the second communication channel when the signal strength is lower than the second strength threshold and the duration of the signal strength is higher than the preset time threshold.
[0124] According to one embodiment of the present invention, the server repair device 10 described above further includes: a second judgment unit and an expansion unit.
[0125] The second judgment unit is used to determine whether to receive the extended script corresponding to the new repair strategy.
[0126] The extension unit is used to perform sandbox verification on the extension script corresponding to the new repair strategy when it receives the extension script, and to update the preset self-healing strategy library using the extension script if the extension script passes the sandbox verification.
[0127] According to one embodiment of the present invention, the hardware monitoring data includes at least one of temperature, voltage and fan speed, and the software monitoring data includes at least one of software process status, application log keywords and network traffic indicators.
[0128] According to one embodiment of the present invention, after locating the root cause event of a fault based on a preset fault knowledge graph and hardware events and software events, the prediction module 200 includes: an alert unit.
[0129] The reminder unit is used to generate a fault reminder command based on the fault time point and the fault root cause event, and to provide acoustic fault reminders and / or optical fault reminders based on the fault reminder command.
[0130] In summary, the description of the features in the embodiments corresponding to the server repair device can be found in the relevant descriptions of the embodiments corresponding to the server repair method, and will not be repeated here.
[0131] Embodiments of the present invention also provide an electronic device, which may include:
[0132] The memory 701, the processor 702, and the computer program stored on the memory 701 and executable on the processor 702.
[0133] When processor 702 executes the program, it implements the server repair method provided in the above embodiments.
[0134] Furthermore, electronic devices also include:
[0135] Communication interface 703 is used for communication between memory 701 and processor 702.
[0136] The memory 701 is used to store computer programs that can run on the processor 702.
[0137] The memory 701 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0138] If the memory 701, processor 702, and communication interface 703 are implemented independently, then the communication interface 703, memory 701, and processor 702 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0139] Optionally, in a specific implementation, if the memory 701, processor 702, and communication interface 703 are integrated on a single chip, then the memory 701, processor 702, and communication interface 703 can communicate with each other through an internal interface.
[0140] The processor 702 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.
[0141] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program configured to execute the steps in any of the above-described server repair method embodiments when run.
[0142] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0143] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described server repair method embodiments.
[0144] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0145] The above provides a detailed description of a server repair method provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make various improvements and modifications to the present invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for repairing a server, characterized in that, Includes the following steps: Acquire hardware monitoring data and software monitoring data; The hardware monitoring data and the software monitoring data are input into a preset fault prediction model to obtain the fault time point. If the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset association conditions, the root cause event of the fault is located based on the preset fault knowledge graph and the hardware event and the software event. The preset fault prediction model is obtained by training a target neural network from historical monitoring data. Based on the fault time point and the fault root cause event, a repair action is determined from a preset self-healing strategy library. The repair action is then sandbox-tested. If the repair action passes the sandbox test, an execution signal is generated based on the repair action, and the repair action is executed based on the execution signal. The feedback result after executing the repair action is then sent to a preset terminal. The timing features of the hardware monitoring data and the software monitoring data are extracted by the preset fault prediction model and then fused to output the fault time point. Before determining that the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet a preset association condition, the process includes: assigning a first weight to the hardware event and assigning a second weight to the software event; determining whether the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than a first preset threshold; if the sum of the product of the hardware event and the first weight and the product of the software event and the second weight is greater than the first preset threshold, then determining that the hardware event determined by the hardware monitoring data and the software event determined by the software monitoring data meet the preset association condition. The step of inputting the hardware monitoring data and the software monitoring data into a preset fault prediction model to obtain the fault time point includes: inputting the hardware monitoring data and the software monitoring data into the preset fault prediction model to obtain the fault time point and prediction confidence; and outputting the fault time point when the prediction confidence is higher than the confidence threshold. The step of sending the feedback result after performing the repair action to a preset terminal includes: acquiring the signal strength and the duration of the signal strength; sending the feedback result to the preset terminal using a first communication channel when the signal strength is higher than a first strength threshold; and sending the feedback result to the preset terminal using a second communication channel when the signal strength is lower than a second strength threshold and the duration of the signal strength is higher than a preset time threshold.
2. The method according to claim 1, characterized in that, Before inputting the hardware monitoring data and the software monitoring data into the preset fault prediction model, the method further includes: Acquire the historical monitoring data and construct a dataset based on the historical monitoring data; Based on a preset partitioning ratio, the dataset is divided into a training set, a validation set, and a test set. Construct the target neural network by inputting the training set into the target neural network for training to obtain the initial model parameters; Based on the initial model parameters, the validation set is input into the target neural network for performance evaluation, and the initial model parameters are adjusted according to the performance evaluation results until the joint loss function of the validation set converges to obtain the optimal model parameters. Based on the optimal model parameters, the test set is input into the target neural network for model testing, and when the test results meet the preset requirements, the preset fault prediction model is obtained.
3. The method according to claim 1, characterized in that, include: The sandbox verification of the repair action includes: In a pre-set simulation container, the target server environment is simulated, and the repair action is performed on the target server to obtain the sandbox verification result.
4. The method according to claim 1, characterized in that, Also includes: Determine whether to accept the extended script corresponding to the new repair strategy; If an extended script corresponding to the new repair strategy is received, the extended script is sandbox verified, and if the extended script passes the sandbox verification, the preset self-healing strategy library is updated using the extended script.
5. The method according to claim 1, characterized in that, The hardware monitoring data includes at least one of temperature, voltage, and fan speed, and the software monitoring data includes at least one of software process status, application log keywords, and network traffic metrics.
6. The method according to claim 1, characterized in that, Also includes: After locating the root cause event of the fault based on the preset fault knowledge graph, according to the hardware events and the software events, the process includes: A fault alert instruction is generated based on the fault time point and the fault root cause event, and an acoustic fault alert and / or optical fault alert is performed based on the fault alert instruction.
7. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the steps of the server repair method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Fault prediction method of server, electronic equipment, medium and product
CN120560900A
Micro-service intelligent operation and maintenance method and device based on deep learning, and medium
CN120579041A