Server fault detection method, device and system and baseboard management controller
By implementing local fault detection and collaborative fault detection of low-latency servers on the server's baseboard management controller (BMC), delay and security issues in server fault monitoring and early warning are solved, and efficient, reliable and secure fault warning services are achieved.
Patent Information
- Application Number
- CN202510124979.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The prior art has latency and security problems in server fault monitoring and early warning, especially when the number of servers increases, resulting in an increase in the latency of fault warning and an increase in the risk of privacy data leakage.
By implementing a fault detection method on the substrate management controller (BMC), local training model is used to perform local fault detection, and the results are sent to the remote server, and status data is sent to the low-latency server for further fault detection, reducing communication delays and improving the reliability and security of fault warnings.
It realizes low-latency server fault monitoring and early warning, improves the timeliness, reliability and security of fault warning, and meets the business needs of strong real-time and large computing volume.
Smart Images

Figure CN119938387A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of servers, and in particular to a server fault detection method, device, system, baseboard management controller and computer-readable storage medium. Background Art
[0002] As the underlying hardware infrastructure architecture in the fields of Internet cloud computing, servers carry the functions of supporting the operation of key businesses and iterative optimization of various technologies. They inevitably have excellent requirements for the overall performance of the machine, and have much higher standards for system maintainability than home PCs (computers). They have more stringent specifications for operational stability, so general servers need to have high performance, high availability, and high reliability. Even if the server has undergone early data simulations, multiple PCB (printed circuit board) return board verifications, and rigorous testing under various extreme scenarios during the development stage, it still cannot be guaranteed to be foolproof. To ensure that the server can be effectively managed during operation and that faults can be diagnosed in a timely manner, this relies on the key component for managing and monitoring the server: the Baseboard Management Controller (BMC).
[0003] In the related art, the traditional BMC management system monitors and alarms the relevant data collected in real time through a monitoring management server deployed at a remote end; Figure 1 As shown in the figure, the BMC firmware is only responsible for the status data collection. A monitoring and management module is deployed on the remote server to check the collected and reported status data in real time and monitor and alarm the abnormal status that exceeds the preset threshold value. As the number of servers increases, a single remote server cannot bear a large amount of status data and fault detection operations. In some solutions, the BMC device will collect status data and report it to the edge server and cloud server for fault detection to reduce the workload of the remote server. However, as the number of servers that need to be monitored continues to grow at a high speed, the data reporting of the BMC devices in these solutions will increase the computing pressure and network transmission pressure of the edge server and cloud server, thereby causing serious fault warning delays, and cannot meet the fault warning requirements of real-time and computationally intensive services, and there is a certain risk of privacy leakage in data reporting.
[0004] Therefore, how to provide low-latency server fault monitoring and early warning services and improve the reliability and security of fault early warning is an urgent problem to be solved today. Summary of the invention
[0005] The purpose of the present invention is to provide a server fault detection method, device, system, baseboard management controller and computer-readable storage medium to provide low-latency server fault monitoring and early warning services and improve the reliability and safety of fault warning.
[0006] In order to solve the above technical problems, the present invention provides a server fault detection method, which is applied to a baseboard management controller, comprising:
[0007] Collect status data of each component in the server to be monitored;
[0008] Using the local trained model of each local detection service, respectively perform fault detection on the status data of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes a privacy data service and / or a time-sensitive service;
[0009] The status data of each peer detection service is sent to a low-latency server, so as to utilize the trained models of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by detection to the remote server; wherein, the low-latency server is an edge server and some servers in the cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
[0010] On the other hand, sending the fault detection result to a remote server includes:
[0011] According to the respective confidence levels of the fault detection results of the local detection services, detecting the local feedback services in the local detection services; wherein the local feedback services include the local detection services whose confidence levels of the fault detection results are greater than the corresponding confidence thresholds;
[0012] The fault detection result of the feedback service of the local end is sent to the remote server; wherein the opposite end detection service includes the services to be detected other than the feedback service of the local end in all the services to be detected of the server to be monitored.
[0013] On the other hand, the method of using the trained local model of each local detection service to perform fault detection on the status data of each local feedback service respectively, obtaining the fault detection result of each local detection service, and sending the fault detection result to the remote server includes:
[0014] Using the local trained models of all the services to be detected, respectively, to perform fault detection on the status data of each service to be detected, and obtain the fault detection results of each service to be detected;
[0015] If the current service to be detected is the privacy data service or the time-sensitive service, determining whether the confidence of the fault detection result of the current service to be detected is greater than a first confidence threshold; wherein the current service to be detected is any of the services to be detected;
[0016] If it is greater than the first confidence threshold, determining that the current service to be detected is a local feedback service;
[0017] If it is not greater than the first confidence threshold, determining that the current service to be detected is the opposite end detection service;
[0018] If the current service to be detected is not the privacy data service or the time-sensitive service, determining whether the confidence of the fault detection result of the current service to be detected is greater than a second confidence threshold; wherein the second confidence threshold is greater than the first confidence threshold;
[0019] If it is greater than the second confidence threshold, determining that the current service to be detected is the local feedback service;
[0020] If it is not greater than the second confidence threshold, determining that the current service to be detected is the opposite end detection service;
[0021] The fault detection results of the local feedback services among all the services to be detected are sent to the remote server.
[0022] On the other hand, the trained model on the local end and the trained model in the edge server are both models sent down after training by the cloud server.
[0023] On the other hand, the method of using the trained local model of each local detection service to perform fault detection on the status data of each local detection service respectively to obtain the fault detection result of each local detection service further includes:
[0024] Receive an initial trained model corresponding to the current local detection service sent by the cloud server; wherein the current local detection service is any of the local trained models, and the initial trained model is a model trained by the cloud server using sample data of at least two service types;
[0025] The initial trained model is migrated and trained by using the sample data of the current local detection service to obtain the local trained model of the current local detection service.
[0026] In another aspect, the method further comprises:
[0027] Receiving a model adjustment instruction sent by the cloud server; wherein the model adjustment instruction includes the neuron parameters to be adjusted of the trained model of the target local end, and the neuron parameters to be adjusted are the neuron parameters whose change amount is greater than the update threshold determined by the cloud server based on the user feedback information of the fault detection result collected by the remote server;
[0028] According to the neuron parameters to be adjusted, the corresponding neuron parameters in the trained model of the target local end are adjusted.
[0029] The present invention also provides a server fault detection device, which is applied to a baseboard management controller, comprising:
[0030] A collection module, used to collect status data of each component in the server to be monitored;
[0031] A local detection module, used to perform fault detection on the status data of each local detection service respectively using the local trained model of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes a privacy data service and / or a time-sensitive service;
[0032] A sending detection module is used to send the status data of each peer detection service to a low-latency server, so as to use the trained models of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by detection to the remote server; wherein, the low-latency server is an edge server and some servers in the cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
[0033] The present invention also provides a baseboard management controller, comprising:
[0034] Memory for storing computer programs;
[0035] A processor is used to implement the steps of the server fault detection method as described above when executing the computer program.
[0036] The present invention also provides a server fault detection system, comprising: an edge server, a cloud server and a baseboard management controller as described above;
[0037] Among them, the baseboard management controller is communicatively connected with the edge server and the cloud server; the edge server and the cloud server are both used to utilize their respective trained models of peer detection services to perform fault detection on the status data of each of the peer detection services sent by the baseboard management controller, and send the fault detection results of each of the peer detection services obtained by the detection to the remote server.
[0038] In addition, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the server fault detection method as described above are implemented.
[0039] A server fault detection method provided by the present invention is applied to a baseboard management controller, comprising: collecting status data of each component in a server to be monitored; using a local-end trained model of each local-end detection service to perform fault detection on the status data of each local-end detection service respectively, to obtain a fault detection result of each local-end detection service, and sending the fault detection result to a remote server; wherein the local-end detection service includes a privacy data service and / or a time-sensitive service; sending the status data of each opposite-end detection service to a low-latency server, so as to use a trained model of each opposite-end detection service in the low-latency server to perform fault detection on the status data of each opposite-end detection service respectively, and sending the fault detection result of each opposite-end detection service obtained by detection to a remote server; wherein the low-latency server is an edge server and some servers in a cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
[0040] It can be seen that the present invention uses the trained model of each local end detection service to perform fault detection on the status data of each local end detection service, obtains the fault detection result of each local end detection service, and sends the fault detection result to the remote server, which can identify the fault of privacy data and time-sensitive data related services on the BMC, avoid the leakage of privacy information in the server as much as possible, and improve the efficiency of fault identification in time-sensitive services, thereby improving the timeliness, reliability and security of fault warning; and by sending the status data of each opposite end detection service to a low-latency server, using the edge server or cloud server with a smaller current communication delay to predict the fault of the corresponding service, it can meet the fault warning requirements of services with strong real-time performance and large computational complexity. In addition, the present invention also provides a server fault detection device, system, baseboard management controller and computer-readable storage medium, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0042] Figure 1 It is a schematic diagram of a traditional BMC management system;
[0043] Figure 2 A flow chart of a server fault detection method provided by an embodiment of the present invention;
[0044] Figure 3 A schematic diagram of the system architecture of a server fault detection method provided by an embodiment of the present invention;
[0045] Figure 4 A structural block diagram of a server fault detection device provided by an embodiment of the present invention;
[0046] Figure 5 A schematic diagram of the structure of a baseboard management controller provided by an embodiment of the present invention;
[0047] Figure 6 A schematic diagram of the structure of a server fault detection system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] Please refer to Figure 2 , Figure 2 A flow chart of a server fault detection method provided by an embodiment of the present invention. The method is applied to a baseboard management controller (BMC) and may include:
[0050] Step 101: Collect status data of each component in the server to be monitored.
[0051] It is understandable that the server to be monitored in this embodiment may be a server that needs to use BMC to detect and collect status data of each component for fault alarm, such as a data center server. The BMC in the server to be monitored in this embodiment can monitor and collect status data of each component to be monitored in the server to be monitored.
[0052] Correspondingly, such as Figure 3 As shown, in this embodiment, the BMC (BMC firmware) can communicate with the edge server, cloud server and remote server; for example, the data center server monitored by the BMC can be networked with the edge server, cloud server and remote server respectively, and the edge server and cloud server can be networked with the remote server respectively to establish a communication connection for data transmission. Each BMC in this embodiment can be a SOC (System on Chip) system, which is divided into two levels: BMC chip and BMC firmware; BMC can be embedded in the motherboard of the data center server (i.e., the server to be monitored). With the rise of the CPU (central processing unit) + GPU (graphics processing unit) + DPU (data processing unit) cloud computing architecture in recent years, the embedding of BMC can also be extended from the motherboard level to the component level. A single server platform can be configured with multiple BMCs, that is, in the networking, multiple BMCs can be distributed and parallel to collect the status data of each component of a monitored server.
[0053] Accordingly, the specific quantity and type of status data monitored and collected by the BMC in this embodiment can be set by the designer according to practical scenarios and user needs. For example, when the server to be monitored is a data center server, the status data collected by the BMC may include all hardware information and operating system (OS) level information of the data center server; for example, the hardware information may include the voltage, temperature, fan speed and power status of each component; in order to timely warn of failures of core components of the server to ensure the normal operation of the core components of the server and reduce the losses caused by failures of core components of the server, the voltage, temperature and fan speed of the above-mentioned components may at least include the mainboard voltage, power status, chassis temperature and fan speed of the data center server; wherein the mainboard voltage may include the power input voltage, the core chip input voltage and the standby voltage, and the chassis temperature may include the air inlet temperature, the air outlet temperature and the core chip temperature.
[0054] Step 102: Use the trained model of each local detection service to perform fault detection on the status data of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes privacy data service and / or time-sensitive service.
[0055] It is understandable that as the number of servers connected to the network continues to grow at a high speed, the number of servers and components on the servers that the BMC needs to monitor also continues to increase at a high speed. The amount of data obtained after the BMC collects the status data of each component on the server is massive. If all the massive data collected is transmitted to the edge server and cloud server for fault warning at this time, it will inevitably cause serious delays due to the transmission of all data, and both the edge server and the cloud server need to provide fault warnings for all data, resulting in repetitive processing of status data, wasting computing resources and communication resources, and also causing serious delays, which cannot meet the fault warning needs of businesses with strong real-time performance and high computing volume (such as autonomous driving and interactive virtual reality or augmented reality games).
[0056] In order to reduce the server fault warning delay and improve the fault warning security and reliability of systems involving privacy data (such as military information processing systems and high-precision technology research and development systems, etc.), in this embodiment, the BMC can be used to perform fault detection on the status data of privacy data services and / or time-sensitive services at the local end to ensure the security and reliability of fault warning. Among them, the privacy data service can be a service that needs to use privacy data for fault detection, and the time-sensitive service can be a service that needs to use time-sensitive data for fault detection.
[0057] Accordingly, the local detection service in this embodiment may be a service that needs to be detected in the BMC (i.e., the service to be detected). The local trained model in this embodiment may be a trained model required for fault prediction in the BMC (e.g., Figure 3 For example, to adapt to the computing power and memory of BMC and edge servers, the trained model on the local side can use an efficient network model with fewer parameters and computing power, such as SquezeNet (a lightweight convolutional neural network model), MobileNet (a lightweight neural network model), and GhostNet (a lightweight deep learning model).
[0058] Correspondingly, the specific method in which the BMC uses the trained model of each local detection service to perform fault detection on the status data of each local detection service respectively, obtains the fault detection result of each local detection service, and sends the fault detection result to the remote server in this step can be set by the designer according to practical scenarios and user needs. For example, the BMC can use the trained model of each local detection service to perform fault detection on the status data of each local detection service respectively, and directly send the fault detection result of each local detection service obtained by the detection to the remote server, so as to use the BMC's fault detection when the local detection service includes a small part of all the services to be detected of the server to be monitored (such as privacy data services and time-sensitive services), so as to ensure the security and reliability of fault warning; that is, the opposite-end detection service in step 103 can be a service other than the local detection service in all the services to be detected.
[0059] Correspondingly, after using the trained local model of each local detection service to perform fault detection on the status data of each local detection service, the BMC can also screen out the local feedback services with confidence levels higher than the corresponding confidence threshold according to the confidence levels of the fault detection results of each local detection service, and send the fault detection results of the local feedback services to the remote server, so as to directly output the fault detection results corresponding to the status data with high confidence levels to the remote server; that is, the opposite-end detection service in step 103 can be a service other than the local feedback service in all services to be detected, so as to send the status data with low confidence levels to the low-latency server for more accurate fault detection.
[0060] For example, the process of sending the fault detection result to the remote server in this step may include: detecting the local feedback service in the local detection service according to the confidence of the fault detection result of each local detection service; wherein the local feedback service includes the local detection service whose confidence of the fault detection result is greater than the corresponding confidence threshold; sending the fault detection result of the local feedback service to the remote server; wherein the opposite-end detection service includes the services to be detected other than the local feedback service in all the services to be detected of the server to be monitored.
[0061] For example, the local detection service may include all services to be detected. In this step, the BMC may use the local trained models of all services to be detected to perform fault detection on the status data of each service to be detected, and obtain the fault detection results of each service to be detected; if the current service to be detected is a privacy data service or a time-sensitive service, determine whether the confidence of the fault detection result of the current service to be detected is greater than the first confidence threshold; wherein the current service to be detected is any service to be detected; if it is greater than the first confidence threshold, determine that the current service to be detected is a feedback service of the local end; if it is not greater than the first confidence threshold , then determine that the current service to be detected is a peer detection service, so as to use the corresponding trained model in the low-latency server to perform more accurate fault detection through step 103; if the current service to be detected is not a privacy data service or a time-sensitive service, then determine whether the confidence of the fault detection result of the current service to be detected is greater than the second confidence threshold; if it is greater than the second confidence threshold, then determine that the current service to be detected is a local feedback service; if it is not greater than the second confidence threshold, then determine that the current service to be detected is a peer detection service; send the fault detection results of the local feedback service in all services to be detected to the remote server. Among them, the second confidence threshold can be greater than the first confidence threshold, so as to use the setting of a smaller first confidence threshold to make the privacy data service and time-sensitive service these two types of services perform fault detection in the BMC as much as possible, so as to ensure the timeliness, reliability and security of fault warning.
[0062] The present embodiment does not limit the specific content of the fault detection result, such as the fault detection result may include any one or more of the fault type, fault occurrence time, fault component, component location, and fault handling solution. For example, in order to enable the remote server to quickly locate the fault location, understand the fault cause and the handling solution, the fault detection result may include the fault type, fault occurrence time, fault component, component location, and fault handling solution.
[0063] Step 103: Send the status data of each peer detection service to the low-latency server, use the trained model of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by the detection to the remote server.
[0064] Among them, the low-latency servers are some servers in the edge servers and cloud servers that are connected to the BMC for communication. The communication delay between the low-latency servers in the edge servers and cloud servers and the BMC is smaller than that between other servers.
[0065] It is understandable that the peer detection service in this embodiment can be a service that needs to perform fault detection on a low-latency server among all services to be detected of the server to be monitored, such as services other than the services to be detected (i.e., local feedback services) for which the BMC sends corresponding fault detection results to the remote server. The low-latency server in this embodiment can be some servers with relatively low current communication delay among the edge servers and cloud servers connected to the BMC for communication, such as a server with the lowest communication delay.
[0066] Accordingly, in this step, the BMC sends the status data of each peer detection service to the low-latency server, so that the low-latency server (such as an edge server or a cloud server) can use its own trained models of each peer detection service to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by the detection to the remote server, so that the low-latency server with a smaller communication delay with the BMC can accurately detect the status data of the peer detection service (such as the status data with unclear BMC fault detection results). Compared with the solution in the related art that the BMC transmits all data to the edge server and the cloud server, it can significantly reduce the network transmission pressure, and can reduce the amount of calculation of the BMC for the data processing process, and make full use of the excellent computing power of the edge server or the cloud server, so that the overall fault warning effect of the system is optimal, and can provide low-latency server fault monitoring and warning services when the number of servers connected to the network continues to grow at a high speed, and can meet the fault warning needs of services with strong real-time performance and large computing volume.
[0067] Correspondingly, the method provided in this embodiment may also include a process in which the BMC detects communication delays with the edge server and the cloud server respectively; for example, network detection, path tracking and / or network diagnostic tools may be used to detect communication delays with the edge server and the cloud server.
[0068] It should be noted that the trained model in the BMC (i.e., the trained model on the local side) and the trained models in the edge server and the cloud server can all be models trained by the cloud server, that is, the trained model on the local side and the trained model in the edge server can all be models sent down after training by the cloud server, so as to utilize the cloud server to be responsible for the training process of the model and reduce the amount of calculation of the BMC and the edge server. For example, in some embodiments, the BMC can directly receive the model sent down after training using the cloud server (i.e., the trained model on the local side).
[0069] In other embodiments, an initial trained model corresponding to the current local detection service sent by the cloud server is received; the initial trained model is transferred and trained using the sample data of the current local detection service to obtain a local trained model of the current local detection service; wherein the current local detection service is any local trained model, and the initial trained model is a model obtained by the cloud server using sample data of at least two types of services. In other words, since server fault warnings of different service types are all performed through the status data collected by the BMC; therefore, in the field of server fault detection based on the status data collected by the BMC, it can be considered that the status data of each component of the server collected by the BMC under different service types obeys the same distribution and comes from the same feature space. Based on this, in this embodiment, the cloud server can perform multi-task learning based on the status data under multiple (i.e., at least two) service types, and can find and understand the correlation and hierarchy of different tasks, thereby improving the generalization ability of the model.
[0070] Correspondingly, the process of the edge server obtaining the trained model corresponding to each peer detection service is similar to the process of the BMC obtaining the local trained model of the current local detection service, which will not be repeated here.
[0071] For example, the cloud server can train the training model with sample data under two business types; the cloud server sends the trained model to the BMC. Since the status data of the server components under different business types collected by the BMC are correlated, in order to make the model obtained by the cloud server through multi-task learning more suitable for fault warning under the current business type of the server, the learned model parameters can be shared with the BMC local model through transfer learning, thereby accelerating and optimizing the learning efficiency of the model, without having to learn from scratch like most networks; that is, the BMC can transfer the model sent by the cloud server with sample data under the current business type to obtain the local training model.
[0072] Correspondingly, in this embodiment, it is taken into account that the feature extractor in the model trained on the cloud server may not have processed a part of the data in the current business, resulting in the output target features may not belong to the source data distribution, so that the model cannot accurately predict this type of features; in this embodiment, a business classifier is added to the model. For example, the cloud server can add a classifier after each intermediate layer of the initial model to obtain the model to be trained, so that during the model training process, the business classifier can be used to determine the business to which the features extracted by the feature extractor belong, so that the feature extractor can extract the features corresponding to the business from the sample data under the current business type, thereby realizing the training of the initial trained models of each business to be detected.
[0073] Among them, the process of the cloud server training the to-be-trained model through sample data under multiple business types may include: obtaining the status data of each component of the data center server under different business types as well as the fault cause, fault location and fault handling solution as sample data, and establishing a multidimensional data set; setting corresponding labels for the data in the multidimensional data set according to whether the component is normal or faulty; dividing the sample data in the multidimensional data set to obtain a training set and a verification set; inputting the data of the training set into the to-be-trained model to train the to-be-trained model; verifying the generalization ability of the to-be-trained model through the data of the verification set, and adjusting the hyperparameters of the to-be-trained model according to the model performance to obtain an initial trained model.
[0074] It should be noted that in this embodiment, the remote server can display the received fault detection results of each service to the user, and the remote server can also support the user to confirm and adjust the displayed fault warning results. In order to improve the accuracy of fault warning, in this embodiment, the user's feedback information on the fault monitoring results (i.e., the user feedback information of the fault detection results) can also be collected to dynamically update the model; considering that using the user feedback information of the fault detection results to adjust the model will generate a large amount of calculation, the remote server can send the user feedback information of the fault detection results to the cloud server, and use the super computing power of the cloud server to dynamically update the model. For example, the cloud server can obtain the user feedback information of the fault detection results sent by the remote server; increase the input of the trained model corresponding to the user feedback information of the fault detection results to update the training of the trained model.
[0075] Since the models of BMC, cloud server, and edge server can all be trained based on the same initial trained model and the same sample data under multiple business types, the model parameters are basically the same, at least the parameters of the neurons that learn common features in the model are basically the same. Based on this, the cloud server can first filter the updated model parameters in the cloud server, and not send the parameters of neurons with smaller changes (i.e., neuron parameters) to the BMC or edge server, and only send the parameters of neurons with larger changes, so that the BMC or edge server can maintain the parameters of other neurons when updating the local model, and only adjust the corresponding neuron parameters according to the sent parameters, which can reduce the communication overhead caused by sending parameters and reduce the training computing power overhead of the local model of the BMC or edge server. In other words, the cloud server can filter the parameters (i.e., neuron parameters) whose changes are greater than the update threshold from the updated trained model; and send the filtered parameters to the edge server and BMC to adjust the corresponding parameters in their respective models.
[0076] Accordingly, the method provided in this embodiment may further include: the BMC receives a model adjustment instruction sent by the cloud server; according to the neuron parameters to be adjusted, the corresponding neuron parameters in the trained model of the target local end are adjusted; wherein the model adjustment instruction includes the neuron parameters to be adjusted of the trained model of the target local end, and the neuron parameters to be adjusted are the neuron parameters whose change amount is greater than the update threshold, determined by the cloud server based on the user feedback information of the fault detection results collected by the remote server; the trained model of the target local end may be a trained model of the local end whose neuron parameters need to be adjusted.
[0077] On this basis, the cloud server updates the model based on the user feedback information of the fault detection results received each time, and sends the neuron parameters with changes greater than the update threshold to the BMC or edge server. Due to the limited computing power of the BMC or edge server, if the cloud server receives multiple user feedback information of fault detection results sent by the remote server within the preset time, in order to alleviate the gradient outdated problem of the local model of the BMC or edge server, the BMC or edge server may not update the local model parameters immediately after receiving the neuron parameters sent by the cloud server, but will update the model parameters after reaching the preset time. During this period, the model of the BMC or edge server still performs fault detection according to the previous parameters. In other words, the BMC can adjust the corresponding neuron parameters in the trained model of the target end according to the neuron parameters to be adjusted after receiving the model adjustment instruction at the preset time.
[0078] In this embodiment, the embodiment of the present invention performs fault detection on the status data of each local detection service by using the local trained model of each local detection service, obtains the fault detection result of each local detection service, and sends the fault detection result to the remote server, so as to perform fault identification of services related to privacy data and time-sensitive data on the BMC, avoid the leakage of privacy information in the server as much as possible, and improve the efficiency of fault identification in time-sensitive services, thereby improving the timeliness, reliability and security of fault warning; and by sending the status data of each opposite-end detection service to a low-latency server, using the edge server or cloud server with a smaller current communication delay to predict the fault of the corresponding service, it can meet the fault warning requirements of services with strong real-time performance and large computational complexity.
[0079] Corresponding to the above method embodiment, an embodiment of the present invention further provides a server fault detection device. The server fault detection device described below and the server fault detection method described above can refer to each other.
[0080] Please refer to Figure 4 , Figure 4 This is a structural block diagram of a server fault detection device provided by an embodiment of the present invention. The device is applied to BMC and may include:
[0081] The acquisition module 10 is used to acquire the status data of each component in the server to be monitored;
[0082] The local detection module 20 is used to perform fault detection on the status data of each local detection service respectively using the local trained model of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes a privacy data service and / or a time-sensitive service;
[0083] The sending detection module 30 is used to send the status data of each peer detection service to the low-latency server, so as to use the trained model of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection result of each peer detection service obtained by the detection to the remote server; wherein, the low-latency server is an edge server and some servers in the cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
[0084] In some embodiments, the sending detection module 30 may include:
[0085] A confidence detection submodule, used to detect the local feedback service in the local detection service according to the confidence of each fault detection result of each local detection service; wherein the local feedback service includes the local detection service whose confidence of the fault detection result is greater than the corresponding confidence threshold;
[0086] The result generation submodule is used to send the fault detection result of the feedback service of the local end to the remote server; wherein the detection service of the opposite end includes the services to be detected other than the feedback service of the local end among all the services to be detected of the server to be monitored.
[0087] In some embodiments, the sending detection module 30 may include:
[0088] The fault detection submodule is used to perform fault detection on the status data of each service to be detected by using the local trained models of all services to be detected to obtain the fault detection results of each service to be detected;
[0089] The first judgment submodule is used to judge whether the confidence of the fault detection result of the current service to be detected is greater than a first confidence threshold if the current service to be detected is a privacy data service or a time-sensitive service; wherein the current service to be detected is any service to be detected; if it is greater than the first confidence threshold, the current service to be detected is determined to be a local feedback service; if it is not greater than the first confidence threshold, the current service to be detected is determined to be a peer detection service;
[0090] The second judgment submodule is used to judge whether the confidence of the fault detection result of the current service to be detected is greater than a second confidence threshold if the current service to be detected is not a privacy data service or a time-sensitive service; wherein the second confidence threshold is greater than the first confidence threshold; if it is greater than the second confidence threshold, it is determined that the current service to be detected is a local feedback service; if it is not greater than the second confidence threshold, it is determined that the current service to be detected is a peer detection service;
[0091] The sending submodule is used to send the fault detection result of the local feedback service among all the services to be detected to the remote server.
[0092] In some embodiments, the trained model on the local end and the trained model in the edge server are both models sent down after training by the cloud server.
[0093] In some embodiments, the apparatus may further include:
[0094] A model receiving module, used to receive an initial trained model corresponding to the current local detection service sent by the cloud server; wherein the current local detection service is any local trained model, and the initial trained model is a model trained by the cloud server using sample data of at least two service types;
[0095] The migration training module is used to use the sample data of the current local detection service to perform migration training on the initial trained model to obtain the local trained model of the current local detection service.
[0096] In some embodiments, the apparatus may further include:
[0097] The adjustment receiving module is used to receive the model adjustment instruction sent by the cloud server; wherein the model adjustment instruction includes the neuron parameters to be adjusted of the trained model of the target local end, and the neuron parameters to be adjusted are the neuron parameters whose change amount is greater than the update threshold determined by the cloud server based on the user feedback information of the fault detection result collected by the remote server;
[0098] The model updating module is used to adjust the corresponding neuron parameters in the trained model of the target end according to the neuron parameters to be adjusted.
[0099] In this embodiment, the embodiment of the present invention uses the local detection module 20 to use the local trained model of each local detection service to perform fault detection on the status data of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server, so as to perform fault identification of services related to privacy data and time-sensitive data on the BMC, avoid the leakage of privacy information in the server as much as possible, and improve the efficiency of fault identification in time-sensitive services, thereby improving the timeliness, reliability and security of fault warning; and send the status data of each opposite-end detection service to the low-latency server by sending the detection module 30, and use the edge server or cloud server with smaller current communication delay to predict the fault of the corresponding service, so as to meet the fault warning requirements of services with strong real-time performance and large computational workload.
[0100] Corresponding to the above method embodiment, the embodiment of the present invention further provides a baseboard management controller. The baseboard management controller described below and the server fault detection method described above can be referred to each other.
[0101] Please refer to Figure 5 , Figure 5 A schematic diagram of the structure of a baseboard management controller provided by an embodiment of the present invention. The server may include:
[0102] A memory D1, for storing computer programs;
[0103] The processor D2 is used to implement the steps of the server fault detection method provided in the above embodiment when executing the computer program.
[0104] Corresponding to the above method embodiment, an embodiment of the present invention further provides a server fault detection system. The server fault detection system described below and the server fault detection method described above can refer to each other.
[0105] Please refer to Figure 6 , Figure 6 A schematic diagram of a server fault detection system provided in an embodiment of the present invention. The system may include: an edge server 100, a cloud server 200, and a baseboard management controller 300 provided in the above embodiment;
[0106] Among them, the baseboard management controller 300 is communicatively connected with the edge server 100 and the cloud server 200; the edge server 100 and the cloud server 200 are both used to use the trained models of their respective peer detection services to perform fault detection on the status data of each peer detection service sent by the baseboard management controller 300, and send the fault detection results of each peer detection service obtained by the detection to the remote server.
[0107] That is, when the edge server 100 and the cloud server 200 are used as the low-latency servers in the above embodiment, they can perform fault detection on the status data of the peer detection service sent by the baseboard management controller 300 .
[0108] In other embodiments, the cloud server 200 can also be used to train and send the trained model on the local side to the baseboard management controller 300; train and send the trained model to the edge server.
[0109] In other embodiments, the cloud server 200 can be specifically used to use sample data of at least two business types to train an initial trained model corresponding to the current local detection business; send the initial trained model to the baseboard management controller 300, so that the baseboard management controller 300 uses the sample data of the current local detection business to perform migration training on the initial trained model to obtain the local trained model of the current local detection business.
[0110] In other embodiments, the cloud server 200 can also be used to receive user feedback information on fault detection results collected by a remote server; train and update the corresponding trained model based on the user feedback information on the fault detection results; if the change in the updated neuron parameters is greater than the update threshold, the updated neuron parameters are used to generate corresponding model adjustment instructions.
[0111] Corresponding to the above method embodiment, the embodiment of the present invention further provides a computer program product. The computer program product described below and the server fault detection method described above can be referred to each other.
[0112] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps of the server fault detection method provided in the above embodiment.
[0113] Corresponding to the above method embodiment, an embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium described below and the server fault detection method described above can be referred to each other.
[0114] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the server fault detection method provided in the above embodiment are implemented.
[0115] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the device, BMC, system, computer program product and computer-readable storage medium disclosed in the embodiment, since they correspond to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0116] The above is a detailed introduction to a server fault detection method, device, system, baseboard management controller and computer-readable storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, the present invention can also be improved and modified, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A method for detecting a server failure, characterized in that: Applicable to baseboard management controllers, including: Collect status data of each component in the server to be monitored; Using the local trained model of each local detection service, respectively perform fault detection on the status data of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes a privacy data service and / or a time-sensitive service; The status data of each peer detection service is sent to a low-latency server, so as to utilize the trained models of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by detection to the remote server; wherein, the low-latency server is an edge server and some servers in the cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
2. The server fault detection method according to claim 1, characterized in that: The step of sending the fault detection result to a remote server comprises: According to the respective confidence levels of the fault detection results of the local detection services, detecting the local feedback services in the local detection services; wherein the local feedback services include the local detection services whose confidence levels of the fault detection results are greater than the corresponding confidence thresholds; The fault detection result of the feedback service of the local end is sent to the remote server; wherein the opposite end detection service includes the services to be detected other than the feedback service of the local end in all the services to be detected of the server to be monitored.
3. The server fault detection method according to claim 1, characterized in that: The method of using the local trained model of each local detection service to respectively perform fault detection on the status data of each local feedback service, obtaining the fault detection result of each local detection service, and sending the fault detection result to the remote server includes: Using the local trained models of all the services to be detected, respectively, to perform fault detection on the status data of each service to be detected, and obtain the fault detection result of each service to be detected; If the current service to be detected is the privacy data service or the time-sensitive service, determining whether the confidence of the fault detection result of the current service to be detected is greater than a first confidence threshold; wherein the current service to be detected is any of the services to be detected; If it is greater than the first confidence threshold, determining that the current service to be detected is a local feedback service; If it is not greater than the first confidence threshold, determining that the current service to be detected is the opposite end detection service; If the current service to be detected is not the privacy data service or the time-sensitive service, determining whether the confidence of the fault detection result of the current service to be detected is greater than a second confidence threshold; wherein the second confidence threshold is greater than the first confidence threshold; If it is greater than the second confidence threshold, determining that the current service to be detected is the local feedback service; If it is not greater than the second confidence threshold, determining that the current service to be detected is the opposite end detection service; The fault detection results of the local feedback services among all the services to be detected are sent to the remote server.
4. The server fault detection method according to any one of claims 1 to 3, characterized in that: The trained model on the local end and the trained model in the edge server are both models sent down after training by the cloud server.
5. The server fault detection method according to claim 4, characterized in that: Before the method of using the trained local model of each local detection service to respectively perform fault detection on the status data of each local detection service to obtain the fault detection result of each local detection service, the method further includes: Receive an initial trained model corresponding to the current local detection service sent by the cloud server; wherein the current local detection service is any of the local trained models, and the initial trained model is a model trained by the cloud server using sample data of at least two service types; The initial trained model is migrated and trained by using the sample data of the current local detection service to obtain the local trained model of the current local detection service.
6. The server fault detection method according to claim 4, characterized in that: Also includes: Receiving a model adjustment instruction sent by the cloud server; wherein the model adjustment instruction includes the neuron parameters to be adjusted of the trained model of the target local end, and the neuron parameters to be adjusted are the neuron parameters whose change amount is greater than the update threshold determined by the cloud server based on the user feedback information of the fault detection result collected by the remote server; According to the neuron parameters to be adjusted, the corresponding neuron parameters in the trained model of the target local end are adjusted.
7. A server fault detection device, characterized in that: Applicable to baseboard management controllers, including: A collection module, used to collect status data of each component in the server to be monitored; A local detection module, used to perform fault detection on the status data of each local detection service respectively using the local trained model of each local detection service, obtain the fault detection result of each local detection service, and send the fault detection result to the remote server; wherein the local detection service includes a privacy data service and / or a time-sensitive service; A sending detection module is used to send the status data of each peer detection service to a low-latency server, so as to use the trained models of each peer detection service in the low-latency server to perform fault detection on the status data of each peer detection service respectively, and send the fault detection results of each peer detection service obtained by detection to the remote server; wherein, the low-latency server is an edge server and some servers in the cloud server that are communicatively connected to the baseboard management controller, and the communication delay between the low-latency server in the edge server and the cloud server and the baseboard management controller is less than that of other servers.
8. A baseboard management controller, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the server fault detection method according to any one of claims 1 to 6 when executing the computer program.
9. A server fault detection system, characterized in that: include: An edge server, a cloud server, and a baseboard management controller as claimed in claim 8; Wherein, the baseboard management controller is communicatively connected with the edge server and the cloud server; The edge server and the cloud server are both used to use their respective trained models of peer detection services to perform fault detection on the status data of each peer detection service sent by the baseboard management controller, and send the fault detection results of each peer detection service obtained by the detection to the remote server.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the server fault detection method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Server fault monitoring method and system based on neural network
CN111143173A
Server fault detection method and system and computer readable storage medium
CN113032218A
Hardware state display method and device, storage medium and mobile terminal
CN115801615A
Server hardware fault detection method and device and medium
CN118708413A
Server fault prediction method and system based on BMC module edge calculation
CN118733317A