A Service Fault Location System and Method Based on BMC

By using a service fault location system based on BMC, server restart testing has been automated and made continuous, solving the problems of cumbersome operation and incomplete data in existing testing methods, and improving the efficiency and accuracy of server stability testing.

CN121301114BActive Publication Date: 2026-05-05CHANGSHA XIANGJI HAIDUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHANGSHA XIANGJI HAIDUN TECH CO LTD
Filing Date
2025-12-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing server restart testing methods are cumbersome to operate, difficult to precisely control the number of restarts, greatly affected by human factors, prone to test interruption, have high hardware resource requirements, incomplete data collection, difficult to process redundant data, and untimely on-site data collection, resulting in difficulties in troubleshooting and low efficiency.

Method used

A service fault location system based on BMC is adopted. A control link is established between BMC and CPLD to collect hardware component data in real time, automatically identify anomalies and control hard reboot or no reboot, persistently store key data, and use IPMI custom protocol to obtain BIOS and operating system data to realize an automated and continuous testing process.

Benefits of technology

It reduces testing costs, improves testing efficiency and continuity, accurately records abnormal field data, reduces redundant data, enhances data processing efficiency and fault diagnosis accuracy, and is suitable for various testing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301114B_ABST
    Figure CN121301114B_ABST
Patent Text Reader

Abstract

This invention relates to a service fault location system and method based on a BMC (Browser Control Center). The system includes: establishing a communication link between the BMC and the server under test (DUT) using a customized IPMI protocol; configuring test parameters and controlling the DUT to power on and start a restart test program; the BMC collecting current hardware-level data; the BMC performing temporary storage and determining whether an anomaly occurred during the current test round; when an anomaly occurs during the current test round, the BMC receives the current customized IPMI message, obtains the anomaly level information, and controls the CPLD (Content Management Logic Controller) to perform a hard restart or not perform a hard restart based on the anomaly level information, and persistently stores the on-site data. No additional test equipment or server hardware / software modifications are required, resulting in lower testing costs and broad applicability; in the event of an abnormal downtime, the BMC can directly operate the CPLD to perform a hard restart, greatly improving testing efficiency and continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer technology, and in particular relates to a service fault location system and method based on BMC. Background Technology

[0002] In the field of server performance testing, stability is a key indicator for measuring server quality. Existing testing methods have very limited ability to detect intermittent hardware problems. Because these problems occur infrequently and randomly, traditional testing methods struggle to capture them. Reboot testing, however, is an effective method that simulates multiple server restarts, increasing the likelihood of intermittent hardware problems occurring and thus providing a comprehensive assessment of server stability.

[0003] Currently, restart testing primarily relies on manual operation or traditional script settings. Manual restarting is cumbersome, the number of restarts is difficult to control precisely, it is greatly affected by human factors, and its repeatability and accuracy are poor. While script testing can achieve good automation, it also has serious limitations. If the server system crashes during testing, the test cannot continue, and the testing process is immediately interrupted. Furthermore, if manual intervention is required during script testing, it must be completed before the operating system restarts, which is quite difficult.

[0004] Traditional restart testing services also suffer from high hardware resource requirements and incomplete data collection. To obtain relatively comprehensive test data, it is often necessary to add additional test computers and connect extra data sensors, which not only increases testing costs but also fails to provide comprehensive reference data that fully covers all key aspects of server operation and potential problems.

[0005] Large-scale testing generates massive amounts of data. Due to a lack of effective data filtering and management mechanisms, this vast amount of redundant data makes manual analysis extremely difficult. Technicians need to spend a significant amount of time and effort locating the error site to obtain relevant data, which undoubtedly increases the workload of problem analysis and reduces efficiency. If only the data at the time of the problem could be saved, redundancy could be greatly reduced from the source, improving the precision and accuracy of problem analysis.

[0006] Existing testing systems struggle to quickly and comprehensively collect detailed on-site data when problems occur. They often only begin collecting data after a problem has already occurred, by which time some critical data may have been lost or overwritten. This makes it difficult for technicians to accurately analyze the cause of the problem, posing a significant challenge to server troubleshooting and performance optimization.

[0007] In summary, existing server restart testing methods have many problems that urgently need to be solved in terms of ease of operation, test continuity, impact of human intervention, data processing, detection of occasional hardware problems, and on-site data collection. There is an urgent need for a more scientific and efficient testing scheme to improve the quality and efficiency of server stability testing. Summary of the Invention

[0008] To address the aforementioned technical problems, this invention provides a service fault location system and method based on BMC.

[0009] The technical solution adopted in this invention is:

[0010] Firstly, a service fault location system based on BMC is provided, including:

[0011] BMC and server under test, the server under test includes CPLD and multiple hardware components;

[0012] The BMC establishes a control link with the CPLD, establishes a communication link with the server under test using the IPMI custom protocol, and connects to multiple hardware components through standard hardware interfaces.

[0013] BMC configures test parameters and controls the power-on of the server under test to start and restart the test program;

[0014] BMC collects real-time hardware-level data from multiple hardware components in the current round of testing through standard hardware interfaces.

[0015] BMC temporarily stores the current hardware-level data and determines whether any anomalies occur during the current round of testing based on the current hardware-level data.

[0016] When an anomaly occurs during the current round of testing, the BMC receives the current IPMI customized message through the communication link of the IPMI customized protocol, obtains the anomaly level information based on the current IPMI customized message, controls the CPLD to perform a hard reboot or not perform a hard reboot based on the anomaly level information, and obtains the anomaly field data based on the current hardware level data and the current IPMI customized message, and persists the field data.

[0017] Furthermore, the test parameters include the number of test rounds to restart, the maximum number of exceptions recorded, the exception data backtracking time limit, and the test requirement data collection rules.

[0018] Furthermore, BMC includes:

[0019] The data filtering module is used to filter the current hardware-level data according to the data collection rules for testing requirements to obtain the necessary test data;

[0020] The anomaly detection module is used to determine whether an anomaly has occurred on the server under test during the current round of testing, based on the necessary test data.

[0021] Furthermore, the IPMI custom protocol proactively pushes a customized IPMI message containing BIOS boot information and operating system-level data to the BMC when the server under test malfunctions.

[0022] Furthermore, BMC also includes:

[0023] The IPMI interface is used to receive the current IPMI customized message actively pushed by the server under test through the communication link of the IPMI customized protocol when an anomaly occurs during the current round of testing. The current IPMI customized message includes the current BIOS boot information and the current operating system level data.

[0024] The abnormal data storage module is used to record the time point of the abnormal occurrence, obtain the backtracking period before the abnormal occurrence time point based on the abnormal data backtracking time limit, and extract the abnormal on-site data from the current hardware level data, the current BIOS boot information and the current operating system level data according to the backtracking period, and persistently store the on-site data.

[0025] Furthermore, BMC also includes:

[0026] The anomaly classification module is used to determine the anomaly level information based on the current BIOS boot information and the current operating system level data. The anomaly level information includes non-critical anomaly level and critical anomaly level.

[0027] The hard reboot control module generates control commands when the anomaly level is severe, and controls the CPLD to perform a hard reboot on the server under test, allowing the current round of testing to continue. When the anomaly level is non-severe, no control commands are generated.

[0028] Furthermore, BMC also includes:

[0029] The exception count module is used to record the total number of exceptions n in all previous test rounds. When an exception occurs, the current total number of exceptions is n+1.

[0030] The test result summary module is used to end the restart test program when the current total number of exceptions n+1 equals the maximum number of exceptions recorded N, to collect the test information corresponding to the restart test program, and to generate a field data packet for all exceptions in the restart test program.

[0031] The test round counting module is used to continue executing the current test round until it ends when the current total number of exceptions n+1 is less than the maximum number of exceptions recorded N, and to record the current test round as i.

[0032] The test result summary module is also used to end the restart test program when the current test round i is equal to the restart test round number I, to collect the test information corresponding to the restart test program and generate the field data packets of all abnormalities in the restart test program; when the current test round i is less than the restart test round number I, it controls the server under test to perform the next test round i+1.

[0033] Furthermore, BMC also includes:

[0034] The operating system entry judgment module is used to determine whether the server under test has entered the operating system normally when no abnormality occurs during the current round of testing.

[0035] The abnormal data storage module is also used to persistently store the on-site data if the server under test fails to enter the operating system normally.

[0036] The extended testing module is used to control the server under test to mount virtual media and execute a preset test script under the operating system if the server under test enters the operating system normally, in order to simulate the current round of testing.

[0037] Secondly, a service fault location method based on BMC is provided, which is applied to the BMC-based service fault location system of the first aspect. The service fault location method includes:

[0038] Configure the test parameters, control the power-on of the server under test, and start the restart test program;

[0039] Real-time data collection at the current hardware level from multiple hardware components in the current round of testing is achieved through standard hardware interfaces;

[0040] The current hardware-level data is temporarily stored, and the current hardware-level data is used to determine whether any anomalies occur during the current round of testing.

[0041] When an anomaly occurs during the current round of testing, the current IPMI customized message is received through the communication link of the IPMI customized protocol;

[0042] Based on the current IPMI customized message, obtain the abnormality level information, and control the CPLD to perform a hard reboot or not perform a hard reboot on the server under test according to the abnormality level information;

[0043] Based on the current hardware data and the current IPMI customized message, abnormal field data is obtained and persistently stored.

[0044] The beneficial effects achieved by this invention are as follows:

[0045] The BMC-based service fault location system includes a BMC and a server under test (DUT). The DUT includes a CPLD and multiple hardware components. The BMC establishes a control link with the CPLD and a communication link with the DUT using a custom IPMI protocol. The BMC connects to the multiple hardware components via standard hardware interfaces. The BMC configures test parameters, powers on the DUT, and initiates the restart test program. The BMC collects real-time hardware-level data from the multiple hardware components during the current test round via the standard hardware interface. The BMC temporarily stores the current hardware-level data and determines whether an anomaly has occurred during the current test round based on this data. When an anomaly occurs during the current test round, the BMC receives the current IPMI custom message via the IPMI custom protocol communication link, obtains the anomaly level information based on the message, controls the CPLD to perform a hard reboot or not based on the anomaly level information, and obtains the anomaly's on-site data based on the current hardware-level data and the IPMI custom message, which is then persistently stored.

[0046] Using BMC and the standard sensors of the server under test, rich and detailed on-site data collection can be completed without the need for additional testing equipment, server hardware modification, or additional data collection equipment, thus reducing testing costs.

[0047] The server under test can be any general-purpose server that supports BMC, which means that there is no need to make large-scale modifications to the server hardware or software, and it has wide applicability.

[0048] BMC establishes a customized IPMI protocol with the server under test, enabling it to obtain abnormal status information actively pushed by the BIOS and operating system of the server under test, and thus accurately record the abnormal situation. BMC can control the mounting of media on the server under test and use scripts to execute test operations under the operating system, building a complete test process from the hardware bottom layer to the operating system top layer. It can not only comprehensively detect the operating status of the server under test at different levels, but also flexibly expand the test content and methods according to actual needs to adapt to diverse test scenarios.

[0049] When the server under test crashes during the restart test, BMC can directly operate the CPLD to perform a hard restart of the server under test without relying on the operating system, and automatically resume the test process. This effectively overcomes the drawback of traditional script testing, which requires manual intervention after the system crashes, and greatly improves test efficiency and continuity.

[0050] By adopting an intelligent logging strategy, the traditional comprehensive logging method is abandoned. By setting reasonable logging trigger conditions, only key log information related to hardware failure is recorded. While ensuring that key information is not lost, the amount of log redundancy is greatly reduced, and the feasibility of large-scale restart testing and data processing efficiency are improved. Attached Figure Description

[0051] Figure 1 This is a structural diagram of the BMC-based service fault location system of the present invention;

[0052] Figure 2 This is a structural diagram of the BMC of the present invention;

[0053] Figure 3 This is a flowchart of the service fault location method based on BMC of the present invention;

[0054] Figure 4 This is a flowchart illustrating the process of handling abnormalities according to the anomaly level in this invention. Detailed Implementation

[0055] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0056] like Figure 1 As shown, this embodiment of the invention provides a service fault location system based on BMC, including:

[0057] BMC101 and server under test 102, server under test 102 includes CPLD1021 and M hardware components 1022; M is a positive integer greater than 2. Specifically, hardware components 1022 include CPU sensors, PSU power supply, chassis fan, PHY physical chip and other types of controllers; CPLD is a programmable logic device with non-volatile configuration memory, usually based on EEPROM technology, ready to use upon power-on, requiring no external configuration chip, and has deterministic latency, suitable for interface control scenarios with strict timing requirements;

[0058] BMC101 establishes a control link with CPLD1021 through the GPIO interface, establishes a communication link with the server under test 102 through the IPMI custom protocol, and connects with multiple hardware components 1022 through standard hardware interfaces.

[0059] Standard hardware interfaces are those that are currently known and include PHY interface, PWM interface, ADC interface, I2C bus interface, etc.

[0060] The Intelligent Platform Management Interface (IPMI) is an open standard for server hardware management that allows for remote monitoring and control of the server's physical status over a network, independent of the operating system (OS), and enables operations such as powering on / off and reinstalling the system. The core of IPMI is the Baseboard Management Controller (BMC), which is an independent subsystem that can still function normally even if the server's OS is not running.

[0061] The specific content of the IPMI custom protocol is as follows: When an anomaly occurs during the testing of the server under test, the BIOS boot information and operating system-level data at the time of the anomaly will be processed into an IPMI custom message through the pre-agreed IPMB custom protocol and actively pushed to the BMC.

[0062] BMC101 configures the test parameters and controls the power-on of the server under test 102 to start the restart test program; the test parameters include the number of restart test rounds, the maximum number of exceptions recorded, the exception data backtracking time limit, and the test requirement data collection rules;

[0063] The BMC101 collects real-time hardware-level data from multiple hardware components 1022 during the current test round through a standard hardware interface. The hardware-level data includes the voltage and current output status of the PSU power supply, CPU data (voltage, temperature and current data of the core sensors), onboard temperature sensor information, chassis temperature sensor information, fan status, and the status of specific GPIOs.

[0064] The BMC101 temporarily stores the current hardware-level data and determines whether any anomalies have occurred during the current round of testing based on the current hardware-level data. The temporary storage can be the BMC's built-in FLASH memory card or an external TF card.

[0065] When an anomaly occurs during the current round of testing, the BMC101 receives the current IPMI customized message through the communication link of the IPMI customized protocol, and obtains the anomaly level information based on the current IPMI customized message. The anomaly level information generally includes non-serious anomaly level and serious anomaly level.

[0066] Based on the anomaly level information, the CPLD1021 is controlled to perform a hard reboot or not on the server under test 102. It also obtains the anomaly's on-site data based on current hardware-level data and current IPMI customized messages, and persistently stores this on-site data. Persistent storage typically involves allocating a separate storage area in memory.

[0067] Based on the above Figure 1 The embodiments shown are preferred embodiments of the present invention, such as... Figure 2 As shown, BMC101 includes:

[0068] The data filtering module 201 is used to filter the current hardware-level data according to the data collection rules for test requirements to obtain the necessary test data;

[0069] The anomaly detection module 202 is used to determine whether the server under test 102 has encountered an anomaly during the current round of testing based on the necessary test data. The anomaly may be a startup failure, system crash, or other similar situations.

[0070] In this embodiment, it is shown that the data filtering module 201 can simplify the current hardware-level data by using the test requirement data collection rules, retaining only the necessary data for judging anomalies during the restart test process, thereby reducing the amount of data processing.

[0071] Preferably, in some embodiments of the present invention, such as Figure 2 As shown, BMC101 also includes:

[0072] IPMI interface 203 is used to receive the current IPMI customized message actively pushed by the server under test 102 through the communication link of the IPMI customized protocol when an abnormality occurs during the current round of testing. The current IPMI customized message includes the current BIOS boot information and the current operating system level data.

[0073] The abnormal data storage module 204 is used to record the time point of the abnormal occurrence, obtain the backtracking period before the time point of the abnormal occurrence based on the abnormal data backtracking time limit, extract the abnormal on-site data from the current hardware level data, the current BIOS boot information and the current operating system level data according to the backtracking period, and persistently store the on-site data.

[0074] In this embodiment, the current BIOS boot information, current operating system level data, and current hardware level data received by the IPMI interface 203 together constitute a multi-dimensional expression of abnormal situation. Furthermore, the on-site data is extracted according to the abnormal data backtracking time limit and persistently stored, so that the abnormal situation can be understood more completely according to the on-site data later.

[0075] Preferably, in some embodiments of the present invention, such as Figure 2 As shown, BMC101 also includes:

[0076] The anomaly classification module 205 is used to determine the anomaly level information based on the current BIOS boot information and the current operating system level data. The anomaly level information includes non-critical anomaly level and critical anomaly level.

[0077] The hard reboot control module 206 is used to generate control commands when the anomaly level information is a severe anomaly level, and control CPLD1021 to perform a hard reboot on the server 102 under test according to the control commands, so that the current round of testing can continue; when the anomaly level information is a non-severe anomaly level, no control commands are generated.

[0078] In this embodiment, the anomaly classification module 205 classifies anomalies and employs different handling methods for different anomaly levels. Non-critical anomalies will not interrupt the restart test process, therefore a hard restart is not required; only the current anomaly data needs to be persistently stored. Critical anomalies will cause the current round of restart testing to be interrupted, requiring a hard restart, and the current data must also be persistently stored.

[0079] Preferably, in some embodiments of the present invention, such as Figure 2 As shown, BMC101 also includes:

[0080] The exception count module 207 is used to record the total number of exceptions n in all rounds of testing before the current round of testing. When an exception occurs, the current total number of exceptions is n+1.

[0081] The test result summary module 208 is used to end the restart test program when the current total number of exceptions n+1 equals the maximum number of exceptions recorded N, to count the test information corresponding to the restart test program, and to generate a field data packet for all exceptions in the restart test program; N is a positive integer greater than 2;

[0082] The test round counting module 209 is used to continue executing the current round of testing until the end when the current total number of exceptions n+1 is less than the maximum number of exceptions recorded N, and to record the current round of testing as i.

[0083] The test result summary module 208 is also used to end the restart test program when the current test round i is equal to the restart test round number I, to count the test information corresponding to the restart test program and generate the field data packets of all abnormalities in the restart test program; when the current test round i is less than the restart test round number I, it controls the server under test to perform the next test round i+1; I is a positive integer greater than 2.

[0084] In this embodiment, during the restart test program, it is also necessary to count the number of exceptions and the number of test rounds. When the number of exceptions is equal to the maximum number of exceptions N, the restart test program ends. When the current test round i ends and the number of restart test rounds I is equal to the number of restart test rounds, the restart test program ends.

[0085] Preferably, in some embodiments of the present invention, such as Figure 2 As shown, BMC101 also includes:

[0086] Furthermore, BMC also includes:

[0087] Operating system entry judgment module 210 is used to determine whether the server under test has entered the operating system normally when no abnormality occurs during the current round of testing.

[0088] The abnormal data storage module 204 is also used to persistently store the on-site data if the server under test 102 fails to enter the operating system normally.

[0089] The extended testing module 211 is used to control the server under test 102 to mount virtual media and execute a preset test script under the operating system if the server under test 102 successfully enters the operating system, in order to simulate the current round of testing. During the simulation of the current round of testing, it is also necessary to collect test data and determine whether there are any abnormalities during the simulation.

[0090] In this embodiment, after the restart test program is started, if the startup is successful, the operating system entry judgment module 210 needs to determine whether the operating system has been entered normally. If the operating system has not been entered, it indicates a serious abnormality level, requiring a hard restart. The abnormal data storage module 204 will persistently store the on-site data. If the operating system has been entered normally, the extended test module 211 can control the server under test 102 to mount virtual media, that is, mount the storage medium of the test script designed by the tester to the server under test, and execute the preset test script under the operating system to simulate the current round of testing. It can simulate various actual operating scenarios, such as high-load computing and large data transmission, thereby improving the completeness and adaptability of the test.

[0091] The above embodiments describe a service fault location system based on BMC. The following embodiments illustrate a service fault location method based on BMC applied to the service fault location system based on BMC.

[0092] like Figure 3 As shown, this embodiment of the invention provides a service fault location method based on BMC, including:

[0093] 301. Configure test parameters and power on the server under test to start the restart test program;

[0094] 302, collects real-time hardware-level data of multiple hardware components in the current round of testing through standard hardware interfaces;

[0095] 303, temporarily store the current hardware-level data, and determine whether an anomaly occurred during the current round of testing based on the current hardware-level data;

[0096] If an exception occurs during the current round of testing, proceed to step 304; if no exception occurs, continue to step 302.

[0097] 304, Receive the current IPMI custom message through the communication link of the IPMI custom protocol;

[0098] 305. Based on the current IPMI customized message, obtain the abnormality level information, and control the CPLD to perform a hard reboot or not perform a hard reboot on the server under test according to the abnormality level information.

[0099] 306. Based on the current hardware-level data and the current IPMI customized message, abnormal field data is obtained and the field data is persistently stored.

[0100] Based on the above description of the BMC-based service fault location system, combined with Figure 3 The BMC-based service fault location method shown is, preferably, in some embodiments of the present invention, such as... Figure 4 As shown, this implementation is a process of handling different cases according to the anomaly level, including:

[0101] 401. Configure test parameters and power on the server under test to start the restart test program;

[0102] 402, collects real-time hardware-level data of multiple hardware components in the current round of testing through standard hardware interfaces;

[0103] 403. Based on the test requirement data collection rules, the current hardware-level data is filtered to obtain the necessary test data;

[0104] 404 indicates that necessary test data is temporarily stored, and the test data is used to determine whether the server under test has encountered any abnormalities during the current round of testing.

[0105] If an exception occurs during the current round of testing, proceed to step 405; if no exception occurs, continue to step 402.

[0106] 405. Receive the current IPMI custom message through the communication link of the IPMI custom protocol. The current IPMI custom message includes the current BIOS boot information and the current operating system level data.

[0107] 406, record the time point of the anomaly occurrence, obtain the backtracking period before the anomaly occurrence time point based on the backtracking time limit of the anomaly data, and extract the anomaly's on-site data from the current hardware level data, the current BIOS boot information and the current operating system level data according to the backtracking period;

[0108] 407 indicates the level of the anomaly based on the current BIOS boot information and current operating system level data.

[0109] The anomaly level information includes non-severe anomaly level and severe anomaly level. If it is a non-severe anomaly level, proceed to step 408; if it is a severe anomaly level, proceed to step 409.

[0110] 408. Persistently store the field data;

[0111] 409. Generate control commands, and control the CPLD to perform a hard reboot on the server under test according to the control commands, and persistently store the field data.

[0112] Based on the above description of the BMC-based service fault location system and method, this application achieves the following beneficial effects:

[0113] Using BMC and the standard sensors of the server under test, rich and detailed on-site data collection can be completed without the need for additional testing equipment, server hardware modification, or additional data collection equipment, thus reducing testing costs.

[0114] The server under test can be any general-purpose server that supports BMC, which means that there is no need to make large-scale modifications to the server hardware or software, and it has wide applicability.

[0115] BMC establishes a customized IPMI protocol with the server under test, enabling it to obtain abnormal status information actively pushed by the BIOS and operating system of the server under test, and thus accurately record the abnormal situation. BMC can control the mounting of media on the server under test and use scripts to execute test operations under the operating system, building a complete test process from the hardware bottom layer to the operating system top layer. It can not only comprehensively detect the operating status of the server under test at different levels, but also flexibly expand the test content and methods according to actual needs to adapt to diverse test scenarios.

[0116] When the server under test crashes during the restart test, BMC can directly operate the CPLD to perform a hard restart of the server under test without relying on the operating system, and automatically resume the test process. This effectively overcomes the drawback of traditional script testing, which requires manual intervention after the system crashes, and greatly improves test efficiency and continuity.

[0117] By adopting an intelligent logging strategy, the traditional comprehensive logging method is abandoned. By setting reasonable logging trigger conditions, only key log information related to hardware failure is recorded. While ensuring that key information is not lost, the amount of log redundancy is greatly reduced, and the feasibility of large-scale restart testing and data processing efficiency are improved.

[0118] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0119] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0122] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of the claims of the present invention pending approval.

Claims

1. A service fault location system based on BMC, characterized in that, include: BMC and server under test, wherein the server under test includes CPLD and multiple hardware components; The BMC establishes a control link with the CPLD, and the BMC establishes a communication link with the server under test using the IPMI custom protocol. The BMC is connected to multiple hardware components through standard hardware interfaces. The IPMI custom protocol is used to actively push IPMI custom messages containing BIOS boot information and operating system-level data to the BMC when the server under test malfunctions. The BMC configures the test parameters and controls the power-on of the server under test to start the restart test program; The BMC collects real-time hardware-level data of multiple hardware components in the current round of testing through the standard hardware interface. The BMC temporarily stores the current hardware-level data and determines whether an anomaly occurs during the current round of testing based on the current hardware-level data. When an anomaly occurs during the current round of testing, the BMC receives the current IPMI customized message through the communication link of the IPMI customized protocol, obtains the anomaly level information based on the current IPMI customized message, controls the CPLD to perform a hard reboot or not perform a hard reboot on the server under test based on the anomaly level information, and obtains the anomaly's on-site data based on the current hardware-level data and the current IPMI customized message, and persistently stores the on-site data.

2. The service fault location system based on BMC according to claim 1, characterized in that, The test parameters include the number of test rounds to restart, the maximum number of exceptions to record, the exception data backtracking time limit, and the test requirement data collection rules.

3. The service fault location system based on BMC according to claim 2, characterized in that, The BMC includes: The data filtering module is used to filter the current hardware-level data according to the test requirement data collection rules to obtain the necessary test data; The anomaly detection module is used to determine whether the server under test has encountered an anomaly during the current round of testing, based on the necessary test data.

4. The service fault location system based on BMC according to claim 2, characterized in that, The BMC also includes: The IPMI interface is used to receive the current IPMI customized message actively pushed by the server under test through the communication link of the IPMI customized protocol when an anomaly occurs during the current round of testing. The current IPMI customized message includes the current BIOS boot information and the current operating system level data. An abnormal data storage module is used to record the time point of occurrence of the abnormality, obtain the backtracking period before the time point of occurrence of the abnormality based on the abnormal data backtracking time limit, extract the on-site data of the abnormality from the current hardware level data, the current BIOS boot information and the current operating system level data according to the backtracking period, and persistently store the on-site data.

5. The service fault location system based on BMC according to claim 4, characterized in that, The BMC also includes: An anomaly classification module is used to determine the anomaly level information of the anomaly based on the current BIOS boot information and the current operating system level data. The anomaly level information includes a non-serious anomaly level and a serious anomaly level. The hard reboot control module is used to generate a control command when the anomaly level information is the severe anomaly level, and control the CPLD to perform a hard reboot on the server under test according to the control command, so that the current round of testing can continue; when the anomaly level information is the non-severe anomaly level, no control command is generated.

6. The service fault location system based on BMC according to claim 5, characterized in that, The BMC also includes: The exception count module is used to record the total number of exceptions n in all rounds of testing before the current round of testing. When the exception occurs, the current total number of exceptions is n+1. The test result summary module is used to end the restart test program when the current total number of anomalies n+1 is equal to the maximum number of recorded anomalies N, to collect test information corresponding to the restart test program, and to generate a field data packet for all anomalies in the restart test program. The test round counting module is used to continue executing the current round of testing until the end when the current total number of anomalies n+1 is less than the maximum number of recorded anomalies N, and to record the current round of testing as i. The test result aggregation module is also used to terminate the restart test program, collect test information corresponding to the restart test program, and generate all abnormal field data packets in the restart test program when the current test round i is equal to the restart test round number I; and to control the server under test to perform the next test round i+1 when the current test round i is less than the restart test round number I.

7. The service fault location system based on BMC according to claim 6, characterized in that, The BMC also includes: The operating system entry judgment module is used to determine whether the server under test has entered the operating system normally when no abnormality occurs during the current round of testing. The abnormal data storage module is also used to persistently store the on-site data if the server under test fails to enter the operating system normally. An extended testing module is used to control the server under test to mount virtual media and execute a preset test script under the operating system if the server under test enters the operating system normally, in order to simulate the current round of testing.

8. A service fault location method based on BMC, characterized in that, The service failure localization method, applied to the BMC-based service failure localization system according to any one of claims 1-7, comprises: Configure the test parameters, control the power-on of the server under test, and start the restart test program; Real-time data collection at the current hardware level from multiple hardware components in the current round of testing is achieved through standard hardware interfaces; The current hardware-level data is temporarily stored, and an anomaly is determined based on the current hardware-level data during the current round of testing. When an anomaly occurs during the current round of testing, the current IPMI customized message is received through the communication link of the IPMI customized protocol; The anomaly level information is obtained based on the current IPMI customized message, and the CPLD is controlled to perform a hard reboot or not perform a hard reboot on the server under test based on the anomaly level information. The abnormal field data is obtained based on the current hardware-level data and the current IPMI customized message, and the field data is persistently stored.

Citation Information

Patent Citations

  • Loop restart testing device and method for server

    CN116701074A

  • Server hardware equipment monitoring method and device, equipment and storage medium

    CN118796608A