Fault detection method and system, electronic equipment and storage medium
The fault detection device connects the bus on the server motherboard to automatically judge the periodic data integrity of the bus signal and detect the fault type, solving the problems of low fault detection efficiency and poor accuracy in the existing technology, and achieving efficient fault data capture and automatic fault classification.
Patent Information
- Application Number
- CN202510081888.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The prior art has low testing efficiency and poor accuracy in fault detection, and has high interface wear of components and equipment.
The pin of the fault detection device is connected to the pinhole on the server motherboard, and the bus to be tested between the management controller and the sub-device is connected. The data integrity is judged based on the periodic identification of the bus signal, the fault data is determined and the fault type is automatically detected.
It improves the efficiency of the acquisition of fault data, realizes automatic classification of fault types, shortens manual detection and analysis time, and improves the efficiency of fault detection.
Smart Images

Figure CN119938426A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a fault detection method and system, an electronic device, and a storage medium. Background Art
[0002] Peripheral Component Interconnect Express (PCIe) devices are connected to the baseboard management controller (BMC) in the server through a bus to achieve out-of-band management. Since there are many types and manufacturers of PCIe devices and firmware is upgraded from time to time, PCIe device bus hangs and other fault problems often occur.
[0003] At present, the related technology usually welds test points on the bus corresponding to the faulty PCIe device, and requires manual monitoring of the logic analyzer to capture error data, resulting in low fault detection efficiency. Summary of the invention
[0004] The present disclosure provides a fault detection method and system, an electronic device and a storage medium, which are mainly intended to solve the problems of low test efficiency and poor test accuracy of fault detection and high wear of the interface of component devices.
[0005] In view of this, the present application provides a fault detection method, system, storage medium and electronic device, the main purpose of which is to improve the technical problem that the current related technology usually welds measurement points on the bus corresponding to the faulty PCIe device, requires manual monitoring of the logic analyzer to capture error data, and has low fault detection efficiency.
[0006] In a first aspect, the present application provides a fault detection method, which is performed by a fault detection device, wherein a pin of the fault detection device is connected to a pin hole installed on a server mainboard, and there is at least one bus between a management controller and a sub-device of the server mainboard, and the pin hole is connected to the at least one bus, comprising:
[0007] Accessing a bus to be tested in the at least one bus through the pin hole;
[0008] Determining whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested;
[0009] If it is detected that the periodic data is incomplete, the periodic data is determined to be fault data, and the fault type of the bus to be tested is detected according to the fault data.
[0010] In some embodiments, detecting the fault type of the bus to be tested according to the fault data includes:
[0011] Detecting the state of the bus to be tested according to the fault data;
[0012] If it is detected that the bus to be tested is in an idle state, determining that the fault type is a first fault, wherein the first fault includes a task interruption caused by a master device switching;
[0013] If it is detected that the bus to be tested is not in an idle state, detecting a bus signal at a low level;
[0014] If it is detected that the bus signal at a low level is a data signal, the bus switch is used to reset and detect whether the front-stage circuit of the bus switch is restored;
[0015] If it is detected that the previous line is restored, determining that the fault type is a second fault, wherein the second fault includes a sub-device fault;
[0016] If it is detected that the previous-stage line has not been restored, it is determined that the fault type is a third fault, and the third fault includes a management controller fault.
[0017] In some embodiments, detecting the fault type of the bus to be tested according to the fault data further includes:
[0018] If it is detected that the bus signal at a low level is a clock signal, the bus switch is used for resetting, and whether the front-stage circuit of the bus switch is restored is detected;
[0019] If it is detected that the previous line is restored, the fault type is determined to be a fourth fault, and the fourth fault includes that the sub-device does not release the bus normally;
[0020] If it is detected that the previous line has not been restored, it is determined that the fault type is a fifth fault, and the fifth fault includes that the management controller does not normally release the bus.
[0021] In some embodiments, after determining that the fault type is a second fault if the previous line is detected to be restored, the method further includes:
[0022] Reset signals are sent to the sub-devices connected to the bus switch in sequence, sub-devices that have not been successfully restored are determined as target faulty devices, and the target faulty devices are recorded.
[0023] In some embodiments, before resetting the bus switch and detecting whether the front-stage line of the bus switch is restored, the method further includes:
[0024] Detecting whether there is a bus switch on the bus to be tested;
[0025] If it is detected that no bus switch exists on the bus to be tested, a reset signal is directly sent to the sub-device corresponding to the bus to be tested.
[0026] In some embodiments, after determining that the periodic data is fault data, the method further includes:
[0027] The fault data is saved.
[0028] In some embodiments, after determining whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested, the method further includes:
[0029] If it is detected that the periodic data is complete, the periodic data is determined to be normal data, and the normal data is discarded.
[0030] In a second aspect, the present application provides a fault detection system, including: a pinhole and a fault detection device;
[0031] The pin hole is installed on the server mainboard, and there is at least one bus between the management controller of the server mainboard and the sub-device, and the pin hole is connected to the at least one bus;
[0032] The pin hole is used to connect the pin of the fault detection device;
[0033] The fault detection device is used to access the bus to be tested in the at least one bus through the pin hole; determine whether the cycle data of the bus signal is complete based on the cycle identifier of the bus signal of the bus to be tested; if the cycle data is detected to be incomplete, determine that the cycle data is fault data, and detect the fault type of the bus to be tested based on the fault data.
[0034] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0035] at least one processor; and
[0036] a memory communicatively connected to the at least one processor; wherein,
[0037] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the first aspect.
[0038] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the first aspect.
[0039] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method as described in the first aspect above.
[0040] The fault detection method and system, electronic device and storage medium provided by the present disclosure are executed by a fault detection device, wherein the pin of the fault detection device is connected to the pin hole installed on the server mainboard, and there is at least one bus between the management controller and the sub-device of the server mainboard, and the pin hole is connected to the at least one bus, wherein the method comprises: accessing a bus to be tested in the at least one bus through the pin hole; judging whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested; if the cycle data is detected to be incomplete, determining the cycle data as fault data, and detecting the fault type of the bus to be tested according to the fault data. Compared with the current prior art, the present application can use the fault detection device to connect the bus to be tested between the management controller and the sub-device by installing the pin hole of the server mainboard, and then determine the incomplete cycle data as fault data according to the cycle identifier of the bus signal of the bus to be tested, and then determine the fault type of the bus to be tested according to the fault data, so that there is no need to manually capture the fault data of the bus to be tested, which effectively improves the capture efficiency of the fault data, and at the same time realizes the automatic classification of the fault type, shortens the manual detection and analysis time, and thus improves the efficiency of fault detection.
[0041] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0044] Figure 1 A schematic diagram showing an example provided by an embodiment of the present application;
[0045] Figure 2 A schematic diagram showing an example provided by an embodiment of the present application;
[0046] Figure 3A schematic diagram of a process flow of a fault detection method provided in an embodiment of the present application is shown;
[0047] Figure 4 A schematic diagram showing an example provided by an embodiment of the present application;
[0048] Figure 5 A structural schematic diagram of a fault detection system provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0049] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0050] At present, in the related art, the BMC in the server motherboard and each sub-device can be connected by a two-wire serial bus (Inter-Integrated Circuit, I2C), such as Figure 1 As shown, BMC can be connected to multiple buses, such as I2C0, I2C1, I2Cn, etc. Complex Programmable Logic Device (FPGA) can be used to receive data from each bus and send a reset signal (reset) to reset the sub-devices connected to the bus. If there are fewer devices hanging down from a certain I2C bus, and the addresses of each hanging down device are different, there is no need for an I2C switch to control the hanging down devices; if there are more devices hanging down from the I2C bus, and the addresses of the devices hanging down from the I2C bus are the same, an I2C switch can be used to control the hanging down devices. Among them, the chip of the I2C switch (I2CSWITCH) (such as the PCA9548 chip, etc.) can be used to try to restore the corresponding sub-device through the reset signal.
[0051] In specific application scenarios, if the server architecture has I2C hang or other fault problems, the following are usually used: Figure 2 The fault detection method is to weld test points on the faulty I2C bus, connect the logic analyzer, and perform fault detection through the logic analyzer. However, due to the limited storage space of the logic analyzer, it is impossible to store a large amount of bus data (such as waveform data, etc.). It is necessary to rely on manual use of the logic analyzer to monitor the bus data, capture the error data during the bus transmission process, and reproduce and analyze the problem based on the error data, resulting in low error data capture efficiency, which in turn affects the efficiency of fault reproduction and analysis, and consumes a lot of labor costs.
[0052] In order to improve the current related technology, the test points are usually welded on the bus corresponding to the PCIe device with fault, and the manual monitoring logic analyzer is required to capture the error data, resulting in low fault detection efficiency. This embodiment provides a fault detection method, such as Figure 3 As shown, the method includes:
[0053] Step 101: access a bus to be tested in at least one bus through a pin hole.
[0054] In a specific application scenario, the executor of this embodiment may be a fault detection device, the pins of the fault detection device are connected to the pin holes installed on the server motherboard, there is at least one bus between the management controller of the server motherboard and the sub-device, and the pin holes are connected to at least one bus.
[0055] In some examples, the pinhole can be installed near the management controller of the server motherboard, and the management controller can include but is not limited to the BMC, and the wiring length is shortened as much as possible to avoid unnecessary bends and excessively long paths to reduce signal attenuation and interference. Specifically, the pinhole can be used to lead out the bus signal and connect it to the fault detection device. When a fault occurs, the cable can be directly connected to the pin, and the cable can be connected to the fault detection device outside the chassis, without the need to use welding to connect the bus, thereby improving the bus connection efficiency.
[0056] In some examples, the server, as the core device of the information system, has the characteristics of high storage performance and strong input / output (I / O) expansion capability. When a sub-device in the server architecture is found to have a fault, the bus to be tested corresponding to the faulty sub-device can be obtained first, and then the pin of the fault detection device can be connected to the pin hole corresponding to the bus to be tested, so that the fault detection device can be used to connect to the bus to be tested, obtain the data of the bus to be tested, and capture the fault data, thereby improving the efficiency of the fault detection device accessing the bus to be tested.
[0057] In some examples, if multiple sub-devices fail and correspond to different buses, you can first obtain multiple buses to be tested corresponding to the multiple sub-devices, and connect the multiple pins of the fault detection device to the pin holes corresponding to each bus to be tested. The fault detection device can then be used to simultaneously connect the multiple buses to be tested, and fault detection can be performed on the multiple buses to be tested, effectively improving the fault detection efficiency while reducing the cost of manual detection.
[0058] Among them, sub-devices may include but are not limited to PCIe devices, such as: Non Volatile Memory Host Controller Interface Specification (NVME) hard disk devices, Serial Attached SCSI (SAS), Redundant Arrays of Independent Disks (RAID) cards, PCI Express network interface cards, Graphics Processing Unit (GPU) cards, GPU modules, etc., which can be used to realize the main functions of the server such as storage, computing, and communication. The PCIe bus of the PCIe device can be connected to the Central Processing Unit (CPU) to quickly interact with the CPU for data.
[0059] In some examples, in order to ensure stable and reliable operation of PCIe devices, BMC, as an independent system unique to the server with monitoring, control and management functions, usually establishes a connection with a device that supports the MCTP protocol (MCTP device) through the Management Component Transport Protocol (MCTP) over I2C, that is, transmits MCTP messages through the I2C bus, so that the user end can manage and monitor the MCTP device through the BMC.
[0060] Step 102: judging whether the periodic data of the bus signal is complete according to the periodic identifier of the bus signal of the bus to be tested.
[0061] In some examples, a fault detection device can be used to read the bus signal of the bus to be tested in real time, and then determine the boundary of each cycle according to the cycle identifier (such as start bit, stop bit, synchronization character, etc.) defined by the bus protocol, parse the cycle data, and determine whether each cycle data is complete. Exemplarily, it is also possible to determine whether each cycle data is complete by comparing the expected cycle data length, checksum, data format and other information.
[0062] Correspondingly, different bus protocols have different cycle identification methods. For the I2C bus, the start bit (Start) of the cycle data is when the serial clock line (SCL) is high, the serial data line (SDA) changes from high to low, and the stop bit (Stop) is when SCL is high, SDA changes from low to high. The cycle identifier can be identified based on the corresponding signal level changes. When a complete start bit and stop bit are detected, it can be determined as a complete cycle data.
[0063] Step 103: If it is detected that the periodic data is incomplete, the periodic data is determined to be fault data, and the fault type of the bus to be tested is detected according to the fault data.
[0064] In some examples, if the cycle identifiers such as the start bit and stop bit of the bus signal do not appear correctly, or if bus signal abnormalities occur, it can be determined that the cycle data is incomplete. For example, if the start bit of the bus signal is detected, but the stop bit corresponding to the start bit is not detected, it can be determined that the cycle data is incomplete. The cycle data can be determined as fault data, and then the fault type of the bus to be tested is analyzed based on the fault data, thereby realizing automatic classification of fault types, shortening the manual detection and analysis time of the bus to be tested, and improving the fault detection efficiency of the bus to be tested.
[0065] Compared with the current prior art, this embodiment can be executed by a fault detection device, the pin of the fault detection device is connected to the pin hole installed on the server motherboard, there is at least one bus between the management controller of the server motherboard and the sub-device, and the pin hole is connected to the at least one bus, wherein the method includes: accessing the bus to be tested in at least one bus through the pin hole; judging whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested; if the cycle data is detected to be incomplete, the cycle data is determined to be fault data, and the fault type of the bus to be tested is detected according to the fault data. Compared with the current prior art, this embodiment can use the fault detection device to connect the bus to be tested between the management controller and the sub-device by installing the pin hole of the server motherboard, and then determine the incomplete cycle data as fault data according to the cycle identifier of the bus signal of the bus to be tested, and then determine the fault type of the bus to be tested according to the fault data, so that there is no need to manually capture the fault data of the bus to be tested, which effectively improves the capture efficiency of the fault data, and at the same time realizes the automatic classification of the fault type, shortens the manual detection and analysis time, and thus improves the efficiency of fault detection.
[0066] In some embodiments, detecting the fault type of the bus to be tested according to fault data may specifically include: detecting the state of the bus to be tested according to the fault data; if it is detected that the bus to be tested is in an idle state, determining the fault type to be a first fault, the first fault including a task interruption caused by a master device switching; if it is detected that the bus to be tested is not in an idle state, detecting a bus signal at a low level; if it is detected that the bus signal at a low level is a data signal, resetting the bus switch to detect whether the front-stage line of the bus switch has been restored; if it is detected that the front-stage line has been restored, determining the fault type to be a second fault; if it is detected that the front-stage line has not been restored, determining the fault type to be a third fault.
[0067] The first fault may include but is not limited to task interruption caused by master device switching, the second fault may include but is not limited to slave device failure, and the third fault may include but is not limited to manager failure. Accordingly, the bus switch can be used to control the signal path, extend or isolate different bus segments, and dynamically select different channels so that the master device can communicate with multiple slave devices without conflict, and may include but is not limited to I2C SWITCH, SPI switch, GPIO expander, etc.
[0068] Optionally, a reset signal may be sent to a sub-device corresponding to the bus to be tested through a logic device in the fault detection apparatus or a logic device in the server, so as to restore bus communication in a variety of ways.
[0069] For example, Figure 4 As shown, the fault type detection process is shown, and the fault type can be classified and recorded in combination with the state of the bus to be tested and the bus signal. Specifically, if it is detected that a complete cycle ends normally, the next cycle no longer appears completely, and the bus to be tested is in an idle state, the incomplete cycle data can be saved, and the fault type is determined and recorded as the first fault, such as recording the error type as 1.
[0070] For the first fault, R&D personnel can intervene to analyze the master device of the last cycle data, for example, whether the BMC is the master device or the slave device is the master device. After determining the master device, analyze the fault reason why the bus is idle but still does not send data. If the fault reason is that the device that has switched to the master has not sent a complete clock cycle, the specific reason for the interruption of the I2C task of the device that has switched to the master can be analyzed in combination with the device log.
[0071] Correspondingly, if it is detected that the bus to be tested is not in an idle state, the bus signal is further detected to determine whether SDA is pulled low, and whether the clock signal is pulled low or the data signal is pulled low; if it is detected that the data signal is pulled low, and there is an I2C SWITCH link in the entire I2C link of the bus to be tested, the data can be saved and the reset I2C SWITCH command can be executed, and then by detecting whether the BMC side can restore normal access, it is determined whether the I2CSWITCH front-stage line is restored;
[0072] Correspondingly, if it is detected that the front-stage circuit of I2C SWITCH is restored and the BMC side can restore normal access, the BMC can be used to re-read each device. If there is a device that cannot be read successfully, it can be determined that the device is hung, and the fault type is recorded as the second fault, such as the recorded error type is 2; if it is detected that the front-stage circuit of I2C SWITCH is not restored, the BMC re-reads each device and still cannot read successfully, it can be determined that the BMC is hung, and the fault type is recorded as the third fault, such as the recorded error type is 3. When the R&D personnel receive the fault type mark, they can try to reset the I2C controller module in the BMC to check whether it can be recovered.
[0073] In some embodiments, the fault type of the bus to be tested is detected based on fault data, which may specifically include: if the bus signal at a low level is detected to be a clock signal, the bus switch is used to reset, and whether the front-stage line of the bus switch is restored is detected; if the front-stage line is detected to be restored, the fault type is determined to be the fourth fault; if the front-stage line is detected to be not restored, the fault type is determined to be the fifth fault.
[0074] The fourth fault may include but is not limited to the failure of the sub-device to release the bus normally, and the fifth fault may include but is not limited to the failure of the management controller to release the bus normally.
[0075] Exemplarily, if the clock signal is pulled low and there is an I2C SWITCH link fast positioning module in the bus to be tested, it is necessary to save the data and execute the reset I2C SWITCH command, and the BMC re-reads other devices except the problematic device. If other devices can be read successfully and the problematic device cannot be read successfully, it is determined that the problematic device is hung due to failure to release the bus normally (such as bus preemption failure but not releasing the bus), and the fault type is recorded as the fourth fault, such as the recorded error type is 4; if the BMC re-reads other devices and still cannot read successfully, it is determined that the BMC is hung due to failure to release the bus normally, and the fault type is recorded as the fifth fault, such as the recorded error type is 5.
[0076] In this way, the fault type of the bus to be tested can be simply classified according to the fault data, so as to notify the R&D personnel corresponding to the fault type to perform fault analysis and reproduction, and provide the R&D personnel with fault analysis guidance, thereby improving the efficiency of R&D personnel in fault reproduction and thus improving the efficiency of fault solving.
[0077] In some embodiments, if the recovery of the previous line is detected, after determining that the fault type is the second fault, the method of this embodiment may also include: sending reset signals to the sub-devices connected to the bus switch in sequence, determining the sub-devices that have not recovered successfully as target faulty devices, and recording the target faulty devices.
[0078] The target faulty device may include at least one sub-device that has failed.
[0079] Specifically, when the front-end line is restored, a reset signal can be sent to multiple sub-devices in sequence through the bus switch to check whether the sub-devices can restore normal communication. If there are sub-devices that cannot be restored successfully, they can be identified as target faulty devices, thereby pointing out the specific target faulty devices and notifying the relevant personnel of the target faulty devices to handle the fault.
[0080] In some embodiments, before using the bus switch to reset and detecting whether the front-stage line of the bus switch is restored, the method of this embodiment may also include: detecting whether there is a bus switch on the bus to be tested; if it is detected that there is no bus switch on the bus to be tested, directly sending a reset signal to the sub-device corresponding to the bus to be tested.
[0081] For example, the addresses of each sub-device can be obtained based on the sub-devices connected to the bus to be tested, and it can be determined whether a bus switch is required to control multiple sub-devices with the same address. If there is no bus switch on the bus to be tested, a reset signal can be directly sent to the sub-device corresponding to the bus to be tested through the FPGA in the fault detection device or the FPGA in the server to restore the communication of the bus.
[0082] In some embodiments, after determining that the periodic data is fault data, the method of this embodiment may further include: saving the fault data.
[0083] Exemplarily, the fault detection device can save the fault data in the memory when it detects that the periodic data is fault data. If the relevant personnel do not have time to collect the fault data in time, they can read the fault data in the memory later and then reproduce and analyze the fault, providing convenience for the relevant personnel.
[0084] In this way, the present embodiment can automatically identify fault data and save it, without the need for personnel on duty to capture fault data, so that corresponding R&D personnel can quickly reproduce the debug version more specifically, thereby quickly reproducing the problem and quickly solving the problem.
[0085] In some embodiments, after determining whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested, the method of this embodiment may further include: if the cycle data is detected to be complete, determining that the cycle data is normal data and discarding the normal data.
[0086] Exemplarily, the fault detection device can determine whether to save the cycle data based on a basic judgment of whether the cycle data is complete. When accessed normally, the cycle can be determined as normal data and does not need to be saved in the memory of the fault detection device, thereby saving storage space of the fault detection device and further reducing the impact of memory limitations on fault data capture.
[0087] In a specific application scenario, after using the fault detection device to start data reading, it is possible to first determine whether there are complete START bits and STOP bits. When the periodic data corresponding to the complete START bits and STOP bits can be captured, another clock cycle can be captured. If the periodic data corresponding to an entire clock cycle can still be captured, the previous periodic data can be discarded until abnormal data (fault data) appears, and the fault data is saved in the memory. In this way, by judging the complete data cycle, the loss of abnormal data can be avoided, and memory consumption can be reduced by not saving normal data.
[0088] Compared with the prior art, the present embodiment can use the fault detection device to connect the bus to be tested between the management controller and the sub-device by installing the pin hole of the server motherboard, and then determine the incomplete cycle data as fault data according to the cycle identification of the bus signal of the bus to be tested, and then determine the fault type of the bus to be tested according to the fault data, so that there is no need to manually capture the fault data of the bus to be tested, which effectively improves the capture efficiency of the fault data, and realizes the automatic classification of the fault type, shortens the manual detection and analysis time, and thus improves the efficiency of fault detection. In addition, it can also automatically identify the fault type of the bus to be tested according to the judgment conditions such as whether the bus to be tested is idle, the bus signal is at a low level, and whether the front-stage line of the bus switch is restored, and save the fault data and discard the normal data, so as to notify the R&D personnel corresponding to the fault type to perform fault analysis and reproduction, and provide the R&D personnel with the analysis direction of the fault, without the need for personnel to be on duty to capture the fault data, improve the efficiency of the R&D personnel to reproduce the fault, and thus improve the efficiency of fault resolution, and at the same time save the storage space of the fault detection device, thereby reducing the impact of memory limitations on fault data capture.
[0089] In order to further illustrate the specific implementation process of the method of this embodiment, this embodiment provides the following Figure 5 The system shown comprises: a pin hole 11 and a fault detection device 21;
[0090] The pin hole 11 is installed on the server mainboard. There is at least one bus between the management controller of the server mainboard and the sub-device, and the pin hole 11 is connected to the at least one bus.
[0091] The pin hole 11 is used to connect the pin of the fault detection device 21;
[0092] The fault detection device 21 is used to access the bus to be tested in at least one bus through the pin hole 11; based on the cycle identifier of the bus signal of the bus to be tested, determine whether the cycle data of the bus signal is complete; if the cycle data is detected to be incomplete, determine that the cycle data is fault data, and detect the fault type of the bus to be tested based on the fault data.
[0093] Among them, the pin hole 11 (pin) can be installed near the management controller (such as BMC) of the server motherboard to shorten the wiring length as much as possible, thereby reducing signal attenuation and interference. The pin hole 11 can be used to access the fault detection device 21 (quick positioning module), perform fault detection on the bus to be tested, receive the bus signal, determine whether the cycle data is complete according to the cycle identifier of the bus signal, determine the incomplete data as fault data, and then perform fault type detection according to the fault data, thereby improving the convenience of connecting the bus and improving the efficiency of fault detection.
[0094] Accordingly, the bus may include a bus of a standard frame format, such as I2C, UART, RS485, etc., and the bus to be tested may include one or more buses that have a fault. When multiple buses are involved, the pinhole 11 may include a plurality of independent pinholes 11, each of which corresponds to a different bus line or signal line. The appropriate type of pinhole 11 may be selected according to the requirements of the bus, such as a single-row pinhole, a double-row pinhole, etc., and in the printed circuit board (PCB) layout stage, each bus is assigned independent pins for wiring to ensure that there is sufficient spacing between the bus signal lines to avoid signal interference.
[0095] Compared with the current existing technology, this embodiment can use the fault detection device 21 to connect the bus to be tested between the management controller and the sub-device by installing the pin hole 11 of the server motherboard, and then determine the incomplete cycle data as fault data according to the cycle identifier of the bus signal of the bus to be tested, and then determine the fault type of the bus to be tested according to the fault data, thereby eliminating the need to manually capture the fault data of the bus to be tested, effectively improving the efficiency of capturing fault data, and at the same time realizing automatic classification of fault types, shortening the manual detection and analysis time, and thereby improving the efficiency of fault detection.
[0096] Based on the above Figure 3The method shown in the embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned Figure 3 The method shown.
[0097] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0098] Based on the above Figure 3 In order to achieve the above-mentioned purpose, the embodiment of the present application also provides an electronic device, such as a personal computer, a server, which includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to achieve the above-mentioned Figure 3 The method shown.
[0099] In some embodiments, the above-mentioned physical device may also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface may include a display, an input unit such as a keyboard, etc., and the optional user interface may also include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.
[0100] Those skilled in the art will appreciate that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different arrangements of components.
[0101] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical device, and supports the operation of the information processing program and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing physical device.
[0102] Through the description of the above implementation methods, the technical personnel in this field can clearly understand that the present application can be implemented by means of software plus the necessary general hardware platform, or by hardware. By applying the scheme of this embodiment, compared with the current prior art, this embodiment can use the fault detection device to connect the bus to be tested between the management controller and the sub-device by installing the pin hole of the server motherboard, and then determine the incomplete cycle data as fault data according to the cycle identification of the bus signal of the bus to be tested, and then determine the fault type of the bus to be tested according to the fault data, so that there is no need to manually capture the fault data of the bus to be tested, which effectively improves the capture efficiency of the fault data, and at the same time realizes the automatic classification of the fault type, shortens the manual detection and analysis time, and thus improves the efficiency of fault detection. In addition, it can also automatically identify the fault type of the bus to be tested based on judgment conditions such as whether the bus to be tested is idle, the bus signal is at a low level, and whether the front-stage line of the bus switch has recovered, and save the fault data and discard normal data, so as to notify the R&D personnel corresponding to the fault type to perform fault analysis and reproduction, and provide R&D personnel with fault analysis directions. There is no need for personnel to be on duty to capture fault data, which improves the efficiency of R&D personnel in fault reproduction, thereby improving the efficiency of fault solving, and at the same time saving storage space for the fault detection device, thereby reducing the impact of memory limitations on fault data capture.
[0103] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0104] The above is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features applied for herein.
Claims
1. A fault detection method, characterized in that: The method is performed by a fault detection device, wherein a pin of the fault detection device is connected to a pin hole installed on a server mainboard, and at least one bus is provided between a management controller of the server mainboard and a sub-device, and the pin hole is connected to the at least one bus. The method comprises: Accessing a bus to be tested in the at least one bus through the pin hole; Determining whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested; If it is detected that the periodic data is incomplete, the periodic data is determined to be fault data, and the fault type of the bus to be tested is detected according to the fault data.
2. The method according to claim 1, characterized in that The detecting the fault type of the bus to be tested according to the fault data comprises: Detecting the state of the bus to be tested according to the fault data; If it is detected that the bus to be tested is in an idle state, determining that the fault type is a first fault, wherein the first fault includes a task interruption caused by a master device switching; If it is detected that the bus to be tested is not in an idle state, detecting a bus signal at a low level; If it is detected that the bus signal at a low level is a data signal, the bus switch is used to reset and detect whether the front-stage circuit of the bus switch is restored; If it is detected that the previous line is restored, determining that the fault type is a second fault, wherein the second fault includes a sub-device fault; If it is detected that the previous-stage line has not been restored, it is determined that the fault type is a third fault, and the third fault includes a management controller fault.
3. The method according to claim 1 or 2, characterized in that: Detecting the fault type of the bus to be tested according to the fault data also includes: If it is detected that the bus signal at a low level is a clock signal, the bus switch is used for resetting, and whether the front-stage circuit of the bus switch is restored is detected; If it is detected that the previous line is restored, the fault type is determined to be a fourth fault, and the fourth fault includes that the sub-device does not release the bus normally; If it is detected that the previous line has not been restored, it is determined that the fault type is a fifth fault, and the fifth fault includes that the management controller does not normally release the bus.
4. The method according to claim 2, characterized in that: After determining that the fault type is a second fault if the preceding line is detected to be restored, the method further includes: Reset signals are sent to the sub-devices connected to the bus switch in sequence, sub-devices that have not been successfully restored are determined as target faulty devices, and the target faulty devices are recorded.
5. The method according to claim 2, characterized in that: Before resetting the bus switch and detecting whether the front-stage circuit of the bus switch is restored, the method further includes: Detecting whether there is a bus switch on the bus to be tested; If it is detected that no bus switch exists on the bus to be tested, a reset signal is directly sent to the sub-device corresponding to the bus to be tested.
6. The method according to claim 1, characterized in that After determining that the periodic data is fault data, the method further includes: The fault data is saved.
7. The method according to any one of claims 1, characterized in that After determining whether the cycle data of the bus signal is complete according to the cycle identifier of the bus signal of the bus to be tested, the method further includes: If it is detected that the periodic data is complete, the periodic data is determined to be normal data, and the normal data is discarded.
8. A fault detection device, characterized in that: include: Pinhole and fault detection device; The pin hole is installed on the server mainboard, and there is at least one bus between the management controller of the server mainboard and the sub-device, and the pin hole is connected to the at least one bus; The pin hole is used to connect the pin of the fault detection device; The fault detection device is used to access the bus to be tested in the at least one bus through the pin hole; determine whether the cycle data of the bus signal is complete based on the cycle identifier of the bus signal of the bus to be tested; if the cycle data is detected to be incomplete, determine that the cycle data is fault data, and detect the fault type of the bus to be tested based on the fault data.
9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
CAN bus controller test method based on RX and TX
CN111104272A
Server bus fault positioning method and device, electronic equipment and storage medium
CN114090379A
LPC bus signal test method and device, electronic equipment and readable medium
CN117033102A
Bus management card for use in a system for bus monitoring
US6311296B1