Fault processing method and device, BMC, storage medium and computer program product
By using the setting monitor and PCH in the BMC to handle PCIe device failures, the complex problem of the fault handling process in the prior art is solved, and efficient fault handling is achieved.
Patent Information
- Application Number
- CN202510025115.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
AI Technical Summary
When handling PCIe equipment failures, the process is complicated and involves the upper-level system, resulting in low fault handling efficiency.
The fault information and device status information of the PCIe device are saved to the fault log file by using the setting monitor and the platform path controller (PCH) set on the PCIe bus in the substrate management controller (BMC) for fault analysis and repair processing.
Simplify the fault processing process, improve the efficiency of fault processing, and avoid interference from the upper-level system.
Smart Images

Figure CN120045368A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a fault handling method, apparatus, BMC, storage medium, and computer program product. Background Art
[0002] In the related art, in the case of a failure of a Peripheral Component Interconnect Express (PCIe) device, the Basic Input Output System (BIOS) performs fault handling based on the Advanced Error Reporting (AER) information reported by the PCIe. However, since this solution involves the upper-layer system, the fault handling process is relatively complex, reducing the efficiency of fault handling. Summary of the Invention
[0003] To solve the problems in the related art, embodiments of this application provide a fault handling method, apparatus, BMC, storage medium, and computer program product.
[0004] The technical solution of the embodiments of this application is implemented as follows:
[0005] Embodiments of this application provide a fault handling method applied to a Baseboard Management Controller (BMC). The method includes:
[0006] Based on a set monitor on the PCIe bus, saving the fault information of a first PCIe device to a fault log file; and / or, based on a Platform Controller Hub (PCH), saving the device status information of the first PCIe device to the fault log file; the set monitor is used to capture the AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents the PCIe device that has failed;
[0007] Based on the fault log file, performing fault analysis processing to obtain an analysis result;
[0008] Based on the analysis result, performing repair processing and / or warning processing on the fault of the first PCIe device.
[0009] In the above solution, the step of saving the fault information of the first PCIe device to the fault log file based on the set monitor on the PCIe bus includes:
[0010] Based on the set monitor, obtain first information; the first information represents AER information generated by the first PCIe device;
[0011] Analyze the first information to obtain second information, and save the second information to a fault log file; the second information represents the fault information of the first PCIe.
[0012] In the above solution, saving the device status information of the first PCIe device to a fault log file based on the PCH includes:
[0013] Detect one or more PCIe devices based on the PCH to obtain third information of each PCIe device; the third information represents the device status information of the corresponding PCIe device;
[0014] Parse out the first PCIe device among the one or more PCIe devices based on the third information;
[0015] When there is a first PCIe device among the one or more PCIe devices, save the device status information of the first PCIe device to a fault log file.
[0016] In the above solution, the one or more PCIe devices represent: all the present PCIe devices controlled by the BMC, or one or more PCIe devices indicated by a set instruction; the set instruction is used to instruct the BMC to detect PCIe devices.
[0017] In the above solution, performing fault analysis and processing based on the fault log file to obtain an analysis result includes:
[0018] At every set time interval, perform fault analysis and processing based on the fault log file to obtain an analysis result.
[0019] In the above solution, the analysis result includes a fault type; the fault type includes one of the following: correctable fault, uncorrectable non-fatal fault, uncorrectable fatal fault.
[0020] In the above solution, performing repair processing on the fault of the first PCIe device based on the analysis result includes:
[0021] When the analysis result does not indicate that the fault type of the first PCIe device is an uncorrectable fatal fault, perform repair processing on the fault of the first PCIe device based on a set strategy.
[0022] In the above solution, performing early warning processing on the fault of the first PCIe device based on the analysis result includes:
[0023] When the analysis result characterizes that the fault type of the first PCIe device is an irreparable fatal fault, determine whether the fault parameter of the first PCIe device exceeds a first set threshold; the fault parameter is determined based on the fault log file; the first set threshold is determined based on the fault parameter when the first PCIe device fails during a historical period;
[0024] When the fault parameter of the first PCIe device exceeds the first set threshold, perform a warning process on the fault of the first PCIe device.
[0025] In the above solution, performing a warning process on the fault of the first PCIe device includes:
[0026] Sending a warning email; the warning email is used to describe one or more of the following of the first PCIe device: fault type, fault time, the fault log file, device name, device model.
[0027] In the above solution, the method further includes:
[0028] Determine whether the remaining capacity of the security register is lower than a second set threshold; the security register is used to store the fault log file;
[0029] When the remaining capacity of the security register is lower than the second set threshold, migrate the file in the security register to a remote server.
[0030] An embodiment of the present application further provides a fault processing device, which is applied to BMC and includes:
[0031] A saving unit, configured to save the fault information of the first PCIe device to a fault log file based on a set monitor set on the PCIe bus; and / or save the device status information of the first PCIe device to the fault log file based on the PCH; the set monitor is used to capture the AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents the PCIe device that has a fault;
[0032] An analysis unit, configured to perform fault analysis processing based on the fault log file to obtain an analysis result;
[0033] A processing unit, configured to perform a repair process and / or a warning process on the fault of the first PCIe device based on the analysis result.
[0034] An embodiment of the present application further provides a BMC, including a first processor and a first communication interface;
[0035] The first processor is configured to save the fault information of the first PCIe device to a fault log file based on a set monitor provided on the PCIe bus; and / or save the device status information of the first PCIe device to a fault log file based on the PCH; the set monitor is used to capture AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents a faulty PCIe device;
[0036] Perform fault analysis processing based on the fault log file to obtain an analysis result;
[0037] Based on the analysis result, perform repair processing and / or warning processing on the fault of the first PCIe device.
[0038] An embodiment of the present application further provides a BMC, including: a first processor and a first memory for storing a computer program that can run on the processor,
[0039] Wherein, when the first processor is used to run the computer program, it executes the steps of any of the above methods.
[0040] An embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of any of the above methods.
[0041] An embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of any of the above methods.
[0042] In the embodiment of the present application, the BMC saves the fault information of the first PCIe device to a fault log file based on a set monitor provided on the PCIe bus; and / or saves the device status information of the first PCIe device to a fault log file based on the PCH; wherein, the set monitor is used to capture AER information transmitted on the PCIe bus, the PCH is used to detect the device status information of the PCIe device, and the first PCIe device represents a faulty PCIe device; then, the BMC performs fault analysis processing based on the fault log file to obtain an analysis result, and based on the analysis result, performs repair processing and / or warning processing on the fault of the first PCIe device. That is to say, the BMC can perform fault processing on the first PCIe device through the set monitor and / or PCH provided on the PCIe bus. Compared with the related art, the upper-level system is not involved in the process of fault processing, thereby simplifying the fault processing flow and improving the efficiency of fault processing. Description of the Drawings
[0043] Figure 1Schematic diagram of the implementation process of a fault handling method provided by an embodiment of the present application;
[0044] Figure 2 Schematic diagram of the original information of a PCIe device provided by an embodiment of the present application;
[0045] Figure 3 Schematic diagram of a classification corresponding to a fault type provided by an embodiment of the present application;
[0046] Figure 4 Schematic diagram of the architecture of a fault handling system provided by an application embodiment of the present application;
[0047] Figure 5 Schematic diagram of the process of a fault handling method provided by an application embodiment of the present application;
[0048] Figure 6 Schematic diagram of the structure of a fault handling device provided by an embodiment of the present application;
[0049] Figure 7 Schematic diagram of the hardware composition of a BMC provided by an embodiment of the present application. Detailed implementation manners
[0050] A PCIe device refers to a device using a PCIe interface, which has the characteristics of high bandwidth and low latency and can be used to connect various external devices, such as intelligent network cards, storage devices, and data acceleration cards.
[0051] In practical applications, a PCIe device can be regarded as one of the key components of a host hardware device. Exemplarily, the host hardware device can be an Advanced RISC Machine (ARM) server. In the case of a fault in the PCIe device, the performance of the host hardware device may be severely affected, thereby affecting the stability of the service. For example, a data link problem caused by a fault in the PCIe device affects the stability of the data center used by the ARM server, thereby causing serious errors in real-time online services.
[0052] In the related art, in the case of a failure of a PCIe device, the upper-layer system captures and analyzes AER information through the PCIe bus, and then the BIOS calls the reliability, availability, and serviceability (RAS) system based on the analyzed AER information to perform fault handling. Alternatively, in the case of a failure of a PCIe device, the BIOS captures AER information based on the interaction with the upper-layer system, and then performs fault handling based on the AER information. It can be seen that the solutions in the related art involve the upper-layer system, resulting in a relatively complex fault handling process, thereby reducing the efficiency of fault handling.
[0053] Based on this, in the embodiments of the present application, the BMC saves the fault information of the first PCIe device to a fault log file based on a setting monitor set on the PCIe bus; and / or saves the device status information of the first PCIe device to the fault log file based on the PCH; wherein, the setting monitor is used to capture the AER information transmitted on the PCIe bus, the PCH is used to detect the device status information of the PCIe device, and the first PCIe device represents the PCIe device that has failed; then, the BMC performs fault analysis and processing based on the fault log file to obtain an analysis result, and based on the analysis result, repairs and / or warns of the fault of the first PCIe device. That is to say, the BMC can perform fault handling on the first PCIe device through the setting monitor and / or PCH set on the PCIe bus. Compared with the related art, the upper-layer system is not involved in the fault handling process, thereby simplifying the fault handling process and improving the efficiency of fault handling.
[0054] The following further describes the present application in detail with reference to the drawings and embodiments.
[0055] The embodiments of the present application provide a fault handling method, which is applied to the BMC. In practical applications, the BMC can be set in an ARM server.
[0056] See Figure 1 , the fault handling method provided by the embodiments of the present application includes:
[0057] Step 101: Save the fault information of the first PCIe device to a fault log file based on a setting monitor set on the PCIe bus; and / or save the device status information of the first PCIe device to the fault log file based on the PCH.
[0058] Wherein, the setting monitor is used to capture the AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; and the first PCIe device represents the PCIe device that has failed.
[0059] In practical applications, in the case of a failure of a PCIe device, the PCIe device can send AER information to the PCIe bus. The AER information sent by the PCIe device can be regarded as the AER information generated by the PCIe device, or can also be expressed as the AER information of the PCIe device.
[0060] Here, it is set that the monitor is set on the PCIe bus. In practical applications, the set monitor can directly monitor the PCIe bus and capture the information transmitted on the PCIe bus. After the failed PCIe device, that is, the first PCIe device, sends AER information, the set monitor can capture the AER information, and then interact with the BMC to save the fault information generated by the first PCIe device to the fault log file. In this way, the BMC can process the AER information of the first PCIe device without going through the upper-layer system.
[0061] In practical applications, the BMC can detect one or more PCIe devices in real time. The BMC can detect one or more PCIe devices based on the set instructions.
[0062] In practical applications, the set instructions can at least include: a login instruction, and / or, an information collection instruction. The BMC can first log in based on the login instruction, then determine that PCIe device detection is required, and then detect one or more PCIe devices. Exemplarily, the login instruction can be a Secure Shell (SSH) login instruction. The BMC can also detect one or more PCIe devices based on the information collection instruction. Detecting a PCIe device can also be expressed as self-checking a PCIe device.
[0063] During the process of the BMC detecting one or more PCIe devices, it can control the PCH to detect the device status information of these PCIe devices, and then based on the device status information detected by the PCH, parse out the PCIe device that has failed among the detected PCIe devices, that is, the parsed first PCIe information, and then save the device status information of the first PCIe device to the fault log file. In this way, the BMC can process the device status information of the first PCIe device without going through the BIOS or the upper-layer system.
[0064] Step 102: Based on the fault log file, perform fault analysis and processing to obtain an analysis result.
[0065] In practical applications, the processing method corresponding to the fault can be determined based on the analysis result.
[0066] In practical applications, the relationship between the severity of a fault and the degree to which the business device corresponding to the fault cannot operate can be positively correlated, that is, the higher the severity of the fault, the less the business device corresponding to the fault can operate normally. The business device corresponding to the fault may include: the PCIe device corresponding to the fault and / or the host hardware device where the PCIe device corresponding to the fault is located. The fault can affect the normal operation of the corresponding business device in aspects such as the Data Link, Physical, Internal, and Transaction.
[0067] In practical applications, a fault occurring in a PCIe device may not affect the operation of the business device, that is, the fault is not severe. The severity of the fault occurring in the PCIe device can be determined based on the analysis result, and then the method for handling the fault can be determined based on the severity of the fault. For example, when the fault is not severe, the fault is not processed or repaired, and when the fault is severe, a warning is issued for the fault.
[0068] Step 103: Based on the analysis result, perform repair processing and / or warning processing on the fault of the first PCIe device.
[0069] In practical applications, performing repair processing on the fault of the first PCIe device can be understood as: eliminating the fault or reducing the adverse effects caused by the fault.
[0070] Exemplarily, when the fault is a clock drift error, the first PCIe device can be controlled to perform device calibration, thereby eliminating the influence of the clock offset error. When the fault is a data retransmission error, the first PCIe device can be controlled to retransmit the data to reduce the influence brought by the data retransmission error. However, since it cannot be guaranteed that the retransmitted data will not have a transmission error, this error is not eliminated.
[0071] In practical applications, performing warning processing on the fault of the first PCIe device can be understood as: actively notifying relevant personnel of the warning information related to the fault occurring in the first PCIe device. Exemplarily, the relevant personnel may be fault management personnel, and the warning information may at least include one or more of the following information of the first PCIe device: fault type, fault time, fault log file, device name, device model, recommended handling method. After receiving the warning notice, the relevant personnel can determine the further processing method of the fault based on the warning information. For example, based on the warning information, it is determined to replace the first PCIe device with other devices.
[0072] In an embodiment, performing warning processing on the fault of the first PCIe device includes:
[0073] Send a warning email; the warning email is used to describe one or more of the following of the first PCIe device: fault type, fault time, fault log file, device name, device model.
[0074] In practical applications, the warning email can be understood as a warning notification, and the warning email can be used to describe warning information.
[0075] In the embodiment of the present application, the BMC can perform fault handling on the first PCIe device through a set monitor and / or PCH disposed on the PCIe bus. Compared with the related art, the upper-layer system is not involved in the process of fault handling, thereby simplifying the fault handling process and improving the efficiency of fault handling.
[0076] The following further describes the manner in which the BMC saves the fault information and / or device status information of the first PCIe device to the fault log file.
[0077] In one embodiment, based on a set monitor disposed on the PCIe bus, saving the fault information of the first PCIe device to the fault log file includes:
[0078] Based on the set monitor, obtain the first information; the first information represents the AER information generated by the first PCIe device;
[0079] Parse the first information to obtain the second information, and save the second information to the fault log file; the second information represents the fault information of the first PCIe
[0080] In practical applications, after the set monitor captures the AER information generated by the first PCIe device, the AER information can be transmitted to the BMC. In this way, the BMC can obtain the AER information generated by the first PCIe device, that is, the first information.
[0081] In practical applications, after the BMC obtains the first information, it can parse the first information to obtain the fault information of the first PCIe, that is, the second information. In practical applications, the second information can at least include one or more of the following of the first PCIe: source of the fault, fault type, fault occurrence time, status, location.
[0082] In practical applications, the BMC can parse the first information into the second information based on the fault analysis mechanism of the RAS system, that is, the RAS fault analysis mechanism.
[0083] In one embodiment, based on the PCH, saving the device status information of the first PCIe device to the fault log file includes:
[0084] Detect one or more PCIe devices based on the PCH to obtain third information of each PCIe device; the third information characterizes the device status information of the corresponding PCIe device;
[0085] Parse the first PCIe device from the one or more PCIe devices based on the third information;
[0086] When the first PCIe device exists among the one or more PCIe devices, save the device status information of the first PCIe device to a fault log file.
[0087] In practical applications, during the initialization process, the BMC can establish a communication connection with the PCH based on the Serial Peripheral Interface (I2C) channel, and then, the BMC can send commands and / or requests to the PCH through the communication interface with the PCH, so as to control the PCH to detect one or more PCIe devices, that is, start reading the device status information of one or more PCIe devices. During the operation of the BMC, it can perform real-time detection on one or more PCIe devices through the PCH.
[0088] In practical applications, during the process of detecting one or more PCIe devices, the PCH can scan the PCIe bus, identify and enumerate the connected PCIe devices; then, the PCH can access these PCIe devices, read the configuration space of each device, and store the device status information of these PCIe devices in internal registers; afterwards, the BMC can read the device status information stored in the PCH through the communication interface with the PCH, so as to collect the device status information of the one or more PCIe devices detected, that is, the third information of each PCIe device.
[0089] In practical applications, the third information can at least include one or more of the following items of the corresponding PCIe device: module identification number, device identifier, manufacturer identifier, device category, supported bus speed, slot, channel, serial number, connection status, error status, transmission rate.
[0090] In an embodiment, the one or more PCIe devices represent: all the in-position PCIe devices controlled by the BMC, or, one or more PCIe devices indicated by a set instruction; the set instruction is used to instruct the BMC to detect the PCIe device.
[0091] In practical applications, the setting instruction can indicate to detect all the PCIe devices in place controlled by the BMC, or the setting instruction can also indicate to detect some of the PCIe devices in place controlled by the BMC. The setting instruction can indicate to the BMC the PCIe devices to be detected by indicating the identifiers of the PCIe devices. The setting instruction can at least include: a login instruction, and / or, an information collection instruction.
[0092] In practical applications, after the BMC obtains the third information, it can save the third information to the PCIe information original file. Then, the BMC can parse the PCIe information original file based on the information parsing instruction to implement the parsing of the third information, and then determine whether each of one or more PCIe devices is faulty based on the parsing result.
[0093] In practical applications, the third information can be expressed as the original information of the PCIe device, which can be a set of encoded data with low readability. Exemplarily, Figure 2 The third information, that is, the original information of the PCIe device, is shown. The BMC can parse the third information of each PCIe device into the fourth information. Both the fourth information and the third information can represent the device status information of the corresponding PCIe device, but the fourth information has higher readability compared with the third information. The fourth information can be expressed as the parsed information of the PCIe device.
[0094] Table 1 gives an example of the fourth information. Among them, the PCIe slot, the in-place information, the device type, and the operating status can all be understood as device status information.
[0095] Table 1
[0096] PCIe Slot Presence Information Device Type Operating Status 0 Present Intelligent Network Card Operating 1 Present NVME Card Operating 2 Present GPU Operating
[0097] In practical applications, the BMC can compare the fourth information of each PCIe device with the preset normal device status information to obtain the first comparison result, and then determine whether the PCIe device is faulty based on the first comparison result, so as to determine the faulty PCIe device, that is, the first PCIe device.
[0098] In practical applications, when the first comparison result indicates that the corresponding PCIe device is not faulty, the BMC can save the fourth information of the PCIe device to the PCIe information parsing file. When the first comparison result indicates that the corresponding PCIe device is faulty, the BMC can save the fourth information of the PCIe device to the fault log file. The BMC can also further determine the fault information of the corresponding PCIe device based on the fourth information and save the determined fault information to the fault log file.
[0099] In practical applications, during the process of saving information to a fault log file, a PCIe information original file, or a PCIe information parsing file, if the corresponding file does not exist, the corresponding file can be generated based on the information to be saved, so as to save the information to the corresponding file.
[0100] In one embodiment, the fault handling method provided by the embodiments of the present application further includes:
[0101] Determine whether the remaining capacity of the security register is lower than a second set threshold; the security register is used to store the fault log file;
[0102] In the case where the remaining capacity of the security register is lower than the second set threshold, migrate the files in the security register to a remote server.
[0103] In practical applications, the security register can also be used to store the PCIe information original file and the PCIe information parsing file.
[0104] In practical applications, each file can correspond to a set path. After the file is generated, the file can be stored in the set path in the security register.
[0105] In practical applications, the BMC can periodically or in real time determine whether the remaining capacity of the security register is lower than the second set threshold. In the case where the remaining capacity of the security register is lower than the second set threshold, migrate the files in the security register to a remote server, so as to realize the regular maintenance of the security register. The remaining capacity of the security register can be understood as the size of the storage space in the security register that has not been used yet. The second set threshold can be expressed as a set capacity threshold.
[0106] In practical applications, the BMC can migrate the files in the security register to a remote server based on the storage time of the files in the security register. Exemplarily, the BMC can sort the storage times of the files in the security register, and then migrate the set number of files with the earliest storage time in the security register to the remote server.
[0107] In practical applications, the storage space of the remote server is large, which can be much larger than the storage space of the security register. In this way, the storage space of the security register can be saved, and the failure of file storage caused by the low remaining capacity of the security register can also be avoided, thereby improving the stability of fault handling.
[0108] The method for fault analysis and processing based on the fault log file will be further described below.
[0109] In one embodiment, based on the fault log file, perform fault analysis and processing to obtain an analysis result, including:
[0110] At regular intervals, based on the fault log file, perform fault analysis and processing to obtain the analysis results.
[0111] In practical applications, performing fault analysis and processing at regular intervals can also be expressed as: performing fault analysis and processing regularly, or performing fault analysis and processing periodically.
[0112] In practical applications, based on the fault log file, the fault information of the first PCIe device can be compared with the latest device status information of the first PCIe device to obtain a second comparison result. Then, based on the second comparison result, it can be determined whether the fault of the first PCIe device increases or is eliminated, and whether the operating state of the first PCIe device decreases due to the fault, so as to determine whether the fault is upgraded. The analysis results can include: relevant judgment results obtained during the fault analysis and processing.
[0113] Exemplarily, in the case where a fault occurs in the slot of the first PCIe device, the fault will cause partial path blockage, but the first PCIe device can continue to work normally relying on the remaining unblocked paths. Through periodic fault analysis and processing, for example, periodically comparing the fault information with the device status information, it can be determined whether the blocked paths increase and whether the operating state of the first PCIe device decreases, and then it can be determined whether the fault is upgraded.
[0114] In one embodiment, the analysis results include the fault type; the fault type includes one of the following: correctable fault, uncorrectable non-fatal fault, uncorrectable fatal fault.
[0115] In practical applications, the fault type can be used to characterize the severity of the fault. The severity of an uncorrectable fatal fault is higher than that of an uncorrectable non-fatal fault, and the severity of an uncorrectable non-fatal fault can be higher than that of a correctable fault.
[0116] Figure 3 A classification schematic corresponding to the fault type is given. In practical applications, both uncorrectable non-fatal faults and uncorrectable fatal faults can be understood as uncorrectable faults.
[0117] In practical applications, both correctable faults and uncorrectable non-fatal faults can support being repaired.
[0118] In the case where the fault type of the fault is a correctable fault, the fault can be corrected, that is, eliminated. Exemplarily, correctable faults can include: flipping of a single bit or byte, clock drift, and occurrence of incorrect packets in transaction layer packets (TLP, Transaction Layer Packet), etc.
[0119] In the case where the fault type of a fault is an uncorrectable non-fatal fault, the fault cannot be corrected, that is, it cannot be completely eliminated, but the impact caused by the fault can be reduced, thereby resolving the fault. Exemplarily, uncorrectable non-fatal faults may include: data retransmission errors, error checking and correcting (ECC) errors, parity errors, and redundant data errors.
[0120] In practical applications, in the case where the fault type of a fault is an uncorrectable fatal fault, the fault cannot be directly repaired. Such faults can be regarded as relatively serious faults, usually causing partial link interruptions, but may not immediately cause the corresponding service device to go down. Exemplarily, uncorrectable fatal faults may include: slot physical damage and memory faults, etc.
[0121] In practical applications, based on the fault type of the first PCIe device, repair processing and / or warning processing can be performed on the first PCIe device, that is, fault processing is performed.
[0122] In the related art, when a PCIe device fails, a serial management interface (SMI) interrupt is triggered to handle the fault. However, the SMI interrupt will cause delays or downtime of the corresponding service device, affecting service operation, and for faults with a lower severity level, even if the SMI interrupt is not triggered, the fault may be repairable. Therefore, the fault handling solution in the related art has a greater impact on service operation. In the embodiments of the present application, determining the fault type of the faulty PCIe device is equivalent to grading the severity of the fault. Based on this, different processing can be performed based on the fault type, avoiding triggering the SMI interrupt for faults with a lower severity level, and reducing the impact of fault handling on service operation.
[0123] The following further describes the method of fault handling based on the fault type.
[0124] In one embodiment, based on the analysis result, repairing the fault of the first PCIe device includes:
[0125] In the case where the analysis result indicates that the fault type of the first PCIe device is not an uncorrectable fatal fault, based on the set policy, repair processing is performed on the fault of the first PCIe device
[0126] In practical applications, the BMC can call the PCIe fault management mechanism of the RAS system to repair the faults of the first PCIe device. The RAS system can be set with a setting policy, and the setting policy can indicate the operations required to repair the faults of the first PCIe device, that is, the way of indicating the repair process.
[0127] Exemplarily, in the case where the fault is a single-bit or byte flip, the way of repair process can be: controlling the hardware device to correct. In the case where the fault is a transmission error, the way of repair process can be: ignoring the error. In the case where the fault is a redundant data error, the way of repair process can be: resetting the first PCIe device.
[0128] In one embodiment, based on the analysis result, early warning processing is performed on the faults of the first PCIe device, including:
[0129] In the case where the analysis result characterizes that the fault type of the first PCIe device is an uncorrectable fatal fault, it is judged whether the fault parameters of the first PCIe device exceed the first set threshold; the fault parameters are determined based on the fault log file; the first set threshold is determined based on the fault parameters when the first PCIe device had a fault during the historical period;
[0130] In the case where the fault parameters of the first PCIe device exceed the first set threshold, early warning processing is performed on the faults of the first PCIe device.
[0131] In practical applications, judging whether the fault parameters of the first PCIe device exceed the first set threshold can be regarded as: predicting the severity of the faults of the first PCIe device, that is, performing fault prediction. In the case where the fault parameters of the first PCIe device exceed the first set threshold, it can be regarded that the fault is relatively serious, that is, the severity is relatively high, affecting the operation of the corresponding service device, and relevant personnel need to be notified through early warning for fault handling. In the case where the fault parameters of the first PCIe device do not exceed the first set threshold, it can be regarded that the severity of the fault is not high, and the PCIe fault management mechanism of the RAS system can be called to repair the first PCIe device.
[0132] In practical applications, the fault parameters of the first PCIe device can be determined based on fault analysis and processing. The fault parameters can be understood as parameters related to the fault, and the fault parameters can at least include one or more of the following: the number of faults, slot positions, channels, categories, transmission rates, and the proportion of faulty channels. Exemplarily, in the case where the fault is a slot fault, the fault parameters can include the proportion of faulty channels, that is, the proportion of the faulty slot positions in all the slot positions of the first PCIe device.
[0133] In practical applications, the fault parameters of the first PCIe device when a fault occurs in a historical period can be statistically analyzed, and a first set threshold can be determined based on the statistical results. It should be noted that the first PCIe device in the historical period represents the PCIe device that had a fault in the historical period, which may or may not be the same device as the PCIe device currently to be fault-processed, and there is no restriction here.
[0134] Exemplarily, when the first PCIe device is a smart network card, assuming that based on the fault parameters of the smart network card when a fault occurs in the historical period, it is determined that when the fault channel ratio is 10%, that is, 10% of the slots have faults, the channels corresponding to the faulty slots will be blocked, but the remaining slots in the smart network card can still meet the normal operation of the network card. However, when the fault channel ratio exceeds 10%, the transmission rate of the smart network card drops by 50%, resulting in the network card crashing. Then, when making a fault prediction, the predicted fault parameters can include the fault channel ratio, and the first set threshold can be set to 10%.
[0135] In this way, in the case of a relatively serious fault, early warning processing of the fault can be carried out so that relevant personnel can handle the fault in a timely manner, improving the efficiency of fault handling.
[0136] The following further describes the present application in detail in combination with application embodiments.
[0137] The application embodiment of the present application provides a fault handling system. Refer to Figure 4 , this system mainly includes: BMC, PCH, PCIe monitor, and security register. Among them,
[0138] The PCIe monitor is equivalent to the set monitor in the embodiment of the present application, is set on the PCIe bus, and can also be expressed as a PCIe monitoring module. The PCIe monitor can be used to capture AER information on the PCIe bus and transmit the AER information to the BMC.
[0139] The BMC includes: an information collection module, an information analysis module, and a fault prediction module.
[0140] In practical applications, the information collection module can include: a login unit, an information collection unit, and an original information storage unit.
[0141] The login unit can be used to: receive a login instruction, and then determine that the PCIe device needs to be detected.
[0142] The information collection unit can be used to: collect the device information of PCIe devices. In practical applications, the information collection unit can detect the device status information of PCIe devices through the PCH, and then obtain the device status information of the detected PCIe devices.
[0143] The original information storage unit can be used to: store the device status information collected by the information collection unit, that is, the original information, into the PCIe information original file.
[0144] In practical applications, the information parsing module can include: an information parsing and fault judgment unit, a fault information storage unit, and a parsed file storage unit.
[0145] The information parsing and fault judgment unit can be used to: parse the PCIe information original file to obtain highly readable parsed information, and then, based on the parsed information, determine whether the corresponding PCIe device has a fault. The information parsing and fault judgment unit can also be used to parse the obtained AER information to obtain fault information.
[0146] The fault information storage unit can be used to: store the parsed information of the faulty PCIe device into the fault log file for subsequent fault handling. The fault information storage unit can also be used to store the fault information parsed from the AER information into the fault log file.
[0147] The parsed file storage unit can be used to: store the parsed information obtained by parsing the PCIe information original file into the PCIe information parsed file so that relevant personnel can obtain the device status information of the PCIe device.
[0148] In practical applications, the fault prediction module can include: a fault analysis unit, a fault monitoring and output unit, a fault prediction unit, and a fault warning unit.
[0149] The fault analysis unit can be used to: perform fault analysis and processing based on the fault log file to obtain an analysis result.
[0150] The fault monitoring and output unit can be used to: based on the analysis result, determine the faulty PCIe device, that is, determine the first PCIe device, and output the fault type and fault description of the first PCIe device; in the case where the fault type of the first PCIe device is a correctable fault or an uncorrectable non-fatal fault, repair the fault of the first PCIe device.
[0151] The fault prediction unit can be used to: determine whether the fault parameters of the first PCIe device exceed the first set threshold, that is, determine whether the current fault of the first PCIe device affects the normal operation of the corresponding service device.
[0152] The fault warning unit can be used to: when the fault parameter of the first PCIe device exceeds the first set threshold, perform fault warning processing on the fault of the first PCIe device.
[0153] In practical applications, the security register can be used to store the fault log file, the original PCIe information file, and the parsed PCIe information file. The security register can also be referred to as an information storage module or an information storage unit.
[0154] The application embodiment of the present application also provides a fault handling method. The fault handling system can perform fault handling based on this fault handling method. See Figure 5 , the process of the fault handling system performing fault handling mainly includes the following steps:
[0155] Step 1: The BMC login instruction is used to determine that the PCIe device needs to be detected.
[0156] Step 2: The BMC detects the PCIe device based on the PCH to obtain the device status information; and / or, based on the AER information captured by the PCIe, determine the fault information.
[0157] Step 3: The BMC saves the obtained device status information to the original PCIe information file.
[0158] Step 4: The BMC parses the original PCIe information file to determine whether the detected PCIe device has a fault.
[0159] In practical applications, the BMC can determine whether each detected PCIe device has a fault.
[0160] If the detected PCIe device has a fault, this PCIe device is equivalent to the first PCIe device, and step 5 is continued.
[0161] If there is no fault in the detected PCIe device, the device status information of the detected PCIe device is saved to the parsed PCIe information file.
[0162] Step 5: The BMC saves the device status information of the first PCIe device to the fault log file.
[0163] In practical applications, the BMC can also save the fault information of the first PCIe device obtained based on the AER information to the fault log file.
[0164] Step 6: The BMC performs fault analysis processing based on the fault log file to obtain an analysis result.
[0165] Step 7: The BMC performs fault handling on the fault of the first PCIe device based on the analysis result.
[0166] In practical applications, when the fault type of the first PCIe device is an uncorrectable fatal fault, the BMC can determine whether the fault parameters of the first PCIe device exceed the first set threshold. When the fault parameters of the first PCIe device exceed the first set threshold, a fault warning process can be performed. When the fault type of the first PCIe device is not an uncorrectable fatal fault, and / or when the fault parameters of the first PCIe device do not exceed the first set threshold, a fault repair process is performed on the fault.
[0167] In the application embodiment of the present application, the BMC can perform fault processing on the first PCIe device through the PCIe monitor and / or PCH set on the PCIe bus. Compared with the related art, the upper layer system is not involved in the process of fault processing, thereby simplifying the fault processing flow and improving the efficiency of fault processing.
[0168] Based on the above embodiments, the present application further provides a fault processing device, which is applied to the BMC. Refer to Figure 6 and this fault processing device includes:
[0169] A storage unit 61, configured to save the fault information of the first PCIe device to a fault log file based on a set monitor set on the PCIe bus; and / or save the device status information of the first PCIe device to the fault log file based on the PCH; the set monitor is used to capture the AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents the PCIe device that has a fault;
[0170] An analysis unit 62, configured to perform fault analysis processing based on the fault log file to obtain an analysis result;
[0171] A processing unit 63, configured to perform a fault repair process and / or a warning process on the fault of the first PCIe device based on the analysis result.
[0172] In an embodiment, the storage unit 61 saves the fault information of the first PCIe device to a fault log file based on a set monitor set on the PCIe bus, including:
[0173] Obtain a first piece of information based on the set monitor; the first piece of information represents the AER information generated by the first PCIe device;
[0174] Parse the first piece of information to obtain a second piece of information, and save the second piece of information to the fault log file; the second piece of information represents the fault information of the first PCIe.
[0175] In one embodiment, the storage unit 61 stores the device status information of the first PCIe device in a fault log file based on the PCH, including:
[0176] Detect one or more PCIe devices based on the PCH to obtain third information of each PCIe device; the third information characterizes the device status information of the corresponding PCIe device;
[0177] Parse out the first PCIe device among the one or more PCIe devices based on the third information;
[0178] When there is a first PCIe device among the one or more PCIe devices, store the device status information of the first PCIe device in a fault log file.
[0179] In one embodiment, the one or more PCIe devices represent all the in-position PCIe devices controlled by the BMC, or one or more PCIe devices indicated by a setting instruction; the setting instruction is used to instruct the BMC to detect the PCIe device.
[0180] In one embodiment, the analysis unit 62 performs fault analysis processing based on the fault log file to obtain an analysis result, including:
[0181] At every set time interval, perform fault analysis processing based on the fault log file to obtain an analysis result.
[0182] In one embodiment, the analysis result includes a fault type; the fault type includes one of the following: a correctable fault, an uncorrectable non-fatal fault, and an uncorrectable fatal fault.
[0183] In one embodiment, the processing unit 63 repairs the fault of the first PCIe device based on the analysis result, including:
[0184] When the analysis result indicates that the fault type of the first PCIe device is not an uncorrectable fatal fault, repair the fault of the first PCIe device based on a set policy.
[0185] In one embodiment, the processing unit 63 gives a warning about the fault of the first PCIe device based on the analysis result, including:
[0186] When the analysis result indicates that the fault type of the first PCIe device is an uncorrectable fatal fault, determine whether the fault parameter of the first PCIe device exceeds a first set threshold; the fault parameter is determined based on the fault log file; the first set threshold is determined based on the fault parameters when the first PCIe device failed during a historical period;
[0187] When the fault parameter of the first PCIe device exceeds the first set threshold, perform a warning process on the fault of the first PCIe device.
[0188] In one embodiment, the processing unit 63 performs a warning process on the fault of the first PCIe device, including:
[0189] Send a warning email; the warning email is used to describe one or more of the following of the first PCIe device: fault type, fault time, the fault log file, device name, device model.
[0190] In one embodiment, the storage unit 61 is further configured to:
[0191] Determine whether the remaining capacity of the security register is lower than a second set threshold; the security register is used to store the fault log file;
[0192] When the remaining capacity of the security register is lower than the second set threshold, migrate the files in the security register to a remote server.
[0193] In actual application, the storage unit 61, the analysis unit 62, and the processing unit 62 can all be implemented by a processor in the fault processing device.
[0194] It should be noted that: when the fault processing device provided in the above embodiment performs fault processing, only the above division of each program module is used for illustration. In actual application, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the device is divided into different program modules to complete all or part of the above-described processing. In addition, the fault processing device provided in the above embodiment and the fault processing method embodiment belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.
[0195] Based on the hardware implementation of the above program module, and in order to implement the method of the embodiments of the present application, the present application also provides a BMC, referring to Figure 7 and this BMC includes:
[0196] A first communication interface 1, capable of interacting with other devices for information;
[0197] The first processor 2 is connected to the first communication interface 1 to enable information interaction with other devices. When running a computer program, it executes the method provided by one or more technical solutions in the above embodiments. The computer program is stored on the first memory 3.
[0198] Specifically, the first processor 2 is configured to save the fault information of the first PCIe device to a fault log file based on a set monitor disposed on the PCIe bus; and / or save the device status information of the first PCIe device to the fault log file based on the PCH. The set monitor is used to capture the AER information transmitted on the PCIe bus. The PCH is used to detect the device status information of the PCIe device. The first PCIe device represents a PCIe device that has failed.
[0199] Based on the fault log file, perform fault analysis processing to obtain an analysis result.
[0200] Based on the analysis result, perform repair processing and / or warning processing on the fault of the first PCIe device.
[0201] In one embodiment, when the first processor 2 saves the fault information of the first PCIe device to the fault log file based on a set monitor disposed on the PCIe bus, it includes:
[0202] Based on the set monitor, obtain first information. The first information represents the AER information generated by the first PCIe device.
[0203] Parse the first information to obtain second information, and save the second information to the fault log file. The second information represents the fault information of the first PCIe.
[0204] In one embodiment, when the first processor 2 saves the device status information of the first PCIe device to the fault log file based on the PCH, it includes:
[0205] Detect one or more PCIe devices based on the PCH to obtain third information of each PCIe device. The third information represents the device status information of the corresponding PCIe device.
[0206] Based on the third information, parse out the first PCIe device among the one or more PCIe devices.
[0207] In the case where the first PCIe device exists among the one or more PCIe devices, save the device status information of the first PCIe device to the fault log file.
[0208] In one embodiment, the one or more PCIe devices represent: all the PCIe devices in place controlled by the BMC, or one or more PCIe devices indicated by a setting instruction; the setting instruction is used to instruct the BMC to detect the PCIe devices.
[0209] In one embodiment, the first processor 2 performs a fault analysis process based on the fault log file to obtain an analysis result, including:
[0210] At every set time interval, perform a fault analysis process based on the fault log file to obtain an analysis result.
[0211] In one embodiment, the analysis result includes a fault type; the fault type includes one of the following: a correctable fault, an uncorrectable non-fatal fault, and an uncorrectable fatal fault.
[0212] In one embodiment, the first processor 2 performs a repair process on the fault of the first PCIe device based on the analysis result, including:
[0213] When the analysis result indicates that the fault type of the first PCIe device is not an uncorrectable fatal fault, perform a repair process on the fault of the first PCIe device based on a set policy.
[0214] In one embodiment, the first processor 2 performs a warning process on the fault of the first PCIe device based on the analysis result, including:
[0215] When the analysis result indicates that the fault type of the first PCIe device is an uncorrectable fatal fault, determine whether the fault parameter of the first PCIe device exceeds a first set threshold; the fault parameter is determined based on the fault log file; the first set threshold is determined based on the fault parameter when the first PCIe device had a fault during a historical period;
[0216] When the fault parameter of the first PCIe device exceeds the first set threshold, perform a warning process on the fault of the first PCIe device.
[0217] In one embodiment, the first processor 2 performs a warning process on the fault of the first PCIe device, including:
[0218] Send a warning email; the warning email is used to describe one or more of the following of the first PCIe device: fault type, fault time, the fault log file, device name, and device model.
[0219] In one embodiment, the first processor 2 is further configured to:
[0220] Determine whether the remaining capacity of the security register is lower than a second set threshold; the security register is used to store the fault log file;
[0221] In the case where the remaining capacity of the security register is lower than the second set threshold, migrate the files in the security register to a remote server.
[0222] It should be noted that: The specific processing process of the first communication interface 1 can be understood with reference to the above method.
[0223] Of course, in actual application, each component in the BMC is coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 7 all kinds of buses are labeled as the bus system 4.
[0224] The first memory 3 in the embodiments of the present application is used to store various types of data to support the operations in the BMC. Examples of these data include: any computer program for operating on the BMC.
[0225] The method disclosed in the above embodiments of the present application can be applied to the first processor 2 or implemented by the first processor 2. The first processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the first processor 2 or by instructions in the form of software. The above-mentioned first processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The first processor 2 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the first memory 3. The first processor 2 reads the information in the first memory 3 and combines its hardware to complete the steps of the foregoing method.
[0226] In an exemplary embodiment, the BMC can be implemented by one or more ASICs, DSPs, PLDs, CPLDs, FPGAs, general-purpose processors, controllers, MCUs, Microprocessors, or other electronic components for performing the foregoing method.
[0227] It can be understood that the first memory 3 in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (FlashMemory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDRSDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random AccessMemory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random AccessMemory), and a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0228] In an exemplary embodiment, the embodiments of the present application further provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a BMC including a stored computer program. The above computer program can be executed by the first processor 2 of the BMC to complete the steps described in the foregoing method. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.
[0229] In an exemplary embodiment, the embodiments of the present application further provide a computer program product, including a computer program, which can be executed by the first processor 2 of the BMC to complete the steps described in any of the foregoing methods.
[0230] It should be noted that: "first", "second", etc. are used to distinguish similar objects and do not necessarily describe a specific order or sequence.
[0231] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "one or more" in this article represents any one or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0232] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
Claims
1. A fault handling method, characterized in that: Applied to a baseboard management controller BMC, the method comprises: Based on a setting monitor set on the PCIe bus, the fault information of the first PCIe device is saved to a fault log file; and / or, based on a platform path controller PCH, the device status information of the first PCIe device is saved to a fault log file; the setting monitor is used to capture the advanced fault report AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents the PCIe device that has failed; Based on the fault log file, perform fault analysis and obtain analysis results; Based on the analysis result, repair processing and / or early warning processing are performed on the fault of the first PCIe device.
2. The method according to claim 1, characterized in that The method of saving the fault information of the first PCIe device to a fault log file based on a setting monitor set on the PCIe bus includes: Based on the setting monitor, obtaining first information; the first information represents AER information generated by the first PCIe device; The first information is parsed to obtain second information, and the second information is saved in a fault log file; the second information represents the fault information of the first PCIe.
3. The method according to claim 1, characterized in that The step of saving the device status information of the first PCIe device to a fault log file based on the PCH includes: Detect one or more PCIe devices based on the PCH to obtain third information of each PCIe device; the third information represents device status information of the corresponding PCIe device; Parsing a first PCIe device among the one or more PCIe devices based on the third information; In a case where there is a first PCIe device among the one or more PCIe devices, device status information of the first PCIe device is saved to a fault log file.
4. The method according to claim 3, characterized in that The one or more PCIe devices represent: all PCIe devices in place controlled by the BMC, or one or more PCIe devices indicated by a setting instruction; The setting instruction is used to instruct the BMC to detect the PCIe device.
5. The method according to claim 1, characterized in that The performing of fault analysis based on the fault log file to obtain analysis results includes: At set intervals, a fault analysis process is performed based on the fault log file to obtain an analysis result.
6. The method according to claim 1, characterized in that The analysis result includes a fault type; the fault type includes one of the following: a correctable fault, an uncorrectable non-fatal fault, and an uncorrectable fatal fault.
7. The method according to claim 1, characterized in that Based on the analysis result, repairing the fault of the first PCIe device includes: When the analysis result indicates that the fault type of the first PCIe device is not an uncorrectable fatal fault, the fault of the first PCIe device is repaired based on a set strategy.
8. The method according to claim 1, characterized in that Based on the analysis result, early warning processing is performed on the failure of the first PCIe device, including: In the case where the analysis result indicates that the fault type of the first PCIe device is an uncorrectable fatal fault, determining whether a fault parameter of the first PCIe device exceeds a first set threshold; the fault parameter is determined based on the fault log file; and the first set threshold is determined based on the fault parameter of the first PCIe device when a fault occurs in a historical period; When the fault parameter of the first PCIe device exceeds the first set threshold, early warning processing is performed on the fault of the first PCIe device.
9. The method according to claim 1, characterized in that: Performing early warning processing on the failure of the first PCIe device includes: Send an early warning email; the early warning email is used to describe one or more of the following of the first PCIe device: fault type, fault time, the fault log file, device name, and device model.
10. The method according to claim 1, characterized in that The method further comprises: Determining whether the remaining capacity of the security register is lower than a second set threshold; the security register is used to store the fault log file; When the remaining capacity of the security register is lower than the second set threshold, the files in the security register are migrated to a remote server.
11. A fault handling device, characterized in that: Applied to BMC, including: A saving unit is used to save the fault information of the first PCIe device to a fault log file based on a setting monitor set on the PCIe bus; and / or, based on the PCH, save the device status information of the first PCIe device to a fault log file; the setting monitor is used to capture the AER information transmitted on the PCIe bus; the PCH is used to detect the device status information of the PCIe device; the first PCIe device represents the PCIe device that has failed; An analysis unit, used to perform fault analysis based on the fault log file to obtain an analysis result; A processing unit is used to perform repair processing and / or early warning processing on the fault of the first PCIe device based on the analysis result.
12. A BMC, characterized in that: include: A first processor and a first communication interface; The first processor is used to save the fault information of the first PCIe device to a fault log file based on a setting monitor set on the PCIe bus; and / or, based on the PCH, saving the device status information of the first PCIe device to a fault log file; The setting monitor is used to capture the AER information transmitted on the PCIe bus; The PCH is used to detect device status information of the PCIe device; The first PCIe device represents a failed PCIe device; Based on the fault log file, perform fault analysis and obtain analysis results; Based on the analysis result, repair processing and / or early warning processing are performed on the fault of the first PCIe device.
13. A BMC, characterized in that: include: a first processor and a first memory for storing a computer program executable on the processor, Wherein, when the first processor is used to run the computer program, the steps of the method described in any one of claims 1 to 10 are executed.
14. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Equipment fault processing method, management firmware, electronic equipment and storage medium
CN121233386A