Fault positioning method, product, equipment and storage medium
By monitoring and analyzing fault signals in PCIe device links and using a combined model for fault prediction, the problem of inaccurate fault location in PCIe devices has been solved, achieving accurate fault location and prediction, reducing operation and maintenance costs, and ensuring the stable operation of the system.
Patent Information
- Application Number
- CN202511071637.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-12-05
AI Technical Summary
In servers, the fault location of PCIe devices cannot be accurately pinpointed to the specific slot, resulting in high maintenance costs. Existing technologies cannot effectively solve the problem of fault identification and reporting for PCIe devices.
By monitoring the fault signals of each device link through the basic input/output system, the fault type is determined. When the fault type is the first fault type and the number of errors reaches the threshold, the combined model is used to predict the fault, and the specific device is determined by combining the device mapping table to generate early warning information.
It enables precise location and prediction of PCIe device faults, reduces operation and maintenance costs, and ensures system stability and service quality.
Smart Images

Figure CN121070656A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of hardware, in particular, to a fault positioning method, product, device and storage medium. BACKGROUND
[0002] In a server, a central processing unit supports connection of multiple PCIe (Peripheral Component Interconnect Express) devices. When the server is running, fault identification and reporting of the PCIe devices are needed. In related technologies, the basic input / output system (BIOS) is used to report faults of the PCIe devices.
[0003] In related technologies, due to the complex topology of the PCIe devices, the fault positioning of the PCIe devices cannot be accurately positioned to a specific slot. For example, professional analysis of the registers of the PCIe advanced error report is needed to analyze the faults, which increases the operation and maintenance cost of the server. SUMMARY
[0004] Embodiments of the present application provide a fault positioning method, product, device and storage medium, aiming to accurately position faults of PCIe devices.
[0005] A first aspect of embodiments of the present application provides a fault positioning method, which comprises: monitoring, by a basic input / output system, a fault signal of each device link; determining, according to the fault signal, a fault type corresponding to the device link; in a case where the fault type corresponding to the device link is a first fault type and the number of errors reaches a fault prediction threshold, performing fault prediction on the device link by using a preset combination model; determining, according to a pre-set device mapping table, a peripheral component interconnect express device corresponding to the device link; generating, according to a fault prediction result, early warning information corresponding to the peripheral component interconnect express device.
[0006] Optionally, before monitoring, by the basic input / output system, the fault signal of each device link, the method further comprises scanning, by the basic input / output system, each peripheral component interconnect express bus, bridge and endpoint device; allocating a corresponding function number to a function corresponding to each peripheral component interconnect express device, the function number at least including a bus number, a device number and a function number corresponding to the peripheral component interconnect express device; establishing a corresponding topology structure tree for all the peripheral component interconnect express devices; Send the function number of each high-speed serial bus device in the topology tree to a baseboard management controller through a predefined interface.
[0007] Optionally, the method further comprises: In a case where the fault type corresponding to the device link is a second fault type, determining the high-speed serial bus device corresponding to the device link according to the device mapping table; Send the alarm information corresponding to the high-speed serial bus device to a baseboard management controller.
[0008] Optionally, the determining of the fault type corresponding to the device link according to the fault signal comprises: Reading the fault signal to determine the fault information of the device link; Determining whether the fault of the device link belongs to a correctable error according to the fault information; In a case where the fault belongs to a correctable error, determining that the fault type corresponding to the device link is a first fault type; In a case where the fault does not belong to a correctable error, determining that the fault type corresponding to the device link is a second fault type.
[0009] Optionally, the fault prediction of the device link through a preset combination model comprises: Statistically analyzing the error rate and error period of the device link to obtain the time characteristics of the fault corresponding to the device link; Statistically analyzing the error of the device link to obtain the intensity characteristics of the fault corresponding to the device link; Statistically analyzing the burst error of the device link to obtain the mode characteristics of the fault corresponding to the device link; Analyzing the error correlation of a plurality of device links to obtain the correlation characteristics of the fault corresponding to the device link; Inputting the time characteristics, the intensity characteristics, the mode characteristics, and the correlation characteristics into the preset combination model; Obtaining a plurality of fault prediction information through each preset model in the prediction model combination; According to a preset weight, the plurality of fault prediction information is weighted combined to obtain a corresponding fault prediction result.
[0010] Optionally, the method further comprises: In a case where the fault type corresponding to the device link is a first fault type and the number of errors does not reach the fault prediction threshold, recording the fault information corresponding to the device link through a log.
[0011] Optionally, the method further comprises: determining, according to the device mapping table, a high-speed serial bus device corresponding to the device link when the fault type corresponding to the device link is a first fault type and the number of errors reaches an alarm threshold; generating alarm information corresponding to the high-speed serial bus device; sending the alarm information to a baseboard management controller.
[0012] A second aspect of the embodiment of the application provides a fault monitoring device, and the device comprises: a fault signal monitoring module configured to monitor a fault signal of each device link through a basic input / output system; a fault type determining module configured to determine a fault type corresponding to the device link according to the fault signal; a fault prediction module configured to perform fault prediction on the device link through a preset combination model when the fault type corresponding to the device link is a first fault type and the number of errors reaches a fault prediction threshold; a first device determining module configured to determine a high-speed serial bus device corresponding to the device link according to a preset device mapping table; a pre-warning information generating module configured to generate pre-warning information corresponding to the high-speed serial bus device according to a fault prediction result.
[0013] Optionally, the device further comprises a device scanning module configured to scan each high-speed serial bus, bridge and endpoint device through the basic input / output system; a function number module configured to assign a corresponding function number to a function corresponding to each high-speed serial bus device, wherein the function number at least comprises a bus number, a device number and a function number corresponding to the high-speed serial bus device; a topology tree establishing module configured to establish a corresponding topology tree for all the high-speed serial bus devices; a function number sending module configured to send the function number of each high-speed serial bus device in the topology tree to a baseboard management controller through a predefined interface.
[0014] Optionally, the device further comprises: a second device determining module configured to determine a high-speed serial bus device corresponding to the device link according to the device mapping table when the fault type corresponding to the device link is a second fault type; a first alarm information sending module configured to send alarm information corresponding to the high-speed serial bus device to a baseboard management controller.
[0015] Optionally, the fault type determining module comprises: a fault information determination submodule, configured to read the fault signal and determine fault information of the device link; an error type judgment submodule, configured to determine whether the fault of the device link belongs to a correctable error according to the fault information; a first fault type determination submodule, configured to determine that the fault type corresponding to the device link is a first fault type in a case where the fault belongs to a correctable error; a second fault type determination submodule, configured to determine that the fault type corresponding to the device link is a second fault type in a case where the fault does not belong to a correctable error.
[0016] Optionally, the fault prediction module comprises: a first feature determination submodule, configured to statistically determine an error rate and an error period of the device link to obtain a time feature of the fault corresponding to the device link; a second feature determination submodule, configured to statistically determine errors of the device link to obtain an intensity feature of the fault corresponding to the device link; a third feature determination submodule, configured to statistically determine burst errors of the device link to obtain a mode feature of the fault corresponding to the device link; a fourth feature determination submodule, configured to analyze error correlations of a plurality of the device links to obtain a correlation feature of the fault corresponding to the device link; a feature input submodule, configured to input the time feature, the intensity feature, the mode feature, and the correlation feature into the preset combination model; a fault information obtaining submodule, configured to obtain a plurality of fault prediction information through each preset model in the prediction model combination; a prediction result obtaining submodule, configured to combine the plurality of fault prediction information according to a preset weight to obtain a corresponding fault prediction result.
[0017] Optionally, the apparatus further comprises: a fault information recording module, configured to record, through a log, the fault information corresponding to the device link in a case where the fault type corresponding to the device link is the first fault type and the number of errors does not reach the fault prediction threshold.
[0018] Optionally, the apparatus further comprises: a second device determination module, configured to determine, according to the device mapping table, the high-speed serial bus device corresponding to the device link in a case where the fault type corresponding to the device link is the first fault type and the number of errors reaches an alarm threshold. a third alarm information generation module, configured to generate alarm information corresponding to the high-speed serial bus device. The second alarm information sending module is configured to send the alarm information to a baseboard management controller.
[0019] The third aspect of the embodiments of the present application provides a product, which comprises a computer program / instruction, and the computer program / instruction is executed by a processor to implement the steps in any of the methods provided in the first aspect of the present application.
[0020] The fourth aspect of the embodiments of the present application provides a readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps in the method provided in the first aspect of the present application.
[0021] The fifth aspect of the embodiments of the present application provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps in the method provided in the first aspect of the present application.
[0022] The fault positioning method provided in the present application monitors the fault signals of each device link through a basic input / output system; determines the fault type corresponding to the device link according to the fault signals; in the case that the fault type corresponding to the device link is a first fault type and the number of errors reaches a fault prediction threshold, performs fault prediction on the device link through a preset combination model; determines the high-speed serial bus device corresponding to the device link according to a pre-set device mapping table; and generates the early warning information corresponding to the high-speed serial bus device according to the fault prediction result.
[0023] In the method, the fault signals of each device link are received through a basic input / output system, in the case that the type of the fault signals is a first fault type and the number of errors reaches a fault prediction threshold, the device link is subjected to fault prediction through a preset combination model, and the specific device corresponding to the device link is determined according to a pre-set device mapping table, and the corresponding early warning information is generated, thereby realizing accurate positioning and fault prediction of the device. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0025] Figure 1 is a fault positioning method architecture schematic diagram proposed by an embodiment of the present application; Figure 2 is a decision logic schematic diagram proposed by an embodiment of the present application. Figure 3 is a schematic diagram of a fault locating process according to an embodiment of the present application; Figure 4 is a schematic diagram of a fault locating device according to an embodiment of the present application; Figure 5 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0026] The technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0027] Reference Figure 5 , Figure 5 is a flowchart of a fault locating method according to an embodiment of the present application. As shown in Figure 5 , the method comprises the following steps: S11: monitoring a fault signal of each device link through a basic input and output system.
[0028] In this embodiment, the basic input and output system is a firmware program that is first run when a computer is started, responsible for initializing hardware devices and loading an operating system. It is stored in a ROM chip on the motherboard, providing the lowest level of hardware control interface, ensuring effective communication between computer hardware and software. The fault signal includes the specific device where the fault occurs and the specific information of the fault.
[0029] In this embodiment, the baseboard management controller detects the fault signal of each device link in the server through the basic input and output system, and the PCIe device fault signal is reported to the basic input and output system through the AER (Advanced Error Reporting) mechanism.
[0030] For example, the fault signal can be a data link layer CRC check error, such as a PCIe link disconnection.
[0031] S12: determining a fault type corresponding to the device link according to the fault signal.
[0032] In this embodiment, the fault type includes a first fault type, Correctable Error (correctable error) and a second fault type, Uncorrectable Error (uncorrectable error). The correctable error is an error that can be repaired by the device itself, and the uncorrectable error is an error that cannot be repaired by the device itself and needs external help to replace the component.
[0033] In this embodiment, after obtaining the fault signal, according to the specific fault information contained in the fault signal, the corresponding fault type is determined, the correctable error is taken as the first fault type, and the uncorrectable error is taken as the second fault type.
[0034] For example, the data link layer CRC check error is the first fault type, which can be repaired by retransmission check, etc., and the PCIe link disconnection is the second fault type, which must be replaced or repaired.
[0035] S13: In the case that the fault type corresponding to the device link is the first fault type and the number of errors reaches the fault prediction threshold, the device link is predicted for fault through a preset combination model.
[0036] In this embodiment, the preset combination model is composed of multiple models and is used for predicting the possible fault of the device. The combination model includes an LSTM (Long Short-Term Memory, long short-term memory network), an isolated forest model, and an exponential smoothing model. The LSTM network learns the long-term error pattern dependence according to the historical errors, and then predicts the fault. The isolated forest model detects abnormal error burst events, and the exponential smoothing model is used for predicting the short-term error growth trend.
[0037] In this embodiment, in the case that the fault type corresponding to the device link is the first fault type, i.e., the correctable error, and the number of errors reaches the fault prediction threshold, the corresponding fault feature is obtained by statistical analysis on the known fault signal, and then the fault feature is input into the combination model for fault prediction.
[0038] In this embodiment, if the number of correctable errors is small, it is a normal phenomenon. If the number of correctable errors exceeds the preset fault prediction threshold, it means that the device may fail, and at this time, the possible fault needs to be predicted to obtain when and what type of fault the device may occur.
[0039] For example, the fault prediction threshold is 100 / h, and when the number of faults occurring per hour exceeds 100, fault prediction is needed, and the prediction result is that the device will fail in 12 hours.
[0040] S14: According to a preset device mapping table, a high-speed serial bus device corresponding to the device link is determined.
[0041] In this embodiment, the device mapping table is a data table recording the mapping relationship between each high-speed serial bus device and a physical slot.
[0042] In this embodiment, the pre-set device mapping table records the bus number, device number and function number corresponding to each link, and according to the mapping relationship, the high-speed serial bus device corresponding to the device link can be determined.
[0043] For example, the BIOS pushes the device of a certain BDF (Bus number, Device number, Function number) at 20.01.2 (BDF number), and the BMC receives the BDF and parses the path as the physical slot corresponding to FRU (Field Replaceable Unit) identified as NVMe Slot0 (non-volatile read-only storage medium) under G1 Port of CPU (Central Processing Unit) 0.
[0044] S15: According to the fault prediction result, the pre-warning information corresponding to the high-speed serial bus device is generated.
[0045] In this embodiment, after obtaining the fault prediction result, the pre-warning information corresponding to the high-speed serial bus device is generated according to the fault prediction result.
[0046] For example, the pre-warning information is that device A is expected to fail after 12 hours.
[0047] In this embodiment, by collecting and analyzing the fault signals of each device, the combination model is used for fault prediction of the device, the accurate positioning of the fault is accurately realized, the potential hardware fault is identified in advance, the operation and maintenance team can replace the risk device in time, the stability of the system operation is ensured, and the service quality of the system is ensured.
[0048] In another embodiment of the application, before monitoring the fault signals of each device link through the basic input and output system, the method further comprises: S21: scanning each high-speed serial bus, bridge and endpoint device through the basic input and output system.
[0049] In this embodiment, the basic input and output system starts from the Host Bridge and Root Complex (core hub connecting processor, memory and PCIe device), recursively scans all high-speed serial (PCIe) buses, bridges and high-speed serial bus devices.
[0050] S22: Assigning a corresponding function number to each function corresponding to the high-speed serial bus device, the function number at least including the bus number, device number and function number corresponding to the high-speed serial bus device.
[0051] In the embodiment, in the scanning process, each high-speed serial bus device is assigned a corresponding function number (BDF address), including bus number, representing the bus where the device is located, device number, representing the device on the bus, function number, representing the logical function on the device, and one physical device can contain multiple logical functions (for example, each port of a multi-port network card can be an independent function).
[0052] For example, the BDF number of a certain device is 20.01.2.
[0053] S23: Establish a corresponding topology tree for all high-speed serial bus devices.
[0054] In the embodiment, the topology tree is used to represent the structure of each device in the system, effectively representing the relationship between each device.
[0055] In the embodiment, after the basic input / output system completes enumeration and scanning of PCIe devices, a completed PCIe topology tree (PCIe Tree) is constructed, which can represent all PCIe Root Ports (Root Complex physical interface components that carry downstream PCIe device connection links), Switches, Bridges, all endpoint devices, their connection relationships (parent-child relationships), BDF addresses and key configuration information of each device / function in the system.
[0056] S24: Send the function number of each high-speed serial bus device in the topology tree to the baseboard management controller through a predefined interface.
[0057] In the embodiment, the basic input / output system sends the function number of each high-speed serial bus device in the topology tree to the baseboard management controller through a predefined interface. The baseboard management controller can determine the corresponding high-speed serial bus device of each device link according to the topology tree.
[0058] In the embodiment, by pre-establishing the topology tree, a mapping relationship between the logical position and the physical slot of each PCIe device is established. When fault positioning is performed, the corresponding fault device can be determined according to the link problem, thereby improving the speed and accuracy of fault positioning.
[0059] In another embodiment of the present application, the method further comprises: S31: In the case where the fault type corresponding to the device link is a second fault type, determining the high-speed serial bus device corresponding to the device link according to the device mapping table.
[0060] In the embodiment, the second fault type is uncorrectable error, and in a case where it is determined that the fault type corresponding to the device link is the second fault type, the high-speed serial bus device corresponding to the device link is determined according to the device mapping table.
[0061] S32: Send the alarm information corresponding to the high-speed serial bus device to the baseboard management controller.
[0062] In the embodiment, after it is determined that the fault is the second fault type and the corresponding high-speed serial bus device is determined, the corresponding alarm information is generated, and the alarm information is directly reported to the baseboard management controller. The baseboard management controller starts an alarm process and informs the operation and maintenance personnel that the device has uncorrectable error.
[0063] In another embodiment of the present application, the determination of the fault type corresponding to the device link according to the fault signal comprises: S41: Read the fault signal to determine the fault information of the device link.
[0064] In the embodiment, the fault information includes what kind of fault occurs in the device link.
[0065] In the embodiment, the basic input / output system reads the fault signal to determine the fault information of the device link and knows the specific fault of the device link.
[0066] S42: Determine whether the fault of the device link belongs to correctable error according to the fault information.
[0067] In the embodiment, the basic input / output system reads the fault information to determine whether the fault of the device link belongs to correctable error.
[0068] For example, a data transmission error in the link is correctable error, which can be verified by retransmission, and the link disconnection is uncorrectable error, which needs to be repaired or replaced by manual operation.
[0069] S43: In a case where the fault belongs to correctable error, determine that the fault type corresponding to the device link is the first fault type.
[0070] In the embodiment, in a case where the fault belongs to correctable error, it is determined that the fault type corresponding to the device link is the first fault type.
[0071] S44: In a case where the fault does not belong to correctable error, determine that the fault type corresponding to the device link is the second fault type.
[0072] In the embodiment, in a case where the fault belongs to uncorrectable error, it is determined that the fault type corresponding to the device link is the second fault type.
[0073] In this embodiment, the fault type is classified according to the fault signal, which is beneficial for subsequent processing according to the fault type and accurate positioning of the equipment fault.
[0074] In another embodiment of the present application, the device link is fault predicted through a preset combination model, including: S51: The error rate and error period of the device link are counted to obtain the time characteristics of the fault corresponding to the device link.
[0075] In this embodiment, the data residual layer obtains the fault information of the device through system logs and corresponding interfaces. For some key monitored parameters, the error rate and error period of the device link are counted to obtain the time characteristics of the fault corresponding to the device link.
[0076] For example, the error rate of a correctable error is 100 times / hour, and the period is 0:00-2:00.
[0077] S52: The errors of the device link are weighted and counted to obtain the intensity characteristics of the fault corresponding to the device link.
[0078] In this embodiment, the errors of the device link are weighted and counted, and the weight of a fatal error (FATAL error) is multiplied by 2 to obtain the intensity characteristics of the device link.
[0079] For example, if the link occurs 3 fatal errors + 5 non-fatal errors → weighted count = (3x2) + 5 = 11 S53: The burst errors of the device link are counted to obtain the mode characteristics of the fault corresponding to the device link.
[0080] In this embodiment, the burst errors of the device link are counted, and the standard deviation analysis of the occurrence of errors is performed to obtain the mode characteristics of the fault corresponding to the device link.
[0081] S54: The error correlation of a plurality of device links is analyzed to obtain the correlation characteristics of the fault corresponding to the device link.
[0082] In this embodiment, the error correlation of a plurality of device links is analyzed to obtain the correlation characteristics of the fault corresponding to the device link.
[0083] S55: The time characteristics, the intensity characteristics, the mode characteristics, and the correlation characteristics are input into the preset combination model.
[0084] In this embodiment, after obtaining the time feature, the intensity feature, the pattern feature and the correlation feature, the features are input into a preset combination model.
[0085] S56: Obtain a plurality of fault prediction information by each preset model in the prediction model combination.
[0086] In this embodiment, the preset combination model analyzes possible faults of the corresponding high-speed serial bus device in a future period of time according to the obtained features, and each model obtains a result, and then a plurality of results are obtained.
[0087] For example, the combination model includes an LSTM network, an isolation forest model and an exponential smoothing model, and three fault information are obtained.
[0088] S57: The plurality of fault prediction information is weighted and combined according to a preset weight to obtain a corresponding fault prediction result.
[0089] In this embodiment, the fault information obtained by each model is weighted and combined according to a preset weight to obtain a corresponding fault prediction result.
[0090] In this embodiment, the combination model analyzes the fault information from various aspects to predict possible faults of the device, and the weight of each model is different. For example, the proportion of a model with strong prediction ability such as a neural network is larger. Finally, the prediction results obtained by each model are weighted and combined to obtain the final fault prediction result.
[0091] In another embodiment of the present application, the method further comprises: S61: In the case that the fault type corresponding to the device link is a first fault type and the number of errors does not reach the fault prediction threshold, recording the fault information corresponding to the device link by logging.
[0092] In this embodiment, in the case that the fault type corresponding to the device link is a first fault type and the number of errors does not reach the preset fault prediction threshold, the number of correctable errors of the device link is small, which is not enough to cause the device to fail. At this time, the fault information corresponding to the device link is recorded by logging.
[0093] In this embodiment, in the case that the fault type is a first fault type and the number of errors does not reach the fault prediction threshold, the fault information is recorded in the log, which is convenient for the staff to understand the faults that have occurred.
[0094] In another embodiment of the present application, the method further comprises: S71: In the case that the fault type corresponding to the device link is the first fault type and the number of errors reaches the alarm threshold, determining the high-speed serial bus device corresponding to the device link according to the device mapping table.
[0095] In the embodiment, in the case that the fault type corresponding to the device link is the first fault type and the number of errors reaches the alarm threshold, it is indicated that the number of errors has reached a certain scale, a large number of correctable errors have occurred, which is also an abnormal phenomenon. At this time, the high-speed serial bus device corresponding to the device link needs to be determined according to the device mapping table. According to the device link, the corresponding device number can be found, and then the corresponding device can be determined.
[0096] S72: Generating alarm information corresponding to the high-speed serial bus device.
[0097] In the embodiment, after the high-speed serial bus device is determined, the alarm information corresponding to the high-speed serial bus device is generated according to the specific information of the first fault type.
[0098] S73: Sending the alarm information to the baseboard management controller.
[0099] In the embodiment, after the alarm information is generated, the alarm information is sent to the baseboard management controller to start the alarm process.
[0100] For example, the number of correctable errors is counted by using a CE funnel mechanism (Correctable Errors Funnel Mechanism), that is, the number of errors of the first error type is counted. The mechanism is a fault-tolerant technology for correctable error storm (CE storm) of a server memory. For a plurality of correctable errors occurring in a short time, the correctable error is filtered, and the reporting frequency of the error is adjusted. By dynamically filtering the error reporting frequency, the system is prevented from being interrupted due to high-frequency interruption, thereby avoiding business interruption.
[0101] When the number of errors does not reach the alarm threshold, the CE Mask mechanism is used to shield interference, that is, the high-frequency correctable error is shielded, the system performance is prevented from being reduced due to redundant errors, the alarm information (including BDF, error type, timestamp, etc.) is reported to the BMC for recording, the fault prediction algorithm of the BMC is used for diagnosis and fault warning, and if the alarm threshold is reached, the alarm is directly performed.
[0102] In the embodiment, when the number of errors of the first error type is large, the alarm information is directly generated to alarm the baseboard management controller, thereby ensuring stable operation of the system.
[0103] Reference Figure 2 , Figure 2is a decision logic diagram proposed by an embodiment of the present application, as shown in Figure 2 When a new error event occurs, the error event is updated to the buffer, it is determined whether the number of error events reaches a prediction threshold, if the number of errors exceeds the prediction threshold, detailed analysis is triggered, fault prediction is started, mixed prediction is performed through multiple models, when the probability of failure of the device is greater than a certain value, for example, 85%, a corresponding early warning is generated, if the number of errors does not exceed the prediction threshold, or the probability of failure is less than a certain value, monitoring is continued.
[0104] Referring to Figure 3 , Figure 3 is a fault positioning flow diagram proposed by an embodiment of the present application, as shown in Figure 3 First, the PCIe device is enumerated through the BIOS to construct a PCIe tree, and the BDF of the device is pushed to the BMC through the BIOS, and the BMC constructs a PCIe device mapping table according to the whole machine PCIe topology and device BDF information. When the BIOS detects the fault signal of the PCIe device, it is determined whether it is a correctable error or an uncorrectable error. When it is a correctable error, the BIOS records the error device BDF and uses the CE funnel mechanism to record the error number. When the alarm threshold is not reached, the BIOS uses the CE Mask mechanism and reports the device BDF information to the BMC for recording. The BMC counts the error types and times of the CE error device, combines the fault prediction algorithm, and performs early warning. When the alarm threshold is reached and the error is an uncorrectable error, the BIOS reports the device BDF information to the BMC for alarm.
[0105] In the above embodiments of the present application, in the case of obtaining the fault information of the device link, for correctable errors, the combined model is used for analysis to predict the possible failure of the device, and for uncorrectable errors, the device is directly reported to the baseboard management controller. The alarm of real-time failure and the early warning of possible future errors are achieved, the intelligent positioning of the PCIe device failure is realized, the operation and maintenance workload of the PCIe device is saved, and the stable operation of the server is effectively ensured, which provides intelligent protection for the reliability of the infrastructure.
[0106] Based on the same inventive concept, an embodiment of the present application provides a fault positioning device. Referring to Figure 4 , Figure 4 is a schematic diagram of a fault positioning device 400 proposed by an embodiment of the present application. As shown in Figure 4 The device comprises: a fault signal monitoring module 401, configured to monitor the fault signal of each device link through a basic input / output system; a fault type determination module 402, configured to determine the fault type corresponding to the device link according to the fault signal; The fault prediction module 403 is configured to perform fault prediction on the device link by using a preset combination model when the fault type corresponding to the device link is a first fault type and the number of errors reaches a fault prediction threshold. The first device determination module 404 is configured to determine the high-speed serial bus device corresponding to the device link according to a preset device mapping table. The early warning information generation module 405 is configured to generate early warning information corresponding to the high-speed serial bus device according to the fault prediction result.
[0107] Optionally, the apparatus further comprises a device scanning module configured to scan each high-speed serial bus, bridge and endpoint device by using a basic input / output system. The function number module is configured to assign a corresponding function number to each function of the high-speed serial bus device, wherein the function number at least includes a bus number, a device number and a function number corresponding to the high-speed serial bus device. The topology tree establishment module is configured to establish a corresponding topology tree for all the high-speed serial bus devices. The function number sending module is configured to send the function number of each high-speed serial bus device in the topology tree to a baseboard management controller through a predefined interface.
[0108] Optionally, the apparatus further comprises: The second device determination module is configured to determine the high-speed serial bus device corresponding to the device link according to the device mapping table when the fault type corresponding to the device link is a second fault type. The first alarm information sending module is configured to send the alarm information corresponding to the high-speed serial bus device to a baseboard management controller.
[0109] Optionally, the fault type determination module comprises: The fault information determination submodule is configured to read the fault signal and determine the fault information of the device link. The error type judgment submodule is configured to determine whether the fault of the device link belongs to a correctable error according to the fault information. The first fault type determination submodule is configured to determine that the fault type corresponding to the device link is a first fault type when the fault belongs to a correctable error. The second fault type determination submodule is configured to determine that the fault type corresponding to the device link is a second fault type when the fault does not belong to a correctable error.
[0110] Optionally, the fault prediction module comprises: a first feature determination submodule configured to statistically determine an error rate and an error period of the device link, and obtain a time feature of a fault corresponding to the device link; a second feature determination submodule configured to statistically determine errors of the device link, and obtain an intensity feature of the fault corresponding to the device link; a third feature determination submodule configured to statistically determine burst errors of the device link, and obtain a mode feature of the fault corresponding to the device link; a fourth feature determination submodule configured to analyze error correlations of a plurality of the device links, and obtain a correlation feature of the fault corresponding to the device link; a feature input submodule configured to input the time feature, the intensity feature, the mode feature, and the correlation feature into the preset combination model; a fault information obtaining submodule configured to obtain a plurality of fault prediction information by each preset model in the prediction model combination; a prediction result obtaining submodule configured to combine the plurality of fault prediction information according to a preset weight, and obtain a corresponding fault prediction result.
[0111] Optionally, the apparatus further includes: a fault information recording module configured to record, by a log, fault information corresponding to the device link in a case where a fault type corresponding to the device link is a first fault type and a number of errors does not reach the fault prediction threshold.
[0112] Optionally, the apparatus further includes: a second device determination module configured to determine, according to the device mapping table, a high-speed serial bus device corresponding to the device link in a case where a fault type corresponding to the device link is a first fault type and a number of errors reaches an alarm threshold. a third alarm information generation module configured to generate alarm information corresponding to the high-speed serial bus device; a second alarm information sending module configured to send the alarm information to a baseboard management controller.
[0113] Based on the same inventive concept, another embodiment of the present application provides a product including a computer program / instruction, which, when executed by a processor, implements steps in the signal line design method according to any of the above embodiments of the present application.
[0114] Based on the same inventive concept, another embodiment of the present application provides a readable storage medium having a computer program stored thereon, which, when executed by a processor, implements steps in the fault positioning method according to any of the above embodiments of the present application.
[0115] Based on the same inventive concept, another embodiment of the present application provides an electronic device, Figure 5 FIG. 1 is a schematic diagram of an electronic device 500 according to an embodiment of the present application, which comprises a memory 502, a processor 501, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the fault locating method according to any of the above embodiments of the present application when executed.
[0116] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts are described in the part of the method embodiment.
[0117] Each embodiment in the present specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts between each embodiment can be referred to each other.
[0118] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can be in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0119] The embodiments of the present application are described with reference to flowcharts and / or block diagrams according to the method, terminal device (system), and computer program product of the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the computer or other programmable data processing terminal device produce a device implemented in the flowcharts and / or block diagrams. Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks.
[0120] These computer program instructions can also be stored in a computer-readable storage medium that can guide the computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable storage medium produce a manufactured product comprising instruction devices that implement the flowcharts and / or block diagrams. Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks. Figure 1 The device that implements the function specified in one flow or multiple flows and / or blocks.
[0121] These computer program instructions can also be loaded into a computer or other programmable data processing terminal device, so that a series of operational steps are performed on the computer or other programmable terminal device to generate a computer-implemented process, thus the instructions executed on the computer or other programmable terminal device provide a process for implementing the functions specified in the flowchart Figure 1 one flow or multiple flows and / or the functions specified in the block Figure 1 one block or multiple blocks.
[0122] Although the preferred embodiments of the present application have been described, those skilled in the art who understand the basic inventive concept after getting to know the present application can make additional changes and modifications to the embodiments. Therefore, the appended claims are intended to cover the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0123] Finally, it should also be noted that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or terminal device including the element.
[0124] The above describes in detail the fault positioning method, product, device and storage medium provided by the present application. The principles and implementation manners of the present application are described by applying specific examples in this document. The above description of the embodiments is only for helping to understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application range; in summary, the content of the present description should not be understood as a limitation of the present application.
Claims
1. A fault location method characterized by, The method comprises: monitoring a fault signal of each device link through a basic input / output system; determining a fault type corresponding to the device link according to the fault signal; when the fault type corresponding to the device link is a first fault type and the number of errors reaches a fault prediction threshold, performing fault prediction on the device link through a preset combination model; determining a high-speed serial bus device corresponding to the device link according to a preset device mapping table; generating early warning information corresponding to the high-speed serial bus device according to the fault prediction result.
2. The fault locating method of claim 1, wherein, Before monitoring the fault signal of each device link through the basic input / output system, the method further comprises: scanning each high-speed serial bus, bridge and endpoint device through the basic input / output system; allocating a corresponding function number to a function corresponding to each high-speed serial bus device, wherein the function number at least comprises a bus number, a device number and a function number corresponding to the high-speed serial bus device; establishing a corresponding topology structure tree for all the high-speed serial bus devices; sending the function number of each high-speed serial bus device in the topology structure tree to a baseboard management controller through a predefined interface.
3. The fault locating method of claim 1, wherein, The method further comprises: when the fault type corresponding to the device link is a second fault type, determining a high-speed serial bus device corresponding to the device link according to the device mapping table; sending the alarm information corresponding to the high-speed serial bus device to the baseboard management controller.
4. The fault locating method of claim 1, wherein, The determination of the fault type corresponding to the device link according to the fault signal comprises: reading the fault signal to determine fault information of the device link; determining whether the fault of the device link belongs to a correctable error according to the fault information; when the fault belongs to a correctable error, determining that the fault type corresponding to the device link is a first fault type; when the fault does not belong to a correctable error, determining that the fault type corresponding to the device link is a second fault type.
5. The fault locating method of claim 1, wherein, The fault prediction on the device link through the preset combination model comprises: statistically obtaining time characteristics of the fault corresponding to the device link by counting error rates and error time periods of the device link; statistically obtaining intensity characteristics of the fault corresponding to the device link by weighting the errors of the device link; statistically obtaining mode characteristics of the fault corresponding to the device link by counting burst errors of the device link; obtaining correlation characteristics of the fault corresponding to the device link by analyzing error correlations of multiple device links; inputting the time characteristics, the intensity characteristics, the mode characteristics and the correlation characteristics into the preset combination model; obtaining multiple fault prediction information through each preset model in the prediction model combination; weighting and combining the multiple fault prediction information according to a preset weight to obtain a corresponding fault prediction result.
6. The fault locating method of claim 1, wherein, The method further comprises: when the fault type corresponding to the device link is the first fault type and the number of errors does not reach the fault prediction threshold, recording fault information corresponding to the device link through a log.
7. The fault locating method of claim 6, wherein, The method further comprises: in a case where the fault type corresponding to the device link is a first fault type and the number of errors reaches an alarm threshold, determining a high-speed serial bus device corresponding to the device link according to the device mapping table; generating alarm information corresponding to the high-speed serial bus device; sending the alarm information to a baseboard management controller.
8. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instruction is executed by a processor to implement the steps in the method of any one of claims 1 to 7.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to implement the steps in the method of any one of claims 1 to 7.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps in the method of any one of claims 1 to 7.