Error Determination Method, Apparatus, Device, and Storage Medium for PCIe Link
By detecting error events in the PCIe link and performing layer-by-layer repair processing, the problem of low accuracy of the traditional PCIe early warning mechanism is solved, and more accurate error prediction and wider error early warning coverage are achieved.
Patent Information
- Application Number
- CN202211265086.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-10-14
AI Technical Summary
The traditional PCIe early warning mechanism has low accuracy and small coverage, making it difficult to effectively identify and repair non-fatal errors in PCIe links.
By adding an error event count every time an event that meets the preset condition is detected, and determining whether to perform repair processing based on the preset bit error rate period and error event count, the repair processing is carried out layer by layer, and the error type of PCIe link is finally determined.
Improve the accuracy of error prediction, expand the scope of error warning, and cover more error types in PCIe links.
Smart Images

Figure CN115904772B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of communication technologies, and in particular, to a method, apparatus, device, and storage medium for determining errors in a PCIe link. Background Art
[0002] In recent years, PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard) devices have been widely used in the fields of servers and storage. Various non-fatal errors (usually manifested as packet errors) may occur in the PCIe link between the motherboard and the PCIe device. Generally speaking, it is necessary to give early warnings about the errors in the PCIe link. The traditional PCIe early warning mechanism is a very simple repairable error alarm, with low early warning accuracy and small coverage. Summary of the Invention
[0003] The present disclosure provides a method for determining errors in a PCIe link, which is applied to a root node device in the PCIe link. The method includes:
[0004] When an event that meets a preset condition is detected each time, increasing the error event count; the preset condition is that an error event occurs in a target clock cycle within the current window period;
[0005] Based on a preset bit error rate period and the error event count, determining whether to perform a repair process;
[0006] If so, performing at least one level of repair process on the PCIe link in sequence; wherein, after each repair process is executed, repeating the step of increasing the error event count when an event that meets the preset condition is detected each time, and each time the repair process targets one level, and the repair process of the next level is determined based on the cumulative number of times of the repair process of the previous level, and different levels correspond to different objects in the PCIe link;
[0007] When a repair cut-off condition is met, determining the error type of the PCIe link based on the repair process of the last level.
[0008] Optionally, the determining whether to perform a repair process based on a preset bit error rate period and the error event count includes:
[0009] Detecting whether the preset bit error rate period is reached currently; wherein, the bit error rate period is greater than the window period;
[0010] If so, reducing the error event count by a preset difference;
[0011] If the reduced error event count reaches a preset first threshold, it is determined to perform repair processing on the PCIe link in sequence;
[0012] If the reduced error event count does not reach the preset first threshold, it is determined not to perform repair processing on the PCIe link.
[0013] Optionally, the performing at least one level of repair processing on the PCIe link where it is located includes:
[0014] Performing first-level repair processing on the signal transmission level of the PCIe link;
[0015] When the cumulative repair times of the first-level repair processing are greater than or equal to a preset second threshold, perform second-level repair processing, and the second level includes the hardware level.
[0016] Optionally, the performing first-level repair processing on the signal transmission level of the PCIe link includes:
[0017] On the signal transmission level, perform repair processing on at least one object of the PCIe link respectively; wherein, the at least one object includes: signal quality and / or signal transmission speed.
[0018] Optionally, on the signal transmission level, performing repair processing on at least one object of the PCIe link respectively includes:
[0019] For each of the objects of the PCIe link, perform cross-repair processing, wherein the previous repair processing is different from the subsequent repair processing;
[0020] Or, in one repair, perform repair processing on at least one of the objects of the PCIe link;
[0021] Or, perform repair processing on the current object of the PCIe link. If the number of times of repair processing on the current object reaches a preset number of times, perform repair processing on the next object of the current object until the repair is successful or the cumulative repair times reach the preset second threshold.
[0022] Optionally, the first-level repair processing includes speed reduction processing and / or rebalancing processing.
[0023] Optionally, the root node device is connected to multiple terminal devices. When the cumulative repair times of the first repair processing reach a preset second threshold, performing second-level repair processing includes:
[0024] When the cumulative repair times of the first repair processing reach the preset second threshold, locate the target terminal device with a fault;
[0025] Perform at least one hardware repair process on the target terminal device; wherein, different hardware repair processes are used to recover the target terminal device.
[0026] Optionally, performing at least one hardware repair process on the port where the target terminal device is located includes:
[0027] Perform a reset process on the port where the target terminal device is located;
[0028] When the cumulative number of reset times in the reset process exceeds a preset third threshold, perform a power-off and then power-on process on the target terminal device.
[0029] Optionally, the window period includes a plurality of clock cycles, and the target clock cycle is the first clock cycle within the window period.
[0030] Optionally, the method further includes:
[0031] Detect whether the preset bit error rate period is reached currently; wherein, the bit error rate period is greater than the window period;
[0032] If so, reduce the error event count according to a preset difference;
[0033] When the error event count is greater than or equal to a preset first threshold, perform at least one level of repair process on the PCIe link in sequence, including:
[0034] When the reduced error event count is greater than or equal to the preset first threshold, perform at least one level of repair process on the PCIe link in sequence.
[0035] Optionally, determining the error type of the PCIe link based on the last level of repair process includes:
[0036] Determine the error type of the PCIe link based on the cumulative number of repair times of the last level of repair process, and / or based on the repair result of the last level of repair process.
[0037] The present disclosure also provides an error determination device for a PCIe link, which is applied to a root node device in the PCIe link. The method includes:
[0038] A counting module, configured to increase the error event count when detecting that an event satisfying a preset condition occurs; the preset condition is that an error event occurs in the target clock cycle within the current window period;
[0039] A repair processing module, configured to perform at least one level of repair processing on the PCIe link in sequence when the error event count is greater than or equal to a preset first threshold; wherein, after each repair processing is performed, the step of increasing the error event count when an event satisfying a preset condition is detected each time is repeated, and each repair processing targets one level, and the repair processing of the next level is determined based on the cumulative number of times of the repair processing of the previous level, and different levels correspond to different objects in the PCIe link;
[0040] An error determination module, configured to determine the error type of the PCIe link based on the repair processing of the last level when a repair cut-off condition is satisfied.
[0041] The present disclosure further provides an electronic device, and a computer program stored therein causes a processor to execute the error determination method of the PCIe link.
[0042] The present disclosure further provides a computer-readable storage medium, and a computer program stored therein causes a processor to execute the error determination method of the PCIe link
[0043] By adopting the technical solution of the embodiment of the present application, the error event count can be increased when an event satisfying a preset condition is detected each time; then, based on the error event count and the preset bit error rate period, it is determined whether repair processing needs to be performed. If so, at least one level of repair processing is performed on the PCIe link in sequence to perform repairs at different levels on the PCIe link; and based on the repair processing of the last level, the error type of the PCIe link is determined.
[0044] Since after each repair processing is performed, the step of increasing the error event count when an event satisfying a preset condition is detected each time is repeated, and each repair processing targets one level, and the repair processing of the next level is determined based on the cumulative number of times of the repair processing of the previous level, and different levels correspond to different objects in the PCIe link. In this way, when it is determined that the error event count reaches a certain threshold in combination with the bit error rate period, different levels of repair are performed on the PCIe link. Since each level targets an object in the PCIe link, thus, based on the results of these different levels of repair processing, it can be determined which object in the PCIe link has failed, that is, the root cause of the error event can be determined, and thus the error type can be accurately located. In this way, not only can the accuracy of error prediction be improved, but also the scope of error warning can be expanded, and more error types in the PCIe link can be covered.
[0045] The above description is only an overview of the technical solution of the present disclosure. In order to understand the technical means of the present disclosure more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present disclosure more obvious and understandable, the following specific embodiments of the present disclosure are specifically given. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. It should be noted that the sizes and shapes of the various figures in the drawings do not reflect the actual proportions, and the purpose is only to schematically illustrate the content of the present invention. The same or similar reference numerals in the drawings represent the same or similar elements or elements with the same or similar functions.
[0047] Figure 1 Schematically shows a schematic diagram of the communication environment of the present application;
[0048] Figure 2 Schematically shows a schematic diagram of the hardware environment of the method for determining errors in the PCIe link supported by the present application;
[0049] Figure 3 Schematically shows a flowchart of the steps of the method for determining errors in the PCIe link;
[0050] Figure 4 Schematically shows an exemplary flowchart of the method for determining errors in the PCIe link;
[0051] Figure 5 Schematically shows a schematic diagram of the structure of an apparatus for determining errors in a PCIe link of the present application. Detailed Embodiments
[0052] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present disclosure.
[0053] In the related art, server failures can be classified into two categories: downtime failures and non-downtime failures. Downtime failures are mainly reflected in two parts: downtime during the startup process and downtime during operation. Non-downtime failures include abnormal monitoring of power supply temperature indicators, abnormal monitoring of motherboard fans, statistical monitoring of repairable and non-fatal failures of CPU / memory / GPU / storage devices / network devices / PCIe external plug-in devices, and monitoring of the health status of components and links.
[0054] After the server is started, various non-downtime errors are inevitable during operation. In terms of early warning for PCIe external plug-in devices, generally, it is a very simple repairable error alarm. Repairable errors are a common type of non-downtime error. In the related art, the CE (Correctable Error) threshold limit strategy is used to determine whether a PCIe device has a hardware failure.
[0055] However, this strategy cannot accurately predict whether the PCIe component actually has a hardware failure, and the related art only sets repair operations such as reset when a fatal error occurs in the PCIe device. In this way, accurate prediction of the repairable error CE cannot be performed.
[0056] In view of this, the present application proposes an error early warning mechanism for PCIe. When a PCIe device frequently has repairable errors, a series of repair mechanisms can be used for gradient repair to sequentially repair the PCIe link at different levels. If an error continues to be reported, when the repair mechanism reaches the threshold, it indicates a fatal error such as a hardware failure. If no error continues to be reported, it indicates a repairable error. Therefore, by counting error events and the number of repair times of the gradient repair mechanism, the accuracy of error determination for the PCIe link can be improved.
[0057] Refer to Figure 1 As shown, a schematic diagram of the communication environment of the present application is shown. As Figure 1 shown, it includes: the CPU, GPU, BMC (Baseboard Management Controller) of the server, multiple root node devices, and multiple terminal devices connected to each root node device. Among them, the root node device is also called the Root port device. In practice, it can be a PCIe Switch (PCIe switch), and this root node device can divide the PCIe link into more segments and finally reach the terminal device.
[0058] In this embodiment, the data transmission link where a Root port device is located can be referred to as a PCIe link. Generally speaking, a Root port device has multiple downstream expansion ports, that is to say, a Root port device can expand the number of PCIe ports, so that more terminal devices can be plugged in ( Figure 1 the EP in
[0059] ). That is to say, in this PCIe link, there are a Root port device and multiple terminal devices. Different terminal devices can be called a link segment in this PCIe link.
[0060] Among them, since a Root port device plugs in multiple terminal devices, when these terminal devices receive or send data, they will reach the corresponding addresses of the CPU and GPU via the Root port device. Therefore, the Root port device can perform error warnings.
[0061] Refer to Figure 2 As shown, it shows a schematic diagram of the hardware environment of the method for determining errors in a PCIe link supported by this application. Among them, RX is the receiving end and TX is the sending end. Based on the PCIe mechanism, a device can be both a sending end and a receiving end. That is to say, in one case, device A is the sending end and device B is the receiving end. In another case, device A is the receiving end and device B is the sending end. Among them, device A and device B communicate through a PCIe link. In this PCIe link communication, signal routing needs to be carried out via the Root port device.
[0062] As Figure 2 shown, in this application, the Root port device can include registers. The registers are the burst protection mechanism of PCIe, which can effectively prevent too many errors from occurring instantaneously. Only one error is allowed to be recorded within a specified time to avoid the generation of error storms. Therefore, it is possible to determine whether to perform multiple gradient repair processes based on the error events recorded in the registers, and determine the type of error based on the cumulative count of the repair processes recorded in the registers.
[0063] Among them, the method for determining errors in the PCIe link of this application can be applied to the root node device in the PCIe link. Among them, this root node device can also be called a Root port device. A root node device can connect multiple terminal devices through the expanded downstream ports, and the upstream port of a root node device can connect to a BMC device.
[0064] Next, in combination with Figure 1The communication environment shown and Figure 2 the hardware environment shown are introduced for the error determination method of the PCIe link of this application. Refer to Figure 3 shown, a step flowchart of the error determination method of the PCIe link of this application is shown. As Figure 3 shown, it may specifically include the following steps:
[0065] Step S301: When an event that meets a preset condition is detected each time, increase the error event count; the preset condition is that an error event occurs in the target clock cycle within the current window period;
[0066] Step S302: Based on the preset bit error rate period and the error event count, determine whether to perform a repair process;
[0067] Step S303: In the case of performing a repair process, perform at least one level of repair process on the PCIe link in sequence;
[0068] Specifically, after each execution of the repair process, step S301 can be repeated; among them, each execution of the repair process corresponds to one level, different repair processes are used to perform different levels of repair on the PCIe link, different levels correspond to different objects in the PCIe link; and the next level of repair process is determined based on the cumulative number of times of the previous level of repair process.
[0069] Step S304: When the repair cut-off condition is triggered, based on the cumulative repair times of the last level of repair process, determine the error type of the PCIe link, and the error type includes non-fatal type and fatal type.
[0070] In this embodiment, generally there are multiple window periods, and one window period can include multiple clock cycles. Among them, the number of clock cycles can be determined according to the working frequency of the PCIe link. The clock cycle is a quantity of time. The higher the working frequency, the smaller the clock cycle. Among them, the window period can be set by the user according to the actual situation and is not limited here.
[0071] Among them, when each window period arrives, it can be detected whether an error event occurs in the target clock cycle within this window period. If an error event occurs, the error event count is increased. For example, after an error event occurs, the error event count is incremented by 1. In some cases, the target clock cycle can also be multiple, and these multiple target clock cycles can be consecutive clock cycles or can be non-consecutive clock cycles. Specifically, in order to avoid an error storm caused by recording multiple error events in one window period, the target clock cycle is preferably one clock cycle.
[0072] More specifically, the target clock cycle can be selected as any one of the clock cycles within the window period. For example, it can be the first clock cycle, the second clock cycle, etc.
[0073] Among them, in each window period, as long as an error event is detected in the target clock cycle, the error event count is incremented by 1, and so on in a loop. That is to say, as long as an event that meets the preset condition (an error event occurs in the target clock cycle within the current window period) is detected, the error event count will be increased.
[0074] Among them, through the bit error rate period, some events caused by bit errors can be excluded from being counted as error events. Therefore, the bit error rate period can be set. When the bit error rate period arrives, the error event count can be subtracted. Thus, the final error event count can be determined according to the preset bit error rate period and the current error event count. When the final error event count exceeds the threshold, it indicates that substantial error events have occurred in multiple window periods. Therefore, it is necessary to give a warning to the PCIe link to determine which type of error it is. In this application, as long as it is determined that repair processing is required, a multi-level repair strategy for the PCIe link will be started. In this multi-level repair strategy, various repair processes can be performed on the PCIe link, and different levels correspond to different types of repair processes.
[0075] Specifically, when performing at least one level of repair processing, in each repair processing, the step of incrementing the error event count when an event that meets the preset condition is detected can be repeated. That is to say, every time a repair processing is executed, it returns to step S301 to continue detecting the error event count. As long as the error event count exceeds the preset first threshold again, another repair processing will be executed.
[0076] In specific implementation, each repair processing can be only for one level. In this application, the levels can include the physical level and the software level. The software level can execute repair through a relatively soft strategy, and the physical level can execute repair through a relatively physical operation strategy. Among them, the objects that the software level can target in the PCIe link can include: transmission signals, data packets, transmission protocols, etc., and the physical level can target the physical devices existing in the PCIe link. For example, it can be for the terminal device or the downstream expansion port of the root node device.
[0077] Since each repair processing targets one level, in this way, the error type can be located through a multi-gradient repair strategy.
[0078] In a specific implementation, whether to enter the repair process at the next level and which level of repair process to enter can be determined by the cumulative number of times of the repair process at the previous level. That is to say, the cumulative number of times of performing a repair process at a certain level on the PCIe link can determine whether to enter the repair process at the next level. In some cases, when entering the repair process at the next level, it can also determine which specific level of repair process to enter.
[0079] Exemplarily, after performing the repair process on level A multiple times, when the cumulative number of times of the repair process on level A reaches a certain threshold, it indicates that the repair process at the next level needs to be entered. When the condition that the next error event count exceeds the preset first threshold arrives, the repair process at the next level is entered. Among them, which specific level it is can be determined according to the cumulative number of times. If the cumulative number of times reaches a certain threshold 1, the repair process at the next level can be the repair process on level B. If the cumulative number of times reaches a certain threshold 2, the repair process at the next level can be the repair process on level C. When the cumulative number of times of the repair process on level B or C reaches the corresponding threshold again, the repair process at the next level D can be entered.
[0080] Among them, the threshold 1 and the threshold 2 can be set by the user himself. It should be noted that when entering the repair process at the next level, the repair process at the previous level can be paused.
[0081] Among them, after each repair process, the above step S301 can be repeatedly executed. Thus, based on the repair process at the last level, the error type of the PCIe link can be determined. Specifically, the error type can include a correctable type and a fatal type. The error of the correctable type is also called a repairable error, which is a common non-crashing type of error. The error of the fatal type is also called a fatal error, which is an error caused by a hardware failure.
[0082] Of course, in some embodiments, the error type may include other types according to requirements, such as a type including abnormal temperature, a type including abnormal fan, etc. It should be noted that in practice, which levels of repair processes to perform can be determined according to the error type to be detected. For example, if the error type to be detected is a repairable type and a fatal type, at least one level of repair process to be performed can include a software-level repair process and a hardware-level repair process.
[0083] In this embodiment, since the repair process of the next layer is determined based on the cumulative number of times of the repair process of the previous layer, that is, whether to enter the repair process of the next layer needs to be determined based on the repair process of the previous layer. Thus, it can be understood as a progressive repair. Each repair targets a possible fault on a certain layer. Therefore, when the repair termination condition is triggered, it can be determined whether the fault on this layer is repaired through the repair process of the last layer. Thus, based on the repair process of the last layer, the error type can be determined.
[0084] Among them, the repair termination condition can be the following conditions: the repair processes of all layers have been performed, and the repair process of the last layer has been executed a preset number of times, or the result of the repair process is a successful repair, or the error event count changes from exceeding a preset first threshold to being lower than the preset first threshold; when any of the above conditions is triggered, it can indicate that the repair termination condition is met.
[0085] Adopting the technical solution of the embodiment of the present application, since after each repair process is executed, the step of increasing the error event count is repeated every time an event that meets the preset conditions is detected, and each repair process targets a certain layer, and the repair process of the next layer is determined based on the cumulative number of times of the repair process of the previous layer. Different layers correspond to different objects in the PCIe link. In this way, when the error event count reaches a certain threshold, by performing repairs on different layers of the PCIe link, since each layer targets an object in the PCIe link, the object in the PCIe link where the fault occurs can be determined based on the results of these different layer repair processes. That is to say, the root cause of the error event can be determined, so as to accurately locate the error type. In this way, not only can the accuracy of error prediction be improved, but also the scope of error warning can be expanded, covering more error types in the PCIe link.
[0086] In an alternative example, when determining the error type of the PCIe link based on the repair process of the last layer, the error type of the PCIe link can be determined based on the cumulative repair times of the repair process of the last layer; and / or the error type of the PCIe link can be determined based on the repair result of the repair process of the last layer.
[0087] In one case, the cumulative number of repair processes at the last level of repair processing can be compared with a preset threshold to determine the error type. For example, if the cumulative number of repair processes at the last level of repair processing reaches the preset threshold, it can be determined that the repair processing at this level has been executed multiple times. Whether the error is repaired or not, the error type can be determined. For example, the last level of repair processing is the hardware level, and the repair processing has been executed 3 times at this hardware level. Whether the error is repaired or not, it can be determined that its error type is a fatal type of error.
[0088] In another case, the error type can be determined based on the repair result of the last level of repair processing. Specifically, if the repair result is a successful repair, it can be determined that its error type is a repairable error. If the repair result is an unsuccessful repair, it can be determined that its error type is a fatal error. In this case, the last level of repair processing can be executed only once, and the error type can be determined according to the result of one repair processing. Thus, the number of repair processes can be reduced, and a high load can be avoided being brought to the root node device.
[0089] In still another case, the cumulative number of repair processes and the repair result of the last level of repair processing can be combined to determine the error type. Specifically, when the cumulative number of repair processes at the last level of repair processing reaches the preset threshold, the error type can be determined according to the result of the repair processing. For example, when the cumulative number of repair processes at the last level of repair processing reaches the preset threshold and the repair result is a successful repair, the error type can be characterized as a repairable error. If the repair result is an unsuccessful repair, the error type can be characterized as a fatal error. In this case, the accuracy of determining the error type can be improved.
[0090] In an optional example, the judgment process of determining whether to perform repair processing based on a preset bit error rate period and the error event count can be as follows:
[0091] It can be detected whether the current time reaches the preset bit error rate period; if so, the error event count is reduced according to a preset difference. Then, when the reduced error event count is greater than or equal to a preset first threshold, it is determined to perform at least one level of repair processing on the PCIe link in sequence.
[0092] In this embodiment, the bit error rate period can be greater than the window period. That is to say, during the operation of the PCIe link, it can be determined whether an event satisfying a preset condition occurs according to the arriving window period. If it occurs, the error event count is incremented by 1. If it does not occur, the error count is not incremented. However, when the bit error rate period arrives, the error event count is subtracted.
[0093] Among them, the preset difference for reducing the error event count can be 1, or it can be other values.
[0094] Among them, when the error event count is increased, it can be detected whether the bit error rate period is reached. If it is reached, the error event count is subtracted. After that, the value after subtracting the error event count is compared with the preset first threshold. If it is greater than or equal to the preset first threshold, a multi-level repair process is entered. If it is less than the preset first threshold, the multi-level repair process is not entered, and step S301 is continued to be repeated.
[0095] Of course, it can also be regularly detected whether the bit error rate period is reached, without having to detect only when the error event count is increased. In the case of regular detection, as long as the bit error rate period is reached, the error event count is subtracted. In this way, after subtracting the error event count, the error event count can be changed according to whether an event that meets the preset conditions occurs.
[0096] In this way, the problem of continuous accumulation of error event counts caused by the bit error rate can be avoided, thereby improving the accuracy of error type determination in this application.
[0097] Based on the above embodiments, how to perform multi-gradient repair processing is described.
[0098] When performing multi-gradient repair processing, in an alternative embodiment, the PCIe link can be repaired at two levels. One level is the signal transmission level, and the other level is the hardware level. Among them, the signal transmission level can be understood as the software level, that is, the error event is repaired by relatively soft means; the hardware level can be understood as the physical level, that is, the error event is repaired by physical means, such as unplugging the terminal device, resetting the port, etc.
[0099] Among them, when the error event is repaired by relatively soft means, objects such as the rate, quality, and bandwidth of the transmission signal in the PCIe link can be repaired to repair errors caused by software, operating systems, etc.; when the error event is repaired by physical means, objects such as the terminal device and port in the PCIe link can be repaired to repair errors caused by abnormal temperature, abnormal power supply, etc.
[0100] Among them, in this alternative embodiment, when the error event count after correction in the preset bit error rate period is greater than or equal to the preset first threshold, the error of the PCIe link is first repaired at the signal transmission level by relatively soft means. If error events still occur and the errors cannot be repaired by relatively soft means, the error of the PCIe link is then repaired at the hardware level by physical means, so as to determine the error type of the PCIe link based on the repair results at the signal transmission level and the hardware level.
[0101] In specific implementation, in the above step S303, when the error event count after being corrected through a preset bit error rate period is greater than or equal to a preset first threshold, a first-level repair process at the signal transmission level can be performed. When the cumulative repair count of the first-level repair process is greater than or equal to a preset second threshold, a second-level repair process is performed.
[0102] Among them, the first level includes the signal transmission level, and the second level includes the hardware level.
[0103] In this embodiment, the preset second threshold can be determined according to the actual situation and no special setting is made here. As described above, since each repair process targets one level, in practice, there may be multiple consecutive repair processes for one level. That is to say, there may be multiple consecutive first-level repair processes. When the cumulative repair count of the first-level repair process is greater than or equal to the preset second threshold, it indicates that the first-level repair process cannot avoid the occurrence of error events. That is to say, the cause of the error event may not be at the first level. In this case, the second-level repair process can be entered.
[0104] Among them, the signal transmission level can correspond to software setting objects of the PCIe link, such as the bandwidth setting, transmission speed setting, signal processing, etc. of the PCIe link. In this case, the first-level repair process is mainly used to repair these objects, such as adjusting the bandwidth setting, transmission speed setting, and signal processing, etc.
[0105] Among them, the hardware level can correspond to physical objects in the PCIe link, such as terminal devices, physical ports on the device, such as the downstream expansion port on the root node device, etc. In this case, the second-level repair process is mainly used to repair these physical objects, such as power on / off, reset, initialization, etc.
[0106] Among them, after the first-level repair process is performed, if it shows that the repair is successful, the second-level repair process may not be performed, that is, if the repair cut-off condition is met, the error type of the PCIe link can be determined as a repairable type based on the first-level repair process, that is, the error of the PCIe link is a repairable error.
[0107] Among them, after the repair process of the first layer is performed, if it is shown that the repair is unsuccessful, in this case, the error count will continuously accumulate, and the repair process of the first layer will be continuously performed. When the cumulative repair times of the repair process of the first layer exceed the preset second threshold, the repair of the second layer will be entered. If it is shown that the repair is successful after the repair of the second layer, that is, the repair cut-off condition is met, then based on the repair process of the second layer, it can be determined that the error type of the PCIe link is a repairable type (although the error is caused by a hardware reason, it is still repairable). If the repair of the second layer is performed multiple times and it is still shown that the repair is unsuccessful, when the cumulative repair times exceed a certain threshold, the repair cut-off condition is met. At this time, it can be determined that the error type of the PCIe link is a fatal type, that is, an irreparable type.
[0108] By adopting the technical solution of the embodiment of the present application, when the error event count exceeds the preset first threshold, the PCIe link can be repaired separately from the signal transmission layer and the hardware layer. When the repair process of the first layer reaches the preset second threshold, the repair process of the second layer will be performed. In this way, the software object and the physical object in the PCIe link can be repaired in sequence to repair the error while investigating the cause of the error, so as to cover the error types on the software and hardware in the PCIe link.
[0109] Among them, when performing the repair process of the first layer, that is, when performing the repair process of the signal transmission layer, the repair process can be performed on the signal transmission speed and the quality of the transmitted signal. In this way, the repair process can be performed on at least one object of the PCIe link; among them, the repair process of the signal transmission speed includes a speed reduction process.
[0110] Among them, the above at least one object includes: signal quality and / or signal transmission speed.
[0111] In one case, the repair process of the signal quality can include rebalancing processing, and the repair process of the signal transmission speed can include a speed reduction process. The speed reduction process is used to reduce the data transmission speed of the PCIe link, and the rebalancing process is used to improve the quality of the transmitted signal.
[0112] Among them, in the speed reduction process, the current signal transmission speed can be obtained, and then the current signal transmission speed can be reduced to the signal transmission speed of the next level. For example, if the current signal transmission speed is Gen5, the signal transmission speed can be reduced from Gen5 to Gen4 or Gen3.
[0113] Of course, in an optional example, if the cumulative number of repair times for the speed reduction process reaches a preset number, the level of the speed reduction process can be increased. Among them, the higher the level of the speed reduction process, the lower the signal transmission speed after reduction. For example, if the signal transmission speed is reduced from Gen5 to Gen4, and still caused by an error event after 3 cumulative executions, the signal transmission speed can be reduced from Gen4 to Gen3, and so on in a cycle until the cumulative repair process at the first level reaches a preset second threshold.
[0114] Of course, in some other cases, the repair process at the first level can also include bandwidth reduction processing, etc., which is not limited here.
[0115] In a specific implementation, the repair process for at least one object can be carried out in the following ways:
[0116] <The repair process at the first level>
[0117] Method 1: For each object of the PCIe link, perform cross-repair processing, where the previous repair process is different from the subsequent repair process.
[0118] In this Method 1, when performing the repair process at the first level, since multiple objects may be involved, such as signal quality and signal transmission speed, during the repair, cross-repair can be performed on multiple objects. For example, when performing the first repair, the signal quality is repaired, and when performing the second repair, the signal transmission speed can be repaired, and so on in a cross cycle; or, when performing the first repair, the signal transmission speed is repaired, and when performing the second repair, the signal quality can be repaired, and so on in a cross cycle.
[0119] By adopting this Implementation Method 1, since cross-repair processing is performed on each object, the probability of successfully repairing errors can be increased, and when a repair is successful at a certain time, the error type can be more accurately determined and the error cause can be located more precisely.
[0120] Method 2: In one repair, perform repair processing on each object of the PCIe link.
[0121] In this Method 2, in one repair process, repair processing can be performed on each object. Specifically, in sequential repairs, both the signal quality can be repaired and the signal transmission speed can be repaired. That is to say, both rebalancing processing and speed reduction processing can be performed. Of course, in this case, in one repair process, the rebalancing processing and the speed reduction processing can be performed separately, such as rebalancing processing first and then speed reduction processing, or speed reduction processing first and then rebalancing processing.
[0122] By adopting the second implementation method, since the repair process is performed on each object in one repair process, the probability of successfully repairing errors can be increased accordingly.
[0123] Method 3: Perform a repair process on the current object of the PCIe link. If the number of times of the repair process on the current object reaches a preset first number, perform a repair process on the next object of the current object until the repair is successful or the cumulative number of repair times reaches the preset second threshold.
[0124] In this Method 3, for a certain object, a repair process can be performed first. When the number of times of the repair process on this object reaches the preset number, a repair process is performed on the next object of this object, and so on in a loop until the repair is successful or the cumulative number of repair times reaches the preset second threshold.
[0125] During specific implementation, a repair process can be performed on the signal quality first, such as performing a re - equalization process. If the number of times of the re - equalization process reaches the preset number, a repair process can be performed on the signal transmission speed, such as performing a speed - reduction process. If the number of times of the speed - reduction process reaches the preset number again, a repair process is performed on the signal quality again. In this way, until the repair is successful or the cumulative number of repair times reaches the preset second threshold. It should be noted that once either the repair is successful or the cumulative number of repair times reaches the preset second threshold, the repair process at the first level can be ended.
[0126] It should be noted that in this Method 3, the preset number is less than the preset second threshold. For example, if the second preset threshold can be set to 6, the preset number can be set to 2 or 3. In practice, the preset number can be set according to actual requirements and will not be particularly limited here.
[0127] By adopting this Method 3, through continuously performing multiple repair processes on different objects, each object is independently and continuously repaired. Thus, the probability of successfully repairing errors can be increased, and the error type can be more accurately located.
[0128] <Repair process at the second level>
[0129] Among them, when performing the repair process at the second level, that is, when performing the repair process at the hardware level, at least one hardware repair process can be adopted for repair.
[0130] During specific implementation, when the cumulative number of repair times of the first repair process reaches the preset second threshold, the target terminal device with a fault can be located; and at least one hardware repair process is performed on the target terminal device.
[0131] Among them, different hardware repair processes are used to recover the target terminal device.
[0132] Among them, the hardware repair process may refer to performing physical operations on the target terminal device. For example, sending an electrical signal to the target terminal device to power on or power off the target terminal. In this way, various physical means can be used to perform hardware repair on the faulty target terminal device. Thus, it is possible to ensure as much as possible that the target terminal device can be repaired, so that the physical correctable errors of the target terminal can be repaired. Therefore, errors such as poor power contact and temporary faults caused by temperature rise can be repaired.
[0133] Among them, when positioning the target terminal device, the implementation method of the positioning is not limited in this embodiment. For example, when an error event occurs, the B / D / F number in the error event can be read, and the device corresponding to the B / D / F number can be determined as the faulty device. In this embodiment, only the above positioning method is taken as an example for introduction, and other implementation methods can refer to the introduction of this embodiment and are not limited here.
[0134] Among them, in one implementation manner, in order to more accurately predict the error type of the target terminal device, the target terminal device can be restored by using two hardware repair processes. Specifically, during implementation, the port where the target terminal device is located can be reset. When the cumulative number of reset operations during the reset process exceeds a preset third threshold, the target terminal device is powered off and then powered on again.
[0135] In this embodiment, the port where the target terminal device is located may refer to the port in the downstream expansion port of the root node device where the target terminal device is plugged in. In practice, the port can be reset, that is, an initialization operation is performed on the port to make the port reconnect to the target terminal device. When the cumulative number of times of performing the reset process exceeds the preset third threshold, it indicates that the reset process of the port can no longer repair the error. At this time, the target terminal device can be powered off and then powered on again. Specifically, the target terminal device can be removed from the root node device, that is, the target terminal is powered off, and then, with the cooperation of the operating system, the target terminal device is connected to the root node device again, that is, the target terminal can be powered on, so as to restore the connection to the target terminal.
[0136] Among them, when the cumulative number of times of performing the power-off and then power-on process exceeds a certain threshold, it can be determined that the error type is a fatal type, that is, a fatal error. When the error no longer accumulates after the power-off and then power-on process, it can be determined that the error type is a repairable error. In this way, not only can various hardware repair means be used to repair possible hardware errors, but also a more fine-grained error type can be determined based on the results of each hardware repair process.
[0137] In an alternative example, a window period includes a plurality of clock cycles, and the target clock cycle is the first clock cycle within the window period. Among them, if an error event occurs in the first clock cycle, the error event count is incremented by 1. If no error event occurs in the first clock cycle, no error event counting is performed. Among them, in the case where no error event occurs in the first clock cycle, regardless of whether an error event occurs in subsequent clock cycles, no error event counting is performed. In this way, only one error can be recorded within a specified time, and an error storm can be avoided.
[0138] Next, in combination with a specific example, an exemplary description of the method for determining errors in the PCIe link of the present application will be given:
[0139] Refer to Figure 4 As shown, an exemplary flowchart of the method for determining errors in the PCIe link of the present application is shown. As Figure 4 shown and Figure 2 shown, it includes the following processes:
[0140] S1: The root node device on the PCIe link is located between the receiving end and the sending end, and the root node device has registers.
[0141] S2: The register is set to the case where ErrorEvent = 1 / Aggr_cnt = 0 in the first clock cycle within a window period. As Figure 4 shown, if an error event occurs, ERR_Count + 1 is obtained to get ERR_Cnt;
[0142] S3: Compare the current value of ERR_Cnt with a preset first threshold, and determine whether the error rate period has been reached. If the error rate period has been reached, then ERR_Count - 1. Compare the value of ERR_Count after subtracting 1 with the preset first threshold. If the preset first threshold is not exceeded, do not enter the multi-level repair process and return to step S2.
[0143] S4: If the preset first threshold is exceeded, perform a multi-level repair process. Specifically, it includes:
[0144] S41: As Figure 2As shown, set G3LBEEN to 1 and set DEGRADEEN to 1, that is, enable G3LBEEN and DEGRADEEN to perform rebalancing processing and speed reduction processing simultaneously. After each rebalancing processing and speed reduction processing, return to step S2. Among them, if the error is repaired and no error event will be generated in this case, the error event count will no longer be incremented. It should be noted that since the error event count is determined whether to be incremented only in the first clock cycle of each window period, when the error repair is successful, the error event count will not be incremented, and at this time, the rebalancing processing and speed reduction processing will not continue.
[0145] S42: When the cumulative count of the rebalancing processing and the speed reduction processing exceeds the preset second threshold, enter the second-level repair processing, such as Figure 4 the EDPC processing shown in, that is, the problematic target terminal device can be removed and a recovery action can be performed with the cooperation of the operating system.
[0146] S43: If the cumulative repair times of the second-level repair processing exceed the preset third threshold, generate a warning message and report it to the BMC for processing. If the cumulative repair times of the second-level repair processing do not exceed the preset third threshold, continue with the EDPC processing. Each time it is processed, EDPC times + 1 is performed, and an error message is generated and reported to the operating system to indicate that it is a hardware type of error.
[0147] Among them, as Figure 4 shown, regardless of whether the second-level repair processing exceeds the preset third threshold, it can be determined that it is an error caused by a hardware device. Therefore, an alarm message of a hardware error can be reported to the operating system, and the error type can be determined as the error type of the hardware device.
[0148] Adopting the technical solution of the embodiment of the present application has the following advantages:
[0149] 1. When the error event count reaches a certain threshold, different levels of repair can be performed on the PCIe link. Since each level targets an object in the PCIe link, based on the results of these different levels of repair processing, it can be determined which object in the PCIe link has failed, that is, the root cause of the error event can be determined, so as to accurately locate the error type. In this way, not only can the accuracy of error prediction be improved, but also the scope of error warning can be expanded, covering more error types in the PCIe link.
[0150] 2. Since the error event count can be subtracted every bit error rate cycle, the error counting caused by the bit error rate can be avoided, thereby improving the accuracy of error type determination.
[0151] 3. Since error event counting is only performed when an error event occurs in the first clock cycle of each window period, and error events occurring in other clock cycles within the window period are not counted, the problem of error storms can be avoided, and the faults where errors truly occur can be detected. Thus, the accuracy of error type determination is improved.
[0152] Based on the same inventive concept, the present application also provides an error determination device for a PCIe link. Referring to Figure 5 as shown, a schematic structural diagram of the device is shown. As Figure 5 shown, the device may specifically include the following modules:
[0153] A counting module 501, configured to increase the error event count when each event that meets a preset condition is detected; the preset condition is that an error event occurs in a target clock cycle within the current window period.
[0154] A judgment module 502, configured to determine whether to perform a repair process based on a preset error code rate period and the error event count.
[0155] A repair processing module 503, configured to perform at least one level of repair processing on the PCIe link in sequence when the error event count is greater than or equal to a preset first threshold; wherein, after each repair process is executed, the step of increasing the error event count when each event that meets the preset condition is detected is repeated, and each repair process targets one level, and the repair process of the next level is determined based on the cumulative number of times of the repair process of the previous level, and different levels correspond to different objects in the PCIe link.
[0156] An error determination module 504, configured to determine the error type of the PCIe link based on the repair process of the last level when a repair cut-off condition is met.
[0157] Optionally, the repair processing module 503 includes:
[0158] A first processing unit, configured to perform a first-level repair process on the signal transmission level of the PCIe link.
[0159] A second processing unit, configured to perform a second-level repair process when the cumulative number of repair times of the first-level repair process is greater than or equal to a preset second threshold, and the second level includes a hardware level.
[0160] Optionally, the first processing unit is specifically configured to perform repair processing on at least one object of the PCIe link respectively at the signal transmission level.
[0161] Wherein, the at least one object includes: signal quality and / or signal transmission speed.
[0162] Optionally, the first processing unit is specifically configured to:
[0163] Perform cross-repair processing on each of the objects of the PCIe link, wherein the previous repair processing is different from the subsequent repair processing;
[0164] Or, in one repair, perform repair processing on at least one of the objects of the PCIe link;
[0165] Or, perform repair processing on the current object of the PCIe link. If the number of times of repair processing on the current object reaches a preset number of times, perform repair processing on the next object of the current object until the repair is successful or the cumulative number of repair times reaches the preset second threshold.
[0166] Optionally, the repair processing at the first level includes speed reduction processing and / or rebalancing processing.
[0167] Optionally, the root node device is connected to multiple terminal devices. The second processing unit includes:
[0168] A positioning unit, configured to locate a target terminal device with a fault when the cumulative number of repair times of the first repair processing reaches the preset second threshold;
[0169] A recovery unit, configured to perform at least one hardware repair processing on the target terminal device; wherein different hardware repair processings are used to recover the target terminal device.
[0170] Optionally, the recovery unit includes:
[0171] A first recovery subunit, configured to reset the port where the target terminal device is located;
[0172] A second recovery subunit, configured to perform a power-off and then power-on processing on the target terminal device when the cumulative number of reset times of the reset processing exceeds a preset third threshold.
[0173] Optionally, the window period includes multiple clock cycles, and the target clock cycle is the first clock cycle within the window period.
[0174] Optionally, the judgment module 502 includes:
[0175] A detection unit, configured to detect whether a preset error rate period is reached currently; wherein the error rate period is greater than the window period;
[0176] A count reduction unit, which is used to reduce the error event count by a preset difference when a preset error code rate period is reached;
[0177] A determination unit, which is used to determine to perform repair processing on the PCIe link in sequence when the reduced error event count reaches a preset first threshold.
[0178] Optionally, the error determination module 504 is specifically used for:
[0179] Determine the error type of the PCIe link based on the cumulative repair times of the repair processing at the last level, and / or based on the repair result of the repair processing at the last level.
[0180] Based on the same inventive concept, the present application also provides a computer-readable storage medium, and the computer program stored therein enables a processor to execute the error determination method for the PCIe link.
[0181] Based on the same inventive concept, the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, the error determination method for the PCIe link is implemented.
[0182] Finally, it should also be noted that, unless otherwise defined, the "first", "second" and similar terms used in this article do not represent any order, quantity or importance, but are only used to distinguish different components. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of other identical elements in the process, method, commodity or device including the element. "Connection" or "connected" and other similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0183] The above has introduced in detail an error determination method, device, equipment and storage medium for a PCIe link provided by the present disclosure. Specific examples are used in this article to elaborate on the principle and implementation manner of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
[0184] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0185] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
[0186] As used herein, the terms "one embodiment", "an embodiment", or "one or more embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. In addition, it should be noted that the examples of the phrase "in one embodiment" herein do not necessarily all refer to the same embodiment.
[0187] In the specification provided herein, a number of specific details are set forth. However, it can be understood that embodiments of the present disclosure may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0188] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present disclosure can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices may be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, not to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A method for error determination of a PCIe link, characterized in that, applied to the root node device in the PCIe link, the method includes: When each event that meets the preset condition is detected, increasing the error event count; the preset condition is that an error event occurs in the target clock cycle within the current window period; Based on the preset bit error rate period and the error event count, determine whether to perform repair processing; If so, perform at least one level of repair processing on the PCIe link in sequence; wherein, after each repair processing is executed, repeat the step of increasing the error event count when each event that meets the preset condition is detected, and each repair processing targets one level, and the next level of repair processing is determined based on the cumulative number of times of the previous level of repair processing; When the repair cut-off condition is met, determine the error type of the PCIe link based on the cumulative repair times of the last level of repair processing; and / or, determine the error type of the PCIe link based on the repair result of the last level of repair processing; Wherein, the levels include a physical level and a software level, and the objects in the software level for the PCIe link include: transmission signals, data packets, and transmission protocols, and the physical level targets the physical devices existing in the PCIe link; the next level is different from the previous level.
2. The method according to claim 1, characterized in that, The determining whether to perform repair processing based on the preset bit error rate period and the error event count includes: Detecting whether the preset bit error rate period is reached currently; wherein, the bit error rate period is greater than the window period; If so, reducing the error event count by a preset difference; If the reduced error event count reaches a preset first threshold, determine to perform repair processing on the PCIe link in sequence.
3. The method according to claim 1, characterized in that, The performing at least one level of repair processing on the PCIe link where it is located in sequence includes: Performing the first-level repair processing on the signal transmission level of the PCIe link; When the cumulative repair times of the first-level repair processing are greater than or equal to a preset second threshold, perform the second-level repair processing, and the second level includes a hardware level.
4. The method according to claim 3, characterized in that, The performing the first-level repair processing on the signal transmission level of the PCIe link includes: On the signal transmission level, performing repair processing on at least one object of the PCIe link respectively; wherein, the at least one object includes: signal quality and / or signal transmission speed.
5. The method according to claim 4, characterized in that, On the signal transmission level, performing repair processing on at least one object of the PCIe link respectively includes: For each object of the PCIe link, performing cross repair processing, wherein the previous repair processing is different from the subsequent repair processing; Alternatively, in one repair, repair processing is performed on at least one of the objects of the PCIe link; Alternatively, repair processing is performed on the current object of the PCIe link. If the number of times of repair processing on the current object reaches a preset number of times, repair processing is performed on the next object of the current object until the repair is successful or the cumulative number of repair times reaches the preset second threshold.
6. The method according to any one of claims 3-5, characterized in that, the repair processing at the first level includes speed reduction processing and / or rebalancing processing.
7. The method according to any one of claims 3-5, characterized in that, the root node device is connected to multiple terminal devices. When the cumulative number of repair times of the repair processing at the first level reaches a preset second threshold, second-level repair processing is performed, including: When the cumulative number of repair times of the first repair processing reaches the preset second threshold, locate the target terminal device with a fault; perform at least one hardware repair processing on the target terminal device; wherein, different hardware repair processings are used to recover the target terminal device.
8. The method according to claim 7, characterized in that, performing at least one hardware repair processing on the port where the target terminal device is located includes: performing a reset processing on the port where the target terminal device is located; when the cumulative number of reset times of the reset processing exceeds a preset third threshold, perform a power-down and then power-up processing on the target terminal device.
9. The method according to claim 1, characterized in that, the window period includes multiple clock cycles, and the target clock cycle is the first clock cycle within the window period.
10. An error determination device for a PCIe link, characterized in that, applied to the root node device in the PCIe link, the device includes: a counting module, configured to increase the error event count when each event satisfying a preset condition occurs; the preset condition is that an error event occurs in the target clock cycle within the current window period; a judgment module, configured to determine whether to perform repair processing based on a preset error code rate period and the error event count; a repair processing module, configured to perform at least one level of repair processing on the PCIe link in sequence when repair processing is performed; wherein, after each repair processing is executed, repeat the step of increasing the error event count when each event satisfying a preset condition occurs, and each time the repair processing is for one level, and the repair processing of the next level is determined based on the cumulative number of times of the repair processing of the previous level, and different levels correspond to different objects in the PCIe link; an error determination module, configured to determine the error type of the PCIe link based on the repair processing of the last level when the repair cut-off condition is satisfied.
11. An electronic device, characterized in that, the stored computer program enables the processor to execute the error determination method of the PCIe link according to any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The stored computer program causes the processor to execute the method for error determination of a PCIe link according to any one of claims 1-9.
Citation Information
Patent Citations
High-speed traffic link transmission quality detection method, device and system
CN110601920A
Fault processing method and device, electronic equipment and storage medium
CN113918375A