Centralized storage device PCIe link fault repair method and device

Through multi-dimensional monitoring and scenario-based repair strategies, the problems of PCIe link status monitoring misjudgment and secondary failure in hot-swap scenarios are solved, the accuracy of link fault detection and the repair success rate are improved, and the stable operation of centralized storage devices is ensured.

CN120560893BActive Publication Date: 2025-09-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511055321.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-26
Estimated Expiration
2045-07-30

AI Technical Summary

Technical Problem

In the existing technology, PCIe link status monitoring relies on a single register state, resulting in a high misjudgment rate, a lack of scenario-based repair strategies, and improper handling in hot-swap scenarios, which can easily lead to secondary failures.

Method used

By obtaining the detection data and status bit values ​​of the PCIe link, combined with the presence signal and heartbeat network, the threshold is dynamically adjusted, the link status is monitored in multiple dimensions, a scenario-based repair strategy is formulated, and the PCIe port is started and stopped to repair the link failure.

Benefits of technology

It improves the accuracy of link fault detection, reduces the misjudgment rate and secondary failure risk, optimizes the repair timing, ensures the stability and availability of data communication, and reduces the impact of link failures on data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120560893B_ABST
    Figure CN120560893B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for repairing PCIe link faults of a centralized storage device, which relates to the technical field of centralized storage devices, and includes: obtaining detection data and status bit values ​​of a PCIe link; determining whether the bit value is a first value; if so, determining an interruption, obtaining in-place signals of corresponding controller nodes respectively, and starting and stopping ports based on the in-place signals until the bit value is not the first value and / or the number of starts and stops reaches a certain number; if not, detecting whether the data meets certain conditions, and if not, determining degradation, starting and stopping ports, until the bit value is not the first value and the detection data meets certain conditions, or the number of starts and stops reaches a certain number, thereby solving the technical problems of high misjudgment rate, lack of scenario-based distinction, and easy occurrence of secondary faults, achieving multi-dimensional monitoring, formulating scenario-based repair strategies, improving detection accuracy and success rate, reducing misjudgment rate and secondary fault risks, and ensuring stable operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of centralized storage devices, and in particular to a method and apparatus for repairing a PCIe (Peripheral Component Interconnect express) link failure in a centralized storage device. Background Art

[0002] In the related technology, after a PCIe node device detects that a link between it and a first downstream PCIe node device has failed, the data transmission between the PCIe node device and each downstream PCIe node device can be stopped, the first PCIe controller in the first downstream PCIe node device can be reset, and the second PCIe controller corresponding to the first PCIe controller in the PCIe node device can be reset, and then the data transmission between the PCIe node device and each downstream PCIe node device can be restored; or when a PCIe node device detects that a channel of the link between it and a downstream PCIe node device has failed, a message interrupt message can be sent to the central processing unit, and then data transmission can be performed according to the number of channels and the number of available channels in the received message interrupt message.

[0003] However, in related technologies, link status monitoring relies only on a single register state, resulting in a high misjudgment rate. In addition, the repair strategy lacks scenario-based differentiation, and the handling in hot-plug scenarios is not appropriate, which can easily cause secondary failures and urgently needs improvement. Summary of the Invention

[0004] The present application provides a method and apparatus for repairing PCIe link faults in a centralized storage device, to at least address the problems in the related art where link status monitoring relies solely on a single register status, resulting in a high misjudgment rate, and the repair strategy lacks scenario-based differentiation. The handling in hot-plug scenarios is not appropriate, and secondary failures are very likely to occur.

[0005] The present application provides a method for repairing a PCIe link failure of a centralized storage device, wherein the centralized storage device includes at least two controller nodes, and different controller nodes are connected through the PCIe link, wherein the method includes the following steps: obtaining detection data corresponding to at least one detection indicator in the PCIe link and the bit value of the PCIe link status bit; based on the PCIe link, determining whether the bit value is the first bit value; if the bit value is the first bit value, determining that the PCIe link is interrupted, and based on the interrupted PCIe link, obtaining a first in-place signal and a second controller status signal of the PCIe link corresponding to the first controller node, respectively; The second in-place signal of the node starts and stops the PCIe port corresponding to the PCIe link based on the first in-place signal and the second in-place signal, until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number; if the bit value is not the first bit value, detect whether the detection data meets the preset condition, and if the detection data does not meet the preset condition, determine that the PCIe link is degraded, and based on the degraded PCIe link, start and stop the PCIe port corresponding to the PCIe link until the bit value is not the first bit value and the detection data meets the preset condition, or the number of starts and stops reaches the preset number.

[0006] The present application also provides a centralized storage device PCIe link fault repair device, the centralized storage device includes at least two controller nodes, and different controller nodes are connected through the PCIe link, wherein the device includes: a first acquisition module, used to obtain detection data corresponding to at least one detection indicator in the PCIe link and the bit value of the PCIe link status bit; a first judgment module, used to judge whether the bit value is the first bit value based on the PCIe link; a first start-stop module, used to determine that the PCIe link is interrupted when the bit value is the first bit value, and based on the interrupted PCIe link, respectively obtain the first in-position status of the first controller node corresponding to the PCIe link; signal and a second in-place signal of a second controller node, starting and stopping the PCIe port corresponding to the PCIe link based on the first in-place signal and the second in-place signal, until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number; a second start-stop module, used to detect whether the detection data meets a preset condition when the bit value is not the first bit value, and if the detection data does not meet the preset condition, determine that the PCIe link is degraded, and start and stop the PCIe port corresponding to the PCIe link based on the degraded PCIe link, until the bit value is not the first bit value and the detection data meets the preset condition, or the number of starts and stops reaches the preset number.

[0007] The present application also provides a storage device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned centralized storage device PCIe link failure repair methods when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned centralized storage device PCIe link failure repair methods are implemented.

[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned centralized storage device PCIe link failure repair methods when executed by a processor.

[0010] Through the present application, the detection data and status bit value corresponding to the PCIe link can be obtained, and when the bit value is the first value, the PCIe link is determined to be interrupted, and based on the interrupted PCIe link, the in-position signals of the dual controller nodes corresponding to the PCIe link are respectively obtained, and then the PCIe port is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a certain number; and when the bit value is not the first value, whether the data meets certain conditions, and when the certain conditions are not met, it is determined that the PCIe link is degraded, and the PCIe port is started and stopped based on the degraded PCIe link until the bit value is not the first value and the detection data meets certain conditions, or, The number of starts and stops reaches a certain number. Therefore, it can solve the technical problems that link status monitoring relies only on a single register status, the misjudgment rate is high, the repair strategy lacks scenario-based differentiation, and the hot-swap scenario is not handled properly, which is very likely to cause secondary failures. It can achieve multi-dimensional monitoring, formulate scenario-based repair strategies, improve the accuracy of link fault detection and the success rate of repair, and reduce the misjudgment rate and the risk of secondary failures. At the same time, it gives priority to ensuring data communication, reasonably arranges the repair time, and sets a limited number of repair and alarm mechanisms to minimize the impact of link failures on data transmission, ensure the stable operation of centralized storage devices, and improve the technical effects of availability and maintainability. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a block diagram of a centralized storage device provided according to one embodiment of the present application;

[0013] Figure 2A flowchart of a method for repairing a PCIe link failure in a centralized storage device provided in an embodiment of the present application;

[0014] Figure 3 A flowchart of repairing a first usage scenario provided according to an embodiment of the present application;

[0015] Figure 4 A flowchart of repairing a second usage scenario provided according to another embodiment of the present application;

[0016] Figure 5 This is a block diagram of a centralized storage device PCIe link fault repair device provided according to an embodiment of the present application.

[0017] Reference numerals:

[0018] Among them, 101-first processor (Central Processing Unit, referred to as CPU), 102-first complex programmable logic device (Complex Programmable Logic Device, referred to as CPLD), 103-first heartbeat network, 104-first in-place circuit, 105-second processor, 106-second complex programmable logic device, 107-second heartbeat network, 108-second in-place circuit, 109-backplane; 10-centralized storage device PCIe link fault repair device; 100-first acquisition module, 200-first judgment module, 300-first start-stop module and 400-second start-stop module. DETAILED DESCRIPTION

[0019] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0021] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0022] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the centralized storage device PCIe link fault repair method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0023] Specifically, Figure 1 A block diagram of a centralized storage device provided according to an embodiment of the present application.

[0024] like Figure 1 As shown, the centralized storage device includes a first processor 101, a first complex programmable logic device 102, a first heartbeat network 103, a first presence circuit 104, a second processor 105, a second complex programmable logic device 106, a second heartbeat network 107, a second presence circuit 108 and a backplane 109.

[0025] The first processor 101 includes a first controller node, the second processor 105 includes a second controller node, and the first controller node and the second controller node are connected via a PCIe link for data transmission.

[0026] Furthermore, in the embodiment of the present application, in-place signals can be designed on the first controller node and the second controller node respectively, wherein the first controller node is connected to the first complex programmable logic device 102 and the second controller node is connected to the second complex programmable logic device 106 .

[0027] The first complex programmable logic device 102 transmits the real-time collected presence signal to the operating system via the I2C (Inter-Integrated Circuit) bus. Furthermore, the second complex programmable logic device 106 transmits the real-time collected presence signal to the operating system via the I2C bus. A monitoring program running in the operating system monitors the presence signal in real time. Upon detecting a change in the operating status of a controller node (e.g., from an in-place state to an out-of-place state, or vice versa, not specifically limited in this application), the monitoring program immediately records the time of the change, including the node ID, in a log file and performs PCIe link repair.

[0028] The first controller node and the second controller node build a heartbeat network to implement a heartbeat mechanism between the two controller nodes, thereby monitoring the operating status of the other controller node.

[0029] The first heartbeat network 103 and the second heartbeat network 107 use the TCP / IP (Transmission Control Protocol / Internet Protocol) protocol, and set the heartbeat packet sending interval to 1 second. The heartbeat packet contains basic status information of the controller node (such as CPU usage, memory usage, etc., which is not specifically limited in this application).

[0030] Furthermore, the embodiment of the present application can determine whether the heartbeat network is interrupted by receiving the heartbeat packet of the other controller node in real time, and record the interruption information in the corresponding log file. For example, if the heartbeat packet of the other node is not received within 5 consecutive heartbeat packet intervals, it is determined that the heartbeat network is interrupted, and the interruption information is recorded in the corresponding log file.

[0031] An embodiment of the present application provides a method for repairing a PCIe link failure of a centralized storage device. The method is described in detail in conjunction with the execution flow of the method for repairing a PCIe link failure of a centralized storage device.

[0032] Specifically, Figure 2 This is a flowchart of a method for repairing a PCIe link failure in a centralized storage device provided according to an embodiment of the present application.

[0033] like Figure 2 As shown, the centralized storage device PCIe link fault repair method includes at least two controller nodes, and different controller nodes are connected through PCIe links. The method includes the following steps:

[0034] In step S201, detection data corresponding to at least one detection indicator in a PCIe link and a bit value of a PCIe link status bit are obtained.

[0035] It is understood that in the embodiments of the present application, the detection indicators of the PCIe link may include, but are not limited to, speed and bandwidth, and this application does not impose any specific restrictions. The detection indicators can be set based on the specific model of the storage device, the version of the PCIe bus, and normal operation requirements, and stored in the controller's configuration file or database, and this application does not impose any specific restrictions.

[0036] Furthermore, the embodiment of the present application can obtain the actual rate data and actual bandwidth data of the PCIe link at a certain time interval (for example, every 100 ms, which can be set by technicians in this field according to actual conditions and is not specifically limited by this application).

[0037] In addition, in the embodiment of the present application, the value of the PCIe link status bit can be obtained through the active status register. The value of the PCIe link status bit can be 0, indicating that the PCIe link is interrupted, or 1, indicating that the PCIe link is not interrupted. The specific setting of the PCIe link status bit can be determined by those skilled in the art according to actual conditions and is not limited by this application.

[0038] As a possible implementation method, the embodiment of the present application can obtain the detection data corresponding to the PCIe link under different detection indicators, and obtain the bit value of the PCIe link status bit.

[0039] Exemplarily, in an embodiment of the present application, when the PCIe link detection indicators are rate and bandwidth, the actual rate data and actual bandwidth data of the PCIe link are obtained every 100 ms, and the bit value of the PCIe link status bit is obtained through the active status register.

[0040] In step S202 , based on the PCIe link, it is determined whether the bit value is the first bit value.

[0041] In actual implementation, the embodiment of the present application can determine whether the bit value of the PCIe link is the first value. In this embodiment of the present application, the first value can be 0 to determine whether the PCIe link is interrupted, or it can be other values. The specific setting can be made by those skilled in the art according to actual conditions and is not specifically limited by this application.

[0042] Optionally, in one embodiment of the present application, before determining that the PCIe link is interrupted, it also includes: counting the total number of times that the bit value is the first value within a preset time; judging whether the total number of times is greater than a preset number of times; if the total number of times is greater than the preset number of times, determining that the PCIe link is interrupted, otherwise, determining that the PCIe link is not interrupted.

[0043] It is understood that the embodiment of the present application can obtain the value of the PCIe link status bit multiple times within a certain period of time to determine whether the PCIe link is interrupted. Among them, the certain period of time can be set by those skilled in the art according to actual conditions and is not specifically limited by this application.

[0044] In some embodiments, embodiments of the present application can, when the bit value is the first value, re-acquire the bit value of the PCIe link status bit at a certain interval, then count the total number of times the PCIe link status bit value is the first value within a certain period of time, and determine whether the total number of times is greater than a certain number. If so, the PCIe link is determined to be interrupted; otherwise, the PCIe link is not interrupted. The certain interval and the certain number of times can be set by those skilled in the art according to actual circumstances and are not specifically limited by this application.

[0045] For example, in an embodiment of the present application, when the read value of the active status register is 0, it is read again after an interval of 200ms. If the read value is 0, it is determined that the PCIe link is interrupted, the interruption information is recorded in the log file, and the repair process is started; if the read value is 1, it indicates that the PCIe link may have just jittered for a moment or the negotiation has not been completed yet, and it is determined that the PCIe link is not interrupted, and normal monitoring can continue.

[0046] The embodiment of the present application can perform multiple statistical calculations on the situation where the bit value is the first value before determining that the PCIe link is interrupted, and then determine that the link is interrupted when the number of times is greater than a certain value, thereby avoiding misjudgment due to a single error, improving the accuracy of the judgment, dynamically adjusting the threshold, adapting to different application scenarios, optimizing resource utilization, and reducing costs.

[0047] In step S203, if the bit value is the first value, the PCIe link is determined to be interrupted, and based on the interrupted PCIe link, a first in-position signal of the first controller node and a second in-position signal of the second controller node corresponding to the PCIe link are respectively obtained, and based on the first in-position signal and the second in-position signal, the PCIe port corresponding to the PCIe link is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0048] It is understandable that the embodiment of the present application can repair the interrupted PCIe link by starting and stopping the PCIe port corresponding to the PCIe link.

[0049] Furthermore, the embodiment of the present application is Figure 1 It can be seen that each controller node corresponds to a presence signal, and thus the embodiment of the present application can repair the PCIe link through the presence signal when the PCIe link is interrupted.

[0050] In some embodiments, the embodiments of the present application can determine that the PCIe link is interrupted when the bit value is the first value, and based on the interrupted PCIe link, obtain the first in-position signal of the first controller node and the second in-position signal of the second controller node corresponding to the PCIe link, and then start and stop the PCIe port corresponding to the PCIe link based on the first in-position signal and the second in-position signal until the bit value is not the first value.

[0051] Exemplarily, an embodiment of the present application can determine that the PCIe link is interrupted when the bit value is the first value, and based on the in-bit signal, by operating the disable / enable register in the PCIe link, first stop the PCIe port and then turn it on again, thereby achieving rapid recovery of the link failure, prompting the PCIe bus to renegotiate to re-establish the connection. If the bit value of the PCIe link becomes 1, it means that the PCIe link has been repaired, the repair is stopped, and the repair information is reported in a timely manner.

[0052] In some embodiments, embodiments of the present application can determine that a PCIe link is interrupted when the bit value is the first value, and based on the interrupted PCIe link, obtain a first presence signal from a first controller node and a second presence signal from a second controller node corresponding to the PCIe link, respectively. Furthermore, based on the first presence signal and the second presence signal, the PCIe port corresponding to the PCIe link is started and stopped until the number of starts and stops reaches a certain number. The certain number can be set by those skilled in the art based on actual circumstances and is not specifically limited by this application.

[0053] Illustratively, an embodiment of the present application can determine that the PCIe link is interrupted when the bit value is the first value, and based on the in-bit signal, by operating the disable / enable register in the PCIe link, first stop the PCIe port and then turn it on again, prompting the PCIe bus to renegotiate to re-establish the connection. If the bit value of the PCIe link is 0, it means that the PCIe link has not been repaired, and it is repaired again until the number of starts and stops reaches 3 times (this application does not impose specific restrictions). At this time, the embodiment of the present application stops repairing regardless of whether the bit value of the PCIe link is still 0, and reports the repair information in a timely manner.

[0054] Optionally, in one embodiment of the present application, based on the first in-place signal and the second in-place signal, the PCIe port corresponding to the PCIe link is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a preset number, including: determining the communication protocol between the first controller node and the second controller node based on the first in-place signal and the second in-place signal; obtaining the heartbeat packet between the first controller node and the second controller node based on the communication protocol; starting and stopping the PCIe port based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0055] It can be understood that the communication protocol in the embodiment of the present application can be a TCP / IP protocol, and the present application does not impose any specific restrictions; the presence signal can be obtained by detecting a sensor installed on the corresponding controller node, or by other means, and the present application does not impose any specific restrictions; the sending interval of the heartbeat packet is a certain value, such as 1 second, and the present application does not impose any specific restrictions.

[0056] It can be understood by those skilled in the art that the embodiment of the present application can determine the communication protocol between the first controller node and the second controller node through the first in-place signal and the second in-place signal, and then obtain the heartbeat packet between the first controller node and the second controller node, thereby starting and stopping the PCIe port until the bit value is not the first bit value.

[0057] For example, combined Figure 1As shown, the embodiment of the present application can obtain the corresponding presence signal through the sensor installed on the controller node, and perform data transmission between the first controller node and the second controller node corresponding to the PCIe link through the TCP / IP protocol, and set the sending interval of the heartbeat packet to 1 second. If the number of consecutive heartbeat packets not received reaches 5 times (the specific setting can be made by technical personnel in this field according to actual conditions, and this application does not make specific restrictions), it is determined that the heartbeat network is interrupted, and then the PCIe port is started and stopped until the bit value is not the first value.

[0058] In addition, the embodiment of the present application can determine the communication protocol between the first controller node and the second controller node through the first in-place signal and the second in-place signal, and then obtain the heartbeat packet between the first controller node and the second controller node, thereby starting and stopping the PCIe port until the number of starts and stops reaches a certain number.

[0059] The embodiment of the present application determines the communication protocol between the first controller node and the second controller node based on the first in-position signal and the second in-position signal, obtains a heartbeat packet, and then starts and stops the PCIe port until the bit value is not the first value and / or the number of starts and stops reaches a certain number, dynamically adapts the protocol, improves compatibility, monitors heartbeat packets, enhances reliability, avoids link freezes caused by driver problems, and makes multi-dimensional judgments to reduce misoperations.

[0060] Optionally, in one embodiment of the present application, based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet, the PCIe port is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a preset number, including: determining the first operating state of the first controller node and the second operating state of the second controller node based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet; based on the first operating state and the second operating state, starting and stopping the PCIe port until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0061] It can be understood that the embodiment of the present application can determine the operating status of the controller node corresponding to the PCIe link through the first presence signal, the second presence signal, the communication protocol and the heartbeat packet.

[0062] Furthermore, the embodiment of the present application can start and stop the PCIe port through the first operating state and the second operating state until the bit value is not the first value.

[0063] Furthermore, the embodiment of the present application can start and stop the PCIe port through the first operating state and the second operating state until the number of starts and stops reaches a certain number.

[0064] The embodiment of the present application can determine the operating status of the corresponding controller node based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet, and then start and stop the PCIe port until the bit value is not the first value and / or the number of starts and stops reaches a preset number. Multi-dimensional status monitoring improves judgment accuracy, significantly reduces the misjudgment rate, quickly locates faults, shortens troubleshooting time, and reduces labor costs.

[0065] Optionally, in one embodiment of the present application, based on the first operating state and the second operating state, the PCIe port is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a preset number, including: based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet, respectively judging whether the first operating state and the second operating state are in-place states; if the first operating state and the second operating state are both in-place states, determining that the PCIe link is in the first usage scenario, and based on the first usage scenario, starting and stopping the PCIe port until the bit value is not the first value and / or the number of starts and stops reaches a preset number; if the first operating state is not in-place state, or the second operating state is not in-place state, determining that the PCIe link is in the second usage scenario, and based on the second usage scenario, starting and stopping the PCIe port of the controller in the in-place state, and adjusting the controller that is not in-place state to the in-place state, and starting and stopping the PCIe port of the adjusted controller until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0066] It can be understood that the embodiment of the present application can determine whether the operating status of the controller node is in the in-place state through the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet of the controller node.

[0067] For example, combined Figure 1 As shown, the first controller node and the second controller node corresponding to the PCIe link in the embodiment of the present application can detect each other's presence signals, and after detecting that the presence signal of the other controller node disappears, the corresponding heartbeat network is interrupted, thereby determining that the operating status of the controller node that has disappeared the presence signal is not the presence state.

[0068] In addition, in an embodiment of the present application, when both the presence signal of the first controller node and the presence signal of the second controller node exist, if one controller node fails to receive a heartbeat packet for 5 consecutive times, it can be determined that the heartbeat network is interrupted, and then the operating status of the corresponding controller node is determined to be not in the presence state.

[0069] In some embodiments, the embodiments of the present application can determine that the usage scenario of the PCIe link is the first usage scenario, that is, the normal scenario, when the first operating state and the second operating state are both in the in-bit state, and then start and stop the PCIe port based on the normal scenario until the bit value is not the first bit value.

[0070] For example, in a normal scenario, when the PCIe link is interrupted, the embodiment of the present application can first stop the PCIe port and then turn it on again based on the in-bit signal by operating the disable / enable register in the PCIe link, prompting the PCIe bus to renegotiate to re-establish the connection. If the bit value of the PCIe link becomes 1, it means that the PCIe link has been repaired. At this time, the embodiment of the present application can stop the repair and report the repair information in a timely manner.

[0071] In some embodiments, the embodiments of the present application can determine that the usage scenario of the PCIe link is a normal scenario when the first operating state and the second operating state are both in the in-place state, and then start and stop the PCIe port based on the normal scenario until the number of starts and stops reaches a preset number.

[0072] For example, in a normal scenario, when the PCIe link is interrupted, the embodiment of the present application can first stop the PCIe port and then turn it on again based on the in-bit signal by operating the disable / enable register in the PCIe link, prompting the PCIe bus to renegotiate to re-establish the connection. If the bit value of the PCIe link is 0, it means that the PCIe link has not been repaired, and it will be repaired again until the number of starts and stops reaches 3 times. At this time, the embodiment of the present application stops the repair regardless of whether the bit value of the PCIe link is still 0, and reports the repair information in a timely manner.

[0073] Furthermore, in some embodiments, the embodiments of the present application determine that the usage scenario of the PCIe link is the second usage scenario, i.e., the hot plug scenario, when the first operating state is not the in-place state, or the second operating state is not the in-place state, and then based on the hot plug scenario, start and stop the PCIe port of the controller in the in-place state until the bit value is not the first value, and adjust the controller that is not in the in-place state to the in-place state, and start and stop the PCIe port of the adjusted controller until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0074] For example, combined Figure 1 As shown, in a hot-swap scenario, two controller nodes in the embodiment of the present application detect each other's controller node's presence signal, which is connected to the CPLD, and the CPLD transmits the detected presence information to the operating system via the I2C bus.

[0075] A heartbeat network channel is built between the two controller nodes to obtain the heartbeat packets of the other controller node and then monitor the running status of the other controller node.

[0076] For example, in an embodiment of the present application, upon detecting that the second controller node has been unplugged (i.e., the second operating state of the second controller node is no longer in-place), the first controller node retrieves PCIe port information from a log file based on changes in the in-place signal and heartbeat network interruption. It then sends a disable command to the link control register of the PCIe link, shutting down the PCIe port. After waiting for a certain period of time (e.g., 100ms, not specifically limited in this application), it then sends an enable command to reopen the PCIe port, prompting the PCIe bus to renegotiate and establish a connection. Throughout this entire process, the system monitors the port status and negotiation process in real time to ensure the accuracy and effectiveness of the operation.

[0077] When the second controller node is reinserted, that is, the second operating state of the second controller node is the in-place state, the first controller node detects that the in-place signal is restored and the heartbeat network returns to normal, obtains relevant information from the log file again, and performs the same PCIe port closing and reopening operation, prompting the PCIe bus to renegotiate and establish a connection.

[0078] After the repair, the activity status register is checked to confirm whether the connection was successfully established. If the connection is normal, the normal data synchronization process is resumed. If the PCIe link is still interrupted, the repair is performed again, and the number of times is also set to three. If the connection cannot be established after three repair attempts, the repair is stopped and an alarm result is reported.

[0079] In an embodiment of the present application, when the first operating state and the second operating state are both in-place states, it can be determined that the PCIe link is in the first usage scenario, and then start and stop the PCIe port until the bit value is not the first value and / or the number of starts and stops reaches a certain number; and when the first operating state is not in-place state, or the second operating state is not in-place state, it can be determined that it is in the second usage scenario, and then start and stop the PCIe port of the controller in the in-place state, and adjust the controller that is not in-place state to the in-place state, and start and stop the PCIe port of the adjusted controller until the bit value is not the first value and / or the number of starts and stops reaches a certain number, thereby refining the scenario division and improving the policy targeting.

[0080] Optionally, in one embodiment of the present application, before starting and stopping the PCIe port corresponding to the PCIe link based on the degraded PCIe link, it also includes: obtaining the data transmission volume of the degraded PCIe link; determining whether the data transmission volume is less than a preset transmission value; if the data transmission volume is less than the preset transmission value, allowing the PCIe port to be started and stopped; if the data transmission volume is greater than or equal to the preset transmission value, adjusting the data transmission volume based on the degraded PCIe link until the data transmission volume is less than the preset transmission value, and allowing the PCIe port to be started and stopped.

[0081] It is understood that, before repairing a degraded PCIe link, embodiments of the present application may first obtain the data transmission volume of the degraded PCIe link, and then perform repairs when the data transmission volume is less than a certain transmission value. The certain transmission value can be set by those skilled in the art based on actual circumstances and is not specifically limited by this application.

[0082] In some embodiments, the embodiments of the present application can obtain the data transmission volume of the degraded PCIe link and determine whether the data transmission volume is less than a certain transmission value. If it is less than, the PCIe link can be repaired, that is, the PCIe port is started and stopped; and if it is greater than or equal to, the data transmission volume is adjusted until the data transmission volume is less than a certain transmission value.

[0083] For example, in an embodiment of the present application, a certain transmission value may be set to 250MB / s, and then when the data transmission rate of the degraded PCIe link is less than 250MB / s, the PCIe port is allowed to be started and stopped; otherwise, the PCIe port is not started and stopped.

[0084] Before starting and stopping the downgraded PCIe link port, the embodiment of the present application first determines whether the data transmission volume is less than a certain transmission value, and allows starting and stopping when it is less than a certain value, thereby avoiding performance interruption caused by starting and stopping under high load, optimizing resource utilization, and improving system stability and reliability.

[0085] Optionally, in one embodiment of the present application, before starting and stopping the PCIe port corresponding to the PCIe link based on the degraded PCIe link, it also includes: obtaining the link operation status of the degraded PCIe link; detecting whether the link operation status is an idle state; and allowing the PCIe port to be started and stopped when the link operation status is an idle state.

[0086] It is understood that, before repairing a degraded PCIe link, embodiments of the present application may first obtain the link operating status of the degraded PCIe link, and then perform repairs when the link operating status is idle. The idle state may be understood as an operating state in which the link bandwidth utilization is less than 30% and lasts for more than 5 minutes. The specific setting can be made by those skilled in the art based on actual conditions and is not specifically limited by this application.

[0087] Illustratively, before repairing a degraded PCIe link, an embodiment of the present application analyzes the data traffic of the PCIe link in real time. When the link bandwidth utilization is lower than 30% and the duration exceeds 5 minutes, the link operation state is determined to be idle, and then the relevant information of the corresponding PCIe link is retrieved from the log file, including the PCIe port number, link parameters, etc. This application does not impose specific restrictions, thereby preparing for the upcoming repair.

[0088] Before starting or stopping the downgraded PCIe link port, the embodiment of the present application can first detect whether the link operation status is idle, and if it is idle, allow the PCIe port to be started or stopped, thereby ensuring data integrity, avoiding data loss, and improving system stability and reliability.

[0089] In step S204, if the bit value is not the first value, the data is detected to see whether it meets the preset conditions. If the detection data does not meet the preset conditions, it is determined that the PCIe link is degraded, and based on the degraded PCIe link, the PCIe port corresponding to the PCIe link is started and stopped until the bit value is not the first value and the detection data meets the preset conditions, or the number of starts and stops reaches a preset number.

[0090] In some embodiments, when the bit value is 1 instead of 0, the present application continues to detect whether the data meets certain conditions. If not, the PCIe link is determined to be degraded, and the PCIe port corresponding to the PCIe link is started or stopped until the bit value is not the first value. The certain conditions can be set by those skilled in the art according to actual circumstances and are not specifically limited in this application.

[0091] In some embodiments, when the bit value is not 0 but 1, the embodiment of the present application continues to detect whether the data meets certain conditions, and if it does not meet the conditions, it determines that the PCIe link is degraded, and then starts and stops the PCIe port corresponding to the PCIe link until the number of starts and stops reaches a certain number.

[0092] For example, an embodiment of the present application can determine that the PCIe link is degraded and needs to be repaired when the actual rate data and actual bandwidth data in the detection data do not meet certain conditions and the active status register read value is 1. At this time, the embodiment of the present application can record the time when the PCIe link degradation occurs, the current actual rate data and actual bandwidth data, etc. in a log file to facilitate subsequent analysis of the degradation cause and assessment of the impact on data transmission. At the same time, it is prepared to start degradation repair, but the repair is not performed immediately. Instead, the data transmission volume and link operation status of the degraded PCIe link are continuously monitored, and then repair is performed when the data transmission volume is less than a certain transmission value or the link operation status is idle.

[0093] Furthermore, during the repair process, embodiments of the present application send a disable command to the link control register of the PCIe link, shutting down the PCIe port. During the port shutdown process, the port status is monitored in real time to ensure the successful completion of the shutdown operation. After the shutdown, a predetermined period of time (e.g., 100ms, not specifically limited in this application) is waited, and then an enable command is sent to the link control register to reopen the PCIe port, prompting the PCIe bus to renegotiate and reestablish the connection. During the renegotiation process, embodiments of the present application can closely monitor the negotiation status and changes in related parameters to ensure the smooth progress of the negotiation process.

[0094] After the repair, the embodiment of the present application can re-obtain the read value of the active status register to confirm whether the link has been successfully restored. If the read value of the active status register is 1, and the actual rate data and actual bandwidth data meet certain conditions, the PCIe link is determined to be successfully repaired, the system resumes the normal data synchronization process, and records the successful repair information in the log file, while clearing the previously recorded degradation-related information.

[0095] After the repair, if the read value of the active status register is still 0 or the actual rate data and actual bandwidth data do not meet certain conditions, the PCIe link repair is determined to have failed, and the repair failure information is recorded in the log file, and the repair count is increased. The certain number of times is set to 3 times. When the start and stop times reach 3 times, the repair is stopped, and the alarm result that the PCIe link failure cannot be repaired is reported to the management platform or operation and maintenance personnel so that further measures can be taken in time.

[0096] Optionally, in one embodiment of the present application, it also includes: when the bit value is the first value, based on the PCIe link, determining the interruption information or degradation information of the PCIe link; based on the interruption information, or, degradation information and detection data, determining the alarm strategy of the PCIe link; based on the alarm strategy, obtaining the alarm result of the PCIe link.

[0097] It can be understood that in the embodiment of the present application, when the bit value is the first value, the corresponding alarm strategy may include an alarm strategy after starting and stopping the PCIe port; an alarm strategy when the detection data does not meet certain conditions; the alarm strategy in the first usage scenario and the alarm strategy in the second usage scenario. The specific settings can be made by technical personnel in this field according to actual conditions, and this application does not impose any specific restrictions.

[0098] As a possible implementation method, the embodiment of the present application can determine the interruption information or degradation information of the PCIe link based on the PCIe link when the bit value is the first value, and then determine the alarm strategy of the PCIe link, thereby obtaining the alarm result of the PCIe link.

[0099] Illustratively, in a first usage scenario, an embodiment of the present application can record the detection data and bit value of the PCIe link to a corresponding log file before starting or stopping the PCIe port. This triggers a first alarm strategy and issues a voice alarm to the user: "PCIe link interrupted, please pay attention." Furthermore, an embodiment of the present application repairs the PCIe link by continuously starting and stopping the PCIe port. If the PCIe link is still interrupted after a certain number of starts and stops, the second alarm strategy is triggered and issues a voice alarm to the user: "PCIe link interrupted, repair unsuccessful, please pay attention." If the PCIe link is repaired, the third alarm strategy is triggered and issues a voice alarm to the user: "PCIe link interrupted, repair successful."

[0100] Furthermore, in the first usage scenario, in the embodiment of the present application, the PCIe link is degraded, and the data transmission volume and link operation status of the degraded PCIe link continue to be monitored, and then repair is performed when the data transmission volume is less than a certain transmission value, or the link operation status is idle. At this time, the fourth alarm strategy can be triggered to issue a voice alarm to the user: "The degraded PCIe link can be repaired."

[0101] Furthermore, in the second usage scenario of the embodiment of the present application, if the PCIe link is interrupted, the fifth alarm strategy can be triggered to issue a voice alarm to the user: "PCIe link is interrupted, please pay attention."

[0102] The above content is only an example. The specific content can be set by technicians in this field according to actual conditions, and this application does not impose any specific limitations.

[0103] The embodiment of the present application can determine the alarm strategy of the PCIe link based on interrupt information, or degradation information and detection data when the bit value is the first value, and then obtain the corresponding alarm result, accurately identify the fault type, avoid false alarms and missed alarms, adapt to the needs of different scenarios, ensure the stable operation of the centralized storage device, and improve availability and maintainability.

[0104] The working principle of the centralized storage device PCIe link failure repair method proposed in the embodiments of the present application is introduced below in combination with multiple embodiments.

[0105] Example 1:

[0106] in, Figure 3 This is a flowchart of repairing a first usage scenario provided according to an embodiment of the present application.

[0107] It should be noted that the first usage scenario is a common scenario, that is, the first operating state of the first controller node of the PCIe link and the second operating state of the second controller node corresponding to the PCIe link are both in-place states.

[0108] Step S301: Acquire detection data and status bit values ​​of a PCIe link.

[0109] Among them, in the embodiment of the present application, when the PCIe link detection indicators are rate and bandwidth, the actual rate data and actual bandwidth data of the PCIe link are obtained every 100ms, and the bit value of the PCIe link status bit is obtained through the active status register.

[0110] Furthermore, in the embodiment of the present application, when the bit value is 1 and the actual rate data and the actual bandwidth data meet certain conditions, step S302 is executed; when the bit value is 0, step S303 is executed; when the bit value is 1 and the actual rate data and the actual bandwidth data do not meet certain conditions, step S305 is executed.

[0111] Step S302: The PCIe link is normal, and the process ends.

[0112] Step S303: Re-acquire the bit value.

[0113] In this embodiment of the present application, if the bit value is 1, step S302 is executed; otherwise, step S304 is executed.

[0114] Step S304: Determine whether the PCIe link is interrupted.

[0115] Step S305: Determine whether the PCIe link is degraded.

[0116] Step S306: Determine whether certain conditions are met.

[0117] In this embodiment of the present application, when the actual rate data and the actual bandwidth data meet certain conditions, step S307 is executed; otherwise, step S308 is executed.

[0118] Step S307: Continue monitoring.

[0119] Step S308: Start and stop the PCIe port.

[0120] Step S309: Verify the bit value.

[0121] In this embodiment of the present application, if the bit value is 1, step S302 is executed; otherwise, step S310 is executed.

[0122] Step S310: Determine whether a certain number of times has been reached.

[0123] In this embodiment of the present application, when the number of starts and stops reaches a certain number, step S311 is executed; otherwise, step S308 is executed.

[0124] Step S311: reporting the repair result.

[0125] Example 2:

[0126] in, Figure 4 This is a flowchart of repairing a second usage scenario provided according to another embodiment of the present application.

[0127] It should be noted that the second usage scenario is a hot-swap scenario, that is, the running status of one controller node is not in the in-place state, and the running status of the other controller node is in the in-place state.

[0128] Step S401: Monitor the operating status of the controller node.

[0129] In this embodiment of the present application, when the operating status of the controller node is in the in-place state, step S402 is executed; otherwise, step S405 is executed.

[0130] Step S402: Immediately close the PCIe port.

[0131] Step S403: Enable the PCIe port after a delay.

[0132] Step S404: Wait for another controller node to be inserted.

[0133] Step S405: Wait for the running status of the controller node to be in the in-position state.

[0134] Step S406: Close the PCIe port.

[0135] Step S407: Enable the PCIe port after a delay.

[0136] Step S408: Verify the PCIe link.

[0137] In this embodiment of the present application, when the PCIe link is normal, step S409 is executed; otherwise, step S410 is executed.

[0138] Step S409: Resume data synchronization.

[0139] Step S410: Start and stop the PCIe port.

[0140] Step S411: Determine whether the number of starts and stops reaches a certain number.

[0141] In this embodiment of the present application, if the condition is not met, step S410 is executed; otherwise, step S412 is executed.

[0142] Step S412: reporting the repair result.

[0143] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0144] According to the centralized storage device PCIe link fault repair method proposed in the embodiment of the present application, the detection data and status bit value corresponding to the PCIe link can be obtained, and when the bit value is the first value, the PCIe link is determined to be interrupted, and based on the interrupted PCIe link, the in-place signals of the dual controller nodes corresponding to the PCIe link are respectively obtained, and then the PCIe port is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a certain number; and when the bit value is not the first value, whether the data meets certain conditions, and if the certain conditions are not met, it is determined that the PCIe link is degraded, and the PCIe port is started and stopped based on the degraded PCIe link until the bit value is not the first value and The detection data meets certain conditions, or the number of starts and stops reaches a certain number. Therefore, it can solve the technical problems that link status monitoring relies only on a single register state, the misjudgment rate is high, the repair strategy lacks scenario-based differentiation, and the hot-swap scenario is not handled properly, which is very likely to cause secondary failures. It achieves multi-dimensional monitoring, formulates scenario-based repair strategies, improves the accuracy of link fault detection and the success rate of repair, reduces the misjudgment rate and the risk of secondary failures; at the same time, it gives priority to ensuring data communication, reasonably arranges the repair time, and sets a limited number of repair and alarm mechanisms to minimize the impact of link failures on data transmission, ensure the stable operation of centralized storage devices, and improve the technical effects of availability and maintainability.

[0145] An embodiment of the present application also provides a centralized storage device PCIe link fault repair device.

[0146] Figure 5 This is a block diagram of a centralized storage device PCIe link fault repair device provided according to an embodiment of the present application.

[0147] like Figure 5 As shown, the centralized storage device PCIe link fault repair device 10, the centralized storage device includes at least two controller nodes, and different controller nodes are connected through PCIe links, wherein the centralized storage device PCIe link fault repair device 10 includes: a first acquisition module 100, a first judgment module 200, a first start-stop module 300 and a second start-stop module 400.

[0148] The first acquisition module 100 is configured to acquire detection data corresponding to at least one detection indicator in a PCIe link and a bit value of a PCIe link status bit.

[0149] The first judgment module 200 is configured to judge whether the bit value is the first bit value based on the PCIe link.

[0150] The first start-stop module 300 is configured to determine that a PCIe link is interrupted when the bit value is the first value, and obtain, based on the interrupted PCIe link, a first in-position signal of a first controller node and a second in-position signal of a second controller node corresponding to the PCIe link, respectively, and start and stop the PCIe port corresponding to the PCIe link based on the first in-position signal and the second in-position signal until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0151] The second start-stop module 400 is used to detect whether the data meets the preset conditions when the bit value is not the first value, and if the detection data does not meet the preset conditions, determine that the PCIe link is degraded, and start and stop the PCIe port corresponding to the PCIe link based on the degraded PCIe link until the bit value is not the first value and the detection data meets the preset conditions, or the number of starts and stops reaches a preset number.

[0152] Optionally, in one embodiment of the present application, it further includes: a statistical module, a second judgment module and a determination module.

[0153] The statistical module is used to count the total number of times the first value is the first value within a preset time before determining that the PCIe link is interrupted.

[0154] The second judgment module is used to judge whether the total number of times is greater than a preset number of times.

[0155] The determination module is configured to determine that the PCIe link is interrupted when the total number of times is greater than a preset number of times; otherwise, determine that the PCIe link is not interrupted.

[0156] Optionally, in one embodiment of the present application, the first start-stop module 300 includes: a determination unit, an acquisition unit, and a start-stop unit.

[0157] The determining unit is configured to determine a communication protocol between the first controller node and the second controller node based on the first presence signal and the second presence signal;

[0158] an acquiring unit, configured to acquire a heartbeat packet between the first controller node and the second controller node based on a communication protocol;

[0159] The start-stop unit is used to start and stop the PCIe port based on the first in-position signal, the second in-position signal, the communication protocol and the heartbeat packet until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0160] Optionally, in one embodiment of the present application, it is characterized in that the start-stop unit includes: a determination subunit and a start-stop subunit.

[0161] The determination subunit is configured to determine a first operating state of the first controller node and a second operating state of the second controller node based on the first presence signal, the second presence signal, the communication protocol, and the heartbeat packet.

[0162] The start-stop subunit is used to start and stop the PCIe port based on the first operating state and the second operating state until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0163] Optionally, in one embodiment of the present application, the start-stop subunit includes: a judgment subcomponent, a first start-stop subcomponent, and a second start-stop subcomponent.

[0164] The judgment component is used to judge whether the first operating state and the second operating state are in-place states based on the first in-place signal, the second in-place signal, the communication protocol and the heartbeat packet.

[0165] The first start-stop sub-component is used to determine that the PCIe link is in a first usage scenario when the first operating state and the second operating state are both in the in-position state, and to start and stop the PCIe port based on the first usage scenario until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0166] The second start-stop sub-component is used to determine that the PCIe link is in a second usage scenario when the first operating state is not the in-place state, or the second operating state is not the in-place state, and based on the second usage scenario, start and stop the PCIe port of the controller in the in-place state, adjust the controller that is not in the in-place state to the in-place state, and start and stop the PCIe port of the adjusted controller until the bit value is not the first value and / or the number of starts and stops reaches a preset number.

[0167] Optionally, in one embodiment of the present application, it further includes: a second acquisition module, a third judgment module, a third start-stop module and a fourth start-stop module.

[0168] The second acquisition module is configured to acquire the data transmission volume of the degraded PCIe link before starting or stopping the PCIe port corresponding to the PCIe link based on the degraded PCIe link.

[0169] The third judgment module is used to judge whether the data transmission amount is less than a preset transmission value.

[0170] The third start-stop module is used to allow the PCIe port to be started and stopped when the data transmission amount is less than a preset transmission value.

[0171] The fourth start-stop module is used to adjust the data transmission volume based on the degraded PCIe link when the data transmission volume is greater than or equal to the preset transmission value until the data transmission volume is less than the preset transmission value, and allow the PCIe port to be started and stopped.

[0172] Optionally, in one embodiment of the present application, it further includes: a third acquisition module, a detection module and a fifth start-stop module.

[0173] The third acquisition module is configured to acquire the link operation status of the degraded PCIe link before starting or stopping the PCIe port corresponding to the PCIe link based on the degraded PCIe link.

[0174] The detection module is used to detect whether the link operation state is idle.

[0175] The fifth start / stop module is configured to allow the PCIe port to be started or stopped when the link operation state is an idle state.

[0176] Optionally, in one embodiment of the present application, the first start-stop module 300 includes: a second determining unit, an acquiring unit, and a second start-stop unit.

[0177] The second determining unit is configured to determine a communication protocol between the first controller node and the second controller node based on the first presence signal and the second presence signal.

[0178] The acquiring unit is configured to acquire a heartbeat packet between the first controller node and the second controller node based on a communication protocol.

[0179] The second start / stop unit is configured to start / stop the PCIe port based on the first presence signal, the second presence signal, the communication protocol, and the heartbeat packet.

[0180] Optionally, in one embodiment of the present application, it further includes: a first determination module, a second determination module and a generation module.

[0181] The first determining module is configured to determine, based on the PCIe link, the interruption information or degradation information of the PCIe link when the bit value is the first value.

[0182] The second determining module is configured to determine an alarm strategy for the PCIe link based on the interruption information, or the degradation information and the detection data.

[0183] The generation module is used to obtain the alarm result of the PCIe link based on the alarm policy.

[0184] For the description of the features in the embodiment corresponding to the apparatus for repairing a PCIe link failure of a centralized storage device, reference may be made to the relevant description of the embodiment corresponding to the method for repairing a PCIe link failure of a centralized storage device, which will not be described in detail here.

[0185] According to the centralized storage device PCIe link fault repair device proposed in the embodiment of the present application, the detection data and status bit value corresponding to the PCIe link can be obtained, and when the bit value is the first value, the PCIe link is determined to be interrupted, and based on the interrupted PCIe link, the in-position signals of the dual controller nodes corresponding to the PCIe link are respectively obtained, and then the PCIe port is started and stopped until the bit value is not the first value and / or the number of starts and stops reaches a certain number; and when the bit value is not the first value, whether the data meets certain conditions, and when the certain conditions are not met, it is determined that the PCIe link is degraded, and the PCIe port is started and stopped based on the degraded PCIe link until the bit value is not the first value and The detection data meets certain conditions, or the number of starts and stops reaches a certain number. Therefore, it can solve the technical problems that link status monitoring relies only on a single register state, the misjudgment rate is high, the repair strategy lacks scenario-based differentiation, and the hot-swap scenario is not handled properly, which is very likely to cause secondary failures. It achieves multi-dimensional monitoring, formulates scenario-based repair strategies, improves the accuracy of link fault detection and the success rate of repair, reduces the misjudgment rate and the risk of secondary failures; at the same time, it gives priority to ensuring data communication, reasonably arranges the repair time, and sets a limited number of repair and alarm mechanisms to minimize the impact of link failures on data transmission, ensure the stable operation of centralized storage devices, and improve the technical effects of availability and maintainability.

[0186] An embodiment of the present application also provides a storage device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned centralized storage device PCIe link failure repair method embodiments.

[0187] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned centralized storage device PCIe link failure repair method embodiments when running.

[0188] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0189] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned centralized storage device PCIe link failure repair method embodiments are implemented.

[0190] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, implementing the steps in any of the above-mentioned centralized storage device PCIe link failure repair method embodiments.

[0191] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0192] The above is a detailed introduction to a centralized storage device PCIe link fault repair method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A method for repairing a PCIe link failure in a centralized storage device, characterized in that: The centralized storage device includes at least two controller nodes, and different controller nodes are connected via the high-speed serial computer expansion bus interface PCIe link, wherein the method includes the following steps: Obtaining detection data corresponding to at least one detection indicator in the PCIe link and a bit value of the PCIe link status bit; Based on the PCIe link, determining whether the bit value is a first bit value; If the bit value is the first bit value, determining that the PCIe link is interrupted, and based on the interrupted PCIe link, respectively obtaining a first presence signal of a first controller node and a second presence signal of a second controller node corresponding to the PCIe link, and starting and stopping the PCIe port corresponding to the PCIe link based on the first presence signal and the second presence signal until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number; If the bit value is not the first bit value, detect whether the detection data meets the preset condition, and if the detection data does not meet the preset condition, determine that the PCIe link is degraded, and based on the degraded PCIe link, start and stop the PCIe port corresponding to the PCIe link until the bit value is not the first bit value and the detection data meets the preset condition, or the number of starts and stops reaches the preset number.

2. The centralized storage device PCIe link failure repair method according to claim 1, characterized in that: Before determining that the PCIe link is interrupted, the method further includes: Counting the total number of times the place value is the first place value within a preset time; Determine whether the total number of times is greater than a preset number of times; If the total number of times is greater than the preset number of times, it is determined that the PCIe link is interrupted; otherwise, it is determined that the PCIe link is not interrupted.

3. The centralized storage device PCIe link failure repair method according to claim 1, characterized in that: The step of starting and stopping the PCIe port corresponding to the PCIe link based on the first in-position signal and the second in-position signal until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number includes: determining a communication protocol between the first controller node and the second controller node based on the first presence signal and the second presence signal; Based on the communication protocol, obtaining a heartbeat packet between the first controller node and the second controller node; Based on the first presence signal, the second presence signal, the communication protocol, and the heartbeat packet, the PCIe port is started and stopped until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number.

4. The centralized storage device PCIe link failure repair method according to claim 3, characterized in that: The starting and stopping of the PCIe port based on the first in-position signal, the second in-position signal, the communication protocol, and the heartbeat packet until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number, includes: Determining a first operating state of the first controller node and a second operating state of the second controller node based on the first presence signal, the second presence signal, the communication protocol, and the heartbeat packet; Based on the first operating state and the second operating state, the PCIe port is started and stopped until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number.

5. The centralized storage device PCIe link failure repair method according to claim 4, characterized in that: The starting and stopping of the PCIe port based on the first operating state and the second operating state until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number includes: Based on the first presence signal, the second presence signal, the communication protocol, and the heartbeat packet, respectively determining whether the first operating state and the second operating state are in-place states; If both the first operating state and the second operating state are the in-position state, determining that the PCIe link is in a first usage scenario, and starting and stopping the PCIe port based on the first usage scenario until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number; If the first operating state is not the in-place state, or the second operating state is not the in-place state, it is determined that the PCIe link is in a second usage scenario, and based on the second usage scenario, the PCIe port of the controller in the in-place state is started and stopped, and the controller that is not in the in-place state is adjusted to the in-place state, and the PCIe port of the adjusted controller is started and stopped until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number.

6. The centralized storage device PCIe link failure repair method according to claim 1, characterized in that: Before enabling or disabling the PCIe port corresponding to the degraded PCIe link, the method further includes: Obtaining the data transmission volume of the degraded PCIe link; Determining whether the data transmission amount is less than a preset transmission value; If the data transmission amount is less than the preset transmission value, allowing the PCIe port to be started and stopped; If the data transmission amount is greater than or equal to the preset transmission value, the data transmission amount is adjusted based on the degraded PCIe link until the data transmission amount is less than the preset transmission value, and the PCIe port is allowed to be started and stopped.

7. The centralized storage device PCIe link failure repair method according to claim 1, characterized in that: Before enabling or disabling the PCIe port corresponding to the degraded PCIe link, the method further includes: Obtaining a link operating status of the degraded PCIe link; Detecting whether the link operation state is an idle state; When the link operation state is the idle state, the PCIe port is allowed to be started and stopped.

8. The centralized storage device PCIe link failure repair method according to claim 1, characterized in that: Also includes: When the bit value is the first bit value, determining, based on the PCIe link, interruption information or degradation information of the PCIe link; Determining an alarm strategy for the PCIe link based on the interruption information, or the degradation information and the detection data; Based on the alarm policy, an alarm result of the PCIe link is obtained.

9. A centralized storage device PCIe link fault repair device, characterized in that: The centralized storage device includes at least two controller nodes, and different controller nodes are connected via the PCIe link, wherein the apparatus includes: An acquisition module, configured to acquire detection data corresponding to at least one detection indicator in the PCIe link and a bit value of the PCIe link status bit; A judging module, configured to judge whether the bit value is a first bit value based on the PCIe link; a first start / stop module, configured to determine, when the bit value is the first bit value, that the PCIe link is interrupted, and obtain, based on the interrupted PCIe link, a first presence signal of a first controller node and a second presence signal of a second controller node corresponding to the PCIe link, respectively, and start / stop the PCIe port corresponding to the PCIe link based on the first presence signal and the second presence signal, until the bit value is not the first bit value and / or the number of starts and stops reaches a preset number; A second start-stop module is configured to detect whether the detection data satisfies a preset condition when the bit value is not the first bit value, and determine that the PCIe link is degraded if the detection data does not satisfy the preset condition, and start and stop the PCIe port corresponding to the PCIe link based on the degraded PCIe link until the bit value is not the first bit value and the detection data satisfies the preset condition, or the number of starts and stops reaches the preset number.

10. A storage device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the centralized storage device PCIe link failure repair method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the centralized storage device PCIe link failure repair method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Link connectivity detection system and method

    CN105306306A

  • Peripheral component interconnect express non-transparent bridge (PCIe NTB)-based dual-controller storage high-availability (HA) subsystem

    CN107766181A