A SSD link exception processing system and method

By designing an SSD link exception handling system, using hardware link switching and software collaborative control, the difficulty of processing SSD in the case of link disconnection and abnormal power outage is solved, the hardware design is simplified, and the stability of the main control system is ensured.

CN115473929BActive Publication Date: 2025-05-13SHANDONG SINOCHIP SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211081409.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-06
Publication Date
2025-05-13
Estimated Expiration
2042-09-06

AI Technical Summary

Technical Problem

The processing of SSD in the case of link disconnection and abnormal power outage is difficult, especially the principle that the physical storage unit of the non-volatile storage medium nand flash particles cannot be overwritten, making the processing of abnormal power outage more demanding.

Method used

An SSD link exception handling system is designed, including a link disconnection detection unit, an interrupt control unit and a link loop back unit. Through front-end hardware link switching and software collaborative control, the complexity of hardware design is simplified and the stability of the main control system is ensured.

Benefits of technology

Effectively handle link abnormalities such as link disconnection and abnormal power outage, reduces the complex interaction process between coupling modules and reset processing of hardware units, simplifies hardware design, and ensures the stability of the main control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115473929B_ABST
    Figure CN115473929B_ABST
Patent Text Reader

Abstract

The present invention discloses a SSD link abnormality processing system and method, the system includes a link disconnection detection unit, an interruption control unit and a link loopback unit; the link abnormality detection unit is used to detect whether a PCIe link has an abnormality, and when an abnormality occurs, triggers the jump of the link smlh_link_up signal, and sends the signal jump to the interruption control unit; the interruption control unit generates an interruption signal according to the received signal jump, and sends the interruption signal to the CPU and the link loopback unit; the link loopback unit automatically configures the PCIe link to enter a self-loopback mode after receiving the interruption signal sent by the interruption control unit, and performs abnormality processing. The present invention designs a processing method for front-end hardware link switching and software collaborative control, simplifies the complexity of hardware design, and ensures the stability of the main control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of SSD storage, and in particular to an SSD link exception processing system and method. Background Art

[0002] Due to the performance advantages of solid-state drives (hereinafter referred to as SSDs), they are increasingly used in servers and storage systems as devices for data storage. Therefore, the stability of SSDs is crucial to the system. The PCIe link is the channel for SSDs to interact with the outside world, which is related to the stability of the data and command paths. Link anomalies include link disconnection and abnormal power failure. Link disconnection will cause the originally transmitted data to be interrupted, which brings challenges to the error handling of many hardware acceleration engines designed inside the SSD. Abnormal power failure refers to the sudden loss of power during the normal operation of the SSD, which is a more serious link anomaly. On the one hand, it is necessary to complete the effective data preservation of the internal cache in time after the abnormal power failure occurs; on the other hand, SSDs use non-volatile storage media nand flash particles, which have the principle that physical storage units cannot be overwritten, which has more stringent requirements for the handling of abnormal power failures. Usually, SSD boards have built-in capacitors to support the power needs of the SSD board after an abnormal power failure occurs. Since the moment of abnormal power failure is unpredictable, it brings difficulties to abnormal handling. Summary of the invention

[0003] In view of the defects of the prior art, the present invention provides an SSD link exception processing system and method. For link abnormalities such as link disconnection and abnormal power failure, a processing method of front-end hardware link switching and software collaborative control is designed, which simplifies the complexity of hardware design and ensures the stability of the main control system.

[0004] In order to solve the technical problem, the technical solution adopted by the present invention is: an SSD link abnormality processing system, including a link disconnection detection unit, an interruption control unit and a link loopback unit;

[0005] The link anomaly detection unit is connected to the PCIe module and the interrupt control unit, and is used to detect whether an abnormality occurs in the PCIe link. When an abnormality occurs, the link smlh_link_up signal jump is triggered and the signal jump is sent to the interrupt control unit;

[0006] The interrupt control unit is connected to the link anomaly detection unit, the CPU, and the link loopback unit, and is used to generate an interrupt signal according to the received signal jump, and send the interrupt signal to the CPU and the link loopback unit;

[0007] The link loopback unit is connected to the PCIe module, the interrupt control unit, and the CPU. After receiving the interrupt signal sent by the interrupt control unit, the link loopback unit automatically configures the PCIe link to enter the self-loopback mode and performs exception processing.

[0008] Furthermore, when the link abnormality detection unit detects that the link is abnormally disconnected, the PCIe link enters the self-loop mode and performs the following abnormal processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus.

[0009] Furthermore, when the abnormality detected by the link abnormality detection unit is an abnormal power-off of the link, the PCIe link enters the self-loop mode and performs the following abnormal processing: notifying the CPU core to start the abnormal power-off action and stop sending new DMA requests, polling and detecting the returned completion message of the DMA requests that have been sent, marking the completion of the DMA write request that has been sent before the interrupt occurs, and performing a logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, a transmission failure is returned to the algorithm layer, the host data is not saved, and the cache is released. If the read request is completed, a transmission completion is returned to the algorithm layer and saved. Then, it is determined whether it is an abnormal power-off. If so, the power-off process is completed. If not, the link training and link establishment are re-performed, the DMA is re-initialized, and the NVME controller is re-initialized.

[0010] Further, when an abnormal power failure event occurs, link abnormality processing follows or precedes processing of the abnormal power failure event.

[0011] Furthermore, the system is applicable to single-core or multi-core main control systems.

[0012] The present invention also discloses a SSD link abnormality detection method, comprising the following steps:

[0013] S01), the link abnormality detection unit detects whether the PCIe link has an abnormality. When an abnormality occurs, it triggers the jump of the link smlh_link_up signal and sends the signal jump to the interrupt control unit;

[0014] S02), the interrupt control unit generates an interrupt signal according to the received signal jump, and sends the interrupt signal to the CPU and the link loopback unit;

[0015] S03), after receiving the interrupt signal sent by the interrupt control unit, the link loopback unit automatically configures the PCIe link to enter the self-loopback mode and performs exception processing.

[0016] Furthermore, when the link abnormality detection unit detects that the link is abnormally disconnected, the PCIe link enters the self-loop mode and performs the following abnormal processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus.

[0017] Furthermore, when the abnormality detected by the link abnormality detection unit is an abnormal power-off of the link, the PCIe link enters the self-loop mode and performs the following abnormal processing: notifying the CPU core to start the abnormal power-off action and stop sending new DMA requests, polling and detecting the returned completion message of the DMA request that has been sent, marking the completion for the DMA write request that has been sent before the interrupt occurs, and performing a logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, a transmission failure is returned to the algorithm layer, the host data is not saved, and the cache is released. If the read request is completed, a transmission completion is returned to the algorithm layer and saved. Then, it is determined whether it is an abnormal power-off. If so, the power-off process is completed. If not, the link is retrained and the link is established, the DMA is reinitialized, and the NVMe controller is reinitialized.

[0018] Further, when an abnormal power failure event occurs, link abnormality processing follows or precedes processing of the abnormal power failure event.

[0019] Furthermore, this method is applicable to single-core or multi-core main control systems.

[0020] Beneficial effects of the invention: The present invention application provides a system and method for processing SSD link abnormalities. Compared with traditional methods, the system designed by this patent targets link abnormalities such as link disconnection and abnormal power failure, and designs a processing method for front-end hardware link switching and software collaborative control, which reduces the complex interaction process between coupling modules and the reset processing of hardware units, simplifies the complexity of hardware design, and ensures the stability of the main control system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 This is the principle block diagram of the exception handling system;

[0022] Figure 2 It is a principle block diagram of the link loopback control unit;

[0023] Figure 3 Flowchart of the exception handling method. DETAILED DESCRIPTION

[0024] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation examples.

[0025] Example 1

[0026] This embodiment discloses a SSD link exception processing system. Figure 1As shown, it includes a link disconnection detection unit, an interruption control unit and a link loopback unit; the link anomaly detection unit is connected to the PCIe module and the interruption control unit, and is used to detect whether the PCIe link has an abnormality. When an abnormality occurs, the link smlh_link_up signal jump is triggered, and the signal jump is sent to the interruption control unit; the interruption control unit is connected to the link anomaly detection unit, the CPU, and the link loopback unit, and is used to generate an interrupt signal according to the received signal jump, and send the interrupt signal to the CPU and the link loopback unit; the link loopback unit is connected to the PCIe module, the interruption control unit, and the CPU. After receiving the interruption signal sent by the interruption control unit, the link loopback unit automatically configures the PCIe link to enter the self-loop mode to perform abnormal processing.

[0027] When the link anomaly detection unit detects an abnormal link disconnection, the PCIe link enters the self-loop mode and performs the following exception processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus. Specifically, the write data request that the DMA has sent to the bus and originally expected to be sent to the host through the PCIeTx port is automatically transferred to the RX cache and discarded; the read data request sent to the host automatically returns fixed pattern data, thereby ensuring that all requests will not be stuck on the AXI bus, and the error response corresponding to the AXI bus is returned. This greatly simplifies the exception handling process of each hardware acceleration module inside the SSD and avoids unnecessary reset operations.

[0028] The abnormality detected by the link abnormality detection unit is that the PCIe link enters the self-loop mode when the link is powered off abnormally. The abnormality processing performed is as follows: notify the CPU core to start the abnormal power-off action and stop sending new DMA requests, poll and detect the returned completion message of the DMA request that has been sent, mark the completion for the DMA write request that has been sent before the interrupt occurs, and perform logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, return the transmission failure to the algorithm layer, do not save the host data, and release the cache. If the read request is completed, return the transmission completion to the algorithm layer and save it; then determine whether it is an abnormal power off. If so, the power-off process is completed. If not, re-train the link and establish the link, re-initialize the DMA, and re-initialize the NVMe controller.

[0029] When an abnormal power failure event occurs, link abnormality processing follows or precedes the processing of the abnormal power failure event to ensure the stability of system data.

[0030] This system is applicable to single-core or multi-core main control systems. The CPU of a multi-core system is divided into different functional cores according to different processing tasks. The method provided by the present invention does not rely on the inter-core division of the multi-core system, does not affect the functions of each core, and guarantees the compatibility of the solution to the greatest extent. Specifically, for a multi-core system, the interrupt signal can be sent to multiple CPU cores at the same time, or it can be sent to one CPU core and then forwarded to other CPUs. Therefore, it does not rely on the inter-core division of the multi-core system.

[0031] like Figure 2 As shown, the link loopback control unit includes a signal sampling unit, an event triggering unit, a PCIe IP configuration unit and a loopback mode enabling unit. The signal sampling unit receives an abnormal jump signal smlh_link_up and a clock signal E-clock sent from the outside. The event triggering unit triggers the self-loopback mode according to the received signal. The PCIe IP configuration unit configures the PCIe IP signal required to enter the self-loopback mode. The loopback mode enabling unit sends a self-loopback enable signal PCIe_loop_en, thereby starting the self-loopback.

[0032] Example 2

[0033] This embodiment discloses a method for handling SSD link abnormalities. Figure 3 As shown, the following steps are included:

[0034] S01), the link abnormality detection unit detects whether the PCIe link has an abnormality. When an abnormality occurs, it triggers the jump of the link smlh_link_up signal and sends the signal jump to the interrupt control unit.

[0035] The PCIe link needs to go through a series of link training states before it can be established. After the link is established, data can be sent normally between upstream and downstream devices. This state is called link up. The link disconnection detection logic unit is responsible for designing the link_down logic to indicate the disconnection of the link according to the change of the physical layer signal when the link is disconnected, and is used as an indication switch for abnormal processing.

[0036] S02), the interrupt control unit generates an interrupt signal according to the received signal jump, and sends the interrupt signal to the CPU and the link loopback unit.

[0037] The interrupt control module is used to generate an interrupt signal to the specified CPU core when link_down occurs. When the CPU core receives the interrupt, the interrupt can be cleared. The interrupt control module is designed with an interrupt signal shielding function, which is more flexible. The CPU core is responsible for processing read and write commands, error handling procedures after abnormal link disconnection, power-off procedures after abnormal power failure, etc. For multi-core systems, interrupt signals can be sent to multiple CPU cores at the same time, or to one CPU core, and then forwarded to other CPUs.

[0038] S03) After receiving the interrupt signal sent by the interrupt control unit, the link loopback unit automatically configures the PCIe link to enter the self-loopback mode and performs exception processing. Different requests received from the internal bus are processed differently. Since the link is disconnected at this time, there is no need to wait for the host's response, so all unfinished and newly received read requests return all-zero data, and the status returns a timeout error; all unfinished or write requests in the receiving end cache return completion, and the data is directly looped back to the receiving end and discarded. All completed requests return a completion status. With the characteristics of the PCIe physical layer, no additional logic is required, ensuring the smoothness of the data path and the integrity of the command.

[0039] Specifically as shown in 3a, when the abnormality detected by the link abnormality detection unit is an abnormal link disconnection, the PCIe link enters the self-loop mode and performs the following abnormality processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus.

[0040] As shown in 3b, the abnormality detected by the link abnormality detection unit is that the PCIe link enters the self-loop mode when the link is powered off abnormally. The abnormality processing is as follows: notify the CPU core to start the abnormal power-off action and stop sending new DMA requests, poll and detect the DMA request that has been sent to return the completion message, mark the DMA write request that has been sent before the interrupt occurs as completed, and perform logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, return the transmission failure to the algorithm layer, do not save the host data, and release the cache. If the read request is completed, return the transmission completion to the algorithm layer and save it; then determine whether it is an abnormal power off. If so, the power-off process is completed. If not, re-train the link and establish the link, re-initialize the DMA, and re-initialize the NVMe controller.

[0041] When an abnormal power failure event occurs, link abnormality processing will follow or precede the abnormal power failure event processing to ensure the stability of system data. This method is applicable to single-core or multi-core main control systems.

[0042] The above description is only the basic principle and preferred embodiments of the present invention. Improvements and substitutions made by those skilled in the art based on the present invention belong to the protection scope of the present invention.

Claims

1. A SSD link exception handling system, characterized in that: It includes a link anomaly detection unit, an interruption control unit and a link loopback unit; The link anomaly detection unit is connected to the PCIe module and the interrupt control unit, and is used to detect whether an abnormality occurs in the PCIe link. When an abnormality occurs, the link smlh_link_up signal jump is triggered and the signal jump is sent to the interrupt control unit; The interrupt control unit is connected to the link anomaly detection unit, the CPU, and the link loopback unit, and is used to generate an interrupt signal according to the received signal jump, and send the interrupt signal to the CPU and the link loopback unit; The link loopback unit is connected to the PCIe module, the interrupt control unit, and the CPU. After receiving the interrupt signal sent by the interrupt control unit, the link loopback unit automatically configures the PCIe link to enter the self-loopback mode and performs exception processing. When the link anomaly detection unit detects that the link is abnormally disconnected, the PCIe link enters the self-loopback mode and performs the following exception processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus; when the link anomaly detection unit detects that the link is abnormally powered off, the PCIe link enters the self-loopback mode and performs the following exception processing: Know that the CPU core starts abnormal power-off and stops sending new DMA requests, polls and detects the completion message returned by the DMA requests that have been sent, marks the DMA write request that has been sent before the interrupt occurs as completed, and performs logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, return the transmission failure to the algorithm layer, do not save the host data, and release the cache. If the read request is completed, return the transmission completion to the algorithm layer and save it; then determine whether it is an abnormal power off. If so, the power-off process is completed. If not, re-train the link and establish the link, re-initialize the DMA, and re-initialize the NVMe controller.

2. The SSD link exception handling system according to claim 1, characterized in that: When an abnormal power failure event occurs, link abnormality processing follows or precedes processing of the abnormal power failure event.

3. The SSD link exception handling system according to claim 1, characterized in that: This system is suitable for single-core or multi-core main control systems.

4. A method for detecting anomalies in an SSD link, characterized in that: The following steps are involved: S01), the link abnormality detection unit detects whether the PCIe link has an abnormality. When an abnormality occurs, it triggers the jump of the link smlh_link_up signal and sends the signal jump to the interrupt control unit; S02), the interrupt control unit generates an interrupt signal according to the received signal jump, and sends the interrupt signal to the CPU and the link loopback unit; S03), after receiving the interrupt signal sent by the interrupt control unit, the link loopback unit automatically configures the PCIe link to enter the self-loopback mode and performs exception processing; When the abnormality detected by the link abnormality detection unit is an abnormal link disconnection, the PCIe link enters the self-loop mode and performs the following abnormal processing: the write request data in the queue is directly discarded, and the completion status is returned to the AXI bus; the read request in the queue is terminated, and a timeout error is returned to the AXI bus; when the abnormality detected by the link abnormality detection unit is an abnormal link power-off, the PCIe link enters the self-loop mode and performs the following abnormal processing: notifying the CPU core to start the abnormal power-off action and stop sending new DMA requests, polling and detecting the returned completion message of the DAM request that has been sent, marking the completion for the DMA write request that has been sent before the interrupt occurs, and performing a logical judgment on the return status of the read request that has been sent before the interrupt occurs. If the read request is not completed, a transmission failure is returned to the algorithm layer, the host data is not saved, and the cache is released. If the read request is completed, a transmission completion is returned to the algorithm layer and saved; then it is determined whether it is an abnormal power-off. If so, the power-off process is completed. If not, the link training and link establishment are re-performed, the DMA is re-initialized, and the NVMe controller is re-initialized.

5. The SSD link anomaly detection method according to claim 4, characterized in that: When an abnormal power failure event occurs, link abnormality processing follows or precedes processing of the abnormal power failure event.

6. The SSD link anomaly detection method according to claim 4, characterized in that: This method is applicable to single-core or multi-core master control systems.

Citation Information

Patent Citations

  • Method for shortening abnormal power failure processing time of solid state disk

    CN108255630A

  • Interrupt message generation device, interrupt message generation method and end equipment

    CN111078597A