SSD link state detection method and system based on link information

By acquiring real-time rate values ​​within the observation window and combining them with PCIe error detection information for hierarchical and progressive judgment, the problem of misjudgment in SSD link status detection is solved, improving detection accuracy and reliability, and ensuring the stable operation of storage devices.

CN122019299APending Publication Date: 2026-05-12DAWNING INFORMATION IND (BEIJING) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DAWNING INFORMATION IND (BEIJING) CO LTD
Filing Date
2026-01-08
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies have a high false alarm rate in SSD link status detection, cannot effectively identify brief interference and sudden traffic fluctuations, leading to unnecessary alarms or repair operations, and cannot identify abnormalities in the connection between the frame backplane or controller, affecting the normal operation of SSD.

Method used

By acquiring real-time rate values ​​within the observation window, performing refined classification based on preset thresholds, and combining PCIe error detection information for hierarchical and progressive judgment, an observation window and state percentage mechanism are introduced, and a first abnormal state and a second abnormal state are set for comprehensive judgment, thereby reducing the false alarm rate and improving detection accuracy.

Benefits of technology

It significantly reduces the false alarm rate, avoids unnecessary alarms or repair operations triggered by misjudgments, improves the accuracy and reliability of SSD link status detection, can identify potential hardware failures early, and ensures the stable operation of storage devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019299A_ABST
    Figure CN122019299A_ABST
Patent Text Reader

Abstract

The invention provides an SSD link state detection method and system based on link information, and the method comprises the steps: obtaining a real-time rate value of a target link of an SSD at each moment in a current observation window; determining a transmission rate state of a target link at each moment according to each real-time rate value and a preset threshold value; according to the type ratio of each transmission rate state, determining a current initial link state detection result of the target link, the initial link state detection result being a normal state, a first abnormal state and a second abnormal state; if the initial link state detection result is a second abnormal state, determining that the current link state detection result of the target link is a link abnormality; and if the initial link state detection result is a first abnormal state, judging a current link state detection result of the target link according to PCIe error detection information of the target link in a preset time period, thereby improving the accuracy of SSD link state detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage device link detection technology, and in particular to a storage device link status detection method and system based on link information. Background Technology

[0002] As the storage medium for storage devices, the health and operational status of Solid State Drives (SSDs) have always been a key research focus for storage manufacturers. Although the number of failed drives is gradually increasing due to advancements in redundancy and other technologies, early detection of SSD anomalies and subsequent corrective action remain crucial. When monitoring SSD status, SMART is an indispensable and critical indicator. Provided by SSD manufacturers based on their firmware operation, it offers significant reference value. SMART information is provided by each manufacturer according to a standard protocol based on their firmware. Some manufacturers have extended SMART, providing richer information and offering query methods such as Log Pages. However, SSDs are plugged into the storage enclosure and rely on the inter-controller link for data transmission. The quality of this link directly impacts service delivery, and the information provided by SMART is relatively limited. Therefore, the information mentioned above cannot effectively and directly reflect the SSD's link status.

[0003] The current common practice in link detection is to detect anomalies as soon as they are detected. This means that if an abnormal value or unexpected result is detected, it is considered a link anomaly, and relevant alarms are reported or intervention measures are taken. However, the link itself may be affected by unpredictable interference factors. These effects are usually temporary and non-continuous, and may recover quickly, with the actual impact being within acceptable limits. If a link anomaly is determined solely based on a momentary anomaly detected, and intervention measures are taken accordingly, it may lead to misjudgment and affect the normal operation of the SSD.

[0004] Furthermore, storage devices support cascading various drive enclosures for expansion purposes. However, regardless of whether the drive controller is integrated or separate, the enclosure containing the SSD is directly connected to it. The current common practice is to treat each SSD within the enclosure as an independent unit for testing, handling each individually when an anomaly is detected. However, this strategy cannot identify anomalies in the enclosure backplane or in the connection between the enclosure and the higher-level controller. These types of anomalies are generally common and usually affect more than just one drive within the enclosure. If such anomalies are encountered, relying solely on identifying the anomaly within the SSD itself and taking corresponding measures will result in a certain number of SSDs being identified as abnormal and queued for subsequent operations. This may not be effective and could even be counterproductive. Summary of the Invention

[0005] To address the aforementioned technical issues, this application provides an SSD link status detection method and system based on link information, thereby improving the accuracy of SSD link status detection.

[0006] In a first aspect, embodiments of this application provide an SSD link state detection method based on link information, including: Within the current observation window, obtain the real-time rate values ​​of the target link of the SSD at various times. Based on the real-time rate values ​​and preset thresholds, the transmission rate status of the target link at each time moment is determined, and the transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. Based on the proportion of each of the transmission rate states, the initial link state detection result of the target link is determined, and the initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. If the initial link status detection result is a second abnormal state, then the current link status detection result of the target link is determined to be a link abnormality; if the initial link status detection result is a first abnormal state, then the current link status detection result of the target link is determined based on the PCIe error detection information of the target link within a preset time period.

[0007] This application provides an SSD link status detection method based on link information. First, the real-time rate value of the target link is continuously acquired within an observation window. Then, the transmission rate status at each moment is finely classified (normal, first anomaly, second anomaly) based on a preset threshold. The initial link status detection result is then obtained by comprehensively considering the status proportion within the entire window. Finally, PCIe error detection information is combined for a comprehensive judgment to obtain the final link status detection result. This application abandons the traditional simplistic and crude method of "instantaneous detection and anomaly determination," introducing the concepts of "observation window" and "status proportion." This makes the link status judgment no longer based on a single instantaneous "snapshot," but on a continuous "trend" over a period of time. This judgment logic based on statistical and trend analysis can effectively filter out occasional rate drops caused by instantaneous interference, sudden traffic fluctuations, or brief hardware jitter, thereby significantly reducing the false alarm rate and avoiding unnecessary alarms or repair operations triggered by misjudgments, which could interfere with the normal operation of the SSD. Furthermore, in the initial link status detection results, in addition to the normal state and the second abnormal state, this embodiment also sets an intermediate state of "first abnormal state," and introduces a second judgment mechanism for the intermediate state—combining PCIe error detection information within a preset time period for final adjudication. This layered and progressive judgment strategy combines two dimensions: "speed performance indicators" and "link error indicators," making the detection results more comprehensive, objective, and reliable. It can not only capture performance problems reflected by speed anomalies but also correlate with underlying hardware error information, which is of great value in distinguishing between transient software / environmental interference and persistent hardware link failures. This significantly improves the accuracy of SSD link status detection and provides a more reliable and refined basis for subsequent anomaly handling decisions.

[0008] Furthermore, determining the transmission rate status of the target link at each time point based on the various real-time rate values ​​and preset thresholds includes: If the real-time rate value of the target link at any time is greater than or equal to the first preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a normal rate state. If the real-time rate value of the target link at any time is less than a first preset threshold and greater than a second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a first abnormal rate state. If the real-time rate value of the target link at any time is less than or equal to the second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be the second abnormal rate state.

[0009] This application specifies and quantifies the rules for determining the transmission rate status, clearly defining a "first preset threshold" and a "second preset threshold," dividing the real-time rate value into three distinct intervals, corresponding to normal, first abnormal, and second abnormal rate states, respectively. By setting two thresholds instead of a single threshold, a "buffer zone," i.e., the first abnormal rate state, is created. This interval represents a situation where the rate has decreased but has not yet reached the minimum warning line. This design is highly practical because it distinguishes between situations where "performance fluctuates slightly but may be harmless" and situations where "performance is severely degraded or interrupted." For anomalies within the buffer zone, the system does not immediately "sentence to death" but triggers a more in-depth secondary detection. This reflects the prudence and refinement of the detection strategy and provides a more refined data foundation for subsequently determining the initial link state detection results.

[0010] In one possible implementation, determining the current initial link state detection result of the target link based on the type proportion of each of the various transmission rate states includes: The observation window is divided into several sub-windows on an average basis according to a preset time length. Traverse each of the sub-windows. For any sub-window, if all transmission rate states in the sub-window are in the normal rate state, then determine that the current initial link state detection result of the target link is in the normal state and stop traversing; otherwise, count the type count of each transmission rate state in the sub-window, and calculate the weighted value corresponding to each transmission rate state in the sub-window based on each type count and the corresponding preset weight. Based on the weighted values ​​corresponding to various transmission rate states in each sub-window, the final weighted value corresponding to various transmission rate states in the observation window is calculated. Based on the magnitude of each final weighted value, the current initial link state detection result of the target link is determined; If all the final weighted values ​​are equal, then the current initial link state detection result of the target link is determined to be the second abnormal state.

[0011] This application provides a method for determining the initial link state detection result based on the proportion of transmission rate state types, introducing a mechanism of "sub-window division" and "weighted statistics". First, this embodiment divides the entire observation window into multiple sub-windows on an average basis, and sequentially analyzes the data within each observation window in segments, enabling a more detailed observation of the distribution and change patterns of the link state over time. Even if there are severe anomalies (secondary abnormal states) at individual moments and time periods within the entire window, as long as all sub-windows are in normal states, the link as a whole can be quickly determined to be normal and the traversal can be stopped. This effectively prevents misjudgments caused by occasional, instantaneous severe interference (such as a strong electromagnetic pulse), greatly enhancing the system's ability to resist transient interference and the robustness of link detection. Second, when no sub-window is found that is in a normal state at every moment after traversing all sub-windows, the counts of different state types within each sub-window are statistically analyzed, and a preset weight is introduced for weighted calculation, ultimately obtaining the "final weighted value" of the entire observation window. This weighting mechanism assigns different levels of importance to anomalies of varying severity. For example, the weight of the second anomaly is typically set higher than that of the first, allowing the final result to more accurately reflect the "overall health" and "trend of anomaly severity" of the link during the observation period, rather than simply a "majority vote," thus improving the accuracy of SSD link status detection. Finally, this embodiment also provides a handling rule for cases where weighted values ​​are equal, i.e., directly determining it as the second anomaly. This provides a clear, conservative, and safe decision-making basis for ambiguous boundary situations, ensuring the determinism of the detection process and the clarity of the final result.

[0012] In one possible implementation, if the initial link state detection result is a first abnormal state, then determining the current link state detection result of the target link based on the PCIe error detection information of the target link within a preset time period includes: Obtain the PCIe error detection information of the target link within a preset time period; Based on a first preset duration, the preset time period is divided into several first sub-time periods on average, and a first indicator value is calculated for each first sub-time period. The first indicator value is the number of times an uncorrectable error occurs in the PCIe error detection information during the sub-time period. If the first indicator value of any first sub-time period is greater than or equal to the corresponding third preset threshold, then the current link status detection result of the target link is determined to be an abnormal state.

[0013] This application proposes a specific method for secondary judgment based on PCIe error detection information when the initial detection result is "first abnormal state," i.e., the rate is in a slightly abnormal range. Its core is focusing on the frequency of "uncorrectable errors." This embodiment divides a preset time period into multiple first sub-time periods according to a "first preset duration," and counts the number of uncorrectable errors occurring in each sub-time period as a first indicator value. If the first indicator value of any sub-time period exceeds a third preset threshold, the current link status detection result of the target link is determined to be abnormal. This embodiment provides crucial, hardware-level corroborating information for "suspected abnormal" link statuses, achieving penetrating diagnosis from "performance manifestations" to "hardware root causes." Uncorrectable errors on the PCIe bus usually indicate serious data integrity errors, such as parity errors or link training failures. These errors are often directly related to physical layer connection problems, signal integrity degradation, or controller / SSD firmware defects, and are strong indicators of hardware link failures. When rate detection detects a decline in link performance (first abnormal state), there are many possible causes on a single basis (such as load changes, queue congestion, etc.). However, if uncorrectable errors occur frequently within the same time period, it is almost certain that there is an underlying hardware link problem. This embodiment quantifies the concept of "frequent occurrence" by setting a "third preset threshold". As long as the number of uncorrectable errors exceeds the safety threshold in any first sub-time period, the link is immediately determined to be abnormal. This mechanism can quickly and accurately capture those "hidden" or "early" hardware faults that, although the current rate decline is not obvious, have already shown serious error symptoms, realizing early warning of potential serious problems and further improving the accuracy of SSD link status detection.

[0014] Furthermore, the SSD link state detection method also includes: If the first indicator value is less than the third preset threshold, then based on the second preset duration, the preset time period is divided into several second sub-time periods on average, and the second indicator values ​​of each second sub-time period are counted. The second indicator value is the number of times various correctable errors in the PCIe error detection information occur in the second sub-time period. The first preset duration is greater than the second preset duration. For any second sub-time period, if several second indicator values ​​of the second sub-time period are greater than the fourth preset threshold, then the second sub-time period is determined to be an abnormal time period. If there is a preset number of consecutive abnormal time periods within the preset time period, then the second index values ​​of each second sub-time period are weighted and summed according to the type of correctable error to obtain a comprehensive index value. If the comprehensive index value is greater than the fifth preset threshold, then the current link status detection result of the target link is determined to be abnormal.

[0015] This application further expands the depth and breadth of PCIe error information analysis. For cases where "uncorrectable errors do not exceed the threshold," it introduces further analysis and evaluation of "correctable errors." Its beneficial effect lies in constructing a multi-layered, progressive error analysis system, capable of more sensitively and comprehensively capturing soft faults and chronic degradation trends in the link. Correctable errors, although automatically repaired by hardware without affecting the final data correctness, frequently occur, directly reflecting problems such as link signal quality degradation, clock skew, and minor interference. Therefore, this embodiment first uses a shorter time granularity, specifically a second preset time period shorter than the first preset duration, to divide the second sub-time period for high-frequency monitoring, capturing those brief but frequent error pulses. This embodiment not only counts the total number of errors but also distinguishes various types of correctable errors and sets thresholds, requiring that multiple types of errors exceed the threshold within a time period before it is marked as an "abnormal time period." This avoids accidental fluctuations of individual error types falsely triggering alarms. Secondly, this embodiment introduces the concept of "continuous abnormal time periods," requiring a preset number of consecutive abnormal time periods before subsequent detection actions can be performed. This ensures that the detected error patterns are persistent and trend-based, rather than isolated events. When the above conditions are met, the statistical results of various correctable errors are weighted and summed by type to obtain a "comprehensive index value," which is then compared with a fifth preset threshold. This mechanism allows the system to assign different weights to the predictive severity of link health based on different types of correctable errors, thereby deriving a more scientific and comprehensive score that better reflects the overall risk. This enables the system to identify "sub-healthy" states where uncorrectable errors have not yet occurred, but correctable errors are consistently frequent and link quality is steadily deteriorating. This achieves earlier and more refined prediction and judgment of potential performance bottlenecks and future failure risks, further improving the accuracy of SSD link status detection.

[0016] In one possible implementation, the SSD link state detection method further includes: If the current link status detection result of any target link of the SSD is abnormal, the target link is determined to be an abnormal link, and the SSD is determined to be an abnormal SSD. Within the current observation window, if the number of abnormal SSDs detected in the same storage device exceeds the preset abnormal number threshold, a device-level alarm is reported; otherwise, abnormal repair is performed on each of the abnormal SSDs.

[0017] This embodiment elevates the link status detection results to the SSD device level and storage device system level for comprehensive decision-making and processing. When any target link is determined to be abnormal, the link and its associated SSD are marked as abnormal. Furthermore, within an observation window, the number of abnormal SSDs within the same storage device (e.g., a disk enclosure) is counted. When the number exceeds an abnormality threshold, a device-level alarm is reported, realizing a hierarchical response strategy from "point" (single link) to "surface" (single SSD) to "volume" (the entire storage device or disk enclosure). Specifically, this embodiment introduces an "abnormality threshold" as a critical condition for determining whether to trigger a "device-level alarm," effectively identifying and distinguishing between "individual random faults" and "system-wide common faults." If only one or two SSDs within the same disk enclosure are detected as abnormal, this is likely a problem with an individual SSD or a specific link. The system will choose to perform independent fault repair (e.g., link reset) on each abnormal SSD, a targeted localized processing approach. Conversely, if the number of abnormal SSDs exceeds a preset threshold within a short period, this strongly suggests a higher-level common fault, such as unstable power supply to the drive enclosure's backplane, a faulty connector or link between the enclosure and the upper-level controller, or poor heat dissipation affecting multiple drives in a localized area. In this case, an "device-level alarm message" should be immediately reported to alert the administrator to focus on enclosure-level or system-level issues, rather than blindly performing potentially ineffective or even harmful repair operations on a large number of SSDs one by one. This avoids inefficient and chaotic "point-to-point" handling in the event of systemic risks, ensuring the safety and reliability of the storage device operation.

[0018] Furthermore, the abnormal repair of each of the abnormal SSDs includes: For any abnormal SSD, all links between the abnormal SSD and the controller are simultaneously shut down and then simultaneously reopened to complete one abnormal repair of the abnormal SSD. If the number of times any SSD is repaired for an anomaly exceeds the preset repair count threshold, the corresponding SSD anomaly alarm information will be reported.

[0019] This application provides a method for repairing faulty SSDs. It involves simultaneously shutting down and then simultaneously reopening all links between the faulty SSD and the controller. This operation is commonly referred to as a "link reset" or "port reset." By simultaneously shutting down and reopening all links, temporary error states at both ends of the links can be forcibly cleared, the physical layer and link layer connections can be rebuilt, and negotiation and training can be re-performed, thereby eliminating many communication failures caused by soft errors or transient interference. Secondly, this operation is fast and has relatively little impact on business operations. Compared to powering down and restarting the entire SSD, link reset typically only occurs at the PCIe link layer, allowing for collaborative completion at the operating system and SSD firmware levels. It is quick and has limited impact on the data and status of the SSD itself. Finally, this embodiment also introduces a "repair count threshold" management mechanism. If the number of fault repairs for the same SSD exceeds a preset threshold, repairs will cease, and an "SSD fault alarm message" will be reported. This is an important protection strategy. This means that if an SSD's link frequently malfunctions and requires repeated resets, it's likely not a temporary issue, but rather a hardware defect in the SSD itself (such as an interface controller failure) or a persistent source of interference that's difficult to eliminate. In this case, continued repeated repairs may not fundamentally solve the problem and could even mask the true fault. Reporting SSD-level alarms can prompt administrators to conduct a more in-depth inspection, diagnosis, or consider replacement of the SSD, preventing intermittent failures from becoming permanent and ensuring the long-term stability of the storage pool.

[0020] Furthermore, the SSD link state detection method also includes: During the first processing time period after the abnormal SSD is repaired, if any first abnormal link in the abnormal SSD is detected to be in an abnormal state again, the first abnormal link will be shut down and the number of abnormal shutdowns of the first abnormal link will be accumulated. If, within the next observation window after the first abnormal link is closed, the unclosed link of the abnormal SSD is detected to be in an abnormal state, then the first abnormal link is opened. If, during the second processing period after the first abnormal link is closed, the abnormal SSD is not detected to be in an abnormal state, the first abnormal link is opened and the abnormal SSD is repaired. If the number of abnormal shutdowns of any first link in the SSD exceeds a preset shutdown threshold, then the first link will no longer be opened. The number of link ports of the SSD is greater than or equal to 2.

[0021] This application proposes a more refined and intelligent link fault isolation and recovery management strategy for SSDs with multiple link ports. By fully utilizing the redundant connection capabilities of multi-port SSDs, it maximizes business continuity, minimizes the impact of faults, and optimizes resource utilization. The core idea of ​​this embodiment is to dynamically and conditionally close and enable abnormal links, prioritizing the normal operation of storage devices. First, if a previously repaired abnormal link is detected as abnormal again within a certain period after repair, it is closed individually. This is equivalent to "isolating" the port, and business traffic will automatically be transmitted through other normal ports, thereby avoiding the continuous impact of the faulty port on services and ensuring uninterrupted service access to the SSD. Second, it sets flexible link reopening conditions: if other unclosed links are also detected as abnormal in the next observation window, it indicates that the problem may not be a single port, or that the business pressure requires more bandwidth. In this case, the previously closed links are reopened to provide additional possible channels. This reflects dynamic adaptation to the overall system load and health. Another reopening condition is that if other links remain normal during a second processing period after the link is shut down, the closed link will be reopened and an anomaly repair attempt will be made. This is equivalent to giving the isolated port a "recovery trial period" to see if it can resume normal operation when the system is relatively stable. Finally, this embodiment introduces a "shutdown count threshold" as the final decision-making basis. If a link is abnormally shut down too frequently, exceeding the threshold, the system determines that the port has an inherent problem that is difficult to recover from, and will no longer automatically reopen it, but will isolate it for a long time, which may trigger a higher-level hardware replacement alarm. This embodiment realizes dynamic and intelligent management of the status of multi-port links. Under the premise of ensuring the continuity of core business, it attempts to recover and utilize redundant resources as much as possible, and degrades resources identified as "bad ports", which greatly improves the self-healing capability, availability and operational stability of the multi-path storage system.

[0022] Secondly, embodiments of this application provide an SSD link state detection system based on link information, including an acquisition module, a rate judgment module, a first detection module, and a second detection module; The acquisition module is used to acquire the real-time rate value of the target link of the SSD at each time point within the current observation window. The rate determination module is used to determine the transmission rate status of the target link at each time based on each real-time rate value and a preset threshold. The transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. The first detection module is used to determine the initial link state detection result of the target link based on the proportion of the types of each of the transmission rate states. The initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. The second detection module is used to determine that the current link status detection result of the target link is abnormal if the initial link status detection result is a second abnormal state; and to determine the current link status detection result of the target link based on the PCIe error detection information of the target link within a preset time period if the initial link status detection result is a first abnormal state.

[0023] Furthermore, the SSD link status detection system also includes an anomaly handling module; The anomaly handling module is used to determine that if the current link status detection result of any target link of the SSD is abnormal, the target link is an abnormal link and the SSD is an abnormal SSD; within the current observation window, if the number of abnormal SSDs detected and determined in the same storage device exceeds a preset abnormal number threshold, a device-level alarm message is reported; otherwise, anomaly repair is performed on each of the abnormal SSDs. Attached Figure Description

[0024] Figure 1 A flowchart illustrating an SSD link state detection method based on link information provided in this application embodiment; Figure 2 A detailed flowchart illustrating an SSD link status detection method during implementation, provided for an embodiment of this application; Figure 3 A flowchart illustrating a method for repairing and handling link anomalies provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of an SSD link state detection system based on link information, provided in an embodiment of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0026] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0027] Example 1: like Figure 1 As shown, Embodiment 1 provides an SSD link state detection method based on link information, including steps S1-S4: Step S1: Within the current observation window, obtain the real-time rate value of the target link of the SSD at each time point; Step S2: Determine the transmission rate status of the target link at each time based on each real-time rate value and the preset threshold. The transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. Step S3: Determine the initial link state detection result of the target link based on the type proportion of each of the transmission rate states. The initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. Step S4: If the initial link status detection result is the second abnormal state, then the current link status detection result of the target link is determined to be a link abnormality; if the initial link status detection result is the first abnormal state, then the current link status detection result of the target link is determined based on the PCIe error detection information of the target link within a preset time period.

[0028] This application provides an SSD link status detection method based on link information. First, the real-time rate value of the target link is continuously acquired within an observation window. Then, the transmission rate status at each moment is finely classified (normal, first anomaly, second anomaly) based on a preset threshold. The initial link status detection result is then obtained by comprehensively considering the status proportion within the entire window. Finally, PCIe error detection information is combined for a comprehensive judgment to obtain the final link status detection result. This application abandons the traditional simplistic and crude method of "instantaneous detection and anomaly determination," introducing the concepts of "observation window" and "status proportion." This makes the link status judgment no longer based on a single instantaneous "snapshot," but on a continuous "trend" over a period of time. This judgment logic based on statistical and trend analysis can effectively filter out occasional rate drops caused by instantaneous interference, sudden traffic fluctuations, or brief hardware jitter, thereby significantly reducing the false alarm rate and avoiding unnecessary alarms or repair operations triggered by misjudgments, which could interfere with the normal operation of the SSD. Furthermore, in the initial link status detection results, in addition to the normal state and the second abnormal state, this embodiment also sets an intermediate state of "first abnormal state," and introduces a second judgment mechanism for the intermediate state—combining PCIe error detection information within a preset time period for final adjudication. This layered and progressive judgment strategy combines two dimensions: "speed performance indicators" and "link error indicators," making the detection results more comprehensive, objective, and reliable. It can not only capture performance problems reflected by speed anomalies but also correlate with underlying hardware error information, which is of great value in distinguishing between transient software / environmental interference and persistent hardware link failures. This significantly improves the accuracy of SSD link status detection and provides a more reliable and refined basis for subsequent anomaly handling decisions.

[0029] In this embodiment, PCIe is a bus standard, while NVMe is a protocol specification designed specifically for flash memory. Theoretically, they can exist without relying on PCIe alone, but combining them can fully leverage their respective advantages to achieve better performance. Therefore, NVMe SSDs based on the PCIe bus standard are the mainstream choice for current flash storage manufacturers. The PCIe device configuration space stores basic and configuration-related information about the PCIe device. From this space, you can read the device's speed (GEN1, GEN2, etc.) and bandwidth (X1, X2, etc., usually representing the number of lanes), including the current speed / bandwidth and the expected speed / bandwidth. If the current speed / bandwidth is lower than the expected value, we generally say that the device has slowed down or reduced its lanes. Some degree of slowdown or lane reduction may have a limited impact on device operation. Simultaneously, PCIe has an Advanced Error Detection and Reporting (AER) mechanism, which can select some errors as key observation items for link status.

[0030] Furthermore, in step S2, determining the transmission rate status of the target link at each time point based on the various real-time rate values ​​and preset thresholds includes: If the real-time rate value of the target link at any time is greater than or equal to the first preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a normal rate state. If the real-time rate value of the target link at any time is less than a first preset threshold and greater than a second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a first abnormal rate state. If the real-time rate value of the target link at any time is less than or equal to the second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be the second abnormal rate state.

[0031] This application specifies and quantifies the rules for determining the transmission rate status, clearly defining a "first preset threshold" and a "second preset threshold," dividing the real-time rate value into three distinct intervals, corresponding to normal, first abnormal, and second abnormal rate states, respectively. By setting two thresholds instead of a single threshold, a "buffer zone," i.e., the first abnormal rate state, is created. This interval represents a situation where the rate has decreased but has not yet reached the minimum warning line. This design is highly practical because it distinguishes between situations where "performance fluctuates slightly but may be harmless" and situations where "performance is severely degraded or interrupted." For anomalies within the buffer zone, the system does not immediately "sentence to death" but triggers a more in-depth secondary detection. This reflects the prudence and refinement of the detection strategy and provides a more refined data foundation for subsequently determining the initial link state detection results.

[0032] In a preferred embodiment, depending on different requirements, the storage device may cascade multiple different drive enclosures, or it may be a single device integrating drive controller and disk controller. Regardless of the actual number of enclosures, to facilitate centralized management of SSDs within different enclosures, each enclosure in the storage device is assigned a unique enclosure number. A timer is set to periodically check the SSD link status. To quickly detect anomalies, the detection period should not be too long, generally recommended to be in the order of seconds.

[0033] During the rate detection process, in addition to comparing the real-time rate with a preset threshold, it can also detect link speed reduction / lane reduction. Here, "reduction" refers to a decrease in bandwidth, which could be a reduction from X2 to X1 within the same rate, or a reduction from GEN3 to GEN1 (possibly accompanied by a simultaneous bandwidth decrease). Different levels of reduction have different impacts on SSD operation. Some reduction levels are within acceptable limits and have limited practical impact, while others can cause SSD malfunctions. This embodiment sets a reduction threshold, which can be a bandwidth reduction threshold like X4 to X1 within the same rate, or a reduction across rates (e.g., a reduction from X1 at a certain rate to a certain bandwidth at the next rate). This reduction threshold can be manually adjusted during storage device operation and is dynamic by default, meaning different thresholds are used based on different desired rates, such as threshold a for desired rate GEN3, b for GEN4, etc.

[0034] As mentioned earlier, link detection uses a second-level unit as the cycle. The link's speed reduction / lane reduction is counted for each cycle according to a set reduction threshold. If no speed reduction / lane reduction occurs in the current cycle (i.e., the real-time rate value is greater than or equal to the first preset threshold), the normal count is incremented by 1. If speed reduction / lane reduction occurs, but the reduction level does not exceed the set reduction threshold (i.e., the real-time rate value is less than the first preset threshold but greater than the second preset threshold), the minor anomaly count is incremented by 1. If the reduction level exceeds the set threshold (i.e., the real-time rate value is less than or equal to the second preset threshold), the anomaly count is incremented by 1. The counts within each window are counted in minute-level sub-windows, and the counts for multiple consecutive windows are also calculated to obtain the overall observation window.

[0035] In a preferred embodiment, the first and second preset thresholds are obtained through dynamic calculation and adjustment. Specifically, the first and second preset thresholds are dynamically calculated and adjusted based on the historical performance of the target SSD. Specifically, within a learning cycle (e.g., 24 hours of stable operation) after the SSD is first put into use or after each repair, the system continuously monitors its link rate under typical load conditions and calculates the SSD's "baseline performance curve" using statistical methods. Subsequently, the first preset threshold can be set as a percentage of this baseline performance value (e.g., 80%), and the second preset threshold can be set as a lower percentage (e.g., 50%) or a minimum absolute rate value to ensure basic link communication. When a significant decrease in the SSD's SMART health index is detected, the system can automatically fine-tune the aforementioned percentages, appropriately relaxing the judgment criteria. In this way, the thresholds can adapt to different models and aging levels of SSDs, avoiding the problem of using a globally fixed threshold that is too lenient for high-performance drives or too harsh for low-performance or old drives. While maintaining detection sensitivity, it further reduces misjudgments caused by individual differences.

[0036] In one possible implementation, step S3, determining the current initial link state detection result of the target link based on the type proportion of each of the transmission rate states, includes: The observation window is divided into several sub-windows on an average basis according to a preset time length. Traverse each of the sub-windows. For any sub-window, if all transmission rate states in the sub-window are in the normal rate state, then determine that the current initial link state detection result of the target link is in the normal state and stop traversing; otherwise, count the type count of each transmission rate state in the sub-window, and calculate the weighted value corresponding to each transmission rate state in the sub-window based on each type count and the corresponding preset weight. Based on the weighted values ​​corresponding to various transmission rate states in each sub-window, the final weighted value corresponding to various transmission rate states in the observation window is calculated. Based on the magnitude of each final weighted value, the current initial link state detection result of the target link is determined; If all the final weighted values ​​are equal, then the current initial link state detection result of the target link is determined to be the second abnormal state.

[0037] This application provides a method for determining the initial link state detection result based on the proportion of transmission rate state types, introducing a mechanism of "sub-window division" and "weighted statistics". First, this embodiment divides the entire observation window into multiple sub-windows on an average basis, and sequentially analyzes the data within each observation window in segments, enabling a more detailed observation of the distribution and change patterns of the link state over time. Even if there are severe anomalies (secondary abnormal states) at individual moments and time periods within the entire window, as long as all sub-windows are in normal states, the link as a whole can be quickly determined to be normal and the traversal can be stopped. This effectively prevents misjudgments caused by occasional, instantaneous severe interference (such as a strong electromagnetic pulse), greatly enhancing the system's ability to resist transient interference and the robustness of link detection. Second, when no sub-window is found that is in a normal state at every moment after traversing all sub-windows, the counts of different state types within each sub-window are statistically analyzed, and a preset weight is introduced for weighted calculation, ultimately obtaining the "final weighted value" of the entire observation window. This weighting mechanism assigns different levels of importance to anomalies of varying severity. For example, the weight of the second anomaly is typically set higher than that of the first, allowing the final result to more accurately reflect the "overall health" and "trend of anomaly severity" of the link during the observation period, rather than simply a "majority vote," thus improving the accuracy of SSD link status detection. Finally, this embodiment also provides a handling rule for cases where weighted values ​​are equal, i.e., directly determining it as the second anomaly. This provides a clear, conservative, and safe decision-making basis for ambiguous boundary situations, ensuring the determinism of the detection process and the clarity of the final result.

[0038] In a preferred embodiment, for three different counts, an increasing weighting is assigned according to normal / slightly abnormal / abnormal. An observation window is divided into several minute-level sub-windows. If only one type of count has a value in an observation window, that count type is used as the link status for the current observation window. If two or more types of counts have values ​​in an observation window, they are sorted according to their weighted proportions. The type with the highest proportion is set as the link status for the current period. If the proportions are roughly equal, the most severe count type in the current period is selected as the link status. The weighting increasing from normal / slightly abnormal / abnormal means that normal counts need to have a certain advantage to set the link status for the entire period as normal.

[0039] Furthermore, if a sub-window within the observation window is confirmed to be in an abnormal state after weighted calculation, but other sub-windows (within the observation window) do not show a decrease in speed / lane, then the entire observation window is discarded, and the detection starts again. In other words, the initial link state detection result corresponding to the current observation window is considered normal. If minor anomalies or alternating anomalies occur within the observation window, the final weighted value of all detection cycles within the observation window is used.

[0040] In a preferred embodiment, when counting the types of rate states within a sub-window, not only the "number of occurrences" is recorded, but also the continuous duration of each abnormal state (first abnormality, second abnormality) is accumulated. For example, a second abnormal state lasting 10 seconds may have a higher weight contribution than 10 second abnormal states that occur at intervals and each lasts only 1 second. When calculating the sub-window weighted value, the "type count" is multiplied by the "duration factor of the abnormality of that type," and then a weighted sum is performed. This mechanism makes the evaluation focus more on persistent abnormalities, rather than just the frequency of abnormality occurrences, thereby better distinguishing between brief transient disturbances and stable performance degradation trends, and improving the indicativeness of the initial state detection results to the actual fault.

[0041] In one possible implementation, in step S4, if the initial link state detection result is a first abnormal state, determining the current link state detection result of the target link based on the PCIe error detection information of the target link within a preset time period includes: Obtain the PCIe error detection information of the target link within a preset time period; Based on a first preset duration, the preset time period is divided into several first sub-time periods on average, and a first indicator value is calculated for each first sub-time period. The first indicator value is the number of times an uncorrectable error occurs in the PCIe error detection information during the sub-time period. If the first indicator value of any first sub-time period is greater than or equal to the corresponding third preset threshold, then the current link status detection result of the target link is determined to be an abnormal state.

[0042] This application proposes a specific method for secondary judgment based on PCIe error detection information when the initial detection result is "first abnormal state," i.e., the rate is in a slightly abnormal range. Its core is focusing on the frequency of "uncorrectable errors." This embodiment divides a preset time period into multiple first sub-time periods according to a "first preset duration," and counts the number of uncorrectable errors occurring in each sub-time period as a first indicator value. If the first indicator value of any sub-time period exceeds a third preset threshold, the current link status detection result of the target link is determined to be abnormal. This embodiment provides crucial, hardware-level corroborating information for "suspected abnormal" link statuses, achieving penetrating diagnosis from "performance manifestations" to "hardware root causes." Uncorrectable errors on the PCIe bus usually indicate serious data integrity errors, such as parity errors or link training failures. These errors are often directly related to physical layer connection problems, signal integrity degradation, or controller / SSD firmware defects, and are strong indicators of hardware link failures. When rate detection detects a decline in link performance (first abnormal state), there are many possible causes on a single basis (such as load changes, queue congestion, etc.). However, if uncorrectable errors occur frequently within the same time period, it is almost certain that there is an underlying hardware link problem. This embodiment quantifies the concept of "frequent occurrence" by setting a "third preset threshold". As long as the number of uncorrectable errors exceeds the safety threshold in any first sub-time period, the link is immediately determined to be abnormal. This mechanism can quickly and accurately capture those "hidden" or "early" hardware faults that, although the current rate decline is not obvious, have already shown serious error symptoms, realizing early warning of potential serious problems and further improving the accuracy of SSD link status detection.

[0043] Furthermore, the SSD link state detection method also includes: If the first indicator value is less than the third preset threshold, then based on the second preset duration, the preset time period is divided into several second sub-time periods on average, and the second indicator values ​​of each second sub-time period are counted. The second indicator value is the number of times various correctable errors in the PCIe error detection information occur in the second sub-time period. The first preset duration is greater than the second preset duration. For any second sub-time period, if several second indicator values ​​of the second sub-time period are greater than the fourth preset threshold, then the second sub-time period is determined to be an abnormal time period. If there is a preset number of consecutive abnormal time periods within the preset time period, then the second index values ​​of each second sub-time period are weighted and summed according to the type of correctable error to obtain a comprehensive index value. If the comprehensive index value is greater than the fifth preset threshold, then the current link status detection result of the target link is determined to be abnormal.

[0044] This application further expands the depth and breadth of PCIe error information analysis. For cases where "uncorrectable errors do not exceed the threshold," it introduces further analysis and evaluation of "correctable errors." Its beneficial effect lies in constructing a multi-layered, progressive error analysis system, capable of more sensitively and comprehensively capturing soft faults and chronic degradation trends in the link. Correctable errors, although automatically repaired by hardware without affecting the final data correctness, frequently occur, directly reflecting problems such as link signal quality degradation, clock skew, and minor interference. Therefore, this embodiment first uses a shorter time granularity, specifically a second preset time period shorter than the first preset duration, to divide the second sub-time period for high-frequency monitoring, capturing those brief but frequent error pulses. This embodiment not only counts the total number of errors but also distinguishes various types of correctable errors and sets thresholds, requiring that multiple types of errors exceed the threshold within a time period before it is marked as an "abnormal time period." This avoids accidental fluctuations of individual error types falsely triggering alarms. Secondly, this embodiment introduces the concept of "continuous abnormal time periods," requiring a preset number of consecutive abnormal time periods before subsequent detection actions can be performed. This ensures that the detected error patterns are persistent and trend-based, rather than isolated events. When the above conditions are met, the statistical results of various correctable errors are weighted and summed by type to obtain a "comprehensive index value," which is then compared with a fifth preset threshold. This mechanism allows the system to assign different weights to the predictive severity of link health based on different types of correctable errors, thereby deriving a more scientific and comprehensive score that better reflects the overall risk. This enables the system to identify "sub-healthy" states where uncorrectable errors have not yet occurred, but correctable errors are consistently frequent and link quality is steadily deteriorating. This achieves earlier and more refined prediction and judgment of potential performance bottlenecks and future failure risks, further improving the accuracy of SSD link status detection.

[0045] In a preferred embodiment, after determining that the SSD link has a minor anomaly based on the count in the observation window, another link detection can be combined to further confirm the link status. This detection relies on PCIe's Advanced Error Detection and Reporting (AER) mechanism, which classifies detected errors into correctable errors and uncorrectable errors. Uncorrectable errors are further divided into fatal and non-fatal errors, each with different corresponding scenarios and impacts.

[0046] Specifically, for unrecoverable errors, due to their significant impact, a smaller threshold (single-digit level) and a longer observation period, i.e., the first sub-time period, are selected. If the number of unrecoverable error observations within the observation period exceeds the set threshold, the SSD link is judged to be abnormal. When the number of unrecoverable error observations in each observation period does not exceed the set threshold, correctable errors are further analyzed. This embodiment selects certain types of correctable errors as key observation objects for link status, such as BAD TLP and BAD DLLP among correctable errors. In BAD TLP (Bad Transaction Layer Packet), TLP is a transaction layer packet, the highest-level packet format in the PCIe protocol stack, carrying the actual read / write requests, configuration information, and the data to be transmitted. When the system or device verifies the TLP at the receiving end and finds a problem, it marks it as BAD TLP. In BAD DLLP (Bad Data Link Layer Packet), DLLP is a data link layer packet, one level lower than TLP. It does not carry application data but is used for link management. Similar to TLP, DLLP also has CRC protection. If its CRC check fails during transmission, or if the format does not meet the protocol requirements, it will be marked as BAD DLLP. The reason for choosing the above two types as observation items is that they are very common errors in actual observation, the data volume is considerable, and the accumulated reference samples are sufficient. In addition, correctable errors also include recerive errors, physical layer detects physical signal errors; REPLAY_NUMRollover, data link layer retransmission counter retransmission; Replay Timer Timeout, data link layer packet transmission timeout; Poisoned TLP Received Status, poisoned (data error) packets received by the device; Poisoned TLP EgressBlocked Status, the device's Egress port detects poison TLP packets and intercepts the transmission of poison data packets, etc. Since different errors have different impacts, different observation periods and thresholds need to be set for the selected errors.

[0047] Taking BAD TLP and BAD DLLP as correctable errors as examples, this embodiment treats them as a single observation item, with an observation period set at the minute level, i.e., the second sub-time period. The number of both errors within the period is counted. Since they are correctable errors, the threshold is set to a relatively large magnitude, and continuous monitoring is performed for several periods. Furthermore, correctable errors other than BAD TLP and BAD DLLP can be considered, with an additional threshold set at a smaller magnitude, maintaining the minute-level observation period, and similarly continuously monitoring for several periods. If the number of each correctable error exceeds the corresponding threshold in several consecutive observation periods, a weighted calculation is performed over the entire preset time period, with BAD TLP and BAD DLLP accounting for 25%, and the other five each accounting for 15%. A comprehensive index value is calculated, and an anomaly is determined based on the comprehensive index value.

[0048] In one possible implementation, the SSD link state detection method further includes: If the current link status detection result of any target link of the SSD is abnormal, the target link is determined to be an abnormal link, and the SSD is determined to be an abnormal SSD. Within the current observation window, if the number of abnormal SSDs detected in the same storage device exceeds the preset abnormal number threshold, a device-level alarm is reported; otherwise, abnormal repair is performed on each of the abnormal SSDs.

[0049] This embodiment elevates the link status detection results to the SSD device level and storage device system level for comprehensive decision-making and processing. When any target link is determined to be abnormal, the link and its associated SSD are marked as abnormal. Furthermore, within an observation window, the number of abnormal SSDs within the same storage device (e.g., a disk enclosure) is counted. When the number exceeds an abnormality threshold, a device-level alarm is reported, realizing a hierarchical response strategy from "point" (single link) to "surface" (single SSD) to "volume" (the entire storage device or disk enclosure). Specifically, this embodiment introduces an "abnormality threshold" as a critical condition for determining whether to trigger a "device-level alarm," effectively identifying and distinguishing between "individual random faults" and "system-wide common faults." If only one or two SSDs within the same disk enclosure are detected as abnormal, this is likely a problem with an individual SSD or a specific link. The system will choose to perform independent fault repair (e.g., link reset) on each abnormal SSD, a targeted localized processing approach. Conversely, if the number of abnormal SSDs exceeds a preset threshold within a short period, this strongly suggests a higher-level common fault, such as unstable power supply to the drive enclosure's backplane, a faulty connector or link between the enclosure and the upper-level controller, or poor heat dissipation affecting multiple drives in a localized area. In this case, an "device-level alarm message" should be immediately reported to alert the administrator to focus on enclosure-level or system-level issues, rather than blindly performing potentially ineffective or even harmful repair operations on a large number of SSDs one by one. This avoids inefficient and chaotic "point-to-point" handling in the event of systemic risks, ensuring the safety and reliability of the storage device operation.

[0050] In a preferred embodiment, the SSD link state detection flowchart is as follows: Figure 2 As shown in the diagram. Simultaneously, link status detection is performed on each link of each SSD in the storage device. During each link status detection process, link speed information and PCIe error report information are acquired and statistically analyzed separately. The entire detection process prioritizes speed information, supplemented by PCIe error report information. When a link anomaly can be directly determined through speed information, further analysis of PCIe error report information is unnecessary, and the process proceeds directly to the subsequent alarm and repair procedures. If a link anomaly cannot be directly determined through speed information alone, the final link status detection result is further determined based on unrecoverable and correctable errors in the PCIe error report information. Finally, when an SSD link anomaly is detected, the link status of all SSDs in the same frame is retrieved based on its frame number. If all SSD links in the frame exceeding a set threshold are abnormal, the current frame is considered abnormal, and a frame-level alarm is reported for upper-level intervention. If the number of SSDs with abnormal links in the frame is less than the threshold, it is considered a disk-level anomaly, and the disk is repaired and processed separately.

[0051] Furthermore, the abnormal repair of each of the abnormal SSDs includes: For any abnormal SSD, all links between the abnormal SSD and the controller are simultaneously shut down and then simultaneously reopened to complete one abnormal repair of the abnormal SSD. If the number of times any SSD is repaired for an anomaly exceeds the preset repair count threshold, the corresponding SSD anomaly alarm information will be reported.

[0052] This application provides a method for repairing faulty SSDs. It involves simultaneously shutting down and then simultaneously reopening all links between the faulty SSD and the controller. This operation is commonly referred to as a "link reset" or "port reset." By simultaneously shutting down and reopening all links, temporary error states at both ends of the links can be forcibly cleared, the physical layer and link layer connections can be rebuilt, and negotiation and training can be re-performed, thereby eliminating many communication failures caused by soft errors or transient interference. Secondly, this operation is fast and has relatively little impact on business operations. Compared to powering down and restarting the entire SSD, link reset typically only occurs at the PCIe link layer, allowing for collaborative completion at the operating system and SSD firmware levels. It is quick and has limited impact on the data and status of the SSD itself. Finally, this embodiment also introduces a "repair count threshold" management mechanism. If the number of fault repairs for the same SSD exceeds a preset threshold, repairs will cease, and an "SSD fault alarm message" will be reported. This is an important protection strategy. This means that if an SSD's link frequently malfunctions and requires repeated resets, it's likely not a temporary issue, but rather a hardware defect in the SSD itself (such as an interface controller failure) or a persistent source of interference that's difficult to eliminate. In this case, continued repeated repairs may not fundamentally solve the problem and could even mask the true fault. Reporting SSD-level alarms can prompt administrators to conduct a more in-depth inspection, diagnosis, or consider replacement of the SSD, preventing intermittent failures from becoming permanent and ensuring the long-term stability of the storage pool.

[0053] Furthermore, the SSD link state detection method also includes: During the first processing time period after the abnormal SSD is repaired, if any first abnormal link in the abnormal SSD is detected to be in an abnormal state again, the first abnormal link will be shut down and the number of abnormal shutdowns of the first abnormal link will be accumulated. If, within the next observation window after the first abnormal link is closed, the unclosed link of the abnormal SSD is detected to be in an abnormal state, then the first abnormal link is opened. If, during the second processing period after the first abnormal link is closed, the abnormal SSD is not detected to be in an abnormal state, the first abnormal link is opened and the abnormal SSD is repaired. If the number of abnormal shutdowns of any first link in the SSD exceeds a preset shutdown threshold, then the first link will no longer be opened. The number of link ports of the SSD is greater than or equal to 2.

[0054] This application proposes a more refined and intelligent link fault isolation and recovery management strategy for SSDs with multiple link ports. By fully utilizing the redundant connection capabilities of multi-port SSDs, it maximizes business continuity, minimizes the impact of faults, and optimizes resource utilization. The core idea of ​​this embodiment is to dynamically and conditionally close and enable abnormal links, prioritizing the normal operation of storage devices. First, if a previously repaired abnormal link is detected as abnormal again within a certain period after repair, it is closed individually. This is equivalent to "isolating" the port, and business traffic will automatically be transmitted through other normal ports, thereby avoiding the continuous impact of the faulty port on services and ensuring uninterrupted service access to the SSD. Second, it sets flexible link reopening conditions: if other unclosed links are also detected as abnormal in the next observation window, it indicates that the problem may not be a single port, or that the business pressure requires more bandwidth. In this case, the previously closed links are reopened to provide additional possible channels. This reflects dynamic adaptation to the overall system load and health. Another reopening condition is that if other links remain normal during a second processing period after the link is shut down, the closed link will be reopened and an anomaly repair attempt will be made. This is equivalent to giving the isolated port a "recovery trial period" to see if it can resume normal operation when the system is relatively stable. Finally, this embodiment introduces a "shutdown count threshold" as the final decision-making basis. If a link is abnormally shut down too frequently, exceeding the threshold, the system determines that the port has an inherent problem that is difficult to recover from, and will no longer automatically reopen it, but will isolate it for a long time, which may trigger a higher-level hardware replacement alarm. This embodiment realizes dynamic and intelligent management of the status of multi-port links. Under the premise of ensuring the continuity of core business, it attempts to recover and utilize redundant resources as much as possible, and degrades resources identified as "bad ports", which greatly improves the self-healing capability, availability and operational stability of the multi-path storage system.

[0055] In a preferred embodiment, such as Figure 3As shown, a method for repairing and handling link anomalies is provided. The repair process begins with a rapid power-down and power-up of the disk with the link anomaly, provided that redundancy and other constraints allow. Simultaneously, the existing link between the SSD and the controller is shut down and then reopened. The link status of the SSD is then continuously monitored. Redundancy refers to the ability of the isolated disk to avoid affecting the data in its storage pool. Limited by the storage's RAID level, the number of disks that a storage device can isolate is limited; exceeding this limit will result in data loss.

[0056] Furthermore, if an SSD link anomaly is still detected within the set timeframe after repair, especially if one of the dual ports was abnormal before repair and remains abnormal after repair, then that port link will be shut down. A shutdown record and timestamp will be added upon successful shutdown. If the disk re-enters repair due to a link anomaly or if no link anomaly occurs again after the specified time, then the link will be reopened during repair or normal operation, and the SSD's link status will continue to be monitored. If the number of times the link is shut down due to post-repair link anomalies exceeds a threshold, the link will no longer be manually reopened until the number of repair attempts exceeds the set threshold before an SSD anomaly is reported.

[0057] In addition, if the SSD link is not detected within the time frame set after repair, but is detected later, the above-mentioned policy of shutting down the link after repair will not be triggered. Instead, the SSD will be repaired directly. However, a threshold is set for the repair. If the number of repairs exceeds the threshold, the SSD will be reported as abnormal.

[0058] The above focuses on situations where a link is repaired but then continues to detect anomalies. For scenarios where a link is repaired but another link detects anomalies after the repair, the link is not shut down, but rather treated as another repair. Meanwhile, for scenarios where both links are abnormal and remain abnormal after repair, shutting down both links simultaneously would essentially isolate the SSD. It is recommended to still follow the repair scenario and identify the SSD as abnormal according to the repair threshold strategy.

[0059] In summary, this embodiment has at least the following beneficial effects: 1) This embodiment does not use a detection-based anomaly determination mechanism. Instead, it uses a longer detection period to comprehensively determine the final link status, which can effectively avoid misjudgment.

[0060] 2) In this embodiment, when detecting speed reduction / lane reduction, different reduction thresholds are set according to different expected rates. Based on these thresholds, the speed reduction / lane reduction situation is divided into three levels, corresponding to different severity levels and assigned different weights. The final link status is determined by the weighted situation of each duration period in the multi-window period, which can effectively identify the influence of short-term fluctuations in the external environment and improve detection accuracy.

[0061] 3) In this embodiment, a unique frame number is assigned to each frame in the storage device, and the SSDs are centrally managed by the frame number. When an SSD link abnormality is detected, the link status of all SSDs in the same frame is obtained to determine whether the current abnormality is at the frame level or the disk level, which can effectively improve the accuracy of subsequent operations.

[0062] 4) In this embodiment, while detecting speed reduction / lane reduction, some errors are selected from AER as the observation objects of link status, and different detection mechanisms are set according to error types, which effectively expands the link detection capability.

[0063] 5) This embodiment also provides a repair and processing method after detecting an SSD link anomaly, which repairs and processes the SSD while minimizing the impact of the current link anomaly on the SSD.

[0064] Example 2: like Figure 4 As shown, this application provides an SSD link state detection system based on link information, including an acquisition module 10, a rate judgment module 20, a first detection module 30, and a second detection module 40. The acquisition module 10 is used to acquire the real-time rate value of the target link of the SSD at each time within the current observation window. The rate determination module 20 is used to determine the transmission rate status of the target link at each time based on each real-time rate value and a preset threshold. The transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. The first detection module 30 is used to determine the current initial link state detection result of the target link based on the type ratio of each of the transmission rate states. The initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. The second detection module 40 is used to determine that the current link status detection result of the target link is abnormal if the initial link status detection result is a second abnormal state; and to determine the current link status detection result of the target link based on the PCIe error detection information of the target link within a preset time period if the initial link status detection result is a first abnormal state.

[0065] Furthermore, the rate determination module 20 determines the transmission rate status of the target link at various times based on each of the real-time rate values ​​and a preset threshold, including: If the real-time rate value of the target link at any time is greater than or equal to the first preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a normal rate state. If the real-time rate value of the target link at any time is less than a first preset threshold and greater than a second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a first abnormal rate state. If the real-time rate value of the target link at any time is less than or equal to the second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be the second abnormal rate state.

[0066] In one possible implementation, the first detection module 30 determines the current initial link state detection result of the target link based on the proportion of each of the transmission rate states, including: The observation window is divided into several sub-windows on an average basis according to a preset time length. Traverse each of the sub-windows. For any sub-window, if all transmission rate states in the sub-window are in the normal rate state, then determine that the current initial link state detection result of the target link is in the normal state and stop traversing; otherwise, count the type count of each transmission rate state in the sub-window, and calculate the weighted value corresponding to each transmission rate state in the sub-window based on each type count and the corresponding preset weight. Based on the weighted values ​​corresponding to various transmission rate states in each sub-window, the final weighted value corresponding to various transmission rate states in the observation window is calculated. Based on the magnitude of each final weighted value, the current initial link state detection result of the target link is determined; If all the final weighted values ​​are equal, then the current initial link state detection result of the target link is determined to be the second abnormal state.

[0067] In one possible implementation, if the initial link state detection result is a first abnormal state, the second detection module 40 determines the current link state detection result of the target link based on the PCIe error detection information of the target link within a preset time period, including: Obtain the PCIe error detection information of the target link within a preset time period; Based on a first preset duration, the preset time period is divided into several first sub-time periods on average, and a first indicator value is calculated for each first sub-time period. The first indicator value is the number of times an uncorrectable error occurs in the PCIe error detection information during the sub-time period. If the first indicator value of any first sub-time period is greater than or equal to the corresponding third preset threshold, then the current link status detection result of the target link is determined to be an abnormal state.

[0068] Furthermore, the SSD link status detection system also includes: If the first indicator value is less than the third preset threshold, then based on the second preset duration, the preset time period is divided into several second sub-time periods on average, and the second indicator values ​​of each second sub-time period are counted. The second indicator value is the number of times various correctable errors in the PCIe error detection information occur in the second sub-time period. The first preset duration is greater than the second preset duration. For any second sub-time period, if several second indicator values ​​of the second sub-time period are greater than the fourth preset threshold, then the second sub-time period is determined to be an abnormal time period. If there is a preset number of consecutive abnormal time periods within the preset time period, then the second index values ​​of each second sub-time period are weighted and summed according to the type of correctable error to obtain a comprehensive index value. If the comprehensive index value is greater than the fifth preset threshold, then the current link status detection result of the target link is determined to be abnormal.

[0069] In one possible implementation, the SSD link status detection system further includes an anomaly handling module; The anomaly handling module is used to determine that if the current link status detection result of any target link of the SSD is abnormal, the target link is an abnormal link and the SSD is an abnormal SSD; within the current observation window, if the number of abnormal SSDs detected and determined in the same storage device exceeds a preset abnormal number threshold, a device-level alarm message is reported; otherwise, anomaly repair is performed on each of the abnormal SSDs.

[0070] Furthermore, the abnormal repair of each of the abnormal SSDs includes: For any abnormal SSD, all links between the abnormal SSD and the controller are simultaneously shut down and then simultaneously reopened to complete one abnormal repair of the abnormal SSD. If the number of times any SSD is repaired for an anomaly exceeds the preset repair count threshold, the corresponding SSD anomaly alarm information will be reported.

[0071] Furthermore, the SSD link status detection system also includes: During the first processing time period after the abnormal SSD is repaired, if any first abnormal link in the abnormal SSD is detected to be in an abnormal state again, the first abnormal link will be shut down and the number of abnormal shutdowns of the first abnormal link will be accumulated. If, within the next observation window after the first abnormal link is closed, the unclosed link of the abnormal SSD is detected to be in an abnormal state, then the first abnormal link is opened. If, during the second processing period after the first abnormal link is closed, the abnormal SSD is not detected to be in an abnormal state, the first abnormal link is opened and the abnormal SSD is repaired. If the number of abnormal shutdowns of any first link in the SSD exceeds a preset shutdown threshold, then the first link will no longer be opened. The number of link ports of the SSD is greater than or equal to 2.

[0072] This application provides an SSD link status detection system based on link information. First, it continuously acquires the real-time rate value of the target link within an observation window. Then, based on a preset threshold, it performs fine-grained classification of the transmission rate status at each moment (normal, first anomaly, second anomaly). Next, it synthesizes the status proportion within the entire window to obtain an initial link status detection result. Finally, it combines PCIe error detection information for comprehensive judgment to obtain the final link status detection result. This application abandons the traditional simplistic and crude method of "instantaneous detection and anomaly determination," introducing the concepts of "observation window" and "status proportion." This makes the link status judgment no longer based on a single instantaneous "snapshot," but on a continuous "trend" over a period of time. This judgment logic based on statistical and trend analysis can effectively filter out occasional rate drops caused by instantaneous interference, sudden traffic fluctuations, or brief hardware jitter, thereby significantly reducing the false alarm rate and avoiding unnecessary alarms or repair operations triggered by misjudgments, which could interfere with the normal operation of the SSD. Furthermore, in the initial link status detection results, in addition to the normal state and the second abnormal state, this embodiment also sets an intermediate state of "first abnormal state," and introduces a second judgment mechanism for the intermediate state—combining PCIe error detection information within a preset time period for final adjudication. This layered and progressive judgment strategy combines two dimensions: "speed performance indicators" and "link error indicators," making the detection results more comprehensive, objective, and reliable. It can not only capture performance problems reflected by speed anomalies but also correlate with underlying hardware error information, which is of great value in distinguishing between transient software / environmental interference and persistent hardware link failures. This significantly improves the accuracy of SSD link status detection and provides a more reliable and refined basis for subsequent anomaly handling decisions.

[0073] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.

[0074] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.

Claims

1. A method for SSD link state detection based on link information, characterized in that, include: Within the current observation window, obtain the real-time rate values ​​of the target link of the SSD at various times. Based on the real-time rate values ​​and preset thresholds, the transmission rate status of the target link at each time moment is determined, and the transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. Based on the proportion of each of the transmission rate states, the initial link state detection result of the target link is determined, and the initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. If the initial link state detection result is the second abnormal state, then the current link state detection result of the target link is determined to be an abnormal link. If the initial link status detection result is a first abnormal state, then the current link status detection result of the target link is determined based on the PCIe error detection information of the target link within a preset time period.

2. The SSD link state detection method based on link information as described in claim 1, characterized in that, Determining the transmission rate status of the target link at each time step based on the real-time rate values ​​and preset thresholds includes: If the real-time rate value of the target link at any time is greater than or equal to the first preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a normal rate state. If the real-time rate value of the target link at any time is less than a first preset threshold and greater than a second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be a first abnormal rate state. If the real-time rate value of the target link at any time is less than or equal to the second preset threshold, then the transmission rate state of the target link at the corresponding time is determined to be the second abnormal rate state.

3. The SSD link state detection method based on link information as described in claim 1, characterized in that, The step of determining the initial link state detection result of the target link based on the proportion of each of the transmission rate states includes: The observation window is divided into several sub-windows on an average basis according to a preset time length. Traverse each of the sub-windows. For any sub-window, if all transmission rate states in the sub-window are in the normal rate state, then determine that the current initial link state detection result of the target link is in the normal state and stop traversing; otherwise, count the type count of each transmission rate state in the sub-window, and calculate the weighted value corresponding to each transmission rate state in the sub-window based on each type count and the corresponding preset weight. Based on the weighted values ​​corresponding to various transmission rate states in each sub-window, the final weighted value corresponding to various transmission rate states in the observation window is calculated. Based on the magnitude of each final weighted value, the current initial link state detection result of the target link is determined; If all the final weighted values ​​are equal, then the current initial link state detection result of the target link is determined to be the second abnormal state.

4. The SSD link state detection method based on link information as described in claim 1, characterized in that, If the initial link state detection result is a first abnormal state, then the current link state detection result of the target link is determined based on the PCIe error detection information of the target link within a preset time period, including: Obtain the PCIe error detection information of the target link within a preset time period; Based on a first preset duration, the preset time period is divided into several first sub-time periods on average, and a first indicator value is calculated for each first sub-time period. The first indicator value is the number of times an uncorrectable error occurs in the PCIe error detection information during the sub-time period. If the first indicator value of any first sub-time period is greater than or equal to the corresponding third preset threshold, then the current link status detection result of the target link is determined to be an abnormal state.

5. The SSD link state detection method based on link information as described in claim 4, characterized in that, The SSD link status detection method further includes: If the first indicator value is less than the third preset threshold, then based on the second preset duration, the preset time period is divided into several second sub-time periods on average, and the second indicator values ​​of each second sub-time period are counted. The second indicator value is the number of times various correctable errors in the PCIe error detection information occur in the second sub-time period. The first preset duration is greater than the second preset duration. For any second sub-time period, if several second indicator values ​​of the second sub-time period are greater than the fourth preset threshold, then the second sub-time period is determined to be an abnormal time period. If there is a preset number of consecutive abnormal time periods within the preset time period, then the second index values ​​of each second sub-time period are weighted and summed according to the type of correctable error to obtain a comprehensive index value. If the comprehensive index value is greater than the fifth preset threshold, then the current link status detection result of the target link is determined to be abnormal.

6. The SSD link state detection method based on link information as described in any one of claims 1-5, characterized in that, The SSD link status detection method further includes: If the current link status detection result of any target link of the SSD is abnormal, the target link is determined to be an abnormal link, and the SSD is determined to be an abnormal SSD. Within the current observation window, if the number of abnormal SSDs detected in the same storage device exceeds the preset abnormal number threshold, a device-level alarm is reported; otherwise, abnormal repair is performed on each of the abnormal SSDs.

7. The SSD link state detection method based on link information as described in claim 6, characterized in that, The specific steps for repairing each of the abnormal SSDs include: For any abnormal SSD, all links between the abnormal SSD and the controller are simultaneously shut down and then simultaneously reopened to complete one abnormal repair of the abnormal SSD. If the number of times any SSD is repaired for an anomaly exceeds the preset repair count threshold, the corresponding SSD anomaly alarm information will be reported.

8. The SSD link state detection method based on link information as described in claim 7, characterized in that, The SSD link status detection method further includes: During the first processing time period after the abnormal SSD is repaired, if any first abnormal link in the abnormal SSD is detected to be in an abnormal state again, the first abnormal link will be shut down and the number of abnormal shutdowns of the first abnormal link will be accumulated. If, within the next observation window after the first abnormal link is closed, the unclosed link of the abnormal SSD is detected to be in an abnormal state, then the first abnormal link is opened. If, during the second processing period after the first abnormal link is closed, the abnormal SSD is not detected to be in an abnormal state, the first abnormal link is opened and the abnormal SSD is repaired. If the number of abnormal shutdowns of any first link in the SSD exceeds a preset shutdown threshold, then the first link will no longer be opened. The number of link ports of the SSD is greater than or equal to 2.

9. An SSD link state detection system based on link information, characterized in that, It includes an acquisition module, a rate determination module, a first detection module, and a second detection module; The acquisition module is used to acquire the real-time rate value of the target link of the SSD at each time point within the current observation window. The rate determination module is used to determine the transmission rate status of the target link at each time based on each real-time rate value and a preset threshold. The transmission rate status is a normal rate status, a first abnormal rate status, and a second abnormal rate status. The first detection module is used to determine the initial link state detection result of the target link based on the proportion of the types of each of the transmission rate states. The initial link state detection result is a normal state, a first abnormal state, and a second abnormal state. The second detection module is used to determine that the current link status detection result of the target link is abnormal if the initial link status detection result is a second abnormal state; and to determine the current link status detection result of the target link based on the PCIe error detection information of the target link within a preset time period if the initial link status detection result is a first abnormal state.

10. The SSD link state detection system based on link information as described in claim 9, characterized in that, The SSD link status detection system also includes an anomaly handling module; The anomaly handling module is used to determine that if the current link status detection result of any target link of the SSD is abnormal, the target link is an abnormal link and the SSD is an abnormal SSD; within the current observation window, if the number of abnormal SSDs detected and determined in the same storage device exceeds a preset abnormal number threshold, a device-level alarm message is reported; otherwise, anomaly repair is performed on each of the abnormal SSDs.