A storage device monitoring apparatus and a monitoring method thereof

By working together with digital logic devices, management and control components, and self-recovering devices, the write operation latency of storage devices is monitored in real time, the lifespan is accurately predicted, and mirror recovery is performed when the lifespan is nearing its end. This solves the problems of high cost and low accuracy of storage device monitoring in edge servers, and improves the reliability and stability of the devices.

CN120670253BActive Publication Date: 2025-11-11LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511164118.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-11
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

In existing technologies, the lifespan monitoring of storage devices in edge servers relies on programmable logic devices, which are costly and have low accuracy, making it difficult to meet the reliability requirements of edge servers.

Method used

Digital logic devices are used to detect write operation latency, management and control components count the number of out-of-limit events, and self-recovery devices disconnect and switch to management and control mode to perform mirror recovery operations, thereby realizing real-time monitoring and lifespan prediction of storage devices.

Benefits of technology

It reduces hardware and maintenance costs, improves the reliability and stability of storage device status, and avoids data loss caused by sudden failures, making it especially suitable for edge scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670253B_ABST
    Figure CN120670253B_ABST
Patent Text Reader

Abstract

This invention discloses a storage device monitoring apparatus and method, relating to the field of storage technology. The method includes: when the storage device responds to a write operation request sent by a processing device, a digital logic device detects the actual latency value of the write operation to determine whether an over-limit event has occurred; a management control unit counts the number of over-limit events; if the number exceeds a preset threshold, it predicts the remaining lifespan of the storage device; when the remaining lifespan is lower than a set lifespan threshold, it sends a start command to a self-recovery device; the self-recovery device disconnects the processing device from the storage device and switches to the dominant mode of the management control unit, enabling the management control unit to perform a mirror recovery operation. This eliminates the need for programmable devices, reducing costs, and provides highly accurate lifespan prediction to precisely control the storage device's status, provide early warning of potential faults, and ensure data security and device operational continuity through automatic mirror recovery, making it suitable for edge computing scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage technology, and in particular to a storage device monitoring apparatus and its monitoring method. Background Technology

[0002] With the development of the Internet of Things (IoT), edge servers are often deployed in industrial sites, roadside base stations, supermarkets, and other environments and operate for extended periods. During this process, the devices face high-intensity data read and write operations, easily causing storage devices to reach their lifespan limit. Simultaneously, environmental factors such as temperature, humidity, and operating voltage all affect the lifespan of storage devices. Because edge servers are often used in harsh environments, their internal storage devices have a shorter lifespan compared to ordinary servers operating in data centers. Traditional solutions rely on field-programmable gate arrays (FPGAs) or complex programmable logic devices (CLPs) to poll and monitor the extended configuration registers of the storage devices. However, this increases hardware costs and firmware maintenance complexity, and the accuracy is low, prone to misjudgments, and fails to meet the reliability requirements of edge servers. Summary of the Invention

[0003] This invention provides a storage device monitoring apparatus and method, which at least solves the problems of high cost and low accuracy in related technologies that rely on programmable logic devices to monitor the lifespan of storage devices.

[0004] This invention provides a storage device monitoring apparatus, comprising:

[0005] A digital logic device, connected to the storage device, is used to detect the actual delay value of the write operation and determine whether an over-limit event has occurred when the storage device responds to a write operation request sent by the processing device.

[0006] A management and control component, connected to the digital logic device, is used to count the number of times the over-limit event occurs. If the count exceeds a preset threshold, the remaining lifespan of the storage device is predicted. When the predicted remaining lifespan is lower than a set lifespan threshold, a start command is sent to the self-recovery device.

[0007] The self-recovery device is used to disconnect the processing device from the storage device after receiving the startup command, and switch to the dominant mode of the management and control unit so that the management and control unit can perform a mirror recovery operation.

[0008] The present invention also provides a monitoring method for a storage device monitoring apparatus, comprising:

[0009] When the storage device responds to a write operation request sent by the processing device, a digital logic device is used to detect the actual delay value of the write operation to determine whether an over-limit event has occurred.

[0010] The number of occurrences of the over-limit events is counted using the management and control components. If the counted number exceeds a preset threshold, the remaining lifespan of the storage device is predicted. When the predicted remaining lifespan is lower than a set lifespan threshold, a start command is sent to the self-recovery device.

[0011] The self-recovery device disconnects the processing device from the storage device and switches to the dominant mode of the management control unit, so that the management control unit performs a mirror recovery operation.

[0012] The storage device monitoring device provided by this invention, through the coordinated operation of digital logic devices, management and control components, and self-recovering devices, can monitor the latency of write operations on storage devices in real time, promptly detect and statistically analyze out-of-limit events, accurately predict the remaining lifespan of the storage device when the number of out-of-limit events is too high, and when the remaining lifespan is below a threshold, the self-recovering device quickly disconnects the original connection and switches to the dominant mode of the management and control component to perform a mirror recovery operation. This reduces hardware and maintenance costs without the need for programmable devices, while providing precise control over the storage device's status and early warning of potential failures through highly accurate lifespan prediction. When the storage device's lifespan is about to expire, the self-recovering device can automatically initiate a recovery mechanism, ensuring data security and device operational continuity through automatic mirror recovery, comprehensively improving the device's reliability and stability. It is particularly suitable for edge computing scenarios, preventing data loss due to sudden failures.

[0013] In addition, the present invention also provides a corresponding monitoring method for a storage device monitoring device, which has the same or corresponding technical features as the storage device monitoring device mentioned above, and has the same effect. Attached Figure Description

[0014] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram of the structure of the storage device monitoring device provided in an embodiment of the present invention;

[0016] Figure 2 This is a schematic diagram of the signaling interaction corresponding to the storage device monitoring device provided in the embodiments of the present invention;

[0017] Figure 3 This is a schematic diagram of the process corresponding to the self-recovery device provided in the embodiments of the present invention;

[0018] Figure 4A flowchart illustrating the monitoring method of the storage device monitoring apparatus provided in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0020] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0021] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0022] The specific application environment architecture or specific hardware architecture on which the execution of the storage device monitoring device depends is described herein.

[0023] The present invention provides a storage device monitoring apparatus, and the apparatus is described in detail in conjunction with the relevant execution flow of the storage device monitoring apparatus. Figure 1 This is a schematic diagram of the structure of the storage device monitoring device provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the device includes:

[0024] Digital logic device 1, connected to a storage device, is used to detect the actual latency value of a write operation when the storage device responds to a write operation request sent by a processing device, and to determine whether an over-limit event has occurred; the actual latency value of a write operation refers to the actual time interval elapsed from the start of the write operation request to the completion of the write operation during data storage or transmission.

[0025] The management and control unit 2, connected to the digital logic device 1, is used to count the number of over-limit events. If the count exceeds a preset threshold, the remaining lifespan of the storage device is predicted. When the predicted remaining lifespan is lower than the set lifespan threshold, a start command is sent to the self-recovery device.

[0026] The self-recovery device 3 is used to disconnect the processing device from the storage device after receiving the startup command, and switch to the dominant mode of the management control unit 2 so that the management control unit 2 can perform the mirror recovery operation.

[0027] It should be noted that the processing device is the initiator of data interaction with the storage device. After establishing a connection with the storage device, the processing device can send write operation requests to the storage device. This processing device can be a host, acting as the source of the entire data write operation, driving the storage device to execute the corresponding data storage action by sending write operation requests. The storage device can be an embedded MultiMediaCard (eMMC) or other storage devices, which are not limited here.

[0028] As a key component connecting storage devices, digital logic device 1 plays a crucial role in the storage device's response to write operation requests. On one hand, it can accurately detect the actual latency of write operations in real time, capturing the time difference between the processing device sending the request and the storage device completing its response. On the other hand, based on a preset latency threshold, it judges the detected latency value; if the value exceeds the preset threshold, it is determined that an over-limit event has occurred. This enables digital logic device 1 to promptly perceive performance fluctuations in the storage device during write operations, providing direct and critical raw data support for subsequent assessment of the storage device's status and fault warning.

[0029] As the core component connecting the digital logic device 1 and the self-recovery device 4, the management and control unit 2 can count the number of out-of-limit events detected by the digital logic device 1. When the count exceeds a preset threshold (e.g., 1000 times), it will initiate the remaining lifespan prediction process of the storage device to determine the aging degree and survival status of the device. When the predicted remaining lifespan is lower than the set lifespan threshold, it means that the storage device is nearing the end of its service life and the risk has increased significantly. At this time, the management and control unit 2 can issue a start command to the self-recovery device 3, thereby promoting the activation of subsequent fault response and device protection mechanisms. This process not only realizes the dynamic control of the storage device status, but also lays the foundation for timely intervention in device management and prevention of fault expansion.

[0030] As a key safety component in the device, the self-recovery device 3 can fully function after receiving the start command from the management control unit 2: it disconnects the connection between the processing device and the storage device to prevent data interaction from escalating risks when the storage device is in poor condition; it simultaneously switches to the dominant mode of the management control unit 2, handing over control of the device to the management control unit 2 so that it can perform mirror recovery operations. This process, by promptly isolating the problematic device and transferring control, provides a stable environment for secure data recovery and is a crucial step in ensuring the device's self-repair when the storage device is at risk.

[0031] In the storage device monitoring device provided in this embodiment of the invention, through the coordinated operation of the digital logic device 1, the management and control unit 2, and the self-recovery device 3, the latency of write operations on the storage device can be monitored in real time, and out-of-limit events can be detected and counted in a timely manner. When the number of out-of-limit events is too high, the remaining lifespan of the storage device can be accurately predicted. When the remaining lifespan is lower than the threshold, the self-recovery device 3 can quickly disconnect the original connection and switch to the management and control unit 2 in a dominant mode to perform a mirror recovery operation. In this way, without the need to use programmable devices, hardware and maintenance costs are reduced, and the storage device status can be accurately controlled with high-precision lifespan prediction, providing early warning of potential failures. When the storage device is about to expire, the self-recovery device 3 can automatically start the recovery mechanism to ensure data security and device operation continuity through automatic mirror recovery, comprehensively improving the reliability and stability of the device. It is especially suitable for edge scenarios and can avoid data loss caused by sudden failures.

[0032] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the digital logic device 1 may include a trigger, a timer, and a delay comparator; wherein, the trigger is used to detect the start signal of a write operation request; the timer is used to start timing when the trigger detects the start signal; and to stop timing after receiving a response signal indicating that the write operation is completed, so as to record the actual delay value of the write operation; the delay comparator is used to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; and to determine whether an over-limit event has occurred based on the comparison and analysis results.

[0033] It should be noted that this invention indirectly assesses the lifetime of storage devices (such as eMMC) by utilizing the write operation latency. As the lifetime of storage devices increases, bad blocks accumulate, and remapping operations lead to increased write operation latency. Therefore, this invention can measure the completion time of each write operation using digital logic device 1. If the write operation time exceeds a preset threshold, it may indicate that a remapping (or other error handling) has occurred. A suspicious remapping event is then recorded and accumulated.

[0034] In implementation, digital logic device 1 may include flip-flops, counters, and delay comparators. Flip-flops can also be called write command flip-flops. The timer can be driven by a quartz crystal oscillator with an accuracy of approximately 10 ppm. When the processing device initiates a write operation request to the storage device, the write command flip-flop detects the start of the write command and simultaneously triggers the timer to start counting, stopping when a response indicating the write operation is complete is received. Then, delay comparison and analysis are performed. The delay comparator compares the measured actual delay value with a preset threshold. This threshold has two levels: a non-warning threshold and a warning threshold. The non-warning threshold is the maximum delay based on the storage device's (e.g., eMMC) specifications; that is, the non-warning threshold is the theoretically acceptable maximum delay upper limit for the storage device when performing read and write operations under design and manufacturing standards, representing the delay boundary for stable device operation. The warning threshold can be set to 150% of the non-warning threshold. If the delay exceeds the warning threshold, an over-limit event is determined to have occurred. The digital logic device 1 of the present invention, through the coordinated operation of flip-flops, timers, and delay comparators, can accurately capture the start and completion signals of write operation requests, thereby accurately recording the actual delay value and comparing it with a preset warning threshold to determine whether an over-limit event has occurred. This not only achieves real-time and accurate monitoring of write operation delays of storage devices, providing a reliable basis for subsequent lifetime prediction and fault warning, but also eliminates the need for complex components. While ensuring detection accuracy, it helps to simplify the hardware structure, further reduce costs, and improve the adaptability and practicality of the entire device in applications such as edge computing scenarios.

[0035] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the trigger can be specifically used to capture the start level or protocol frame header of the write operation request through the command line of the storage device; the timer can be specifically used to start timing at the same time as the trigger captures the start level or protocol frame header.

[0036] In implementation, the trigger can accurately capture the start level of the write operation request or the protocol frame header through the command line (CMD) of the storage device, providing a clear and precise timing start point for the timer, i.e. the start time of the write operation. This ensures that the timer starts timing from the moment the write operation starts, effectively avoiding timing deviations and thus improving the accuracy and reliability of the entire delay detection process.

[0037] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the timer can also be used to stop timing after receiving the response code corresponding to the write operation completion status returned by the command line of the storage device, or after receiving the response frame indicating that the write operation is complete.

[0038] In implementation, the timer can stop counting when it receives a write operation completion status response code or response frame from the storage device command line. This allows for precise capture of the end time of the write operation, forming a complete time loop with the start time. This ensures the integrity and accuracy of timing the actual latency value of the write operation, avoids latency calculation deviations caused by inaccurate judgment of the end time, further improves the accuracy of write operation latency detection, and provides more reliable time data support for subsequent over-limit event judgment and lifetime prediction.

[0039] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management control component 2 may include a bad block counter and a management controller; wherein, the bad block counter is used to count a corresponding number of times when an over-limit event is determined to occur, so as to accumulate and record the over-limit events that occur during the operation of the storage device, and send a lifetime detection signal to the management controller; the management controller is used to read the count value of the bad block counter, and analyze the bad block growth trend based on the read count value to predict the remaining lifetime of the storage device.

[0040] In implementation, the management control component 2 may include a bad block counter and a management controller. The bad block counter can accurately count when an over-limit event occurs, accumulate and record the over-limit situation during the operation of the storage device and issue a life detection signal, providing accurate basic data for subsequent analysis. The management controller can analyze the bad block growth trend by reading the count value, thereby predicting the remaining life of the storage device. This trend analysis method based on actual operating data makes the life prediction more in line with the actual state of the device, further improving the accuracy and reliability of the prediction.

[0041] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the count value of the bad block counter can be stored using a non-volatile storage medium; when the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to accumulate from the current value.

[0042] In implementation, each time an out-of-limit event occurs, the bad block counter can increment by a set value (e.g., 1). The bad block counter can use non-volatile storage media to store the count value, ensuring that the count value remains unchanged when the storage device is powered off and can continue to accumulate from the current value after power is restored. Electrically Erasable Programmable Read-Only Memory (EEPROM) can be used as the non-volatile storage media, and there is no specific limitation. This avoids the loss of out-of-limit event counts due to power failure, ensuring the continuity and integrity of the count. This provides a consistent and reliable data foundation for the management controller to analyze bad block growth trends and predict the remaining lifespan of the storage device, further improving the accuracy of lifespan prediction. It also reduces the additional verification or recounting work caused by data interruption, lowering the additional operating costs of the device.

[0043] Figure 2 This is a schematic diagram of the signaling interaction corresponding to the storage device monitoring device provided in an embodiment of the present invention. Figure 2 As shown, a processing device (such as a host) can send a write operation request to a storage device, triggering the data writing process of the storage device (such as an eMMC). When the storage device responds to the write operation request, it feeds back response latency information to digital logic device 1 to detect latency and monitor the response performance of the storage device. Digital logic device 1 can collect latency data and, together with bad block count information, report it to the management controller. The management controller can be a baseboard management controller (BMC). The management controller can perform trend analysis on latency data and the number of bad blocks, and enter different branches based on the analysis results: when the remaining lifetime is not lower than a set lifetime threshold (such as 5%), it returns a normal response to the processing device, and the write operation is completed according to the normal process without triggering a recovery action; when the remaining lifetime is lower than the set lifetime threshold (such as 5%), it determines that the storage device's lifetime is nearing its limit (reaching the recovery trigger condition), triggers the recovery process, and sends a start command to self-recovery device 3. Self-recovery device 3 can perform isolation connection to avoid data conflicts and initiate a read mirror operation request to read a backup of critical data from the storage device. Subsequent operations such as fast erase and image writing, as well as device reset, can be performed to reinitialize critical components and ensure that the device can be re-interacted in a repaired state.

[0044] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management controller can be specifically used to read the count value of the bad block counter according to a preset period, and calculate the number of bad blocks increasing per unit time based on the read count value; use an exponentially weighted moving average algorithm to predict the number of bad blocks increasing, and obtain the bad block growth rate within a future set time period; and calculate the remaining lifespan of the storage device based on the bad block growth rate combined with the current number of bad blocks and the maximum allowable number of bad blocks.

[0045] In implementation, the management controller can read the bad block counter value at preset intervals (e.g., hourly) to calculate the number of bad blocks increasing per unit time (bad blocks / hour). Then, it uses an Exponentially Weighted Moving Average (EWMA) algorithm to predict the number of bad blocks increasing, obtaining the bad block growth rate for a future set time period (e.g., the next 24 hours). Finally, it combines the current number of bad blocks with the maximum allowable number of bad blocks to calculate the remaining lifespan of the storage device. The maximum allowable number of bad blocks can be determined by the storage device model (generally 5% of the total number of chips). This method makes the analysis of bad block growth trends more timely and accurate, and bases the calculation of remaining lifespan on reliable dynamic data, significantly improving the accuracy of lifespan prediction and reflecting the true state of the storage device more promptly and accurately.

[0046] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management controller can be specifically used to obtain the current number of bad blocks from the bad block counter; subtract the sum of the current number of bad blocks and the bad block growth rate from 1, divide the difference obtained by the maximum allowable number of bad blocks, and then multiply the obtained value by 100 to calculate the remaining lifespan of the storage device.

[0047] In implementation, the management controller obtains the current number of bad blocks from the bad block counter and then calculates the remaining lifetime of the storage device. The calculation formula is: Remaining lifetime (%) = 100 * (1 - (Current number of bad blocks + Predicted growth) / Maximum allowable number of bad blocks). This calculation method directly relates to the current bad block status, growth trend, and the device's tolerance limit. It is logically clear and computationally simple, quickly yielding an intuitive remaining lifetime value. This ensures a close match between the calculation results and the actual wear and tear of the device, improves the efficiency of lifetime assessment, and provides a clear and reliable quantitative basis for subsequent decisions on whether to activate the self-recovery mechanism.

[0048] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management controller can also be used to compare the predicted remaining lifespan with a set lifespan threshold after predicting the remaining lifespan of the storage device; if the predicted remaining lifespan is lower than the set lifespan threshold, the activation conditions of the storage device self-recovery mechanism are triggered, and a startup command is sent to the self-recovery device 3.

[0049] In implementation, the management controller can accurately determine whether the device is close to the critical failure state by comparing the predicted remaining lifespan of the storage device with a set lifespan threshold. When the remaining lifespan is lower than the set lifespan threshold, the self-recovery mechanism is immediately triggered and a start command is sent to the self-recovery device 3. This process realizes dynamic monitoring and timely response of the device status, ensuring that the protection mechanism is proactively activated before the storage device's lifespan expires, avoiding data loss due to sudden device failure. At the same time, early intervention reduces the cost of fault diagnosis and on-site maintenance, effectively ensuring the continuous and stable operation of the storage system in edge scenarios.

[0050] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management controller can also be used to trigger an alarm through the web management interface and drive the alarm light to illuminate while sending a start command to the self-recovery device 3, so as to prompt the administrator to back up relevant files.

[0051] During implementation, the management controller can trigger alarms and illuminate alarm lights through the web management interface while sending self-recovery startup commands. This timely alerts administrators of impending storage device failures, reminding them to back up relevant files as soon as possible. This enhances the timeliness of manual intervention, reducing the risk of data loss and ensuring that administrators do not miss critical information through dual alarm mechanisms. This effectively combines automatic protection with manual operation, minimizing potential file loss due to device problems and improving the security and controllability of the storage system in emergency scenarios.

[0052] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the self-recovery device 3 may include an isolating switch. The isolating switch may be a high-speed analog switch (such as PI3A3157), and the switching time may be less than 10 ns. The management controller may be used to send a self-recovery control signal to the isolating switch simultaneously with sending a start command to the self-recovery device 3. The isolating switch may be used to disconnect the connection between the processing device and the storage device after receiving the self-recovery control signal from the management controller, and switch to the dominant mode of the management controller. The management controller may also be used to temporarily store the image data in a hardware buffer, and write it to the storage device after processing by the hardware buffer.

[0053] In implementation, after receiving the self-recovery control signal sent by the management controller, the isolating switch in the self-recovery device 3 can quickly disconnect the connection between the processing device and the storage device and switch to the management controller-dominated mode. At the same time, the management controller temporarily stores the mirror data in the hardware buffer and writes it to the storage device after processing. This not only avoids the additional pressure on the faulty device caused by the continuous interaction of the processing device through isolation operation, but also ensures the stability and security of the mirror data writing with the help of the hardware buffer. With the dominant control of the management controller, it not only ensures the orderly and efficient self-recovery process, but also provides a reliable channel for data migration while automatically recovering, further reducing the risk of data loss, while reducing the complexity of manual intervention and improving the convenience and reliability of the device's emergency handling.

[0054] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the isolating switch can be used to disconnect the command line and data line connection between the processing device and the storage device after receiving the self-recovery control signal from the management controller, and synchronously switch to the corresponding control line between the management controller and the storage device to establish a management channel connection.

[0055] In implementation, upon receiving the self-recovery control signal from the management controller, the isolating switch disconnects the command and data lines between the processing unit and the storage device, completely severing data interaction between them and preventing the faulty device from continuously receiving operational commands and exacerbating the problem. Simultaneously, it synchronously switches to the corresponding control lines of the management controller and the storage device, establishing a management channel connection. This ensures the management controller's independent control over the storage device, providing a dedicated and stable communication path for subsequent image recovery operations. This achieves secure isolation between the faulty device and the original processing unit, establishing a reliable channel for the recovery process led by the management controller, thereby effectively guaranteeing the independence and security of the self-recovery process.

[0056] Furthermore, in a specific implementation, in the storage device monitoring apparatus provided in the embodiments of the present invention, the self-recovery device 3 may include a Serial Peripheral Interface Flash (SPI FLASH); the SPI FLASH is used to store backup image data. The management controller is specifically used to read the pre-stored backup image data from the SPI FLASH, temporarily store the read backup image data in a page buffer, and perform image writing of the storage device through format conversion processing of the page buffer.

[0057] In practice, the serial peripheral interface flash memory in the self-recovery device 3 can stably store backup image data. Figure 3 This is a schematic diagram of the process corresponding to the self-recovery device provided in an embodiment of the present invention. Figure 3As shown, after the management controller reads the backup image data from the serial peripheral interface flash memory, it first temporarily stores it in the page buffer for format conversion processing, and then performs the image write operation on the storage device. The page buffer is a hardware buffer that temporarily stores data after it is read from the serial peripheral interface flash memory and before it is written to the storage device. It is responsible for buffering and format conversion of high-speed data. In this way, the reliable storage characteristics of the serial peripheral interface flash memory ensure the security and accessibility of the backup data, and the format conversion of the page buffer adapts to the requirements of the storage device, ensuring the compatibility and accuracy of the image write process. The entire process realizes a coherent and efficient operation from backup data reading, format adaptation to write recovery, further improving the reliability of the self-recovery mechanism. It can quickly complete data recovery when the storage device is nearing the end of its lifespan, reduce manual intervention costs, and effectively ensure the integrity of data and the continuous and stable operation of the device in edge scenarios.

[0058] Furthermore, in a specific implementation, in the storage device monitoring device provided in the embodiments of the present invention, the management controller can also be used to trigger the reset circuit to perform a reset operation on the server motherboard after the storage device image is written, so that the storage device is reinitialized and the connection with the processing device is restored.

[0059] In implementation, after the storage device image is written, the management controller can trigger a reset circuit to perform a reset operation on the server motherboard, causing the storage device to reinitialize and restore its connection with the processing unit, thus completing the closed-loop self-recovery process. The reset operation ensures the storage device is put back into use in a stable state, restoring normal data interaction with the processing unit. This guarantees rapid recovery of the device after image restoration, avoiding operational failures caused by connection anomalies or improper initialization. Furthermore, it eliminates the need for manual restarts, reducing on-site maintenance intervention costs and further improving the continuity and stability of the edge server after storage device recovery, ensuring that services can quickly return to normal operation.

[0060] Through the above description of the embodiments, those skilled in the art can clearly understand that the device according to the above embodiments can be implemented in hardware. From the implementation perspective, the core components of the device, such as digital logic devices (including flip-flops, timers, and delay comparators), management and control units (integrating bad block counters and management controllers), and self-recovering devices (including isolating switches and serial peripheral interface flash memory), all adopt hardware circuit design, rather than relying on the loading and running of software programs. This pure hardware architecture avoids compatibility issues, code vulnerability risks, and processor resource consumption that software operation may face, and can execute various operations with faster response speeds. For example, the detection of write operation latency by digital logic devices can be completed within nanoseconds, and the counting and judgment of over-limit events by the management and control unit does not require a complex instruction parsing process, significantly improving the real-time performance of the device. At the same time, the hardware implementation is more stable and unaffected by fluctuations in the operating system or operating environment, especially in complex and harsh environments such as edge computing scenarios, maintaining a continuously reliable working state. From a cost and maintenance perspective, the hardware circuitry eliminates the need for expensive programmable chips, reducing hardware costs through streamlined circuit design. Furthermore, the fixed hardware functionality minimizes the maintenance burden caused by software upgrades and vulnerability fixes, further aligning with the edge computing scenario's demand for low-cost, highly stable equipment. In addition, hardware-level collaborative operations (such as rapid switching of disconnect switches and format conversion of page buffers) ensure precise synchronization of operations across all stages, providing a solid physical foundation for core functions such as highly accurate lifetime prediction and reliable self-recovery mechanisms. This makes the device more stable and controllable in terms of data security and fault prevention.

[0061] Based on the same inventive concept, embodiments of the present invention also provide a monitoring method for a storage device monitoring apparatus. Figure 4 This is a flowchart of the monitoring method of the storage device monitoring apparatus provided in an embodiment of the present invention, as shown below. Figure 4 As shown, the method includes:

[0062] S401. When the storage device responds to a write operation request sent by the processing device, a digital logic device is used to detect the actual delay value of the write operation to determine whether an over-limit event has occurred.

[0063] S402. Use the management control component to count the number of over-limit events. If the count exceeds a preset threshold, predict the remaining lifespan of the storage device. When the predicted remaining lifespan is lower than the set lifespan threshold, send a start command to the self-recovery device.

[0064] S403. The connection between the processing device and the storage device is disconnected by a self-recovering device, and the system switches to the dominant mode of the management control unit so that the management control unit can perform a mirror recovery operation.

[0065] In the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, the latency of write operations on the storage device can be monitored in real time, and over-limit events can be detected and counted in a timely manner. When the number of over-limit events is too high, the remaining lifespan of the storage device can be accurately predicted. When the remaining lifespan is lower than the threshold, the original connection can be quickly disconnected by a self-recovering device and the system can switch to the management and control unit-dominated mode to perform a mirror recovery operation. In this way, without the need to use programmable devices, hardware and maintenance costs are reduced, and the storage device status can be accurately controlled by highly accurate lifespan prediction, providing early warning of potential failures. When the lifespan of the storage device is about to expire, the self-recovering device can automatically start the recovery mechanism to ensure data security and device operation continuity through automatic mirror recovery, comprehensively improving the reliability and stability of the device. It is especially suitable for edge scenarios and can avoid data loss caused by sudden failures.

[0066] Since the embodiments of the monitoring method portion of the storage device monitoring device correspond to each other, the descriptions of the features in the embodiment corresponding to the monitoring method of the storage device monitoring device can be found in the relevant descriptions of the embodiment corresponding to the storage device monitoring device, and will not be repeated here. Furthermore, it has the same beneficial effects as the storage device monitoring device mentioned above.

[0067] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, step S401 uses digital logic devices to detect the actual delay value of the write operation and determine whether an over-limit event has occurred. Specifically, it may include: using a trigger to detect the start signal of the write operation request; when the trigger detects the start signal, starting a timer; stopping the timer after receiving a response signal indicating that the write operation is completed, to record the actual delay value of the write operation; using a delay comparator to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; and determining whether an over-limit event has occurred based on the comparison and analysis results.

[0068] Furthermore, in specific implementation, in the above steps, a trigger is used to detect the start signal of the write operation request; when the trigger detects the start signal, a timer starts timing; after receiving the response signal indicating that the write operation is complete, the timing stops. Specifically, this may include: using a trigger to capture the start level or protocol frame header of the write operation request through the command line of the storage device; simultaneously with the trigger capturing the start level or protocol frame header, a timer starts timing; the starting point of the timing is the start time of the write operation; after receiving the response code corresponding to the write operation completion status returned by the command line of the storage device, or after receiving the response frame indicating that the write operation is complete, the timing stops.

[0069] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, step S402 utilizes the management control component to count the number of times an over-limit event occurs. Specifically, this may include: when an over-limit event is determined to have occurred, using a bad block counter to count the corresponding number of times, so as to cumulatively record the over-limit events occurring during the operation of the storage device, and sending a life detection signal to the management controller. In implementation, the count value of the bad block counter is stored using a non-volatile storage medium; when the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to accumulate from the current value.

[0070] If the number of counts exceeds a preset threshold, step S402 predicts the remaining lifespan of the storage device. Specifically, this may include: using a management controller to read the bad block counter's count value, analyzing the bad block growth trend based on the read count value, and predicting the remaining lifespan of the storage device. In implementation, the management controller reads the bad block counter's count value at a preset period, and calculates the number of bad blocks increasing per unit time based on the read count value; an exponentially weighted moving average algorithm is used to predict the number of bad blocks increasing, obtaining the bad block growth rate within a future set time period; and the remaining lifespan of the storage device is calculated based on the bad block growth rate combined with the current number of bad blocks and the maximum allowed number of bad blocks.

[0071] The remaining lifespan of the storage device is calculated based on the bad block growth rate, the current number of bad blocks, and the maximum allowed number of bad blocks. Specifically, this can include: obtaining the current number of bad blocks from the bad block counter; subtracting the sum of the current number of bad blocks and the bad block growth rate from 1, dividing the difference by the maximum allowed number of bad blocks, and then multiplying the result by 100 to calculate the remaining lifespan of the storage device.

[0072] Step S402 When the predicted remaining lifetime is lower than the set lifetime threshold, a start command is sent to the self-recovery device. Specifically, this may include: after predicting the remaining lifetime of the storage device, comparing the predicted remaining lifetime with the set lifetime threshold; if the predicted remaining lifetime is lower than the set lifetime threshold, triggering the start condition of the storage device's self-recovery mechanism, and sending a start command to the self-recovery device.

[0073] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the above embodiments of the present invention, after executing step S402, it may further include: while sending a start command to the self-recovery device, triggering an alarm through a web page management interface using the management controller, and driving the alarm light to illuminate, so as to prompt the administrator to back up relevant files.

[0074] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, step S403 uses a self-recovering device to disconnect the processing device from the storage device and switches to the dominant mode of the management control unit so that the management control unit performs a mirror recovery operation. Specifically, this may include: while sending a start command to the self-recovering device, using the management controller to send a self-recovery control signal to the isolation switch of the self-recovering device; after receiving the self-recovery control signal from the management controller, the isolation switch disconnects the processing device from the storage device and switches to the dominant mode of the management controller; the management controller temporarily stores the mirror data in a hardware buffer, and after processing by the hardware buffer, writes it to the storage device.

[0075] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, after receiving the self-recovery control signal from the management controller, the connection between the processing device and the storage device is cut off, and the mode of the management controller is switched to the dominant mode. Specifically, this may include: after receiving the self-recovery control signal from the management controller, disconnecting the command line and data line connection between the processing device and the storage device, synchronously switching to the corresponding control line between the management controller and the storage device, and establishing a management channel connection.

[0076] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring device provided in the embodiments of the present invention, the image data is temporarily stored in a hardware buffer and then written to the storage device after being processed by the hardware buffer. Specifically, this may include: reading pre-stored backup image data from the serial peripheral interface flash memory of the self-recovering device, temporarily storing the read backup image data in a page buffer, and performing image writing to the storage device through format conversion processing of the page buffer.

[0077] Furthermore, in a specific implementation, the monitoring method of the storage device monitoring device provided in the embodiments of the present invention may further include: after the storage device image is written, using the management controller to trigger the reset circuit to perform a reset operation on the server motherboard, so that the storage device is reinitialized and the connection with the processing device is restored.

[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on these embodiments without creative effort are within the scope of protection of this invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art can still combine, add to, delete from, or otherwise adjust the features of the various embodiments of the present invention as appropriate without conflict or creative effort, thereby obtaining different technical solutions that do not fundamentally depart from the concept of the present invention. These technical solutions also fall within the scope of protection of this invention.

[0079] The storage device monitoring apparatus and monitoring method provided by the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention, and are not intended to limit the scope of protection of the invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A storage device monitoring device, characterized in that, include: A digital logic device, connected to a storage device, is used to detect the actual latency value of a write operation when the storage device responds to a write operation request sent by a processing device, and to determine whether an over-limit event has occurred. A management and control unit, connected to the digital logic device, is used to count the number of times the over-limit event occurs. If the counted number exceeds a preset threshold, the remaining lifespan of the storage device is predicted. When the predicted remaining lifespan is lower than a set lifespan threshold, a start command is sent to the self-recovery device. The management and control unit includes a bad block counter and a management controller. The bad block counter is used to count a corresponding number of times when an over-limit event is detected, so as to accumulate and record the over-limit events that occur during the operation of the storage device, and send a lifetime detection signal to the management controller; the management controller is used to read the count value of the bad block counter according to a preset period, and calculate the number of bad blocks increasing per unit time based on the read count value; use an exponential weighted moving average algorithm to predict the number of bad blocks increasing, and obtain the bad block growth rate in a future set time period; obtain the current number of bad blocks from the bad block counter; calculate the remaining lifetime of the storage device based on the bad block growth rate combined with the current number of bad blocks and the maximum allowable number of bad blocks; the formula for calculating the remaining lifetime of the storage device is: remaining lifetime (%) = 100 * (1 - (current number of bad blocks + predicted growth) / maximum allowable number of bad blocks); The self-recovering device is used to disconnect the processing device from the storage device after receiving the boot command, and switch to the dominant mode of the management and control unit, so that the management and control unit can perform a mirror recovery operation; the self-recovering device includes an isolation switch and a serial peripheral interface flash memory; The management controller is used to send a self-recovery control signal to the disconnecting switch while simultaneously sending a start command to the self-recovery device. The isolating switch is used to disconnect the connection between the processing device and the storage device after receiving the self-recovery control signal from the management controller, and switch to the dominant mode of the management controller; the management controller is also used to temporarily store the image data in a hardware buffer, and write it to the storage device after processing by the hardware buffer. The serial peripheral interface flash memory is used to store backup image data; the management controller is also used to read the pre-stored backup image data from the serial peripheral interface flash memory, temporarily store the read backup image data in a page buffer, and perform image writing of the storage device through the format conversion processing of the page buffer.

2. The storage device monitoring device according to claim 1, characterized in that, The digital logic device includes flip-flops, timers, and delay comparators; The trigger is used to detect the start signal of the write operation request; The timer is configured to start timing when the trigger detects the start signal, and stop timing after receiving a write operation completion response signal to record the actual delay value of the write operation. The delay comparator is used to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; based on the comparison and analysis results, it is determined whether an over-limit event has occurred.

3. The storage device monitoring device according to claim 2, characterized in that, The trigger is used to capture the start level or protocol frame header of the write operation request via the command line of the storage device; The timer is used to start timing when the trigger captures the start level or protocol frame header.

4. The storage device monitoring device according to claim 3, characterized in that, The timer is further configured to stop timing after receiving a response code corresponding to the write operation completion status returned by the command line of the storage device, or after receiving a write operation completion response frame.

5. The storage device monitoring device according to claim 1, characterized in that, The count value of the bad block counter is stored using a non-volatile storage medium; When the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to increment from the current value.

6. The storage device monitoring device according to claim 1, characterized in that, The management controller is further configured to compare the predicted remaining lifespan of the storage device with a set lifespan threshold after predicting the remaining lifespan of the storage device; if the predicted remaining lifespan is lower than the set lifespan threshold, the controller triggers the start condition of the storage device self-recovery mechanism and sends a start command to the self-recovery device.

7. The storage device monitoring device according to claim 1, characterized in that, The management controller is also used to trigger an alarm through the web management interface and drive the alarm light to illuminate while sending a start command to the self-recovering device, so as to prompt the administrator to back up relevant files.

8. The storage device monitoring device according to claim 1, characterized in that, The isolating switch is used to disconnect the command line and data line connection between the processing device and the storage device after receiving the self-recovery control signal from the management controller, and synchronously switch to the corresponding control line between the management controller and the storage device to establish a management channel connection.

9. The storage device monitoring device according to claim 1, characterized in that, The management controller is also configured to trigger a reset circuit to perform a reset operation on the server motherboard after the storage device image is written, so that the storage device is reinitialized and the connection with the processing device is restored.

10. A monitoring method for a storage device monitoring apparatus as described in any one of claims 1 to 9, characterized in that, include: When the storage device responds to a write operation request sent by the processing device, a digital logic device is used to detect the actual delay value of the write operation to determine whether an over-limit event has occurred. The management and control unit counts the number of occurrences of the out-of-limit events. If the count exceeds a preset threshold, the remaining lifespan of the storage device is predicted. When the predicted remaining lifespan is lower than a set lifespan threshold, a startup command is sent to the self-recovery device. The management and control unit includes a bad block counter and a management controller. When an over-limit event is detected, the bad block counter is used to count the corresponding number of times to accumulate and record the over-limit events that occur during the operation of the storage device, and a lifetime detection signal is sent to the management controller. The management controller reads the bad block counter's count value at a preset period and calculates the bad block growth rate per unit time based on the read count value. An exponentially weighted moving average algorithm is used to predict the bad block growth rate over a future set time period. The current bad block count is obtained from the bad block counter. The remaining lifespan of the storage device is calculated based on the bad block growth rate, the current bad block count, and the maximum allowed bad block count. The formula for calculating the remaining lifespan of the storage device is: Remaining lifespan (%) = 100 * (1 - (Current bad block count + Predicted growth) / Maximum allowed bad block count). The self-recovery device disconnects the processing device from the storage device and switches to the dominant mode of the management control unit, so that the management control unit performs a mirror recovery operation; the self-recovery device includes an isolation switch and a serial peripheral interface flash memory; While sending a start command to the self-recovering device, the management controller sends a self-recovery control signal to the isolating switch; after receiving the self-recovery control signal from the management controller, the isolating switch disconnects the processing device from the storage device and switches to the dominant mode of the management controller. The management controller temporarily stores the image data in a hardware buffer, processes it in the hardware buffer, and then writes it into the storage device. The management controller reads the pre-stored backup image data from the serial peripheral interface flash memory, temporarily stores the read backup image data in a page buffer, and performs image writing to the storage device through format conversion processing of the page buffer.

Citation Information

Patent Citations

  • Cloud mobile terminal cooperative fault early warning method, related device and system

    CN109634820A

  • Storage space management method and device, storage medium and electronic equipment

    CN119322587A