Storage equipment monitoring device and monitoring method thereof

By detecting the storage device write operation delay value through digital logic devices, managing and controlling the statistics of over-limit events, and combining with self-recovery devices to perform mirror recovery, the problems of high cost and low accuracy of storage device life monitoring in the existing technology are solved, and efficient and reliable storage device status monitoring and fault warning are achieved.

CN120670253AActive Publication Date: 2025-09-19LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511164118.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-09-19
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Existing technologies rely on programmable logic devices when monitoring the life of storage devices, resulting in high costs and low accuracy, making it difficult to meet the reliability requirements of edge servers.

Method used

Digital logic devices are used to detect the actual delay value of storage device write operations, and the management and control components count the number of over-limit events. Self-recovery devices are used to cut off the connection when the device life is about to expire and perform mirror recovery operations.

Benefits of technology

This reduces hardware and operation and maintenance costs without the need for programmable devices, while improving the accuracy and reliability of storage device life prediction, ensuring data security and continuous and stable operation of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670253A_ABST
    Figure CN120670253A_ABST
Patent Text Reader

Abstract

The invention discloses a storage equipment monitoring device and a monitoring method thereof, and relates to the technical field of storage, and the method comprises the following steps: when storage equipment responds to a write operation request sent by a processing device, a digital logic device detects an actual delay value of write operation, and judges whether an over-limit event occurs or not; the management control part counts the occurrence frequency of the overrun event, and if the frequency exceeds a preset frequency threshold value, the remaining life of the storage device is predicted; when the residual life is lower than a set life threshold value, a starting instruction is sent to the self-recovery device; and the self-recovery device cuts off the connection between the processing device and the storage equipment, and switches to the dominant mode of the management control component, so that the management control component executes the mirror image recovery operation. In this way, a programmable device does not need to be used, cost is reduced, potential faults can be warned in advance by means of high-precision service life prediction and accurate control over the state of the storage device, data safety and device operation continuity can be guaranteed through automatic mirror image recovery, and the method is suitable for edge scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of storage technology, and in particular to a storage device monitoring device and a monitoring method thereof. Background Art

[0002] With the development of the Internet of Things, edge servers are often deployed in environments such as industrial sites, roadside base stations, and supermarkets, and operate for extended periods. During this process, the devices face high-intensity data read and write operations, which can easily cause storage devices to reach their service life limit. Furthermore, factors such as ambient temperature, humidity, and operating voltage can affect the lifespan of storage devices. Because edge servers are often used in harsh environments, their internal storage devices have a shorter service life than standard servers operating in computer rooms. Traditional solutions rely on field programmable gate arrays (FPGAs) or complex programmable logic devices (CPLDs) to poll and monitor the storage device's extended configuration registers. However, this increases hardware costs and firmware maintenance complexity, and its accuracy is low, prone to misjudgments, making it difficult to meet the reliability requirements of edge servers. Summary of the Invention

[0003] The present invention provides a storage device monitoring apparatus and a monitoring method thereof, so as to at least solve the problem in the related art of relying on programmable logic devices to monitor the life of storage devices, which is high in cost and low in accuracy.

[0004] The present invention provides a storage device monitoring device, comprising: a digital logic device connected to the storage device, and configured to detect an actual delay value of a write operation when the storage device responds to a write operation request sent by the processing device, and determine whether an over-limit event occurs; a management control component connected to the digital logic device and configured to count the number of occurrences of the over-limit event, and predict the remaining life of the storage device if the counted number exceeds a preset threshold; and send a startup instruction to the self-recovery device when the predicted remaining life is lower than a set threshold; The self-recovery device is used to cut off the connection between the processing device and the storage device after receiving the startup instruction, and switch to the dominant mode of the management and control component to enable the management and control component to perform the mirror recovery operation.

[0005] The present invention also provides a monitoring method for a storage device monitoring apparatus, comprising: When the storage device responds to the write operation request sent by the processing device, the digital logic device is used to detect the actual delay value of the write operation to determine whether an over-limit event occurs; Utilizing a management control component to count the number of occurrences of the over-limit event, and if the counted number exceeds a preset threshold, predicting the remaining life of the storage device; and sending a startup instruction to the self-recovery device when the predicted remaining life is lower than a set life threshold; The self-recovery device is used to cut off the connection between the processing apparatus and the storage device, and switch to the dominant mode of the management and control component, so that the management and control component performs a mirror recovery operation.

[0006] The storage device monitoring device provided by the present invention can monitor the delay of storage device write operations in real time through the coordinated operation of digital logic devices, management and control components, and self-recovery components, timely detect and count over-limit events, accurately predict the remaining life of the storage device when the number of over-limit events is too many, and when the remaining life is lower than the threshold, quickly cut off the original connection through the self-recovery component and switch to the dominant mode of the management and control component to perform the mirror recovery operation. In this way, without the need to use programmable devices, it not only reduces hardware costs and operation and maintenance costs, but also can accurately control the status of storage devices and provide early warning of potential failures by relying on highly accurate life predictions; when the storage device is about to expire, the self-recovery component can automatically start the recovery mechanism, and ensure data security and device operation continuity through automatic mirror recovery, which comprehensively improves the reliability and stability of the device. It is particularly suitable for edge scenarios and can avoid data loss caused by sudden failures.

[0007] In addition, the present invention also provides a corresponding monitoring method for a storage device monitoring apparatus, which has the same or corresponding technical features as the above-mentioned storage device monitoring apparatus and has the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0009] Figure 1 A schematic diagram of the structure of a storage device monitoring device provided by an embodiment of the present invention; Figure 2 A schematic diagram of signaling interactions corresponding to the storage device monitoring apparatus provided in an embodiment of the present invention; Figure 3 A schematic diagram of a process flow corresponding to a self-recovering device provided in an embodiment of the present invention; Figure 4 This is a flow chart of a monitoring method for a storage device monitoring apparatus provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0011] It should be noted that, in the description of the present invention, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. The terms "first," "second," etc., in the present invention are used to distinguish similar objects, and are not used to describe a particular order or precedence.

[0012] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0013] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the storage device monitoring apparatus depends, the specific application environment architecture or specific hardware architecture is described herein.

[0014] An embodiment of the present invention provides a storage device monitoring apparatus. The apparatus is described in detail in conjunction with relevant execution processes of the storage device monitoring apparatus. Figure 1 A schematic diagram of the structure of a storage device monitoring device provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the device includes: A digital logic device 1 is connected to the storage device and is configured to detect an actual delay value of a write operation when the storage device responds to a write operation request sent by the processing device, and determine whether an over-limit event has occurred; the actual delay value of the write operation refers to the actual time interval from the start of the write operation request to the completion of the write operation during the data storage or transmission process; The management control component 2 is connected to the digital logic device 1 and is used to count the number of times the limit-exceeding event occurs. If the counted number exceeds a preset threshold, the remaining life of the storage device is predicted; when the predicted remaining life is lower than the set life threshold, a startup instruction is sent to the self-recovery device; The self-recovery device 3 is used to cut off the connection between the processing device and the storage device after receiving the startup instruction, and switch to the dominant mode of the management control component 2 to enable the management control component 2 to perform the mirror recovery operation.

[0015] It should be noted that the processing device is the initiator of data interaction with the storage device. After establishing a connection with the storage device, the processing device can send a write operation request to the storage device. The processing device can be the host, serving as the source of the entire data write operation. By sending the write operation request, it drives the storage device to perform the corresponding data storage action. The storage device can be an embedded MultiMediaCard (eMMC) or other storage device, which is not limited here.

[0016] As a key component connecting to a storage device, digital logic device 1 plays a crucial role in the storage device's response to write requests. First, it accurately detects the actual latency of write operations in real time, capturing the time difference between the processing unit sending a request and the storage device completing its response. Second, it determines the actual latency based on a preset latency threshold. If the detected latency exceeds the threshold, it signals an over-limit event. This allows digital logic device 1 to promptly detect fluctuations in the storage device's performance when processing write operations, providing direct and critical raw data support for subsequent storage device status assessments and fault warnings.

[0017] As the core component connecting the digital logic device 1 and the self-recovery device 4, the management and control component 2 can count the number of out-of-limit events detected by the digital logic device 1; when the statistical number exceeds the preset threshold (such as 1000 times), it will start the remaining life prediction process of the storage device to determine the aging degree and survival status of the device; and when the predicted remaining life is lower than the set life threshold, it means that the storage device is approaching the end of its service life and the risk is significantly increased. The management and control component 2 can issue a startup instruction to the self-recovery device 3, thereby promoting the subsequent fault response and device protection mechanism to start. This process not only realizes the dynamic control of the storage device status, but also lays the foundation for timely intervention in equipment management and avoiding the expansion of faults.

[0018] As a key component within the device, self-recovery device 3 fully functions upon receiving a startup command from management and control component 2: it disconnects the processing device from the storage device, preventing data exchange and exacerbating risks even when the storage device is in poor condition; it then simultaneously switches to the dominant mode of management and control component 2, handing control of the device over to it, enabling it to perform image recovery operations. This process, by promptly isolating the problematic device and transferring control, provides a stable environment for secure data recovery and is crucial for ensuring the device's ability to self-repair in the event of a storage device risk.

[0019] In the above-mentioned storage device monitoring device provided by the embodiment of the present invention, through the coordinated operation of the digital logic device 1, the management and control component 2 and the self-recovery device 3, it is possible to monitor the delay of the storage device write operation in real time, timely discover and count the over-limit events, accurately predict the remaining life of the storage device when the number of over-limit events is too many, and when the remaining life is lower than the threshold, the self-recovery device 3 quickly cuts off the original connection and switches to the dominant mode of the management and control component 2 to perform the mirror recovery operation. In this way, without the need to use programmable devices, it not only reduces the hardware cost and operation and maintenance cost, but also can accurately control the status of the storage device and give early warning of potential failures by relying on highly accurate life predictions; when the storage device is about to expire, the self-recovery device 3 can automatically start the recovery mechanism, and ensure data security and device operation continuity through automatic mirror recovery, which comprehensively improves the reliability and stability of the device. It is especially suitable for edge scenarios and can avoid data loss caused by sudden failures.

[0020] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the digital logic device 1 may include a trigger, a timer and a delay comparator; wherein the trigger is used to detect the start signal of a write operation request; the timer is used to start timing when the trigger detects the start signal; after receiving a response signal indicating that the write operation is completed, the timer stops timing to record the actual delay value of the write operation; the delay comparator is used to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; and based on the comparison and analysis results, determine whether an over-limit event has occurred.

[0021] It should be noted that the present invention indirectly assesses the lifespan of storage devices (such as eMMC) by utilizing the write latency. As the storage device's lifespan decreases, the number of bad blocks increases, and remapping operations can lead to increased write latency. Therefore, the present invention measures the completion time of each write operation using digital logic device 1. If the write operation time exceeds a preset threshold, it may indicate that a remapping (or other error handling) has occurred. A suspected remapping event is then recorded and counted.

[0022] In implementation, the digital logic device 1 may include a trigger, a counter, and a delay comparator. The trigger may also be referred to as a write command trigger. The timer may be driven by a quartz crystal oscillator with an accuracy of approximately 10ppm. When the processing device initiates a write operation request to the storage device, the write command trigger detects the start of the write command and simultaneously triggers the timer to start counting. The timer stops counting upon receiving a response indicating the write operation is complete. Delay comparison and analysis is then performed, with the delay comparator comparing the measured actual delay value with a preset threshold. This threshold is divided into two levels: a non-warning threshold and a warning threshold. The non-warning threshold is based on the maximum delay specified in the storage device specification (such as eMMC). Specifically, the non-warning threshold represents the theoretically acceptable upper limit for read and write operations under the storage device's design and production standards, representing the delay boundary for stable device operation. The warning threshold can be set to 150% of the non-warning threshold. If the delay exceeds the warning threshold, an over-limit event is determined. The digital logic device 1 of the present invention can accurately capture the start and completion signals of the write operation request through the coordinated cooperation of the trigger, timer, and delay comparator, thereby accurately recording the actual delay value and comparing it with the preset warning threshold to determine whether an over-limit event has occurred. This not only realizes real-time and accurate monitoring of the write operation delay of the storage device, but also provides a reliable basis for subsequent life prediction and fault warning. It does not need to rely on complex components. While ensuring detection accuracy, it helps to simplify the hardware structure, further reduce costs, and improve the adaptability and practicality of the entire device in applications such as edge scenarios.

[0023] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the trigger can be specifically used to capture the starting level or protocol frame header of the write operation request through the command line of the storage device; the timer can be specifically used to start timing when the trigger captures the starting level or protocol frame header.

[0024] In practice, the trigger can accurately capture the starting level or protocol frame header of the write operation request through the command line (CMD) of the storage device, providing the timer with a clear and precise timing starting point, namely the start time of the write operation. This ensures that the timer starts timing from the moment the write operation is started, effectively avoiding timing deviation, thereby improving the accuracy and reliability of the entire delay detection process.

[0025] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the timer can also be used to stop timing after receiving a response code corresponding to the write operation completion status returned by the command line of the storage device, or after receiving a response frame indicating that the write operation is completed.

[0026] During implementation, the timer can stop timing when it receives the write operation completion status response code or response frame returned by the storage device command line, and can accurately capture the end time of the write operation, forming a complete time closed loop with the start time. This ensures the integrity and accuracy of the timing of the actual delay value of the write operation, avoids the delay calculation deviation caused by inaccurate judgment of the end time, further improves the accuracy of write operation delay detection, and provides more reliable time data support for subsequent over-limit event judgment and life prediction.

[0027] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the management control component 2 may include a bad block counter and a management controller; wherein the bad block counter is used to count a corresponding number of times when it is determined that an over-limit event has occurred, so as to cumulatively record the over-limit events that occur during the operation of the storage device and send a life detection signal to the management controller; the management controller is used to read the count value of the bad block counter, and analyze the bad block growth trend based on the read count value to predict the remaining life of the storage device.

[0028] In implementation, the management and control component 2 may include a bad block counter and a management controller; the bad block counter can accurately count when an over-limit event occurs, cumulatively record the over-limit conditions during the operation of the storage device and send out a life detection signal, providing accurate basic data for subsequent analysis; the management controller can analyze the bad block growth trend by reading the count value, thereby predicting the remaining life of the storage device. This trend analysis method based on actual operation data makes the life prediction more in line with the actual status of the device, further improving the accuracy and reliability of the prediction.

[0029] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the count value of the bad block counter can be saved using a non-volatile storage medium; when the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to accumulate from the current value.

[0030] During implementation, each time an over-limit event occurs, the bad block counter can automatically increment by a set value (e.g., 1). The bad block counter can use a non-volatile storage medium to store the count value, ensuring that the count value remains unchanged when the storage device is powered off, and can continue to accumulate from the current value after power is restored. The non-volatile storage medium can be an Electrically Erasable Programmable Read-Only Memory (EEPROM), which is not limited here. This avoids the loss of over-limit event counts due to power outages, ensures the continuity and integrity of the count, and provides a coherent and reliable data basis for the management controller to analyze the bad block growth trend and predict the remaining life of the storage device, further improving the accuracy of life prediction, reducing the additional verification or re-counting work caused by data interruption, and reducing the additional cost of device operation.

[0031] Figure 2 Schematic diagram of signaling interaction corresponding to the storage device monitoring device provided in an embodiment of the present invention. Figure 2 As shown, a processing device (such as a host) can send a write operation request to a storage device, triggering the data write process in the storage device (such as an eMMC). When the storage device responds to the write operation request, it feeds response delay information back to digital logic device 1 to detect the delay and monitor the storage device's response performance. Digital logic device 1 collects the delay data and combines it with bad block count information to report it to the management controller. The management controller can be a baseboard management controller (BMC). The management controller can perform trend analysis on the delay data and bad block count and, based on the analysis results, initiate different branches. If the remaining life is at least a set lifespan threshold (such as 5%), a normal response is returned to the processing device, and the write operation is completed according to the normal process without triggering recovery. If the remaining lifespan is below the set lifespan threshold (such as 5%), the storage device is determined to be nearing its lifespan limit (meeting the recovery trigger condition), triggering the recovery process and sending a start command to self-recovery device 3. Self-recovery device 3 can isolate the connection to avoid data conflicts and initiate a read mirror operation request to access the backup of the storage device's critical data. Subsequently, you can perform fast erase and image write operations, as well as device reset operations, to reinitialize key components and ensure that they can interact again in the repaired state.

[0032] Furthermore, in specific implementation, in the above-mentioned storage device monitoring device provided in the embodiment of the present invention, the management controller can be specifically used to read the count value of the bad block counter according to a preset period, and calculate the number of bad blocks growing per unit time based on the read count value; use the exponentially weighted moving average algorithm to predict the number of bad blocks growing, and obtain the bad block growth rate in the future set time period; calculate the remaining life of the storage device based on the bad block growth rate combined with the current number of bad blocks and the maximum allowable number of bad blocks.

[0033] In implementation, the management controller reads the bad block counter at a preset interval (e.g., hourly), calculates the number of bad blocks growing per unit time (bad blocks / hour), and then uses an exponentially weighted moving average (EWMA) algorithm to predict the bad block growth rate over a set future time period (e.g., the next 24 hours). Finally, the remaining lifespan of the storage device is calculated by combining the current number of bad blocks with the maximum allowable number of bad blocks. The maximum allowable number of bad blocks is determined by the storage device model (typically 5% of the total number of blocks on the chip). This approach makes analysis of bad block growth trends more timely and accurate, and allows calculation of remaining lifespan to be based on reliable dynamic data. This significantly improves the accuracy of lifespan predictions and provides a more timely and accurate reflection of the actual status of the storage device.

[0034] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the management controller can be specifically used to obtain the current number of bad blocks from the bad block counter; subtract the sum of the current number of bad blocks and the bad block growth rate from 1, divide the obtained difference by the maximum allowable number of bad blocks, and then multiply the obtained value by 100 to calculate the remaining life of the storage device.

[0035] In practice, the management controller obtains the current number of bad blocks from the bad block counter and then calculates the remaining lifespan of the storage device. This calculation formula can be: Remaining Lifespan (%) = 100 * (1 - (Current Number of Bad Blocks + Predicted Growth) / Maximum Allowable Number of Bad Blocks). This calculation method directly links the current bad block status, growth trend, and the device's tolerance limit. Its clear logic and simple calculations allow for a quick and intuitive remaining lifespan value. This ensures that the calculation results closely align with the actual device wear and tear, improves the efficiency of lifespan assessment, and provides a clear and reliable quantitative basis for subsequent decisions on whether to activate the self-recovery mechanism.

[0036] Furthermore, in specific implementation, in the above-mentioned storage device monitoring device provided in the embodiment of the present invention, the management controller can also be used to compare the predicted remaining life with the set life threshold after predicting the remaining life of the storage device; if the predicted remaining life is lower than the set life threshold, the start-up conditions of the storage device self-recovery mechanism are triggered, and a start-up instruction is sent to the self-recovery device 3.

[0037] During implementation, the management controller can accurately determine whether the device is approaching a critical failure state after predicting the remaining life of the storage device by comparing it with the set life threshold: when the remaining life is lower than the set life threshold, the self-recovery mechanism start condition is immediately triggered and a start instruction is sent to the self-recovery device 3. This process realizes dynamic monitoring and timely response to the device status, ensuring that the protection mechanism is actively started before the storage device life is about to expire, avoiding data loss due to sudden failure of the device, and at the same time reducing the cost of troubleshooting and on-site maintenance through early intervention, effectively ensuring the continuous and stable operation of the storage system in edge scenarios.

[0038] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the management controller can also be used to trigger an alarm through the web management interface while sending a startup instruction to the self-recovery device 3, and drive the alarm light to light up to prompt the administrator to back up relevant files.

[0039] During implementation, the management controller can trigger an alarm and turn on the alarm light through the web management interface while sending the self-recovery startup command. This can promptly convey early warning information of impending storage device failure to the manager, reminding him to back up relevant files as soon as possible, further enhancing the timeliness of manual intervention, reducing the risk of data loss, and ensuring that managers will not miss key information through dual alarms, thereby forming an effective coordination between automatic protection and manual operation, minimizing the file loss that may be caused by equipment problems, and improving the security and controllability of the storage system in emergency scenarios.

[0040] Furthermore, in a specific implementation, in the storage device monitoring device provided in an embodiment of the present invention, the self-recovery device 3 may include an isolation switch. The isolation switch may be a high-speed analog switch (such as the PI3A3157), with a switching time of less than 10ns. The management controller may be configured to simultaneously send a self-recovery control signal to the isolation switch while simultaneously sending a startup command to the self-recovery device 3. Upon receiving the self-recovery control signal from the management controller, the isolation switch may be configured to disconnect the processing device from the storage device and switch to the management controller's dominant mode. The management controller may also be configured to temporarily store mirrored data in a hardware buffer, which is then processed by the hardware buffer and written to the storage device.

[0041] During implementation, after receiving the self-recovery control signal sent by the management controller, the isolation switch in the self-recovery device 3 can quickly cut off the connection between the processing device and the storage device and switch to the management controller dominant mode. At the same time, the management controller temporarily stores the mirror data in the hardware buffer and writes it to the storage device after processing. This not only avoids the additional pressure on the faulty device caused by the continuous interaction of the processing device through the isolation operation, but also ensures the stability and security of the mirror data writing with the help of the hardware buffer. In conjunction with the dominant control of the management controller, it not only ensures the order and efficiency of the self-recovery process, but also provides a reliable channel for data migration during automatic recovery, further reducing the risk of data loss, while reducing the complexity of manual intervention and improving the convenience and reliability of emergency processing of the device.

[0042] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the isolation switch can be specifically used to disconnect the command line and data line connections between the processing device and the storage device after receiving the self-recovery control signal of the management controller, and synchronously switch to the corresponding control line between the management controller and the storage device to establish a management channel connection.

[0043] In practice, upon receiving the self-recovery control signal from the management controller, the isolation switch disconnects the command and data lines between the processing unit and the storage device, completely severing data exchange between the two. This prevents the faulty device from continuously receiving operational commands and exacerbating the problem. Simultaneously, it switches to the corresponding control lines between the management controller and the storage device and establishes a management channel connection, ensuring the management controller's sole control of the storage device and providing a dedicated, stable communication path for subsequent image recovery operations. This securely isolates the faulty device from the original processing unit, establishing a reliable channel for the recovery process led by the management controller, and effectively ensuring the independence and security of the self-recovery process.

[0044] Furthermore, in a specific implementation, in the storage device monitoring apparatus provided in an embodiment of the present invention, the self-recovery device 3 may include a serial peripheral interface flash memory (SPI FLASH) configured to store backup image data. The management controller may be configured to read pre-stored backup image data from the SPI FLASH, temporarily store the read backup image data in a page buffer, and perform image writing to the storage device through format conversion processing in the page buffer.

[0045] In implementation, the serial peripheral interface flash memory in the self-recovery device 3 can stably store the backup image data. Figure 3 This is a flow chart corresponding to the self-recovery device provided in the embodiment of the present invention. Figure 3As shown, after the management controller reads the backup image data from the serial peripheral interface flash memory, it is first temporarily stored in the page buffer for format conversion processing, and then the image write operation of the storage device is executed. The page buffer refers to a hardware buffer where data is temporarily stored after being read from the serial peripheral interface flash memory and before being written to the storage device. It is responsible for buffering and format conversion of high-speed data. In this way, the security and accessibility of the backup data are guaranteed by the reliable storage characteristics of the serial peripheral interface flash memory, and the format conversion of the page buffer is adapted to the requirements of the storage device to ensure the compatibility and accuracy of the image writing process. The entire process realizes the coherent and efficient operation of backup data reading, format adaptation to write recovery, further improves the reliability of the self-recovery mechanism, can quickly complete data recovery when the storage device is about to reach the end of its life, reduces the cost of manual intervention, and effectively guarantees the integrity of data and the continuous and stable operation of the device in edge scenarios.

[0046] Furthermore, in a specific implementation, in the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the management controller can also be used to trigger the reset circuit to perform a reset operation on the server mainboard after the storage device image is written, so that the storage device is reinitialized and restored to the connection with the processing device.

[0047] In practice, after the storage device image is written, the management controller triggers the reset circuit to reset the server motherboard, prompting the storage device to reinitialize and restore its connection to the processing unit, completing the self-recovery process. This reset operation ensures that the storage device is put back into use in a stable state and restores normal data interaction with the processing unit. This not only ensures the device's rapid recovery after image restoration, but also avoids operational failures caused by abnormal connections or improper initialization after restoration. It also eliminates the need for manual restarts, reducing on-site maintenance intervention costs, further improving the continuity and stability of edge servers after storage device restoration, and ensuring that services can quickly return to normal operation.

[0048] From the above description of the embodiments, those skilled in the art will clearly understand that the device according to the above embodiments can be implemented using hardware. From an implementation perspective, the core components of the device, such as the digital logic device (including triggers, timers, and delay comparators), the management and control unit (with an integrated bad block counter and management controller), and the self-recovery device (including the isolation switch and serial peripheral interface flash memory), are all designed using hardware circuits, rather than relying on the loading and execution of software programs. This purely hardware architecture avoids the compatibility issues, code vulnerability risks, and processor resource utilization that software may face, enabling faster execution of various operations. For example, the digital logic device can detect write operation delays in nanoseconds, and the management and control unit can count and determine over-limit events without complex instruction parsing, significantly improving the device's real-time performance. Furthermore, the hardware implementation offers greater stability, unaffected by fluctuations in the operating system or operating environment, ensuring continuous and reliable operation, especially in complex and harsh environments such as edge scenarios. From a cost and maintenance perspective, the hardware circuits eliminate the need for expensive programmable chips, reducing hardware costs through streamlined circuit design. Furthermore, the solidified hardware functionality reduces the operational burden associated with software upgrades and bug fixes, further meeting the edge's demand for low-cost, highly stable equipment. Furthermore, hardware-level collaboration (such as rapid switching of isolation switches and format conversion of page buffers) ensures precise synchronization of operations across all links, providing a solid physical foundation for core functions such as highly accurate lifespan prediction and reliable self-recovery mechanisms. This ensures more stable and controllable performance in terms of data security and fault prevention.

[0049] Based on the same inventive concept, an embodiment of the present invention further provides a monitoring method for a storage device monitoring apparatus. Figure 4 A flow chart of a monitoring method for a storage device monitoring apparatus according to an embodiment of the present invention is shown as follows: Figure 4 As shown, the method includes: S401: When the storage device responds to a write operation request sent by a processing device, the storage device uses a digital logic device to detect an actual delay value of the write operation to determine whether an over-limit event occurs.

[0050] S402. Utilize the management control component to count the number of times the over-limit event occurs. If the counted number exceeds a preset threshold, predict the remaining life of the storage device. When the predicted remaining life is lower than the set life threshold, send a startup instruction to the self-recovery device.

[0051] S403: Use the self-recovery device to cut off the connection between the processing apparatus and the storage device, and switch to the dominant mode of the management and control component, so that the management and control component performs the mirror recovery operation.

[0052] In the monitoring method of the above-mentioned storage device monitoring device provided in the embodiment of the present invention, the delay of the storage device write operation can be monitored in real time, the over-limit events can be discovered and counted in time, and the remaining life of the storage device can be accurately predicted when the number of over-limit events is too many. When the remaining life is lower than the threshold, the self-recovery device quickly cuts off the original connection and switches to the management and control component dominant mode to perform the mirror recovery operation. In this way, without the need to use programmable devices, it not only reduces the hardware cost and operation and maintenance cost, but also can accurately control the status of the storage device and give early warning of potential failures by relying on highly accurate life predictions; when the storage device is about to expire, the self-recovery device can automatically start the recovery mechanism, and ensure data security and device operation continuity through automatic mirror recovery, which comprehensively improves the reliability and stability of the device. It is especially suitable for edge scenarios and can avoid data loss caused by sudden failures.

[0053] Since the embodiments of the monitoring method of the storage device monitoring apparatus correspond to the embodiments of the storage device monitoring apparatus, the description of the features of the corresponding embodiments of the monitoring method of the storage device monitoring apparatus can be found in the description of the corresponding embodiments of the storage device monitoring apparatus, and will not be repeated here. The embodiments of the storage device monitoring apparatus and the monitoring method of the storage device monitoring apparatus have the same beneficial effects as the aforementioned storage device monitoring apparatus.

[0054] Furthermore, in a specific implementation, in the monitoring method of the above-mentioned storage device monitoring apparatus provided in an embodiment of the present invention, step S401 uses a digital logic device to detect the actual delay value of the write operation to determine whether an over-limit event occurs, which may specifically include: using a trigger to detect the start signal of the write operation request; when the trigger detects the start signal, starting timing through a timer; after receiving a response signal indicating that the write operation is completed, stopping timing to record the actual delay value of the write operation; using a delay comparator to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; and judging whether an over-limit event occurs based on the comparison and analysis results.

[0055] Furthermore, in a specific implementation, in the above steps, a trigger is used to detect the start signal of the write operation request; when the trigger detects the start signal, the timer starts timing; after receiving the response signal indicating that the write operation is completed, the timing is stopped, which may specifically include: using a trigger to capture the start level or protocol frame header of the write operation request through the command line of the storage device; when the trigger captures the start level or protocol frame header, the timer starts timing; the starting point of the timing is the start time of the write operation; after receiving the response code corresponding to the write operation completion status returned by the command line of the storage device, or after receiving the response frame indicating that the write operation is completed, the timing is stopped.

[0056] Furthermore, in a specific implementation, in the monitoring method of the storage device monitoring apparatus provided in an embodiment of the present invention, step S402 utilizes the management control component to count the number of over-limit events. Specifically, this may include: when an over-limit event is determined to have occurred, using a bad block counter to count the corresponding number of over-limit events to cumulatively record the over-limit events that occurred during the operation of the storage device, and sending a life detection signal to the management controller. In implementation, the count value of the bad block counter is stored using a non-volatile storage medium; when the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to accumulate from the current value.

[0057] If the number of counts exceeds a preset threshold in step S402, then the remaining life of the storage device is predicted. This may specifically include: using the management controller to read the count value of the bad block counter, and analyzing the bad block growth trend based on the read count value to predict the remaining life of the storage device. In implementation, the management controller reads the count value of the bad block counter according to a preset period, and calculates the number of bad blocks that have increased per unit time based on the read count value; uses an exponentially weighted moving average algorithm to predict the number of bad blocks that have increased to obtain the bad block growth rate within a future set time period; and calculates the remaining life of the storage device based on the bad block growth rate, the current number of bad blocks, and the maximum allowable number of bad blocks.

[0058] The remaining life of the storage device is calculated based on the bad block growth rate combined with the current number of bad blocks and the maximum allowable number of bad blocks. Specifically, the method may include: obtaining the current number of bad blocks from the bad block counter; subtracting the sum of the current number of bad blocks and the bad block growth rate from 1, dividing the resulting difference by the maximum allowable number of bad blocks, and then multiplying the resulting value by 100 to calculate the remaining life of the storage device.

[0059] Step S402 sends a startup instruction to the self-recovery device when the predicted remaining life is lower than the set life threshold. Specifically, it may include: after predicting the remaining life of the storage device, comparing the predicted remaining life with the set life threshold; if the predicted remaining life is lower than the set life threshold, triggering the startup condition of the storage device self-recovery mechanism and sending a startup instruction to the self-recovery device.

[0060] Furthermore, in a specific implementation, in the monitoring method of the above-mentioned storage device monitoring device provided in an embodiment of the present invention, after executing step S402, it may also include: while sending a startup instruction to the self-recovery device, using the management controller to trigger an alarm through the web management interface, and driving the alarm light to light up, so as to prompt the administrator to back up relevant files.

[0061] Furthermore, in a specific implementation, in the monitoring method of the above-mentioned storage device monitoring device provided in an embodiment of the present invention, step S403 uses a self-recovery device to cut off the connection between the processing device and the storage device, and switches to the dominant mode of the management control component, so that the management control component performs the mirror recovery operation, which may specifically include: while sending a startup instruction to the self-recovery device, using the management controller to send a self-recovery control signal to the isolation switch of the self-recovery device; after receiving the self-recovery control signal from the management controller, the isolation switch cuts off the connection between the processing device and the storage device, and switches to the dominant mode of the management controller; the mirror data is temporarily stored in the hardware buffer through the management controller, and written to the storage device after being processed by the hardware buffer.

[0062] Furthermore, in specific implementation, in the monitoring method of the above-mentioned storage device monitoring device provided in an embodiment of the present invention, after receiving the self-recovery control signal of the management controller, the connection between the processing device and the storage device is cut off, and the dominant mode of the management controller is switched. Specifically, it may include: after receiving the self-recovery control signal of the management controller, the command line and data line connection between the processing device and the storage device is disconnected, and the corresponding control line between the management controller and the storage device is synchronously switched to establish a management channel connection.

[0063] Furthermore, in a specific implementation, in the monitoring method of the above-mentioned storage device monitoring device provided in an embodiment of the present invention, the mirror data is temporarily stored in a hardware buffer and written to the storage device after being processed by the hardware buffer. Specifically, it can include: reading pre-stored backup mirror data from the serial peripheral interface flash memory of the self-recovery device, and temporarily storing the read backup mirror data in a page buffer, and performing mirror writing to the storage device through format conversion processing of the page buffer.

[0064] Furthermore, in a specific implementation, the monitoring method of the above-mentioned storage device monitoring device provided in an embodiment of the present invention may also include: after the storage device image is written, using the management controller to trigger the reset circuit to perform a reset operation on the server motherboard, so that the storage device is reinitialized and restored to the connection with the processing device.

[0065] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other. The described embodiments are only some embodiments of the present invention, not all embodiments. Based on these embodiments, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in this field can still combine, add, delete or make other adjustments to the features in the various embodiments of the present invention according to the circumstances without conflict and without making creative work, so as to obtain different other technical solutions that do not deviate from the concept of the present invention in essence. These technical solutions also fall within the scope of protection of the present invention.

[0066] The above is a detailed introduction to the storage device monitoring device and monitoring method provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only intended to help understand the method and core ideas of the present invention, and is not intended to limit the scope of protection of the invention. At the same time, for those skilled in the art, based on the ideas of the present invention, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A storage device monitoring device, characterized in that: include: a digital logic device connected to the storage device, and configured to detect an actual delay value of a write operation when the storage device responds to a write operation request sent by the processing device, and determine whether an over-limit event occurs; a management control component connected to the digital logic device and configured to count the number of occurrences of the over-limit event, and predict the remaining life of the storage device if the counted number exceeds a preset threshold; and send a startup instruction to the self-recovery device when the predicted remaining life is lower than a set threshold; The self-recovery device is used to cut off the connection between the processing device and the storage device after receiving the startup instruction, and switch to the dominant mode of the management and control component to enable the management and control component to perform the mirror recovery operation.

2. The storage device monitoring device according to claim 1, wherein: The digital logic device includes a trigger, a timer and a delay comparator; The trigger is used to detect the start signal of the write operation request; The timer is configured to start timing when the trigger detects the start signal; and stop timing after receiving a response signal indicating that the write operation is completed, so as to record an actual delay value of the write operation; The delay comparator is used to compare and analyze the actual delay value recorded by the timer with a preset warning threshold; and determine whether an over-limit event occurs based on the comparison and analysis result.

3. The storage device monitoring device according to claim 2, wherein: The trigger is configured to capture the start level or protocol frame header of the write operation request via a command line of the storage device; The timer is configured to start timing when the trigger captures the start level or the protocol frame header.

4. The storage device monitoring device according to claim 3, wherein: The timer is further configured to stop timing after receiving a response code corresponding to a write operation completion status returned by a command line of the storage device, or after receiving a response frame indicating that the write operation is completed.

5. The storage device monitoring device according to claim 1, wherein: The management and control component includes a bad block counter and a management controller; The bad block counter is configured to count a corresponding number of times when it is determined that an over-limit event has occurred, so as to cumulatively record the over-limit events that have occurred during the operation of the storage device, and to send a life detection signal to the management controller; The management controller is used to read the count value of the bad block counter, and analyze the bad block growth trend according to the read count value to predict the remaining life of the storage device.

6. The storage device monitoring device according to claim 5, characterized in that: The count value of the bad block counter is stored in a non-volatile storage medium; When the storage device is powered off, the count value of the bad block counter remains unchanged; when the storage device is powered on, the count value of the bad block counter continues to accumulate from the current value.

7. The storage device monitoring device according to claim 5, characterized in that: The management controller is used to read the count value of the bad block counter according to a preset period, and calculate the number of bad blocks growing per unit time based on the read count value; use an exponentially weighted moving average algorithm to predict the number of bad blocks growing to obtain the bad block growth rate within a future set time period; and calculate the remaining life of the storage device based on the bad block growth rate combined with the current number of bad blocks and the maximum allowable number of bad blocks.

8. The storage device monitoring device according to claim 7, wherein: The management controller is used to obtain the current number of bad blocks from the bad block counter; subtract the sum of the current number of bad blocks and the bad block growth rate from 1, divide the obtained difference by the maximum allowable number of bad blocks, and then multiply the obtained value by 100 to calculate the remaining life of the storage device.

9. The storage device monitoring device according to claim 5, characterized in that: The management controller is also used to compare the predicted remaining life with a set life threshold after predicting the remaining life of the storage device; if the predicted remaining life is lower than the set life threshold, the start-up conditions of the self-recovery mechanism of the storage device are triggered, and a start-up instruction is sent to the self-recovery device.

10. The storage device monitoring device according to claim 5, wherein: The management controller is further configured to trigger an alarm through a web management interface and drive an alarm light to light up while sending a startup instruction to the self-recovery device, so as to prompt the administrator to back up relevant files.

11. The storage device monitoring device according to claim 5, wherein: The self-restoring device includes an isolating switch; The management controller is configured to send a self-recovery control signal to the isolation switch while sending a start-up instruction to the self-recovery device; The isolation switch is configured to cut off the connection between the processing device and the storage device and switch to the dominant mode of the management controller after receiving the self-recovery control signal from the management controller; The management controller is further configured to temporarily store the mirror data in a hardware buffer, and write the mirror data into the storage device after being processed by the hardware buffer.

12. The storage device monitoring apparatus according to claim 11, wherein: The isolation switch is used to disconnect the command line and data line connections between the processing device and the storage device after receiving the self-recovery control signal from the management controller, and synchronously switch to the corresponding control line between the management controller and the storage device to establish a management channel connection.

13. The storage device monitoring apparatus according to claim 11, wherein: The self-recovery device includes a serial peripheral interface flash memory; The serial peripheral interface flash memory is used to store backup image data; The management controller is used to read pre-stored backup image data from the serial peripheral interface flash memory, temporarily store the read backup image data in a page buffer, and perform image writing on the storage device through format conversion processing of the page buffer.

14. The storage device monitoring apparatus according to claim 13, wherein: The management controller is further configured to trigger a reset circuit to perform a reset operation on the server mainboard after the storage device image is written, so as to reinitialize the storage device and restore the connection with the processing device.

15. A monitoring method for a storage device monitoring apparatus according to any one of claims 1 to 14, characterized in that: include: When the storage device responds to the write operation request sent by the processing device, the digital logic device is used to detect the actual delay value of the write operation to determine whether an over-limit event occurs; Utilizing a management control component to count the number of occurrences of the over-limit event, and if the counted number exceeds a preset threshold, predicting the remaining life of the storage device; and sending a startup instruction to the self-recovery device when the predicted remaining life is lower than a set life threshold; The self-recovery device is used to cut off the connection between the processing apparatus and the storage device, and switch to the dominant mode of the management and control component, so that the management and control component performs a mirror recovery operation.

Citation Information

Patent Citations

  • Cloud mobile terminal cooperative fault early warning method, related device and system

    CN109634820A

  • SSD flash memory life prediction method, device and equipment, and readable medium

    CN111475115A

  • SSD life prediction method and device, equipment and readable storage medium

    CN111950796A

  • Interactive perception data acquisition and storage method based on BGA SSD

    CN118550475A

  • SPI NAND flash memory life prediction method and device, equipment and storage medium

    CN118643280A