Storage Device and Its Control System

By setting up dual control modules and non-volatile memory in the storage device, automatic switching and update of firmware abnormalities is achieved, resource waste and manual intervention problems caused by firmware errors in the prior art are solved, and the automatic repair capability and reliability of the system are improved.

CN114253763BActive Publication Date: 2025-07-08KUNDA COMP TECHKUSN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010992639.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-21
Publication Date
2025-07-08
Estimated Expiration
2040-09-21

AI Technical Summary

Technical Problem

In existing storage systems, when the firmware error causes abnormality in the control module, it is necessary to manually replace the motherboard or increase non-volatile memory, resulting in waste of resources and space consumption, and the exception handling is not automated enough.

Method used

Two control modules and non-volatile memory are set in the storage device. Each module initializes the self-reading firmware program code, automatically switches the mode in case of abnormalities and updates the firmware program code to realize automatic repair and backup functions.

Benefits of technology

It realizes automatic repair and backup of storage equipment when firmware abnormalities are abnormal, reduces manual intervention and resource waste, and improves system reliability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253763B_ABST
    Figure CN114253763B_ABST
Patent Text Reader

Abstract

A storage device includes a control system, and the control system includes two main boards, two control modules, and two non-volatile memories. The two control modules respectively read a firmware program code from the two non-volatile memories to execute a firmware, and respectively operate in a master mode and a slave mode. When an exception occurs during the operation of the control module operating in the master mode, the control module operating in the slave mode is converted to operate in the master mode, and controls the control module where the exception occurs to be converted to operate in a recovery mode, and transmits the firmware program code stored in the corresponding volatile memory to the control module where the exception occurs to update the one stored in the corresponding non-volatile memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a storage device and its control system, and particularly to a storage device and its control system with automatic repair and redundancy functions.

Background Art

[0002] Existing enterprise storage systems belong to a high-availability (HA) system and include a power supply unit, a fan unit, two mainboards, two control modules (Input / Output module, IOM), and at least one hard disk. The two control modules are respectively arranged on the two mainboards and are used to monitor and manage each of the hard disks, the power supply unit, and the fan unit, and operate in an active mode and a passive mode respectively after power-on to serve as a redundancy system. Each control module is initialized to read a firmware program code (such as an Image) from a corresponding non-volatile memory to execute a firmware. When an exception occurs in the control module operating in the active mode, one of the main reasons is that the corresponding firmware has an error.

[0003] To solve the problem of firmware errors, a simple method in a conventional technology is to replace the corresponding mainboard, but it requires a lot of manpower, money, and time. Another method in the conventional technology is to set two non-volatile memories corresponding to the control module on each mainboard, and store the same firmware program code in the two non-volatile memories, that is, one of them is the one that the control module defaults to read after power-on, and the other is used as a backup, or one of them stores an updated version of the firmware program code, and the other stores the operable version of the firmware program code at the time of factory. For example, when an error occurs in the firmware executed by the control module operating in the active mode, the control module reads the backup firmware program code to operate normally. However, this method requires setting two non-volatile memories corresponding to the control module on each mainboard, which not only wastes the cost of non-volatile memory but also occupies the space of the circuit board. Furthermore, another method is that when the firmware program code read and executed by the control module operating in the active mode from the non-volatile memory causes an exception in the control module, another control module is switched from the passive mode to the active mode to maintain the normal operation of the storage system.

[0004] At this time, although the storage system is still operating normally, the firmware on the non-volatile memory that caused the abnormality of the control module still needs to be manually re-burned by engineers or users with other versions of the firmware program code that can operate normally. Therefore, how to handle other firmware abnormalities in enterprise storage systems has become a problem to be solved.

Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a storage device and its control system with automatic repair and redundancy functions.

[0006] To solve the above technical problems, the present invention provides a control system for a storage device, which is applicable to a power supply unit, a fan unit, and at least one hard disk, and includes two main boards, two control modules, and two non-volatile memories. The two control modules are respectively disposed on the two main boards, electrically connected to each other and both electrically connected to the power supply unit, the fan unit, and the at least one hard disk for monitoring and managing the power supply unit, the fan unit, and the at least one hard disk. The two non-volatile memories are respectively disposed on the two main boards, electrically connected to the two control modules respectively, and both store a firmware program code.

[0007] Wherein, the two control modules are respectively initialized to read the firmware program code from the two non-volatile memories respectively to execute a corresponding firmware, so that the two control modules operate in a main control mode and a controlled mode respectively.

[0008] When the control module operating in the main control mode has an abnormality, the control module operating in the controlled mode is converted to operate in the main control mode, and controls the control module originally operating in the main control mode to be converted to operate in a restoration mode, and transmits the firmware program code stored in the volatile memory corresponding to the control module operating in the main control mode to the control module operating in the restoration mode to update the non-volatile memory corresponding to the control module operating in the restoration mode.

[0009] Preferably, each of the control modules uses a health signal to enable the other to determine whether the other has an abnormality.

[0010] Preferably, each control module includes a first bus for outputting the health signal and for receiving the health signal output by the other party. When any one of the control modules operates normally, the logical value of the corresponding output health signal will jump between logic 1 and logic 0. When any one of the control modules operates abnormally, the logical value of the corresponding output health signal will remain at logic 0 or logic 1.

[0011] Preferably, each of the control modules outputs the health signal via a first bus and includes a buffer. When one of the control modules operates normally, the value of the buffer of the other is changed over time by using the corresponding health signal. When one of the control modules operates abnormally, the value of the buffer of the other cannot be changed over time by using the corresponding health signal.

[0012] Preferably, each of the control modules includes a second bus for outputting a restoration signal and for receiving another restoration signal output by the other party. Each of the control modules determines whether to operate in the restoration mode according to the received restoration signal.

[0013] Preferably, each of the control modules includes a third bus for transmitting the firmware program code and for receiving the firmware program code transmitted by the other party.

[0014] Preferably, the control modules are respectively defined as a first control module and a second control module. The first control module is default to operate in the main control mode, and the second control module is default to operate in the controlled mode. When the first control module malfunctions and changes to operate in the restoration mode, and receives the firmware program code from the second control module to update the stored non-volatile memory, and then re-executes the firmware and operates normally, the first control module changes to operate in the main control mode and the second control module changes to operate in the controlled mode.

[0015] To solve the above technical problems, the present invention further provides a storage device, including a power supply unit, a fan unit, at least one hard disk, and a control system. The power supply unit is used to provide operating power. The fan unit is used to provide heat dissipation. The at least one hard disk is used to store data. The control system includes two main boards, two control modules, and two non-volatile memories.

[0016] The two control modules are respectively disposed on the two main boards, are electrically connected to each other, and are both electrically connected to the power supply unit, the fan unit, and the at least one hard disk for monitoring and managing the power supply unit, the fan unit, and the at least one hard disk. The two non-volatile memories are respectively disposed on the two main boards, are respectively electrically connected to the two control modules, and both store a firmware program code.

[0017] Wherein, the two control modules are respectively initialized to read the firmware program code from the two non-volatile memories to execute corresponding firmware, so that the two control modules respectively operate in a main control mode and a controlled mode.

[0018] When an exception occurs in the control module operating in the master mode, the control module operating in the slave mode switches to operate in the master mode, and controls the control module originally operating in the master mode to switch to operate in a recovery mode, and transmits the firmware program code stored in the volatile memory corresponding to the operation in the master mode to the control module operating in the recovery mode to update the non-volatile memory corresponding to the operation in the recovery mode.

[0019] Preferably, each of the control modules uses a health signal to enable the other to determine whether an exception has occurred in the former.

[0020] Preferably, each control module includes a first bus for outputting a recovery signal and for receiving another recovery signal output by the other party. Each control module determines whether to operate in the recovery mode according to the received recovery signal.

[0021] Compared with the prior art, when an exception occurs in the control module operating in the master mode during the startup process, the present invention switches the control module operating in the slave mode to operate in the master mode, controls the control module with the exception to switch to operate in a recovery mode, and transmits the firmware program code stored in the corresponding volatile memory to the control module with the exception to update the non-volatile memory stored correspondingly, so as to be able to repair the abnormal situation of the firmware program code stored in the non-volatile memory.

[0022]

Description of the Drawings

[0023] Other features and effects of the present invention will be clearly presented in the embodiments with reference to the drawings, wherein:

[0024] Figure 1 is a block diagram illustrating an embodiment of the storage device of the present invention.

Detailed Embodiments

[0025] Before the present invention is described in detail, it should be noted that in the following description, similar components are denoted by the same reference numerals.

[0026] Refer to Figure 1, One embodiment of the storage device of the present invention includes a backplane 6, a power supply unit 7, a fan unit 8, a storage unit 9, and a control system 1. The storage device is, for example, a JBOD (Just A Bunch Of Disks); the power supply unit 7 is, for example, a power supply for providing the power required for the operation of the storage device; the fan unit 8 includes, for example, multiple fans for providing heat dissipation capacity to the storage device; the storage unit 9 includes, for example, at least one hard disk for storing data; but it is not limited thereto.

[0027] The control system 1 includes a first main board 21, a second main board 22, a first non-volatile memory 31, a second non-volatile memory 32, a first control module 41, and a second control module 42. The first main board 21 and the second main board 22 are, for example, inserted on the backplane 6. The first non-volatile memory 31 and the second non-volatile memory 32 are respectively disposed on the first main board 21 and the second main board 22, and are respectively electrically connected to the first control module 41 and the second control module 42, and both store a firmware program code.

[0028] The first control module (Input / Output module, IOM) 41 and the second control module 42 are, for example, both integrated circuit chips, and are respectively disposed on the first main board 21 and the second main board 22, and are electrically connected to each other and are both electrically connected to the power supply unit 7, the fan unit 8, and the storage unit 9 for monitoring and managing the power supply unit 7, the fan unit 8, and the storage unit 9. For example, for monitoring and managing the relevant power supply status of the power supply, the rotation speeds of the fans and the corresponding ambient temperatures, and the relevant operation status of each hard disk.

[0029] The first control module 41 and the second control module 42 are electrically connected via a first bus 411, a second bus 412, and a third bus 413 through the backplane 6. And Figure 1 For simplicity of illustration, the actual electrical connection form is not shown. The first bus 411, the second bus 412, and the third bus 413 may be directly electrically connected or indirectly electrically coupled to the first control module 41 and the second control module 42 through the backplane 6.

[0030] When the storage device is powered on, that is, when the control system 1 is powered on, the first control module 41 and the second control module 42 are respectively initialized to read the firmware program code from the first non-volatile memory 31 and the second non-volatile memory 32 to execute a firmware, so that the first control module 41 and the second control module 42 operate in an active mode and a passive mode respectively. For example, the first control module 41 operates in the active mode and the second control module 42 operates in the passive mode.

[0031] The first control module 41 (or the second control module 42) outputs a health signal to the second control module 42 through the first bus 411. When the first control module 41 (or the second control module 42) operates normally, the logical value of the corresponding output health signal will jump between logic 1 and logic 0. When the first control module 41 (or the second control module 42) operates abnormally, the logical value of the corresponding output health signal will remain at logic 0 or logic 1. Therefore, the second control module 42 (or the first control module 41) can determine whether the first control module 41 (or the second control module 42) is abnormal by the received health signal.

[0032] In this embodiment, the first bus 411 is, for example, a bus composed of connection lines between two general-purpose input / output (GPIO) pins of the first control module 41 and the second control module 42 respectively, used to transmit and receive the health signal. The first bus 411 can also be an Inter-Integrated Circuit (I2C) bus between the first control module 41 and the second control module 42 to transmit and receive the health signal. In other embodiments, the first control module 41 and the second control module 42 both further include a buffer. When the first control module 41 (or the second control module 42) operates normally, the value of the buffer of the second control module 42 (or the first control module 41) is changed over time by the corresponding health signal. When the first control module 41 (or the second control module 42) operates abnormally, the value of the buffer of the second control module 42 (or the first control module 41) cannot be changed over time by the corresponding health signal, so that the second control module 42 (or the first control module 41) can determine whether the first control module 41 (or the second control module 42) is abnormal according to the change in the value of the corresponding buffer.

[0033] When an exception occurs during the operation of the first control module 41 operating in the master mode, the second control module 42 operating in the slave mode is converted to operate in the master mode according to the judgment, and controls the first control module 41 originally operating in the master mode to be converted to operate in a recovery mode. More specifically, the first control module 41 (or the second control module 42) outputs the recovery signal to the second control module 42 (or the first control module 41) through the second bus 412. The first control module 41 (or the second control module 42) determines whether to operate in the recovery mode according to the received recovery signal, for example, according to the logical value of the recovery signal. In this embodiment, the second bus 412 is, for example, a connection line between two general-purpose input / output pins of the first control module 41 and the second control module 42 or another integrated circuit (I2C) bus between the first control module 41 and the second control module 42, but not limited thereto.

[0034] Next, the second control module 42 operating in the master mode transmits the firmware program code stored in the corresponding volatile memory 32 to the first control module 41 operating in the recovery mode, so that the first control module 41 in the recovery mode updates (i.e., rewrites) the received firmware program code and stores it in the corresponding non-volatile memory 31, and reloads the updated firmware program code to execute the firmware again, and is converted to operate in the slave mode, and then can operate normally.

[0035] In this embodiment, the first control module 41 (or the second control module 42) transmits or receives the firmware program code through the third bus 413. In this embodiment, the third bus 413 is another integrated circuit (I2C) bus or an Intelligent Platform Management Interface (IPMI) bus, but not limited thereto. In addition, in other embodiments, the first control module 41 may also be defaulted to operate in the master mode, and the second control module 42 is defaulted to operate in the slave mode. When the first control module 41 has an exception and is changed to operate in the recovery mode, and receives the firmware program code from the second control module 42 to update and store it in the corresponding non-volatile memory, and then re-executes the firmware and operates normally, the first control module 41 will be changed back to operate in the master mode, and the second control module 42 that is changed to operate in the master mode due to the exception of the first control module 41 will be changed back to operate in the default slave mode at the same time.

[0036] It should be particularly emphasized that: when the first control module 41 or the second control module 42 malfunctions, the normally operating first control module 41 or the second control module 42 transmits the firmware program code corresponding to the "firmware being executed by itself". In addition, both main boards (i.e., the first main board 21 and the second main board 22) only need to be respectively provided with a single non-volatile memory (i.e., the first non-volatile memory 31 and the second non-volatile memory 32) for storing the firmware program code corresponding to the firmware being executed by itself, and only one copy of the firmware program code needs to be stored in the single non-volatile memory, which can not only be provided for the control modules on the two main boards (i.e., the first control module 41 and the second control module 42) to execute, but also achieve the function of mutually repairing the firmware.

[0037] In addition, it should be particularly supplemented and explained that: the first control module 41 operating in the master mode and the second control module 42 operating in the slave mode will both monitor the power supply unit 7, the fan unit 8, and the storage unit 9 and respectively record and store the monitored information obtained from the monitoring, so as to be able to switch the master mode / the slave mode in real time and provide a redundancy function in real time when one of them malfunctions. However, only the first control module 41 operating in the master mode will manage (i.e., control) the power supply unit 7, the fan unit 8, and the storage unit 9 according to the monitored information obtained from the monitoring.

[0038] Furthermore, it should be particularly emphasized that: the first control module 41 operating in the master mode will also monitor the health signal of the second control module 42. For example, when the first control module 41 detects that the second control module 42 is malfunctioning, it will also transmit the restoration signal to the second control module 42, and then transmit the firmware program code corresponding to the firmware being executed by the first control module 41 itself to the second control module 42. After receiving the firmware program code transmitted by the first control module 41, the second control module 42 updates the firmware program code in its corresponding second non-volatile memory 32 with the received firmware program code, and after the firmware update is completed, it executes the updated firmware again to switch from the restoration mode to the slave mode according to an identification code of the second main board 22. Among them, since the first control module 41 is already operating in the master mode, when the first control module 41 monitors that the second control module 42 is malfunctioning, the first control module 41 does not switch to operate in the slave mode, but still continues to operate in the master mode.

[0039] In summary, when an abnormality occurs during the operation of the first control module 41 (or the second control module 42) in the main control mode, the second control module 42 (or the first control module 41) operating in the controlled mode is used to switch to the main control mode, and the first control module 41 (or the second control module 42) with the abnormality is controlled to switch to the restore mode. Moreover, the firmware program code stored in the corresponding volatile memory is transmitted to the first control module 41 (or the second control module 42) with the abnormality. For the first control module 41 (or the second control module 42) with the abnormality in the restore mode, the received firmware program code is automatically updated and stored in the corresponding first non-volatile memory 31 (or the second non-volatile memory 32). Furthermore, because the first control module 41 (or the second control module 42) with the abnormality is in the restore mode, after storing the received firmware program code, it automatically restarts itself and loads and executes the firmware corresponding to the stored firmware program code to switch the operation to the originally default mode, thus being able to repair the abnormal situation of the firmware program code stored in the first non-volatile memory 31 (or the second non-volatile memory 32). Therefore, the object of the present invention can indeed be achieved.

[0040] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A control system for a storage device, applicable to a power supply unit, a fan unit, and at least one hard disk, characterized in that, and includes: two main boards; two control modules, respectively disposed on the two main boards, electrically connected to each other and both electrically connected to the power supply unit, the fan unit, and the at least one hard disk, for monitoring and managing the power supply unit, the fan unit, and the at least one hard disk; and two non-volatile memories, respectively disposed on the two main boards, respectively electrically connected to the two control modules, and both storing a firmware program code, wherein, the two control modules are respectively initialized to respectively read the firmware program code from the two non-volatile memories to execute a corresponding firmware, so that the two control modules respectively operate in a main control mode and a controlled mode, when the control module operating in the main control mode has an abnormality, the control module operating in the controlled mode is converted to operate in the main control mode, and controls the control module originally operating in the main control mode to be converted to operate in a recovery mode, and transmits the firmware program code stored in the non-volatile memory corresponding to the control module operating in the main control mode to the control module operating in the recovery mode, to update the non-volatile memory stored in the control module corresponding to operating in the recovery mode.

2. The control system of the storage device according to claim 1, characterized in that Each of the two control modules uses a health signal to enable the other to determine whether an abnormality has occurred to each of the foregoing.

3. The control system of the storage device according to claim 2, characterized in that, Each control module includes a first bus for outputting the health signal and for receiving the health signal output by the other party. When any one of the two control modules operates normally, the logical value of the corresponding output health signal will jump between logic 1 and logic 0, and when any one of the two control modules operates abnormally, the logical value of the corresponding output health signal will remain at logic 0 or logic 1.

4. The control system of the storage device according to claim 2, characterized in that, Each control module outputs the health signal through a first bus and both include a buffer. When one of the two control modules operates normally, the value of the buffer of the other is changed over time by using the corresponding health signal, and when one of the two control modules operates abnormally, the value of the buffer of the other cannot be changed over time by using the corresponding health signal.

5. The control system of the storage device according to claim 1, characterized in that, Each control module includes a second bus for outputting a recovery signal and for receiving another recovery signal output by the other party. Each control module determines whether to operate in the recovery mode according to the received recovery signal.

6. The control system of the storage device according to claim 1, wherein, Each control module includes a third bus for transmitting the firmware program code and for receiving the firmware program code transmitted by the other party.

7. The control system of the storage device according to claim 1, characterized in that The two control modules are respectively defined as a first control module and a second control module. The first control module is default to operate in the main control mode, and the second control module is default to operate in the controlled mode. When the first control module has an abnormality and is changed to operate in the recovery mode, and receives the firmware program code from the second control module to update the non-volatile memory stored correspondingly, and then re-executes the firmware and operates normally, the first control module is changed to operate in the main control mode and the second control module is changed to operate in the controlled mode.

8. A storage device, characterized in that, Includes: A power supply unit for providing operating power; A fan unit for providing heat dissipation; At least one hard disk for storing data; And A control system, comprising: Two mainboards; Two control modules respectively disposed on the two mainboards, electrically connected to each other and both electrically connected to the power supply unit, the fan unit, and the at least one hard disk for monitoring and managing the power supply unit, the fan unit, and the at least one hard disk; and Two non-volatile memories respectively disposed on the two mainboards, respectively electrically connected to the two control modules, and both storing a firmware program code, wherein the two control modules are respectively initialized to respectively read the firmware program code from the two non-volatile memories to execute corresponding firmware, such that the two control modules respectively operate in a master mode and a slave mode, when the control module operating in the master mode encounters an abnormality, the control module operating in the slave mode is converted to operate in the master mode, and controls the control module originally operating in the master mode to be converted to operate in a recovery mode, and transmits the firmware program code stored in the non-volatile memory corresponding to the master mode to the control module operating in the recovery mode to update the non-volatile memory corresponding to the recovery mode.

9. The storage device according to claim 8, characterized in that, Each of the two control modules uses a health signal to enable the other to determine whether the foregoing each has an abnormality.

10. The storage device according to claim 8, wherein Each of the control modules includes a first bus for outputting a recovery signal and for receiving another recovery signal output by the other party, and each of the control modules determines whether to operate in the recovery mode according to the received recovery signal.

Citation Information

Patent Citations

  • Fault automatic recovery system and method

    CN101211281A

  • System and electronic device provided with firmware updating function and firmware updating method of system

    CN103176857A