Power supply backup system of storage cluster, control method, server and storage medium
By introducing a third backup power supply and controller into the storage cluster, the problem of unstable power supply of the controller in the event of backup power failure is solved, redundant power supply is achieved, and the reliability and data security of the storage cluster are improved.
Patent Information
- Application Number
- CN202510732770.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing storage cluster cannot effectively ensure the controller's continuous and reliable power supply and cluster performance when the backup power fails, and lacks redundancy and fault tolerance, which poses great safety hazards.
The third backup power supply and the third controller are introduced, and the power supply link between the third backup power supply and the controller is turned on through the switch assembly when the backup power supply fails, thereby realizing redundant power supply and ensuring that at least one controller works normally.
Improve the reliability of the storage cluster and user data security, ensuring the continuous and reliable power supply and cluster performance of the controller.
Smart Images

Figure CN120255682A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technologies, and in particular, to a power supply backup system, a control method, a server, and a storage medium for a storage cluster. Background Art
[0002] The BBU (Battery Backup Unit) on a storage device provides backup power support for each controller in a storage cluster. It can immediately start BBU power supply when the external power supply of the cluster system fails, so as to ensure that the data in the storage system cache is safely written into the disk array when the external power supply fails, improve the high reliability of the storage system, and guarantee the security of user service data. It is an essential component in a storage cluster.
[0003] In existing storage clusters, the BBU module and the controller adopt a one-to-one dedicated power supply mode. Each controller depends on its dedicated BBU to provide services. When the dedicated BBU of a certain controller fails and does not recover for a long time, the controller will be degraded, resulting in the loss of redundancy of the cluster and affecting system stability and data reliability. Although the cluster management system periodically detects the health status of the BBU and reports a first-level alarm at the initial stage of the BBU failure to remind maintenance, if the failure lasts for more than 14 days or another BBU also shows abnormalities, it will be upgraded to a second-level alarm, triggering controller degradation and further exacerbating system risks. The existing mechanism cannot guarantee that at least one controller continues to work in case of dual BBU failures, posing a significant security hazard. Summary of the Invention
[0004] This application provides a power supply backup system, a control method, a server, and a storage medium for a storage cluster to at least solve the technical problem that in the related art, the continuous and reliable power supply of the controller and the cluster performance cannot be effectively guaranteed in case of backup power failure, and there is a lack of redundancy and fault tolerance capabilities.
[0005] This application provides a power supply backup system for a storage cluster. The storage cluster includes a first controller and a second controller. Among them, the system includes: a first backup power supply, a second backup power supply, and a third backup power supply. The first backup power supply is connected to the first controller, the second backup power supply is connected to the second controller, and the third backup power supply is respectively connected to the first controller and the second controller; a switch component for conducting or disconnecting the power supply link of the first controller and the second controller; a third controller for controlling the switch component to conduct the power supply link between the third backup power supply and one of the first controller and the second controller when at least one of the first backup power supply and the second backup power supply fails.
[0006] This application also provides a server including the power supply backup system for the storage cluster described above.
[0007] The present application also provides a control method for a power supply backup system of a storage cluster. The method is applied to a third controller in the power supply backup system of the above storage cluster. The method includes: obtaining the health data of the storage cluster; judging whether the first backup power supply and the second backup power supply are faulty according to the health data; when at least one of the first backup power supply and the second backup power supply is faulty, controlling the switch component to conduct the power supply link between the third backup power supply and one of the first controller and the second controller.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the control method for any one of the above power supply backup systems of the storage cluster are implemented.
[0009] Through the present application, since it includes the first to third backup power supplies, the switch component, and the third controller, when at least one of the first backup power supply and the second backup power supply is faulty, the third controller controls the switch component to conduct the power supply link between the third backup power supply and one of the first controller and the second controller, so as to ensure that one controller can still work normally when one or two backup power supplies are faulty, solving the technical problem that in the related art, the continuous and reliable power supply of the controller and the cluster performance cannot be effectively guaranteed when the backup power supply is faulty, and the lack of redundancy and fault tolerance capabilities, achieving the technical effects of improving the reliability of the storage cluster and the data security of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 FIG. is a schematic structural diagram of a power supply backup system of a storage cluster according to an embodiment of the present application; Figure 2 FIG. is a flow chart of the global hot standby alarm reporting of the BBU of the storage system according to an embodiment of the present application; Figure 3 FIG. is a flow chart for processing a fault of a dedicated BBU according to an embodiment of the present application; Figure 4 FIG. is a flow chart for processing faults of two dedicated BBUs simultaneously according to an embodiment of the present application; Figure 5 FIG. is a schematic flow chart of the control method for the power supply backup system of the storage cluster according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0013] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0014] In the related art, the BBU module in the storage cluster and the controller in the storage cluster adopt a one-to-one service mode. For a dual-controller cluster, each storage chassis provides two BBU modules, which respectively correspond to two controllers, that is, each controller is equipped with a dedicated BBU. When the dedicated BBU of a certain controller fails, the controller will completely lose the BBU service, which poses a risk to the storage cluster.
[0015] At the same time, in order to ensure the stability and reliability of the operation of the storage system, the cluster management system will monitor the health status of the BBU module in real time. When the BBU module shows an abnormality due to external factors or its own factors, the cluster management system will report an alarm for different abnormal situations to remind the user that the dedicated BBU of a certain controller in the storage cluster has an abnormal situation and needs to be checked or replaced in time.
[0016] The management process of the BBU by the cluster management system is as follows: During the operation of the BBU, the storage system periodically detects the health status of the dedicated BBU of the two controllers in the same chassis. When the BBU status is normal, the timestamp of the fault status flag bit of the BBU is 0. If the cluster management system detects that the timestamp of the BBU fault status is 0, it is considered that the BBU health status is normal, and the next cycle of polling detection continues; when it is detected that the BBU (e.g., BBU_A) of a certain controller (e.g., controller A) is in an abnormal state, and the BBU (e.g., BBU_B) of the other controller (e.g., controller B) is in a normal state, the timestamp of the detection moment will be marked on the fault status flag bit of BBU_A to prevent false alarms caused by human operations. For example, if BBU_A is manually unplugged, the cluster management system will not immediately report an alarm after detecting the timestamp of the BBU_A fault status flag bit. If the fault is restored within 300 seconds, the cluster management system will clear the timestamp of the BBU_A fault status flag bit and continue the next cycle of polling detection; if the fault still exists after more than 300 seconds, and at the same time BBU_B is detected to be in a normal state, the cluster will report a first-level alarm for BBU_A. At this time, BBU_A cannot provide services, but the status of controller A is still normal and can provide services to the cluster. Users can check and replace BBU_A without the controller_A being downgraded, which ensures the redundancy of the controller for the storage cluster; if the cluster management system detects that the first-level alarm has not been restored for more than 14 days, or if BBU_B is detected to be out of position or there is an alarm for BBU_B after the first-level alarm is reported, the first-level alarm will be converted into a second-level alarm. At the same time, the cluster management system will downgrade controller A. After downgrading, controller A cannot provide services, and only one controller B provides services to the storage cluster. The storage cluster loses controller redundancy, which poses a great risk to the user's business data. Users must replace BBU_A in time to ensure the controller redundancy of the cluster.
[0017] Therefore, the related technology cannot effectively guarantee the continuous and reliable power supply of the controller and the cluster performance when the backup power supply fails, and lacks redundancy and fault tolerance capabilities.
[0018] In view of the deficiencies of the related technology, the present application provides a power supply backup system, a control method, a server and a storage medium for a storage cluster, so as to realize that when one or two backup power supplies fail, a controller can still work normally, and solve the technical problems that the related technology cannot effectively guarantee the continuous and reliable power supply of the controller and the cluster performance when the backup power supply fails, and lacks redundancy and fault tolerance capabilities. The specific solution will be described in detail below.
[0019] In order to enable those skilled in the art of this technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] Specifically,Figure 1 The structural schematic diagram of a power supply backup system for a storage cluster provided by an embodiment of the present application is as follows Figure 1 As shown in the figure, the storage cluster 10 includes a first controller 101 and a second controller 102. The power supply backup system 11 of the storage cluster includes: a first backup power supply 111, a second backup power supply 112, a third backup power supply 113, a switch assembly 114, and a third controller 115.
[0021] Among them, the first backup power supply 111 is connected to the first controller 101, the second backup power supply 112 is connected to the second controller 102, and the third backup power supply 113 is connected to both the first controller 101 and the second controller 102; the switch assembly 114 is used to conduct or disconnect the power supply link of the first controller 101 and the second controller 102; the third controller 115 is used to control the switch assembly 114 to conduct the power supply link between the third backup power supply 113 and one of the first controller 101 and the second controller 102 when at least one of the first backup power supply 111 and the second backup power supply 112 fails.
[0022] Among them, the first backup power supply 111, the second backup power supply 112, and the third backup power supply 113 are all BBU, which are energy storage modules used to provide temporary power support for devices when the main power supply fails, and are commonly used in storage systems to protect cache data from being lost; the third backup power supply 113 is deployed inside the storage chassis and hot-stands by the first backup power supply 111 and the second backup power supply 112, that is, it is in a state of being available at any time.
[0023] It can be understood that the first backup power supply 111 of the embodiment of the present application is connected to the first controller 101, the second backup power supply 112 is connected to the second controller 102, and the third backup power supply 113 is connected to both the first controller 101 and the second controller 102. As redundant backups, the switch assembly 114 is used to conduct or disconnect the power supply link between the first controller 101 and the second controller 102. When the first backup power supply 111 or the second backup power supply 112 fails, the third controller 115 is responsible for detecting and judging the fault state, and controlling the switch assembly 114 to conduct the power supply link between the third backup power supply 113 and the affected controller, so as to realize continuous power supply to the first controller 101 or the second controller 102, ensuring the high availability and reliability of the system. The third backup power supply 113 is deployed inside the storage chassis and hot-stands by the first backup power supply 111 and the second backup power supply 112 to ensure that the normal operation of the controller can still be maintained when any one of the backup power supplies fails.
[0024] In the embodiment of the present application, the switch assembly 114 includes a first switch, a second switch, a third switch, and a fourth switch. Among them, the first switch is disposed on the first power supply link between the first backup power supply 111 and the first controller 101, the second switch is disposed on the second power supply link between the second backup power supply 112 and the second controller 102, the third switch is disposed on the third power supply link between the third backup power supply 113 and the first controller 101, and the fourth switch is disposed on the fourth power supply link between the third backup power supply 113 and the second controller 102.
[0025] Among them, the first power supply link refers to the power transmission path connecting the first backup power supply 111 and the first controller 101, the second power supply link refers to the power transmission path connecting the second backup power supply 112 and the second controller 102, the third power supply link refers to the power transmission path connecting the third backup power supply 113 and the first controller 101, and the fourth power supply link refers to the power transmission path connecting the third backup power supply 113 and the second controller 102.
[0026] It can be understood that the switch assembly 114 in the embodiment of the present application is composed of the first to fourth switches and is used to control the first to fourth power supply links. Through the coordinated control of these four switches, when any backup power supply fails, the third backup power supply 113 can provide emergency power supply support for any controller, thereby improving the redundancy and reliability of the system.
[0027] In the embodiment of the present application, the third controller 115 is used to control the first switch to conduct the first power supply link and the second switch to conduct the second power supply link when the external power supply of the storage cluster 10 fails.
[0028] It can be understood that when the external power supply of the storage cluster 10 fails in the embodiment of the present application, after the third controller 115 detects the failure, it controls the first switch to conduct the first power supply link, so that the first backup power supply 111 provides power for the first controller 101, and at the same time controls the second switch to conduct the second power supply link, so that the second backup power supply 112 supplies power for the second controller 102, thereby ensuring that the two controllers can continue to operate normally relying on their respective dedicated backup power supplies in the case of a main power interruption.
[0029] In the embodiment of the present application, the third controller 115 is used to, when any one of the first backup power supply 111 and the second backup power supply 112 fails, if the failed backup power supply recovers from the failure, conduct the power supply link between the recovered backup power supply and the corresponding controller, and disconnect the power supply link between the third backup power supply 113 and the corresponding controller.
[0030] Among them, the recovery from the failure means that the originally failed backup power supply has returned to the normal working state and has the ability to continue to supply power to the corresponding controller.
[0031] It can be understood that after any one of the first backup power supply 111 or the second backup power supply 112 fails in the embodiment of the present application, the third controller 115 will intervene in the control, conduct the power supply link between the third backup power supply 113 and the corresponding controller to maintain its continuous operation. If the failed backup power supply is restored later and has the normal power supply capacity, the third controller 115 will control to re-conduct the power supply link between the restored backup power supply and its corresponding controller, and at the same time disconnect the power supply link between the third controller 115 and this controller, so as to realize the automatic switchback from the hot standby power supply to the original backup power supply, restore the original power supply relationship and maintain the system redundancy. For example, if the first backup power supply 111 is manually unplugged and then reinstalled after a period of time, during the period when the first backup power supply 111 is absent, the third controller 115 controls the third backup power supply 113 to supply power to the first controller 101 instead of the first backup power supply 111. When the first backup power supply 111 is reinstalled and the fault is restored, the power supply link between the first backup power supply 111 and the first controller 101 is conducted, and the power supply link between the third backup power supply 113 and the first controller 101 is disconnected.
[0032] In the embodiment of the present application, the third controller 115 is used to conduct the power supply link between the controller corresponding to the backup power supply with an earlier fault time and the third backup power supply 113 when both the first backup power supply 111 and the second backup power supply 112 fail; if the fault times of the first backup power supply 111 and the second backup power supply 112 are the same, the target controller is determined according to the priorities of the first controller 101 and the second controller 102, and the power supply link between the target controller and the third backup power supply 113 is conducted.
[0033] Among them, the priorities of the first controller 101 and the second controller 102 are determined by the configuration node of the storage cluster 10. The configuration node is the core controller node with the highest authority in the storage cluster 10, responsible for global coordination and management, and the controller where the configuration node is located has a higher priority.
[0034] It can be understood that in the embodiments of the present application, when both the first backup power supply 111 and the second backup power supply 112 fail, the third controller 115 determines the power supply object of the third backup power supply 113 according to the difference in the failure time of the two or the controller priority. Specifically, if the failure times of the two backup power supplies are different, the power supply link between the controller corresponding to the backup power supply with the earlier failure time and the third backup power supply 113 is preferentially conducted; if the two backup power supplies fail simultaneously, it is judged according to the priority of the first controller 101 and the second controller 102, and the power supply link of the third backup power supply 113 is conducted to the target controller with the higher priority. For example, when the backup power supplies of the two controllers fail simultaneously and an alarm is triggered, if the first controller 101 includes a configuration node, the third backup power supply 113 is preferentially bound to the first controller 101 to ensure the normal operation of the cluster management function.
[0035] In the embodiments of the present application, the third controller 115 is configured to generate a primary fault alarm for the faulty backup power supply when at least one of the first backup power supply 111 and the second backup power supply 112 fails.
[0036] Among them, the primary fault alarm is an alarm type with a relatively high set fault level, which is used to prompt the user to process the abnormal backup power supply problem in time.
[0037] It can be understood that in the embodiments of the present application, when at least one of the first backup power supply 111 and the second backup power supply 112 fails, the third controller 115 detects the fault status and generates a primary fault alarm for the corresponding faulty backup power supply to remind the user, such as the system administrator, to perform maintenance or replacement operations in time, so as to ensure the backup power supply capacity of the controller and the high availability of the storage system.
[0038] In the embodiments of the present application, the third controller 115 is configured to, when both the first backup power supply 111 and the second backup power supply 112 fail, if the failure duration of any faulty backup power supply is greater than a preset duration, generate a secondary fault alarm for the faulty backup power supply and restrict the service permission of the controller corresponding to the faulty backup power supply of the secondary fault alarm. Among them, the alarm method of the secondary fault alarm is different from that of the primary fault alarm.
[0039] Among them, the preset duration is specifically set according to actual requirements and will not be specifically limited here.
[0040] It can be understood that when both the first backup power supply 111 and the second backup power supply 112 fail in the embodiments of the present application, the third controller 115 continuously monitors the failure time of each failed backup power supply. If the failure time of any one of the failed backup power supplies exceeds a preset duration, the third controller 115 will generate a corresponding secondary fault alarm and impose service permission restrictions on the controller corresponding to the secondary fault alarm to prevent it from continuing to provide services externally to avoid potential data or performance risks. Among them, the alarm method of the secondary fault alarm is different from that of the primary fault alarm, with a higher severity identifier and a more urgent notification mechanism, so that the operation and maintenance personnel can identify and handle key problems in a timely manner.
[0041] In the embodiments of the present application, the third controller 115 is used to periodically detect the health status when the failed backup power supply recovers from the failure. If the health status of the failed backup power supply meets the health conditions after a preset duration, the power supply link between the failed backup power supply and the corresponding controller is turned on, and the power supply link between the corresponding controller and the third backup power supply 113 is disconnected.
[0042] It can be understood that when a certain failed backup power supply recovers from the failure in the embodiments of the present application, the third controller 115 starts to periodically detect the health status of the backup power supply. If after a preset duration, the health status of the failed backup power supply is confirmed to meet the health conditions, the third controller 115 will turn on the power supply link between the failed backup power supply and its corresponding controller, and at the same time disconnect the power supply link between the controller and the third backup power supply 113, so as to realize the automatic switchback from the shared hot standby power supply to the original dedicated backup power supply, restore the original power supply relationship and release the shared resources for other controllers to use.
[0043] According to the power supply backup system of the storage cluster provided by the embodiments of the present application, it includes a first backup power supply 111, a second backup power supply 112, a third backup power supply 113, a switch assembly 114, and a third controller 115. When at least one of the first backup power supply 111 and the second backup power supply 112 fails, the third controller 115 controls the switch assembly 114 to turn on the power supply link between the third backup power supply 113 and one of the first controller 101 and the second controller 102, so as to ensure that one controller can still work normally when one or two backup power supplies fail, achieving the technical effects of improving the reliability of the storage cluster and the data security of users.
[0044] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0045] The power supply backup system of the storage cluster will be further described below through a specific embodiment.
[0046] This embodiment proposes a BBU global hot standby method based on a storage system. In the storage cluster architecture, in addition to the dedicated backup power supplies, the first backup power supply and the second backup power supply, corresponding to the first controller and the second controller respectively, a BBU, that is, the third backup power supply, is introduced. The third backup power supply is deployed in the storage chassis and hot stands by the dedicated BBU of each of the two controllers in the chassis. At the same time, the storage cluster will also perform periodic health status detection on the third backup power supply. When a fault of a dedicated BBU of a certain controller is reported and an alarm is issued, the third backup power supply immediately replaces the faulty BBU to provide services for the controller. At this time, the third backup power supply becomes the dedicated BBU of the controller. When the fault of the dedicated BBU of the controller is recovered, the third backup power supply exits the belonging controller and reverts to the third backup power supply again.
[0047] The third backup power supply is deployed in the storage chassis, independent of the first controller and the second controller, and is connected to the first controller and the second controller through the shared power supply board in the chassis.
[0048] The cluster management system assigns an independent logical identifier (such as a global ID) to the third backup power supply, which is different from the local ID of the dedicated BBU of the controller; a priority policy for the third backup power supply is set in the cluster management system. When the status of the dedicated BBU of the controller is normal, the third backup power supply is not bound to any controller and defaults to the global hot standby state.
[0049] The cluster management system periodically detects the health status of all BBUs (the first backup power supply, the second backup power supply, and the third backup power supply). When a fault of a dedicated BBU of a certain controller is detected and a first-level alarm is reported, the cluster management system immediately marks the dedicated BBU of the controller as an untrusted state, and at the same time issues a BBU switching instruction to bind the logical identifier of the third backup power supply to the controller, so as to realize the replacement of the faulty dedicated BBU by the third backup power supply and provide services for the controller.
[0050] Before the fault of the dedicated BBU is repaired or a new dedicated BBU is replaced, the third backup power supply always provides services for the controller.
[0051] Before the exclusive BBU failure of the failed controller is repaired or a new exclusive BBU is replaced, the first-level alarm for the exclusive BBU failure on the cluster management system will always exist. In this scenario, if the first-level alarm exists for more than 14 days, since there is a third backup power supply to provide services for the controller, the first-level alarm will no longer be converted to a second-level alarm, and the cluster management system will not set the controller to the degraded state, ensuring the controller redundancy of the storage cluster. Until the failure is repaired or a new exclusive BBU is replaced, after the cluster management system detects that the status of the exclusive BBU is normal, the alarm is automatically corrected. The alarm reporting process of the cluster management system for the BBU is as Figure 2 shown.
[0052] As Figure 2 shown, the alarm reporting process of the cluster management system for the BBU is as follows: The cluster management system monitors the health status of the BBU. If a BBU failure is detected, it reads the BBU timestamp. When the timestamp is 0, it starts timing and continues to monitor the BBU health status. If the timing exceeds 300 seconds, and the failed BBU is an exclusive BBU and the peer BBU is normal, then the third backup power supply binds to the controller to which the failed BBU belongs and supplies power to it, while reporting a first-level alarm. If the timing exceeds 14 days, the first-level alarm is upgraded to a second-level alarm.
[0053] After the exclusive BBU failure of the controller is restored, when the cluster management system detects that the status of the exclusive BBU has returned to normal, to prevent false status reports, the cluster management system will not immediately issue a switching instruction. After continuing to periodically detect the health status of all BBUs for 300 seconds and confirming that the health status of the exclusive BBU is still normal, it issues a BBU switching instruction. The third backup power supply exits the exclusive state of the controller and switches to the global hot standby state, no longer providing services for a single controller, and the original controller-exclusive BBU resumes serving the controller. The processing flow after a single exclusive BBU failure is as Figure 3 shown.
[0054] As Figure 3 shown, when a single exclusive BBU fails and the third backup power supply is normal, if the failure exceeds 300 seconds, the failed exclusive BBU is marked as an untrusted state. The third backup power supply switches to the exclusive mode to provide services for the controller and continues to monitor the BBU health status. If the failed exclusive BBU returns to normal, it continues to monitor the BBU health status for 300 seconds to confirm that the failed exclusive BBU has indeed returned to normal and can resume serving the controller. The third backup power supply exits the exclusive mode and returns to the global hot standby state. If a failure is detected within 300 seconds, the failed exclusive BBU is marked as an untrusted state, and the third backup power supply switches to the exclusive mode to provide services for the controller.
[0055] When a dedicated BBU of a controller fails and a first-level alarm is reported, if the dedicated BBU of another controller also fails and reports a first-level alarm when using the third backup power supply to replace the failed dedicated BBU and provide services for the controller, the cluster management system processes it according to the existing logic. That is, the cluster management system first reports a first-level alarm for the dedicated BBU of another controller. If the first-level alarm is repaired within 14 days after reporting, the alarm is eliminated; if the first-level alarm is not repaired within 14 days after reporting, the first-level alarm is converted to a second-level alarm, and the cluster management system degrades the controller.
[0056] When the dedicated BBUs of two controllers fail simultaneously and report alarms, the cluster management system will preferentially bind the third backup power supply to the controller where the cluster configuration node is located (i.e., the BOSS node with the highest authority of the cluster management system) according to the role of the controller in the storage cluster. The cluster management system processes the alarm of the other controller (non-BOSS node) according to the existing logic. The processing flow for the simultaneous failure of two dedicated BBUs is as Figure 4 shown.
[0057] As Figure 4 shown, when two dedicated BBUs fail and the third backup power supply is normal, if the failure exceeds 300 seconds, the two failed dedicated BBUs are marked as untrusted states. The third backup power supply switches to the dedicated mode to provide services for the configuration node and continues to monitor the health status of the BBU. If the dedicated BBU of the failed configuration node returns to normal, continue to monitor the health status of this BBU for 300 seconds. If a failure is detected within 300 seconds, continue to control the third backup power supply to switch to the dedicated mode to provide services for the configuration node. If it is confirmed that the dedicated BBU that has failed has indeed returned to normal after 300 seconds and can provide services for the controller again, after the third backup power supply exits the dedicated mode and returns to the global hot standby state, it switches to the dedicated mode again to provide services for non-configuration nodes; if the dedicated BBU of the non-configuration node returns to normal, continue to monitor the health status of this BBU for 300 seconds. If a failure is detected within 300 seconds, continue to control the third backup power supply to switch to the dedicated mode to provide services for the controller. If it is confirmed that the status of the dedicated BBU of the non-configuration node has returned to normal after 300 seconds and provide services for the non-configuration node again, the status of the third backup power supply remains unchanged and still provides services for the configuration node in the dedicated mode.
[0058] When all three BBUs (the first backup power supply, the second backup power supply, and the all-hot standby BBU) in the storage device fail, the cluster management system will consider that a very serious failure has occurred in the storage cluster, and will perform a downgrade operation on both nodes to ensure the security of the data that has been written and will no longer accept newly written data.
[0059] The BBU global hot standby method based on the storage system proposed in this embodiment introduces a global hot standby BBU in the storage cluster and connects it to two controllers in the storage chassis through a shared power supply board. At the same time, when a dedicated BBU in the storage cluster fails, the global hot standby BBU will automatically replace the faulty BBU to provide services for the controller, ensuring the security of the controller. In the extreme scenario where two dedicated BBUs in the storage cluster fail simultaneously, the global hot standby BBU will automatically replace the dedicated BBU of the cluster BOSS node and provide services for the BOSS node, avoiding the cluster from entering the degraded mode and being unable to provide services to users due to the simultaneous failure of two dedicated BBUs, and improving the high reliability of the storage cluster.
[0060] An embodiment of the present application also provides a server, including the power supply backup system of the above storage cluster.
[0061] An embodiment of the present application also provides a control method for a power supply backup system of a storage cluster. Figure 5 As shown in the flowchart of the control method for the power supply backup system of the storage cluster provided by the embodiment of the present application, this method is applied to the third controller in the power supply backup system of the above storage cluster, as Figure 5 shown, this method includes the following steps: In step S201, obtain the health data of the storage cluster.
[0062] Among them, the health data is a data set used to reflect the operating status, performance indicators, and reliability level of the entire storage cluster or some components. The health data may include, but is not limited to, controller status, backup power supply status, etc., and is usually implemented through monitoring modules, management interfaces, or data collection tools.
[0063] In step S202, determine whether the first backup power supply and the second backup power supply are faulty according to the health data.
[0064] It can be understood that in the embodiment of the present application, by analyzing the obtained health data of the storage cluster, it is determined whether the status indicators related to the first backup power supply and the second backup power supply meet the normal operating conditions, so as to identify whether these two backup power supplies are faulty. If the health data reflects abnormal situations, such as too low voltage, communication failure, insufficient capacity, etc., it is determined that the corresponding backup power supply is in a faulty state.
[0065] In step S203, when at least one of the first backup power supply and the second backup power supply is faulty, control the switch component to conduct the power supply link between the third backup power supply and one of the first controller and the second controller.
[0066] It is understandable that when at least one of the first backup power supply and the second backup power supply fails in the embodiment of the present application, the switch assembly is controlled to selectively conduct the power supply link between the third backup power supply and the failed controller, so that the controller can still obtain power support in the case of the failure of the dedicated backup power supply, thereby maintaining the continuous operation of the storage cluster and the controller redundancy capability.
[0067] According to the control method of the power supply backup system of the storage cluster provided by the embodiment of the present application, including the first to third backup power supplies, a switch assembly, and a third controller, when at least one of the first backup power supply and the second backup power supply fails, the third controller is used to control the switch assembly to conduct the power supply link between the third backup power supply and one of the first controller and the second controller, so as to ensure that one controller can work normally when one or two backup power supplies fail, achieving the technical effects of improving the reliability of the storage cluster and the data security of users.
[0068] For the description of the features in the corresponding embodiment of the control method of the power supply backup system of the storage cluster, reference can be made to the relevant description of the corresponding embodiment of the power supply backup system of the storage cluster, which will not be elaborated here one by one.
[0069] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in the embodiment of the control method of the power supply backup system of the storage cluster when running.
[0070] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0071] The embodiment of the present application also provides a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of the control method of the power supply backup system of the storage cluster.
[0072] The embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of the control method of the power supply backup system of the storage cluster.
[0073] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0074] The above has introduced in detail a power supply backup system, a control method, a server, and a storage medium of a storage cluster provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A power supply backup system for a storage cluster, characterized in that The storage cluster includes a first controller and a second controller. Among them, the system includes: A first backup power supply, a second backup power supply, and a third backup power supply. The first backup power supply is connected to the first controller, the second backup power supply is connected to the second controller, and the third backup power supply is connected to both the first controller and the second controller; A switch assembly for conducting or disconnecting the power supply links of the first controller and the second controller; A third controller for controlling the switch assembly to conduct the power supply link between the third backup power supply and one of the first controller and the second controller when at least one of the first backup power supply and the second backup power supply fails.
2. The power supply backup system of the storage cluster according to claim 1, wherein The switch assembly includes a first switch, a second switch, a third switch, and a fourth switch. Among them, the first switch is arranged on the first power supply link between the first backup power supply and the first controller, the second switch is arranged on the second power supply link between the second backup power supply and the second controller, the third switch is arranged on the third power supply link between the third backup power supply and the first controller, and the fourth switch is arranged on the fourth power supply link between the third backup power supply and the second controller.
3. The power supply backup system of the storage cluster according to claim 2, wherein The third controller is used to control the first switch to conduct the first power supply link and the second switch to conduct the second power supply link when the external power supply of the storage cluster fails.
4. The power supply backup system for a storage cluster according to claim 1, wherein The third controller is used to conduct the power supply link between the recovered backup power supply and the corresponding controller and disconnect the power supply link between the third backup power supply and the corresponding controller when any one of the first backup power supply and the second backup power supply fails and the failure of the faulty backup power supply is recovered.
5. The power supply backup system of the storage cluster according to claim 1, characterized in that, The third controller is used to conduct the power supply link between the controller corresponding to the backup power supply with an earlier failure time and the third backup power supply when both the first backup power supply and the second backup power supply fail and the failure times of the first backup power supply and the second backup power supply are different; If the failure times of the first backup power supply and the second backup power supply are the same, determine the target controller according to the priorities of the first controller and the second controller, and conduct the power supply link between the target controller and the third backup power supply.
6. The power supply backup system of the storage cluster according to claim 1, characterized in that, The third controller is used to generate a primary fault alarm for the faulty backup power supply when at least one of the first backup power supply and the second backup power supply fails.
7. The power supply backup system of the storage cluster according to claim 1, wherein The third controller is used to generate a secondary fault alarm for the faulty backup power supply and limit the service authority of the controller corresponding to the faulty backup power supply with the secondary fault alarm when both the first backup power supply and the second backup power supply fail and the failure duration of any faulty backup power supply is greater than a preset duration. Among them, the alarm method of the secondary fault alarm is different from that of the primary fault alarm.
8. A server, characterized in that, A power supply backup system including the storage cluster according to any one of claims 1-7.
9. A control method for a power supply backup system of a storage cluster, the method being applied to a third controller in the power supply backup system of the storage cluster according to any one of claims 1-7, wherein, The method includes: Obtaining the health data of the storage cluster; Judging whether the first backup power supply and the second backup power supply are faulty according to the health data; When at least one of the first backup power supply and the second backup power supply fails, the control switch assembly conducts the power supply link between the third backup power supply and one of the first controller and the second controller.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the control method of the power supply backup system of the storage cluster as described in claim 9 are implemented.
Citation Information
Patent Citations
High-order autonomous vehicle redundant power supply system and method
CN118589677A
Power supply redundancy backup circuit and display device
CN222423306U
Power failure processing method and system for storage system
WO2015051633A1