Cluster server

By designing the ring topology structure and main server control mechanism of the cluster server, the problem of low storage resource utilization in the existing technology is solved, and higher storage resource utilization and system stability are achieved.

CN113535473BActive Publication Date: 2025-06-27ZHEJIANG HUAQI INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110721354.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-28
Publication Date
2025-06-27
Estimated Expiration
2041-06-28

AI Technical Summary

Technical Problem

The storage resource utilization rate of existing cluster servers is low, and it is impossible to effectively utilize the hard disk resources and storage links of the failed server.

Method used

A cluster server structure is designed, including a switch and at least three servers, the server connects to the disk array through a hard disk controller and forms a ring topology through a disk connector. The master server is used to control each server to acquire or release control of the disk array, ensuring that storage resources can be quickly taken over in the event of a failure.

Benefits of technology

It improves the storage resource utilization rate of cluster servers, ensures that storage resources can be quickly switched and utilized when a failure occurs, and enhances the stability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113535473B_ABST
    Figure CN113535473B_ABST
Patent Text Reader

Abstract

The present application relates to a cluster server, comprising: a switch and at least three servers, the servers being connected to the switch; the servers include storage devices, the storage devices include hard disk controllers and disk arrays, each hard disk controller is connected to the disk arrays of two other servers through a disk connector, and the storage devices of each server are connected in a ring topology; at least three servers include a primary server, and the primary server is used to control each server to acquire or release the control right of the disk array of the current server and / or the disk arrays of at least one other server. Through the present application, the problem of low utilization rate of storage resources of the cluster server in the related art is solved, and the utilization rate of storage resources of the cluster server is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server clusters, and particularly to a cluster server. Background Art

[0002] A server cluster refers to a collection of many servers that work together to provide the same service. To the client, it appears as if there is only one server. A cluster can utilize multiple computers for parallel computing to achieve high computing speeds, or use multiple computers for backup so that the entire system can still operate normally even if any one machine fails.

[0003] Existing cluster servers can usually only achieve clustering at the software system level. That is, when a certain server fails, the applications running on this server will switch to other servers, and the hard disk resources on this faulty machine will no longer be used by the applications, and the storage link transmitted to this server will also be cut off, so the storage content on this server cannot be obtained either, resulting in the underutilization of storage resources. Summary of the Invention

[0004] In this embodiment, a cluster server is provided to solve the problem of low utilization rate of storage resources in related cluster servers.

[0005] In this embodiment, a cluster server is provided, including: a switch and at least three servers, where the servers are connected to the switch;

[0006] Each server includes a storage device, and the storage device includes a hard disk controller and a disk array. Each hard disk controller is connected to the disk arrays of at least one other server through a disk connector;

[0007] The at least three servers include a main server, and the main server is used to control each server to obtain or release the control right of the disk array of the current server and / or the disk arrays of at least one other server.

[0008] In some of these embodiments, each hard disk controller is connected to the disk arrays of two other servers through a disk connector, and the storage devices of each server are connected in a ring topology.

[0009] In some of these embodiments, the server includes a central processing unit, and the central processing unit is connected to the switch; the central processing unit of the main server is used to monitor the online status of other servers, and in the case of other servers going offline, notify adjacent servers to obtain the control right of the disk array of the offline server.

[0010] In some of these embodiments, other servers establish heartbeat connections with the primary server, and the primary server monitors the online status of other servers through heartbeat information.

[0011] In some of these embodiments, the server further includes a baseboard management controller, which is connected to the switch. The baseboard management controller is also connected to the hard disk controller of the current server, and is used to acquire or release the control right of the disk array of the current server and / or the disk array of at least one other server.

[0012] In some of these embodiments,

[0013] The baseboard management controller is used to monitor the operating status of each hardware in the current server and send the operating status to the central processing unit of the current server;

[0014] The central processing unit of other servers is used to release the control right of the disk array and notify the primary server of the abnormal operating status in the case of abnormal operating status;

[0015] The central processing unit of the primary server is used to notify adjacent servers to acquire the control right of the disk array of the server with abnormal operating status after receiving the notification of abnormal operating status.

[0016] In some of these embodiments, other servers are also used to perform self-check and repair on the hardware of the current server after transferring the control right of the disk array of the current server to other servers, and reacquire the control right of the disk array of the current server after the self-check and repair is successful.

[0017] In some of these embodiments, the disk arrays of each server are powered by independent power supplies, and other servers perform self-check and repair by restarting the current server.

[0018] In some of these embodiments, the disk connector is a serial attached small computer system interface connector.

[0019] In some of these embodiments, the storage devices of each server are physically centrally arranged within the server.

[0020] Compared with the related art, the cluster server provided in this embodiment includes: a switch and at least three servers, where the servers are connected to the switch; the servers include storage devices, and the storage devices include hard disk controllers and disk arrays. Each hard disk controller is connected to the disk arrays of at least one other server through a disk connector; the at least three servers include a primary server, and the primary server is used to control each server to acquire or release the control right of the disk array of the current server and / or the disk arrays of at least one other server, which solves the problem of low utilization rate of storage resources of the cluster server in the related art and improves the utilization rate of storage resources of the cluster server.

[0021] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects, and advantages of this application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0023] Figure 1 is a schematic diagram of the server in this embodiment.

[0024] Figure 2 is a schematic structural diagram of the cluster server provided in this embodiment.

[0025] Figure 3 is a schematic diagram of the linear topology structure in this embodiment.

[0026] Figure 4 is a schematic diagram of the ring topology structure in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To understand the purpose, technical solution, and advantages of this application more clearly, the following describes and explains this application with reference to the drawings and embodiments.

[0028] Unless otherwise defined, technical terms or scientific terms involved in this application shall have the ordinary meanings understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "an", "one kind", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "linked", "coupled", etc. involved in this application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The term "plurality" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0029] This embodiment provides a cluster server, and the cluster server includes three or more servers. Figure 1 It is a schematic diagram of the server in this embodiment. The server can also be called a host, such as Figure 1 shown, each server includes a computing part 10 and a storage part 20. Among them, the computing part 10 usually includes a central processing unit 110 (CPU, also called the main controller or master controller); the storage part is usually composed of a storage device 210.

[0030] The storage device 210 includes a disk array 211. It should be noted that the disk array 211 referred to in this embodiment may include only one disk drive, or may be a disk group composed of multiple disk drives. Moreover, the disk drives constituting the disk array are not limited to HDD disk drives or SDD disk drives, and in some embodiments, they may also be a combination of HDD disk drives and SDD disk drives. And the disk array 211 can be a large-capacity disk drive formed by connecting all disk drives in series using JBOD (Just a Bunch Of Disks) technology, or can also be used by the server using technologies such as RAID (Redundant Array of Independent Disks) to improve the disk fault tolerance ability.

[0031] The interface device between the computing part 10 and the disk array 211 is called the hard disk controller 212, which is also known as the disk drive adapter. At the software level, the hard disk controller 212 is used to interpret the commands given by the computing part 10 and send various control signals to the disk drive; detecting the status of the disk drive, or writing data to and reading data from the disk in accordance with the specified disk data format is also controlled by the hard disk controller 212. At the hardware level, the hard disk controller 212 provides one or more physical interfaces for connecting to the disk array 211. The hard disk controller 212 can connect to one or more disk arrays 211 through these physical interfaces, and acquire or release the control right of the disk array 211 connected by the physical interface.

[0032] Each disk array 211 can include one or more physical interfaces for connecting to the hard disk controller 212. For example, a disk array 211 based on SAS (Serial Attached SCSI) technology can be connected to the hard disk controllers 212 of multiple servers to enable multiple servers to share the same disk array 211.

[0033] The computing part 10 and the storage part 20 of each server can be physically centralized, for example, set in the same server chassis. The computing part 10 and the storage part 20 can be set on the same main circuit board, or can be set separately. For example, the storage part 20 is set on the server backplane, and the computing part 10 is set on the main circuit board.

[0034] In addition to the storage part 20 and the computing part 10, the server usually also has two core firmware, namely BIOS (Basic Input / Output System) (not shown in the figure) and BMC (Baseboard Management Controller) (not shown in the figure). Among them, in the computer system, the BIOS has a more underlying and fundamental role than the server's operating system, mainly responsible for detecting, accessing, and debugging underlying hardware resources and allocating them to the operating system to ensure the smooth and safe operation of the entire machine. The BMC is a small operating system independent of the server's operating system, usually integrated on the motherboard, or inserted on the motherboard in the form of PCIe, etc. The external manifestation of the BMC is usually a standard RJ45 network port, and the BMC has an independent IP firmware system. Usually, the server can perform unattended operations using BMC instructions, such as remote management, monitoring, installation, restart, etc. of the server.

[0035] Figure 2 It is a schematic structural diagram of the cluster server provided by this embodiment. In Figure 2 Five servers are taken as an example for illustrative purposes. In other embodiments, the number of servers can be any number greater than three, usually specifically set according to the requirements for the computing resources and storage resources of the cluster server, and the number is not limited in this embodiment.

[0036] As Figure 2 shown, the cluster server includes a switch 40 and five servers. Each server is connected to the switch 40. The hard disk controller 212 of each server is connected to the disk array 211 of the current server and the disk arrays 211 of at least one other server through a disk connector (such as a SAS connector). Herein, the other servers refer to the other servers in the cluster server except the current server.

[0037] Among these five servers, one server is elected as the master server (server A in Figure 2 ) through self-election or user configuration, and the other servers are called slave servers relative to the master server. The master server is used to control each server to acquire or release the control right of the disk array of the current server and / or the disk arrays of at least one other server.

[0038] However, it should be noted that the master and slave servers in this embodiment do not limit that these servers have a master-slave relationship in actual business processing, but only indicate that the master server is in a dominant position when implementing this embodiment. In actual business processing, the master server may have the same status as other slave servers, or may have a lower or higher status than a certain other slave server.

[0039] By taking one server as the master server in the cluster server provided in this embodiment, and connecting the hard disk controllers of each server to the disk array of the current server and the disk arrays of at least one other server through disk connectors, and controlling other servers to acquire or release the control right of the disk array of the current server and / or the disk arrays of at least one other server through the master server, when a certain other server fails, the master server can control other servers to take over the disk array 211 of the failed server, improving the utilization rate of the disk array 211. Compared with the related art that uses expensive SAS switches to implement the sharing of the disk array 211, this embodiment does not require any additional SAS switches, and can directly use the switch 40 for business processing in the cluster server to meet the requirements, greatly reducing the cost.

[0040] The computing part 10 of each server includes a central processing unit 110, and the central processing unit 110 of each server is connected to a switch 40. The central processing unit 110 of each slave server can report the status of the current server to the master server periodically or aperiodically through the switch 40. Once the status reported by the slave server indicates that the slave server cannot normally complete business processing, the master server can notify the slave server to release the control right of the disk array of the current server and notify other slave servers to obtain the control right of the disk array of the faulty slave server. In some embodiments, the slave server can also actively release the control right of the disk array of the current server in case of a fault.

[0041] However, when the slave server drops offline or experiences abnormal power-off, the slave server cannot report the status of the current server to the master server. Therefore, in some embodiments, the master server can actively monitor the status of the slave server, such as monitoring the online status of other slave servers. The central processing unit 110 of the master server, in the case of other servers dropping offline, notifies the server adjacent to the offline slave server to obtain the control right of the disk array 211 of the offline server.

[0042] Among them, the online status of other slave servers can be detected by using a heartbeat connection. That is, a heartbeat connection is established between other slave servers and the master server, and the slave server periodically sends heartbeat information (keep-alive information) to the master server through the heartbeat connection. If the master server does not receive the heartbeat information sent by a certain slave server after exceeding the set time interval, it is considered that the slave server has dropped offline.

[0043] In each server, the central processing unit 110 and the BMC can both be used to control the hard disk controller 212 to obtain or release the control right of the disk array 211. The BMC is connected to the switch 40 through an RJ45 network port, and the BMC is also connected to the hard disk controller 212 of the current server. The BMC is used to control the hard disk controller 212 to obtain or release the control right of the disk array 211 of the current server and / or the disk array 211 of at least one other server. In some cases, if the operating system of the slave server crashes or the CPU fails, resulting in the inability to control the hard disk controller 212 to release the control right of the disk array 211 of the current server, the master server can control the BMC of the slave server through the switch 40 to release the control right of the disk array 211 of the current server. Similarly, the master server can also control the BMC of other slave servers through the switch 40 to obtain the control right of a certain disk array 211.

[0044] Since the BMC is a small operating system independent of the server operating system, even if the server's operating system crashes due to hardware or software failures, the BMC can still work properly to ensure the normal transfer of control of the disk array 211 of the cluster server.

[0045] As an independent third party in the server, the BMC can monitor the hardware information of the entire server, such as the system temperature, power supply voltage, fan speed, etc., and can also monitor the working status of the system network module, user interaction module (such as USB module, display module) or other modules. Once an abnormality that can affect the normal business capabilities of the server occurs in a certain module of the server, and the BMC determines that the server cannot complete the storage function, the BMC will transmit the abnormality information to the central processing unit 110 of the current server, or directly transmit it to the central processing unit of the main server through the switch 40. Therefore, the monitoring of the operating status of the current server by the central processing unit 110 can be achieved through the BMC. For example, the BMC monitors the operating status of each hardware in the current server and sends the operating status to the central processing unit 110 of the current server; when the central processing unit 110 of other servers is in an abnormal operating status, it releases the control of the disk array 211 and notifies the main server of the abnormal operating status; after receiving the notification of the abnormal operating status, the central processing unit 110 of the main server notifies the adjacent server to obtain the control of the disk array 211 of the server with the abnormal operating status.

[0046] To avoid the increased cost caused by using SAS switches to interconnect all the disk arrays 211 in the cluster server, in this embodiment, each hard disk controller 212 is connected to the disk array 211 of the current server and the disk arrays 211 of at least one other server through a disk connector (SAS connector). Through such a connection, the storage devices of each server can form, for example Figure 3 the linear topology structure as shown. In the linear topology structure, when the servers at both ends of the topology structure fail, the storage device can only be taken over by one adjacent server. In the case where the adjacent server has a large computing load, it may cause the adjacent server to malfunction due to further increased load after taking over the storage device, resulting in a reduction in the stability of the cluster server. Or if two adjacent servers at both ends of the topology structure fail continuously, the storage device of the outermost server will not be taken over by any server. Thus, there is still room for improvement in the utilization rate of the storage device.

[0047] Therefore, in some of these embodiments, each hard disk controller 212 is connected to the disk array 211 of the current server and the disk arrays of two other servers through a disk connector (SAS connector), and the storage devices of each server form, for example Figure 4The ring topology structure shown. In such a connection method, in the case of any server failure, there are two adjacent servers that can take over the storage device of the failed server; even if two consecutive adjacent servers fail, it can be ensured that there is one server respectively taking over the disk arrays of these two failed servers; only in the case of three consecutive adjacent servers failing, it is possible to cause the storage device of one server not to be taken over by any server. Thus, adopting the ring topology structure improves the stability of the cluster server and the utilization rate of the storage device.

[0048] The working process of the cluster server in this embodiment will be described below.

[0049] Embodiment 1

[0050] In this embodiment, the main server controls the hard disk controllers 212 of each server to obtain or release the control right of the disk arrays of the current server and / or other servers.

[0051] Referring to Figure 4 which is the topology structure, taking the main server as server A and other servers as slave servers as an example, the working process of the cluster server provided in this embodiment includes the following steps:

[0052] Step 1, both the main and slave servers monitor the running states of each hardware in their respective servers.

[0053] Step 2, in the case of an abnormal running state of server B, control the hard disk controller 212 of server B to release the control right of the disk array 211 of server B.

[0054] Step 3, server B sends a disk array control right handover instruction to server A.

[0055] Among them, the disk array control right handover instruction sent by server B to server A carries the identification information of server B, or carries the identification information of the disk array of server B.

[0056] Step 4, in the case where server A receives the disk array control right handover instruction sent by server B through the switch 40, according to the identification information carried in the disk array control right handover instruction, determine that the adjacent servers of server B corresponding to the identification information are server A and server C respectively by querying the pre-configured mapping table.

[0057] Step 5, if server A confirms that the current server is adjacent to server B, it can directly control the hard disk controller 211 of server A to obtain the control right of the disk array 211 of server B.

[0058] Embodiment 2

[0059] In this embodiment, the main server controls the hard disk controllers 212 of each server to acquire or release the control right of the disk array of the current server and / or other servers.

[0060] Refer to Figure 4 the topology structure, taking the main server as server A and other servers as slave servers as an example. The working process of the cluster server provided in this embodiment includes the following steps:

[0061] Step 1, both the master and slave servers monitor the running status of each hardware in their respective servers.

[0062] Step 2, when the running status of server B is abnormal, control the hard disk controller 212 of server B to release the control right of the disk array 211 of server B.

[0063] Step 3, server B sends a disk array control right handover instruction to server A.

[0064] Among them, the disk array control right handover instruction sent by server B to server A carries the identification information of server B, or carries the identification information of the disk array of server B.

[0065] Step 4, when server A receives the disk array control right handover instruction sent by server B through the switch 40, according to the identification information carried in the disk array control right handover instruction, determine that the adjacent servers of server B corresponding to the identification information are server A and server C respectively by querying the pre-configured mapping table.

[0066] Step 5, server A finds that it is adjacent to server B, but at this time the running status of server A indicates that its workload is already large. At this time, server A sends a disk array control right acquisition instruction to server C, and the disk array control right acquisition instruction carries the identification information of server B or the identification information of the disk array 211 of the faulty server.

[0067] Step 6, after receiving the disk array control right acquisition instruction, server C acquires the control right of the disk array 211 of server B according to the identification information carried therein.

[0068] The adjacent servers can be one or more servers. For example, in a ring topology structure, each server has two adjacent servers. In some other embodiments, in the case where the server adopts a disk array 211 such as SAS technology, two adjacent servers can jointly take over the control right of the disk array 211 of the same faulty server.

[0069] The master server can maintain a mapping table of the physical interfaces of the hard disk controllers 212 and the disk arrays 211 to know the identification information of the disk arrays 211 connected to each physical interface, or the identification information of the servers to which the disk arrays 211 belong; it can also maintain the topology information of the cluster servers to know the adjacent servers of each server. After the slave server obtains the disk array control right acquisition instruction, it determines the physical interface to which the disk array 211 to be taken over is connected according to the identification information carried in the disk array control right acquisition instruction, so as to control the hard disk controller 212 to obtain the control right of the disk array 211 of the faulty server connected to this physical interface.

[0070] Embodiment 3

[0071] In this embodiment, the master server controls the hard disk controllers 212 of each server to obtain or release the control right of the disk arrays of the current server and / or other servers, but each slave server does not need to actively report the operating status of its own server.

[0072] Referring to Figure 4 is the topology structure, taking the master server as server A and other servers as slave servers as an example. The working process of the cluster servers provided in this embodiment includes the following steps:

[0073] Step 2, server A monitors the heartbeat information of each server in the cluster server, and determines that server B is offline according to the heartbeat information.

[0074] Step 3, according to the identification information of server B, server A determines through querying the pre-configured mapping table that the adjacent servers of server B are server A and server C.

[0075] Step 4, server A finds that it is adjacent to server B, but at this time, the operating status of server A indicates that its workload is already large. At this time, server A sends a disk array control right acquisition instruction to server C, and the identification information of server B or the identification information of the disk array 211 of server B is carried in the disk array control right acquisition instruction.

[0076] Step 5, after receiving the disk array control right acquisition instruction, server C obtains the control right of the disk array 211 of server B according to the identification information carried therein.

[0077] The master server can maintain a mapping table of the physical interfaces of the hard disk controller 212 and the disk arrays 211 to know the identification information of the disk arrays 211 connected to each physical interface, or the identification information of the servers to which the disk arrays 211 belong; it can also maintain the topology information of the cluster servers to know the adjacent servers of each server. After the slave server obtains the disk array control right acquisition instruction, it determines the physical interface to which the disk array 211 to be taken over is connected according to the identification information carried in the disk array control right acquisition instruction, so as to control the hard disk controller 212 to obtain the control right of the disk array 211 of the faulty server connected to this physical interface.

[0078] Step 6, after Server B actively or passively transfers the control right of the disk array 211 of Server B to Server C due to faults or disconnections, etc., perform self-check and repair on the hardware of Server B.

[0079] Step 7, if the self-check and repair of Server B is successful, Server B reacquires the control right of the disk array 211 of Server B.

[0080] Among them, when Server B reacquires the control right of the disk array 211 of Server B, it can send a disk array control right acquisition request to Server A. After receiving the disk array control right acquisition request from Server B, Server A sends a disk array control right release instruction to Server C that currently controls the disk array control right of Server B, and after receiving the notification that Server C has successfully released the disk array control right, sends a confirmation message to Server B. When Server B receives this confirmation message, it reacquires the control right of its disk array. Through the above method, the self-check and self-repair of the faulty server are realized.

[0081] Among them, the disk arrays 211 of each server are powered by independent power supplies. The server can perform self-check and repair by restarting the current server, and ensure that the disk arrays 211 of the current server are not powered off and can be taken over and utilized by other servers.

[0082] The cluster server can also include a control node. The control node is connected to the switch 40 and is used to configure each server, such as configuring the control programs of each server, or the identification information of each server, or the mapping table stored in each server. In addition, through the control node, the BMC of each server can also be controlled to realize the remote unattended function, such as remote restart, etc.

[0083] In summary, traditional cluster services usually cut off the services of an abnormal node and cannot call the storage part. This embodiment realizes the cluster service from the hardware aspect, effectively utilizes the storage part of the abnormal device for reuse and retrieves the content of the storage part. This embodiment uses a disk connector to interconnect the disk arrays of multiple servers, making the storage parts of multiple servers an integrated whole capable of transferring control rights, and through one of the servers as the main server to participate in cluster control, greatly improving the stability and security of the cluster solution. Once an abnormality occurs, a quick decision can be made to transfer the control rights of the disk array, greatly improving the stability of the cluster solution.

[0084] It should be understood that the specific embodiments described herein are only used to explain this application and not to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of this application.

[0085] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations based on these drawings without creative work. Additionally, it can be understood that although the work done during this development process may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.

[0086] The term "embodiment" in this application means that the specific features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean being independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.

[0087] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.

Claims

1. A cluster server, characterized in that Including: A switch and at least three servers, where the servers are connected to the switch; The servers include storage devices, and the storage devices include hard disk controllers and disk arrays. Each hard disk controller is connected to the disk arrays of at least one other server through a disk connector; The at least three servers include a primary server, and the primary server is used to control each server to obtain or release the control right of the disk array of the current server and / or the disk arrays of at least one other server; The servers include central processing units, and the central processing units are connected to the switch; the central processing unit of the primary server is used to monitor the online status of other servers, and in the case of other servers going offline, notify the adjacent servers to obtain the control right of the disk array of the offline server; The servers also include baseboard management controllers, the baseboard management controllers are connected to the switch, and the baseboard management controllers are also connected to the hard disk controllers of the current server. The baseboard management controllers are used to obtain or release the control right of the disk array of the current server and / or the disk arrays of at least one other server; The baseboard management controllers are used to monitor the operating status of each hardware in the current server and send the operating status to the central processing unit of the current server; The central processing units of other servers are used to release the disk array control right in the case of abnormal operating status and notify the primary server of the abnormal operating status; The central processing unit of the primary server is used to notify the adjacent servers to obtain the control right of the disk array of the server with abnormal operating status after receiving the notification of the abnormal operating status.

2. The cluster server according to claim 1, wherein Each hard disk controller is connected to the disk arrays of two other servers through a disk connector, and the storage devices of each server are connected in a ring topology.

3. The cluster server according to claim 1, characterized in that Other servers establish a heartbeat connection with the primary server, and the primary server monitors the online status of other servers through heartbeat information.

4. The cluster server according to claim 1, wherein Other servers are also used to perform self-check and repair on the hardware of the current server after transferring the control right of the disk array of the current server to other servers, and re-obtain the control right of the disk array of the current server after the self-check and repair is successful.

5. The cluster server according to claim 4, wherein The disk arrays of each server are powered by independent power supplies, and other servers perform self-check and repair by restarting the current server.

6. The cluster server according to any one of claims 1 to 5, characterized in that, The disk connector is a SAS connector.

7. The cluster server according to any one of claims 1 to 5, characterized in that, The storage device of each server is physically centrally arranged in the server.

Citation Information

Patent Citations

  • Server direct attached storage shared through virtual sas expanders

    CN103095796A

  • Local management console for storage devices

    CN110058803A

  • Cluster system control method and cluster system

    CN111045602A

  • Fail over method through disk take over and computer system having fail over function

    US20060143498A1