Link redundant disk system, fault processing method, server and electronic equipment
By designing a disk system with link redundancy in a disk system, using the design of multi-stage extenders and redundant links, JBOD's shortcomings in data redundancy are solved, and the redundancy of hard disk identification and the stability of data paths are achieved, ensuring data security and system reliability.
Patent Information
- Application Number
- CN202510081607.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-06
Smart Images

Figure CN119938415A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of servers, and in particular to a disk system with link redundancy and a fault handling method, a server, and an electronic device. Background Art
[0002] In today's era of rapid development of information technology, data security and reliability are becoming increasingly important for enterprises and individuals. Data plays a key role in daily operations and life. Once lost, it will bring various losses and many inconveniences to enterprises and individuals.
[0003] Among the current disk management methods, Just a Bunch Of Disks (JBOD) is suitable for scenarios with large storage space requirements but low data redundancy requirements due to its simplicity, flexibility and low cost. JBOD can easily access each hard disk, and can also integrate physical hard disks to form a large-capacity logical volume for storing various data files. However, JBOD has significant deficiencies in data redundancy. Once a hard disk or disk link fails, the data stored on it will face the risk of irrecoverable loss, posing a major risk to data security. Summary of the invention
[0004] The present invention provides a disk system with link redundancy and a fault handling method, a server, and an electronic device, which are mainly intended to solve the problem of hard disk recognition failure and data loss caused by the lack of redundancy in the hard disk connection link.
[0005] According to a first aspect of the present disclosure, there is provided a disk system with link redundancy, the system comprising: an expander group, a hard disk group, a head adapter card, and a fault processing module;
[0006] The expander group includes at least two levels of expanders, the first level expander is connected to the head adapter card via a first data link, the upper level expander is connected to at least two expanders in the lower level via at least two second data links, and the hard disk group is connected to at least two last level expanders via at least two third data links respectively;
[0007] The fault handling module is connected to the head adapter card and is used to identify the faulty expander in the expander group step by step when it is determined that there is a missing hard disk group that cannot be identified; control other expanders at the same level to replace the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link is composed of a first data link, a second data link at different levels, and a third data link.
[0008] In some embodiments, the set of expanders includes a single-level expander;
[0009] The single-level expander is connected to the head adapter card via a first data link, and the hard disk group is connected to the single-level expander via at least two third data links.
[0010] In some embodiments, one end of the first data link is connected to the interface of the head adapter card, and the other end is connected to the uplink port of the first stage expander;
[0011] One end of the second data link is connected to the downlink port of the upper-level expander, and the other end is connected to the uplink port of the lower-level expander;
[0012] One end of the third data link is connected to the downlink port of the last-stage expander, and the other end is connected to the interface of the hard disk group.
[0013] In some embodiments, the system further comprises: a data calling module;
[0014] The data calling module is connected to the head adapter card and is used to send a calling instruction of target data to the hard disk group; in response to the calling instruction, the hard disk group uses one of the at least two third data links to send the target data to the last-level expander, and the remaining third data link is closed;
[0015] The last-stage expander transmits the target data upward to the first-stage expander by using an extended data link, wherein the extended data link is generated by combining one second data link of each level in the expander group level by level, and the remaining second data links are closed;
[0016] The first-level expander sends the target data to the head adapter card through a first data link.
[0017] In some embodiments, when the expander group includes two stages of expanders, the last stage expander of the expander group is a second stage expander;
[0018] The first-stage expander is connected to the head adapter card via a first data link, and the first-stage expander is connected to at least two second-stage expanders via at least two second data links respectively;
[0019] The hard disk group is connected to at least two second-level expanders respectively through at least two third data links.
[0020] In some embodiments, the fault handling module is also used to obtain the operating status information of the expander, and determine whether the expander at the current level has a fault based on the operating status information; if there is a fault, control other expanders at the same level to replace the faulty expander and switch the data link; if there is no fault, continue to check the expanders at the next level until the expander group inspection is completed.
[0021] According to a second aspect of the present disclosure, a method for handling a fault is provided, the method being applied to the disk system with link redundancy in the first aspect, the method comprising:
[0022] Check the status of the hard disk group to determine whether there is an unrecognizable missing hard disk group;
[0023] When it is determined that there is an unrecognizable missing hard disk group, the faulty expanders in the expander group are identified step by step;
[0024] Control other expanders at the same level to take over the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link consists of a first data link, a second data link at different levels, and a third data link.
[0025] In some embodiments, the stepwise identifying of a faulty expander in the expander group includes:
[0026] Acquire the operation status information of the expander, and determine whether the current level expander has a fault according to the operation status information;
[0027] If there is a fault, other expanders at the same level are controlled to take over the faulty expander and switch the data link;
[0028] If there is no fault, continue to check the expander at the next level until the expander group is checked.
[0029] According to a third aspect of the present disclosure, a server is provided, the server comprising the disk system with link redundancy of the first aspect
[0030] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:
[0031] at least one processor; and
[0032] a memory communicatively connected to the at least one processor; wherein,
[0033] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the second aspect.
[0034] According to a fifth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the second aspect.
[0035] According to a sixth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the method described in the second aspect is implemented.
[0036] The present invention provides a disk system with link redundancy and a fault handling method, a server, and an electronic device, wherein the system comprises: an expander group, a hard disk group, a head adapter card, and a fault handling module; the expander group comprises at least two levels of expanders, the first level expander is connected to the head adapter card via a first data link, the upper level expander is connected to at least two expanders in the lower level via at least two second data links, and the hard disk group is respectively connected to at least two last level expanders via at least two third data links; the fault handling module is connected to the head adapter card, and is used to identify the faulty expanders in the expander group step by step when it is determined that there is a missing hard disk group that cannot be identified; and control other expanders at the same level to take over the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link is composed of a first data link, second data links at different levels, and a third data link. Compared with the related art, the disclosed embodiment of the present invention can solve the problem of hard disk recognition and data loss caused by expander failure in the disk system through the expander group architecture and fault handling mechanism; when a failure occurs, the fault handling module can quickly locate the faulty expander, and use redundant links to allow other expanders at the same level to take over the work, ensuring that all hard disks can always be recognized by the head adapter card and maintain the smooth flow of data paths, thereby ensuring the integrity and security of data and reducing the risk of data loss due to hardware failure; the multi-level expander structure and redundant link design enable the system to still operate stably in the face of single or partial expander failures, avoiding the overall paralysis of the system due to local failures, ensuring the continuous availability of storage services, and reducing system downtime caused by maintenance and repair of failures.
[0037] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0039] Figure 1A schematic diagram of the structure of a disk system with link redundancy provided in an embodiment of the present disclosure;
[0040] Figure 2 A schematic diagram of the structure of a disk system with single-level expander link redundancy provided by an embodiment of the present disclosure;
[0041] Figure 3 A schematic diagram of the structure of another disk system with link redundancy provided by an embodiment of the present disclosure;
[0042] Figure 4 A schematic diagram of the structure of a disk system with two-level expander link redundancy provided by an embodiment of the present disclosure;
[0043] Figure 5 A schematic diagram of a typical JBOD architecture;
[0044] Figure 6 A flowchart of a fault handling method provided by an embodiment of the present disclosure;
[0045] Figure 7 A schematic block diagram of an exemplary electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0046] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0047] The following describes a disk system and fault handling method, a server, and an electronic device with link redundancy according to an embodiment of the present disclosure with reference to the accompanying drawings.
[0048] Figure 1 A schematic diagram of the structure of a disk system with link redundancy provided in an embodiment of the present disclosure is provided. The system includes: an expander group 11, a hard disk group 12, a head adapter card 13, and a fault processing module 14.
[0049] The expander group 11 includes at least two levels of expanders, the first level expander is connected to the head adapter card 13 via a first data link, the upper level expander is connected to at least two expanders in the lower level via at least two second data links, and the hard disk group 12 is connected to at least two last level expanders respectively via at least two third data links.
[0050] In the embodiment of the present disclosure, the expander group 11 serves as a key component for connecting the head adapter card 13 and the hard disk group 12, and a complex and orderly hierarchical structure design of at least two levels of expanders is adopted inside. Among them, the first-level expander is connected to the head adapter card 13, and they achieve efficient data transmission and communication with the help of the first data link. This first data link can adopt but is not limited to high-speed cable connection at the physical level, and follows the predetermined communication protocol standard to ensure that data can be exchanged quickly, stably and accurately between the head adapter card 13 and the first-level expander, laying a solid foundation for the data flow of the entire storage system. It should be noted that this embodiment does not limit the use of which communication protocol.
[0051] In terms of the internal hierarchical connection of the expander group 11, the upper-level expander is connected to at least two expanders in the lower level through at least two configured second data links. In order to better realize the link redundancy design, and when the interface of the expander allows, a second data link can be used to connect an upper-level expander to all expanders in the lower level. Similarly, the second data link can also be connected by a high-speed cable but not limited to it at the physical level, and follows a predetermined communication protocol standard. The second data link connected by the upper and lower expanders designed in this way can form a multiplexing technology or a backup link mechanism. When a certain second data link fails, the system can automatically switch the data to other normal second data links for transmission, ensuring the stability and reliability of the hierarchical connection of the entire expander group 11, thereby building a stable and efficient data transmission network architecture.
[0052] The connection between the hard disk group 12 and the expander group 11 also has a rigorous design. The hard disk group 12 is connected to at least two last-level expanders through at least two third data links. These third data links are highly adapted to the interfaces of the hard disk group 12 and the last-level expander in terms of electrical characteristics and interface specifications, which can ensure the accuracy and efficiency of data transmission between the hard disk and the expander. In actual application scenarios, these links may adopt high-speed serial connection technology and be equipped with corresponding signal enhancement and error correction mechanisms to cope with the interference that may be caused by long-distance transmission and complex electromagnetic environments, ensuring that the hard disk data can be safely and quickly received and processed by the expander, thereby achieving smooth operation of the data storage and reading functions of the entire storage system.
[0053] The fault handling module 14 is connected to the head adapter card 13, and is used to identify the faulty expander in the expander group step by step when it is determined that there is a missing hard disk group that cannot be identified; control other expanders at the same level to replace the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card 13, wherein the data link is composed of a first data link, a second data link at different levels, and a third data link.
[0054] In the embodiment of the present disclosure, in the operation architecture of the entire disk storage system, the fault handling module 14 plays a crucial central control role, which is closely connected with the head adapter card 13 to form an efficient fault response and data recovery system.
[0055] When the system is in operation, the fault handling module 14 will continuously and closely monitor the storage status of the system. Once an abnormal situation of a missing hard disk group that cannot be identified is found, the fault handling module 14 will check the expander group step by step according to the preset logic and algorithm to determine which expander or expanders in the expander group have failed, so as to accurately lock the faulty expander.
[0056] After the faulty expander is identified, the fault handling module 14 will send control instructions to other normal expanders at the same level, instructing one of these other normal expanders at the same level to quickly take over the work of the faulty expander. At the same time, it will perform intelligent switching operations on the entire data link. In this process, the first data link involved, the second data links at different levels, and the third data links will be reconfigured and connected under the control of the fault handling module 14. For the first data link, it will ensure that the connection stability and data transmission efficiency between the head adapter card 13 and the newly replaced expander are not affected; for the second data links at different levels, the data flow direction and transmission path will be readjusted so that the data can bypass the faulty expander and flow smoothly in the new expander connection architecture; and for the third data link, it will ensure that it is tightly connected to the hard disk group 13 and the data interaction is normal. Through such a series of complex and orderly operations, the originally unrecognizable missing hard disk group and the head adapter card 13 are finally successfully re-established with a stable and reliable connection relationship, ensuring that the entire storage system can maintain data integrity and availability in the face of expander failure, greatly improving the robustness and reliability of the system, and providing solid protection for users' data storage needs.
[0057] The present invention provides a disk system with link redundancy, the system comprising: an expander group, a hard disk group, a head adapter card, and a fault processing module; the expander group comprises at least two levels of expanders, the first level expander is connected to the head adapter card via a first data link, the upper level expander is connected to at least two expanders in the lower level via at least two second data links, and the hard disk group is respectively connected to at least two last level expanders via at least two third data links; the fault processing module is connected to the head adapter card, and is used to identify the faulty expanders in the expander group step by step when it is determined that there is a missing hard disk group that cannot be identified; control other expanders at the same level to take over the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link is composed of a first data link, second data links of different levels, and a third data link. Compared with the related art, the disclosed embodiment of the present invention can solve the problem of hard disk recognition and data loss caused by expander failure in the disk system through the expander group architecture and fault handling mechanism; when a failure occurs, the fault handling module can quickly locate the faulty expander, and use redundant links to allow other expanders at the same level to take over the work, ensuring that all hard disks can always be recognized by the head adapter card and maintain the smooth flow of data paths, thereby ensuring the integrity and security of data and reducing the risk of data loss due to hardware failure; the multi-level expander structure and redundant link design enable the system to still operate stably in the face of single or partial expander failures, avoiding the overall paralysis of the system due to local failures, ensuring the continuous availability of storage services, and reducing system downtime caused by maintenance and repair of failures.
[0058] Furthermore, in a possible implementation of this embodiment, Figure 2 As shown, the expander group 11 includes a single-level expander;
[0059] The single-level expander is connected to the head adapter card 13 via a first data link, and the hard disk group 12 is connected to the single-level expander via at least two third data links.
[0060] Specifically in the embodiments of the present disclosure, in the architectural design of the hard disk storage system, the expander group 11 has a special configuration form, that is, it includes a single-level expander. This single-level expander plays a key connection role in the entire storage link. Due to certain limitations in the design of the interface of the head adapter card 13, the upper limit of its physical port number or electrical performance determines that more hard disk groups 12 cannot be connected. However, in some specific application scenarios, when the number of hard disk groups 12 is relatively small, it becomes a more ideal choice for the expander group 11 to connect to the hard disk group 12 in the form of a single-level expander. This connection method can not only meet the needs of data storage and transmission, but also has certain advantages in terms of system complexity and cost control. Compared with the architecture of multi-level expanders, single-level expanders reduce the hardware equipment and data transmission paths in the intermediate links, reduce the probability of system failure and maintenance costs, and also improve the system's response speed and data transmission efficiency.
[0061] Specifically, if Figure 2 As shown, the single-level expander is connected to the head adapter card 13 by means of a specially designed first data link. The high speed and stability requirements of data transmission are fully considered during the construction of this first data link, and high-quality cables and adaptive interface technologies that meet specific standards are selected to ensure that data can flow efficiently and accurately between the head adapter card 13 and the single-level expander. Its transmission protocol has been optimized to effectively reduce delays and error rates during data transmission, laying the foundation for the stable operation of the entire storage system. At the same time, the connection between the hard disk group 12 and the single-level expander is achieved through at least two third data links. Figure 2 The middle dashed line indicates that it is closed under normal circumstances. These third data links have been carefully debugged in terms of electrical characteristics, bandwidth allocation, and interface compatibility to adapt to the connection requirements of the hard disk group 12 and the single-level expander. They can accurately transmit data in the hard disk group 12 to the single-level expander, or pass instructions and data from the head adapter card 13 to each hard disk in the hard disk group 12, ensuring the smooth storage and reading of data.
[0062] Further, in a possible implementation of this embodiment, one end of the first data link is connected to the interface of the head adapter card 13, and the other end is connected to the uplink port of the first-level expander;
[0063] One end of the second data link is connected to the downlink port of the upper-level expander, and the other end is connected to the uplink port of the lower-level expander;
[0064] One end of the third data link is connected to the downlink port of the last-stage expander, and the other end is connected to the interface of the hard disk group 12 .
[0065] Specifically in the embodiment of the present disclosure, the head adapter card 13 is a host bus adapter card (Host Bus Adapter Card, HBA card), which is an interface card used to connect a host (such as a server motherboard, etc.) and an external storage device (such as a disk array, etc.) in a computer system. Its function is to enable the host to communicate data with the storage device and realize operations such as reading and writing the storage device (the hard disk in the hard disk group 12).
[0066] In the architecture of the entire storage system, the connection mode of the data link plays a key role in the stable operation and efficient data transmission of the system. At the physical level, one end of the first data link is closely connected to the interface of the head adapter card 13 through a physical link. The other end of the first data link is firmly connected to the uplink port of the first-level expander. For the second data link, it plays a bridge role in the hierarchical architecture of the expander group. One end of it is reliably connected to the downlink port of the upper-level expander, and the downlink port of the upper-level expander is designed to integrate an efficient data distribution mechanism, which can accurately distribute data to the corresponding second data link according to different priorities and target addresses. One end of the third data link is closely connected to the downlink port of the last-level expander, and the other end of the third data link is docked with the interface of the hard disk group 12. The interface of the hard disk group 12 adopts a technical standard adapted to the third data link, which can efficiently receive data from the third data link and accurately write it into the corresponding hard disk storage unit, or when reading data, it can quickly transmit the data in the hard disk back to the expander group through the third data link, thereby realizing the data storage and reading function of the entire storage system. Specifically, the expander, the hard disk group 12, and the head adapter card 13 adopt the Serial Attached SCSI (SAS) standard.
[0067] Furthermore, in a possible implementation of this embodiment, Figure 3 As shown, the system further includes: a data calling module 15;
[0068] The data calling module 15 is connected to the head adapter card 13, and is used to send a calling instruction of target data to the hard disk group 12; in response to the calling instruction, the hard disk group 12 uses one of the at least two third data links to send the target data to the last-level expander, and the remaining third data links are closed;
[0069] The last-stage expander transmits the target data upward to the first-stage expander by using an extended data link, wherein the extended data link is generated by combining one second data link of each level in the expander group level by level, and the remaining second data links are closed;
[0070] The first-level expander sends the target data to the head adapter card 13 through the first data link.
[0071] Specifically, in the embodiment of the present disclosure, in the efficient operation mechanism of the entire storage system, the data call module 15 serves as a control unit for data call and establishes a stable connection with the head adapter card 13. Its core function is to accurately send a call instruction for the target data to the hard disk group 12. When the data call module 15 starts the data call process, it will generate a call instruction in a specific format according to the communication protocol and data addressing method preset by the system, and accurately transmit the instruction to the hard disk group 12 through the connection link with the head adapter card 13. After receiving this call instruction, the hard disk group 12 will quickly start the internal data retrieval and extraction mechanism. At this time, the hard disk group 12 will intelligently select and use one of the at least two third data links to perform the task of sending the target data. This selected third data link will transmit the target data to the last-level expander in a high-speed and stable manner under the coordinated control of hardware and software. In this process, in order to avoid confusion and interference in data transmission, the remaining third data links will be in a closed state to ensure the singleness and accuracy of data transmission.
[0072] When the target data successfully reaches the last-level expander, the expander will use an extended data link to pass the target data upward to the first-level expander according to the system's established hierarchical transmission rules. This extended data link is not a simple single link, but is generated by combining a second data link at each level in the expander group step by step. At the same time, in order to ensure the efficiency and orderliness of data transmission, the remaining second data links will remain closed to prevent data from being diverted or misrouted during transmission.
[0073] Finally, when the target data successfully reaches the first-level expander, the first-level expander will use the first data link between it and the head adapter card 13 to send the target data completely to the head adapter card 13. In this process, the first data link will give full play to its high-speed and stable transmission performance, and adopt advanced signal modulation and error correction technology to ensure that the data can be accurately received and processed by the head adapter card 13, thereby completing the entire data call process, providing the required target data for the upper-layer applications of the system, and ensuring smooth data interaction between the storage system and other system components.
[0074] Furthermore, in a possible implementation of this embodiment, Figure 4 As shown, when the expander group 11 includes two-stage expanders, the last-stage expander of the expander group 11 is the second-stage expander;
[0075] The first-stage expander is connected to the head adapter card 13 via a first data link, and the first-stage expander is connected to at least two second-stage expanders via at least two second data links respectively;
[0076] The hard disk group 12 is connected to at least two second-level expanders respectively through at least two third data links.
[0077] In the related art, the link of the hard disk storage expansion cabinet (Just a Bunch Of Disks, JBOD) does not have a redundant design. Once a SAS expander fails, the corresponding downlink link will immediately become inoperable. The following is an example of a typical JBOD architecture, taking a JBOD that accommodates 120 hard disks as an example. Please refer to Figure 5 Understand. Usually, there are 4 SASExpanders inside the JBOD. The uplink of each SASExpander card is configured as a SASport with a bandwidth of X4. The uplink SASports of these 4 SASExpander cards correspond one-to-one to the 4 X4 SASports ABCD on the HBA card of the head and are directly connected. In order to realize the expansion function of the link, the downlink of each SASExpander card is designed as an X30SASPort, and each downstream port is directly connected to 30 HDD hard drives. In this way, the original X16SAS link of the head is expanded with the help of the 4 SASExpander cards inside the JBOD, and its bandwidth can be improved, which can achieve connection with 120 hard drives to meet the needs of large-scale storage. In this way, the 30 hard drives connected to the faulty SASExpander cannot be recognized by the system, which leads to the risk of data loss stored on these hard drives, seriously affecting the security of the data and the stability of the storage system.
[0078] Specifically, in the embodiment of the present disclosure, when the expander group 11 adopts a two-stage expander structure, the last stage expander of the expander group 11 is the second stage expander. Figure 4 The actual case shown is for reference only. The architecture includes 2 L1-level SAS Expanders as the first-level expanders and 4 L2-level SAS Expanders as the second-level expanders.
[0079] The first-level expander is closely connected to the head adapter card 13 through a carefully constructed first data link. In the actual design, the uplink of the L1-level SAS Expander is 2 X4 SAS ports, and the 4 uplink X4 SAS ports of the 2 L1-level SAS Expanders are directly connected to the 4 X4 SAS ports ABCD of the head HBA card respectively, which is like opening up four high-speed channels for data, ensuring that data can be quickly and stably transmitted between the head adapter card 13 and the first-level expander.
[0080] At the same time, the first-level expander is connected to at least two second-level expanders through at least two second data links. Taking the L1-level SAS Expander as an example, its downlink is 4 X4 SAS ports, of which the 2 X4 SAS ports in the solid line part are normally open and connected to 2 L2-level SAS Expanders respectively; the other 2 X4 SAS ports in the dotted line part are normally closed and are also connected to the remaining 2 L2-level SAS Expanders respectively. This design builds a redundancy mechanism, and when the normal link fails, the spare dotted line link can be quickly put into use.
[0081] Let's look at the connection between the hard disk group 12 and the second-level expander. The hard disk group 12 is connected to at least two second-level expanders through at least two third data links. In the actual solution, the solid line part X30 of the L2-level SAS Expander is directly connected to 30 hard disks; the dotted line part X30 is connected to 30 hard disks of another SAS Expander of the same level L2. For example, the solid line part of SAS L2Expander1 is connected to 30 hard disks, and the dotted line part is connected to 30 hard disks of SAS L2 Expander2; the solid line part of SAS L2Expander2 is connected to 30 hard disks, and the dotted line part is connected to 30 hard disks of SAS L2 Expander1, and the same is true for SAS L2Expander3 and SAS L2 Expander4. Normally, the solid line part is open and the dotted line part is closed. These third data links use high-reliability cables and interfaces in hardware to ensure the stability of physical connections; at the software level, they are equipped with an intelligent data scheduling system that can reasonably allocate data transmission tasks according to the link load conditions. Similarly, the third data link also has redundancy function. When a link fails, the system can quickly switch data to the backup link.
[0082] Through such connection, link redundancy is successfully achieved. Under normal circumstances, the X16 SAS link of the head is expanded to X120 to connect 120 disks through two L1-level SAS Expander cards and four L2-level SAS Expander cards in the JBOD, that is, the head can identify 120 disks through the HBA card. When a certain L1-level Expander card fails, such as SAS L1Expander1, the head cannot detect SAS L1 Expander1, SAS L2 Expander1, and SAS L2Expander2, and the head will detect that 60 disks are missing. However, due to the link redundancy design, the head program shuts down SAS L1Expander1 and urgently activates the two X4 SAS ports in the dotted part of SAS L1 Expander2. Because these two X4 SAS ports in the dotted part are still connected to the uplink X4 SAS ports of SAS L2 Expander1 and SAS L2 Expander2, the head HBA card can still connect to SAS L2 Expander1 and SAS L2Expander2 through the SAS link in the dotted part, which means that the 60 disks mounted under SAS L2Expander1 and SAS L2 Expander2 can still be identified, thus ensuring that the overall 120 disks can still be identified, effectively preventing the loss of hard disks caused by SAS L1 Expander failure. When a certain L2-level Expander card fails, such as SAS L2 Expander1, the head cannot detect SASL2 Expander1, and the head will detect that 30 disks are missing. At this time, the head program shuts down SAS L2 Expander1 and urgently activates the X30 SAS port in the dotted part of SAS L2 Expander2. The dotted line indicates that it is shut down under normal circumstances. Because the X30 SAS port in the dotted part is still connected to the 30 disks mounted under SAS L2 Expander1, the head HBA card can still connect to the 30 disks mounted under SAS L2 Expander1 through the SAS link in the dotted part, so that the entire 120 disks can still be identified, successfully avoiding the loss of hard disks caused by SAS L2 Expander failure. This redundant design greatly improves the stability and reliability of the storage system and ensures the security and integrity of data.
[0083] Furthermore, in a possible implementation of the present embodiment, the fault handling module 14 is also used to obtain operating status information of the expander, and determine whether the current-level expander has a fault based on the operating status information; if there is a fault, control other expanders at the same level to replace the faulty expander and switch the data link; if there is no fault, continue to check the expanders at the next level until the expander group inspection is completed.
[0084] Specifically, in the embodiment of the present disclosure, the fault handling module 14 first continuously obtains the operating status information of the expander through a set of efficient data collection mechanisms. This mechanism can go deep into the hardware bottom layer and software running process of the expander, and comprehensively collect multi-dimensional operating parameters such as the operating temperature, power consumption, data transmission rate, error checking code, etc. of the expander.
[0085] After obtaining the operating status information of the expander, the fault processing module 14 will deeply analyze the information to determine whether the current level expander has a fault. For example, when the operating temperature of the expander exceeds the safety threshold, or a large number of error check codes appear during data transmission, it may be determined that there is a fault.
[0086] Once it is determined through analysis that the current level expander has a fault, the fault handling module 14 will immediately send control instructions to other normal expanders at the same level. These instructions inform other expanders in detail which tasks need to take over the faulty expander and how to switch the data link. When switching data links, the fault handling module 14 will fully consider factors such as the link's bandwidth and transmission stability to ensure that data can be transmitted in the new link in an efficient and stable manner. If it is determined through analysis that the current level expander does not have a fault, the fault handling module 14 will continue to check the expander at the next level, and check each level in turn according to the established hierarchical order. The fault handling module 14 can promptly discover and resolve potential fault problems in the expander group.
[0087] Figure 6 A flowchart of a fault handling method provided in an embodiment of the present disclosure.
[0088] like Figure 6 As shown, the method comprises the following steps:
[0089] Step 201 , detecting the presence status of the hard disk group, and determining whether there is an unrecognizable missing hard disk group.
[0090] Step 202: When it is determined that there is an unidentified missing hard disk group, a faulty expander in the expander group is identified step by step.
[0091] Step 203, control other expanders at the same level to take over the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link consists of a first data link, a second data link at different levels, and a third data link.
[0092] Specifically, in step 201 to step 203, in the daily operation and maintenance of the hard disk storage system, a rigorous and efficient troubleshooting and repair process is essential to ensure stable storage and transmission of data.
[0093] The system first starts a sophisticated hard disk group presence detection mechanism. This mechanism uses a dedicated hardware sensor to work with the underlying driver to continuously scan the physical connection status of each hard disk in the hard disk group and its logical communication status. By sending a specific detection command to each hard disk and waiting for it to return a corresponding response signal, it can accurately determine whether the hard disk is in a normally identifiable presence state. Once it is found that one or more hard disks cannot respond to the detection signal, the system will determine that there is an unrecognizable missing hard disk group.
[0094] When it is determined that there is an unrecognized missing hard disk group, the system will quickly start the fault expander troubleshooting program. Since the expander group adopts a hierarchical architecture, the troubleshooting work will start from the first-level expander and gradually go deeper into each level in the hierarchical order. At each level, the troubleshooting program will send detailed status query instructions to each expander at that level. These instructions not only require the expander to feedback its own working status, such as temperature, power consumption, data transfer rate, etc., but also test its data forwarding function. Once the status information returned by an expander is abnormal, or data loss, errors, etc. occur during the data forwarding test, the system will mark it as a possible faulty expander. Subsequently, further diagnostic procedures, such as repeatedly sending test instructions and performing different types of data transfer tests, are used to finally determine whether the expander is a faulty expander. This process is like following a complex pipeline system, checking each valve and connection section by section to find the fault point that causes the water flow to be interrupted (data transmission interrupted).
[0095] After successfully identifying the faulty expander, the system's fault handling module will immediately issue a series of control instructions to instruct other normal expanders at the same level to quickly take over the work tasks of the faulty expander. In order to ensure the continuity and stability of data transmission, the fault handling module will carefully plan the switching plan of the data link. In this process, the first data link involved, the second data link at different levels, and the third data link will be reconfigured.
[0096] Furthermore, in a possible implementation of this embodiment, the step of gradually identifying a faulty expander in the expander group includes:
[0097] Acquire the operation status information of the expander, and determine whether the current level expander has a fault according to the operation status information;
[0098] If there is a fault, other expanders at the same level are controlled to take over the faulty expander and switch the data link;
[0099] If there is no fault, continue to check the expander at the next level until the expander group is checked.
[0100] Specifically, the system obtains all-round operating status information of the expander through a specially designed data acquisition interface and monitoring program. This information covers many key parameters at the hardware level, such as the operating temperature of the expander chip, the power supply voltage, the current consumption of each port, etc. These parameters directly reflect the health of the expander hardware. At the same time, it also includes operating data at the software level, such as the real-time rate of data transmission, the number of data packets sent and received, and the statistics of error codes that occur during data verification. This information is crucial for judging the working status of the expander during data processing and transmission.
[0101] After obtaining detailed operating status information, the system will conduct an in-depth analysis of this data. The algorithm accurately compares the operating parameters of the current level expander with the standard normal threshold range to determine whether the current level expander is faulty.
[0102] Once it is determined that the current level expander has a fault, the system's fault handling module will send control instructions to other normal expanders at the same level. Inform other expanders of the specific work tasks that need to take over the faulty expander, including key information such as the target address of data forwarding and the priority of data processing, and provide detailed instructions on how to switch data links. When switching data links, the system will comprehensively consider factors such as the link's bandwidth utilization, transmission delay, and signal stability. Prioritize links with sufficient bandwidth and low current load to plan an optimal route for data traffic.
[0103] If the system determines that there is no fault in the current level expander after analysis, the system will not stop troubleshooting. It will continue to check the expander at the next level in the established hierarchy order. The system will continue to repeat the above steps of obtaining operating status information and analyzing and judging faults until every level of the entire expander group has been checked. Through this comprehensive, systematic and step-by-step inspection method, the system can accurately locate the faulty expander in the expander group, providing a solid foundation for subsequent fault repair and system recovery.
[0104] It should be noted that the embodiments of the present disclosure may include multiple steps. For the convenience of description, these steps are numbered, but these numbers do not limit the execution time slots or execution order between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not limit this.
[0105] Furthermore, an embodiment of the present disclosure also provides a server, which includes the disk system with link redundancy in the above embodiment, and processes the fault of the disk system with link redundancy by using the above fault handling method.
[0106] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0107] Figure 7 A schematic block diagram of an example electronic device 300 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0108] like Figure 7 As shown, the device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 302 or a computer program loaded from a storage unit 308 to a RAM (Random Access Memory) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An I / O (Input / Output) interface 305 is also connected to the bus 304.
[0109] A number of components in the device 300 are connected to the I / O interface 305, including: an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0110] The computing unit 301 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as a fault handling method. For example, in some embodiments, the fault handling method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 308. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of the method described above may be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute the aforementioned fault handling method in any other appropriate manner (for example, by means of firmware).
[0111] Various embodiments of the systems and techniques described above herein may be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor that may be a special purpose or general purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0112] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0113] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory) or a flash memory, an optical fiber, a CD-ROM (Compact Dis sc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0115] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0116] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server of a distributed system, or a server combined with a blockchain.
[0117] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0118] The various numerical numbers such as first and second involved in the present disclosure are only for the convenience of description and are not used to limit the scope of the embodiments of the present disclosure, but also indicate the order of precedence.
[0119] At least one in the present disclosure may also be described as one or more, and a plurality may be two, three, four or more, which is not limited in the present disclosure. In the embodiments of the present disclosure, for a technical feature, the technical features in the technical feature are distinguished by "first", "second", "third", "A", "B", "C" and "D", etc., and there is no order of precedence or size between the technical features described by the "first", "second", "third", "A", "B", "C" and "D".
[0120] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0121] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A link redundant disk system, characterized in that: The system includes: an expander group, a hard disk group, a head adapter card, and a fault processing module; The expander group includes at least two levels of expanders, the first level expander is connected to the head adapter card via a first data link, the upper level expander is connected to at least two expanders in the lower level via at least two second data links, and the hard disk group is connected to at least two last level expanders via at least two third data links respectively; The fault handling module is connected to the head adapter card and is used to identify the faulty expander in the expander group step by step when it is determined that there is a missing hard disk group that cannot be identified; control other expanders at the same level to replace the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link is composed of a first data link, a second data link at different levels, and a third data link.
2. The system according to claim 1, characterized in that The expander group includes a single-level expander; The single-level expander is connected to the head adapter card via a first data link, and the hard disk group is connected to the single-level expander via at least two third data links.
3. The system according to claim 1, characterized in that One end of the first data link is connected to the interface of the head adapter card, and the other end is connected to the uplink port of the first stage expander; One end of the second data link is connected to the downlink port of the upper-level expander, and the other end is connected to the uplink port of the lower-level expander; One end of the third data link is connected to the downlink port of the last-stage expander, and the other end is connected to the interface of the hard disk group.
4. The system according to claim 1, characterized in that The system also includes: a data calling module; The data calling module is connected to the head adapter card and is used to send a calling instruction of target data to the hard disk group; in response to the calling instruction, the hard disk group uses one of the at least two third data links to send the target data to the last-level expander, and the remaining third data link is closed; The last-stage expander transmits the target data upward to the first-stage expander by using an extended data link, wherein the extended data link is generated by combining one second data link of each level in the expander group level by level, and the remaining second data links are closed; The first-level expander sends the target data to the head adapter card through a first data link.
5. The system according to claim 1, characterized in that When the expander group includes two stages of expanders, the last stage expander of the expander group is a second stage expander; The first-stage expander is connected to the head adapter card via a first data link, and the first-stage expander is connected to at least two second-stage expanders via at least two second data links respectively; The hard disk group is connected to at least two second-level expanders respectively through at least two third data links.
6. The system according to claim 1, characterized in that The fault handling module is also used to obtain the operating status information of the expander, and determine whether the current level expander has a fault based on the operating status information; if there is a fault, control other expanders at the same level to replace the faulty expander and switch the data link; if there is no fault, continue to check the expander at the next level until the expander group inspection is completed.
7. A method for troubleshooting, characterized in that: The method is applied to the disk system with link redundancy as described in claims 1-6, and the method comprises: Check the status of the hard disk group to determine whether there is an unrecognizable missing hard disk group; When it is determined that there is an unrecognizable missing hard disk group, the faulty expanders in the expander group are identified step by step; Control other expanders at the same level to take over the faulty expander and switch the data link so as to reconnect the missing hard disk group with the head adapter card, wherein the data link consists of a first data link, a second data link at different levels, and a third data link.
8. The method according to claim 7, characterized in that The step of gradually identifying a faulty expander in the expander group that has a fault includes: Acquire the operation status information of the expander, and determine whether the current level expander has a fault according to the operation status information; If there is a fault, other expanders at the same level are controlled to take over the faulty expander and switch the data link; If there is no fault, continue to check the expander at the next level until the expander group is checked.
9. A server, characterized in that: The server comprises: a disk system with link redundancy as claimed in any one of claims 1 to 6.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 7 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 7-8.
Citation Information
Patent Citations
Expansion chip management method and device, storage medium and electronic equipment
CN116150068A
Hard disk backboard device and communication link fault detection method
CN117991870A