Method and system for replacing and testing data storage devices in a data storage array
Through the data storage array system, the problem of expansion and management of large-scale data storage systems is solved, efficient equipment maintenance and data redundancy management are achieved, and the stability and scalability of the system are improved.
Patent Information
- Application Number
- CN202111373120.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-02
- Filing Date
- 2021-11-19
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2041-11-19
AI Technical Summary
The prior art is difficult to efficiently scale and manage data storage devices in large data storage systems, resulting in high maintenance costs and difficult to adapt to the growth of data storage demand.
The data storage array (DSA) system is used to detect faulty data storage devices through the management server, replace and add new devices, and distribute redundant copies on the logical path, and use I/O servers to provide data access to realize automated management and maintenance of devices.
It realizes efficient management and expansion of data storage arrays, reduces maintenance costs, and improves system stability and data storage capabilities.
Smart Images

Figure CN114579357B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to data storage arrays and, more particularly, to the management of data storage arrays. Background Art
[0002] As data storage systems become larger and more complex, the need for robust maintenance tools to extend the capabilities of operators is an important aspect of software-defined data center management.
[0003] Enterprise RAID systems attempt to provide the required functionality for servicing data storage devices (DSDs) such as hard disk storage devices. Conventional methods require large resources to service such systems, making it difficult to scale as the data storage system grows to accommodate the ever-increasing demands for data storage. Summary of the Invention
[0004] According to an embodiment of the present invention, a system is disclosed that includes a data storage array (DSA) that includes a plurality of data storage devices (DSDs) in an enclosure, and the DSA is configured to distribute redundant copies of data among the plurality of DSDs on a logical path of the DSA. The system further includes an I / O server that couples the DSA to a client node and is configured to provide data access between the client node and the DSA. The system further includes a management server coupled to the DSA. In an embodiment, the management server is configured to detect a failed DSD in the DSA. The management server is further configured to store a first log including a first state of the DSA and the I / O server. The management server is also configured to remove the failed DSD from the logical path of the DSA. The management server is further configured to distribute redundant copies of data among the plurality of DSDs in the logical path of the DSA. The management server is further configured to detect a replacement DSD in the enclosure that replaces the failed DSD. The management server is further configured to add the replacement DSD to the logical path of the DSA. The management server is further configured to store a second log including a second state of the DSA and the I / O server. The management server is further configured to compare the first log and the second log and display an indication of the state of the DSA based on the comparison.
[0005] In certain embodiments, a computer program product for data storage device replacement and testing is disclosed. The computer program product has a computer-readable storage medium having computer-readable program code embodied therewith. In certain embodiments, the computer-readable program code is executable by one or more computer processors to detect a failed data storage device (DSD) in a data storage array (DSA). The computer-readable program code is further executable to store a first log including a first state of the DSA and an I / O server. The computer-readable program code is further executable to remove the failed DSD from the logical path of the DSA. The computer-readable program code is further executable to distribute redundant copies of data among a plurality of DSDs in the logical path of the DSA. The computer-readable program code is further executable to detect a replacement DSD that replaces the failed DSD in an enclosure. The computer-readable program code is further executable to add the replacement DSD to the logical path of the DSA. The computer-readable program code is further executable to store a second log including a second state of the DSA and the I / O server. The computer-readable program code is further executable to compare the first log and the second log and display an indication of the state of the DSA based on the comparison.
[0006] In further embodiments, a method for data storage device replacement and testing is disclosed, including: detecting a failed data storage device (DSD) in a data storage array (DSA). The method further includes storing a first log including a first state of the DSA and an I / O server. The method further includes removing the failed DSD from the logical path of the DSA. The method further includes distributing redundant copies of data among a plurality of DSDs in the logical path of the DSA. The method further includes detecting a replacement DSD that replaces the failed DSD in an enclosure. The method further includes adding the replacement DSD to the logical path of the DSA. The method further includes storing a second log including a second state of the DSA and the I / O server. The method further includes comparing the first log and the second log, and displaying an indication of the state of the DSA based on the comparison. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 A system for disk replacement and testing according to certain embodiments is shown.
[0008] Figure 2 A flowchart depicting a process for detecting and replacing a data storage device (DSD) according to certain embodiments is shown.
[0009] Figure 3 A flowchart depicting a process for detecting a replacement DSD by a successful DSD replacement according to the disclosed embodiments is shown.
[0010] Figure 4Describes a method for DSD replacement and testing according to the disclosed embodiments. Detailed Description
[0011] In the following, reference is made to the embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specifically described embodiments. Instead, any combination of the following features and elements, whether or not related to different embodiments, is expected to implement and practice the intended embodiments. Additionally, although the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether a given embodiment achieves a particular advantage does not limit the scope of the present disclosure. Therefore, the following aspects, features, embodiments, and advantages are illustrative only and are not to be considered elements or limitations of the appended claims unless expressly recited in the claims. Similarly, references to "the invention" should not be construed as a generalization of any inventive subject matter disclosed herein and should not be considered an element or limitation of the appended claims unless expressly recited in one or more of the claims.
[0012] Figure 1 Illustrates a system 100 for disk replacement and testing according to certain embodiments. The system 100 includes an enclosure 105 for holding one or more data storage devices (DSDs) such as DSD 110, DSD 115 to DSD N. The enclosure 105 can be any type of physical enclosure capable of holding a DSD, such as a cabinet, drawer, drum, tower, or other physical structure. A DSD can be any type of device capable of storing computer-readable data, such as a hard disk drive, solid-state drive, non-volatile memory, volatile memory, or other device capable of storing data. A group of DSDs such as DSD 110 to DSD N can be configured as a data storage array (DSA) such as DSA120. The data storage array 120 can be a RAID-type array, a de-clustered array, or configured to store data based on a redundant data storage strategy that can be implemented on the DSDs. The DSA can include systems similar to the IBM Spectrum Scale Native GPFS RAID sold by IBM, as well as other commercially available data storage array systems that utilize RAID and de-clustered storage technologies. In some embodiments, the DSA 120 can be configured virtually, while in other embodiments, the DSA 120 can include a physical architecture that manages the redundant storage strategy for each DSD coupled to the DSA 120.
[0013] The system 100 further includes a management server 125 that manages one or more of the components of the system 100. The management server 125 includes a DSA maintenance module 130 for performing in conjunction with Figure 2 and Figure 3Computer-readable instructions for replacing and testing a DSD for a DSA 120, which are discussed in more detail. The management server 125 may include one or more processors and a memory, which may be locally located in a single computer system, while in some embodiments, one or more components of the management server may be remotely located and accessed via a network. The management server 125 manages the interaction (i.e., data storage and retrieval) between one or more I / O servers and client nodes.
[0014] The system 100 further includes a plurality of I / O servers, such as I / O servers 135 and 140, up to I / O server N, which are configured to provide and receive data to and from the DSA 120 based on instructions received from one or more client nodes, such as client node 145. In this context, the client node 145 may be a single computer, a multi-computer, or a computer network that utilizes access to data housed on the DSA 120 via one or more I / O servers.
[0015] Figure 2 Flowchart 200 for detecting and replacing a DSD according to certain embodiments is depicted. The actions of flowchart 200 are performed based on computer-readable instructions received from Figure 1 the DSA maintenance module 130.
[0016] At 205, the DSA maintenance module 130 detects the failure mode of the DSD, and at 210, identifies the enclosure that contains the DSD having the failure mode, such as enclosure 105. At 215, the DSA maintenance module 130 receives a selection of a recovery group that includes the failed DSA, where the recovery group is a part of the DSA that contains the failed DSD, such as DSA120.
[0017] At 225, the DSA maintenance module 130 confirms the existence of at least one failed DSD in the DSA and obtains the status of each DSD in the DSA. If there is no failed DSD, flowchart 200 exits. If there is a failed DSD in the DSA, the DSA maintenance module 130 receives a selection of the failed DSD at 230.
[0018] Once the failed DSD is determined, the DSA maintenance module 130 issues a series of commands for the DSA to rebalance the data stored among the non-selected DSDs at 235, and stores logs of one or more of the DSA (such as DSA120), one or more I / O servers coupled to the DSA (such as I / O servers 135 and I / O server 140), and the management server (such as management server 125).
[0019] Once the log is backed up at 240, the failed DSD is logically released at 245. In this context, "logically released" means removed from the logical path of the DSA that hosts the failed DSD and removed from further communication with another component of the system (such as Figure 1 system 100). At 250, the failed DSD is replaced, and the DSA maintenance module 130 detects at 255 that the replacement of the DSD has occurred.
[0020] Figure 3 FIG. 300 is a flow chart depicting an execution process for detecting a replacement DSD through a successful DSD replacement according to the disclosed embodiments. At 305, Figure 1 the DSA maintenance module 130 (such as Figure 1 DSA 120) detects a replacement DSD in the DSA. When replacing the DSD, it is possible that the replacement DSD has been previously used and is not new. To detect whether the DSD has been used, the DSA maintenance module 130 reads one or more DSD metadata sectors to see if any data has been written. If data has been written in the replacement DSD metadata sector, then at 310, the DSA maintenance module 130 delivers a command to the replacement DSD to erase all its metadata content.
[0021] Once cleared (if needed), at 315, the replacement DSD is placed in the logical path of the DSA such that it can be accessed as part of the DSA. Further, once the replacement DSD is in the logical path of the DSA, the DSA rebalances its data across all DSDs on its logical path at 320 and displays a successful DSD replacement to the user at 325.
[0022] Figure 4 FIG. 400 is a method for DSD replacement and testing according to the disclosed embodiments. In some embodiments, the method 400 is executed as computer-readable instructions by a DSA maintenance module (such as Figure 1 DSA maintenance module 130).
[0023] At 405, the method 400 detects a failed data storage device (DSD) in a data storage array (DSA), and at 410, stores a first log including a first state of the DSA and an I / O server. In some embodiments, detecting a failed DSD in the DSA includes detecting that the DSA has removed a failed disk from the logical definition of the DSA. In some embodiments, the DSA is one of a clustered array and a RAID array. In some embodiments, detecting a failed DSD in the DSA includes placing a good DSD among multiple DSDs in a failed state.
[0024] At 415, method 400 removes a failed DSD from the logical path of the DSA, and at 420, distributes redundant copies of data among multiple DSDs in the logical path of the DSA.
[0025] At 425, method 400 detects a replacement DSD in the enclosure that replaces the failed DSD, and at 430, adds the replacement DSD to the logical path of the DSA. In some embodiments, detecting the replacement DSD in the enclosure includes displaying an indicator that removes the failed DSD on the failed DSD.
[0026] At 435, method 400 stores a second log including the second state of the DSA and the I / O server, compares the first log and the second log at 440, and displays an indicator of the state of the DSA based on the comparison at 445.
[0027] Embodiments of method 400 may further include: after placing the replacement DSD in the logical path of the DSA, distributing redundant copies of data among multiple DSDs in the logical path of the DSA. Certain further embodiments may include reading data in the metadata sector of the replacement DSD and erasing the data in the metadata sector.
[0028] Above, reference is made to the embodiments presented in the present disclosure. However, the scope of the present disclosure is not limited to the specifically described embodiments. Instead, any combination of features and elements, whether or not related to different embodiments, is expected to implement and practice the expected embodiments. In addition, although the embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether a given embodiment achieves a particular advantage does not limit the scope of the present disclosure. Thus, the aspects, features, embodiments, and advantages discussed herein are illustrative only and are not to be considered elements or limitations of the appended claims, unless explicitly recited in the claims. Similarly, references to "the invention" should not be construed as generalizing any inventive subject matter disclosed herein and should not be considered elements or limitations of the appended claims, unless explicitly recited in one or more of the claims.
[0029] Aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software and hardware aspects, which may all be collectively referred to herein as "circuitry", "module" or "system."
[0030] The present invention can be a system, method, and / or computer program product at any possible level of integration of technical details. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0031] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0032] The computer-readable program instructions described herein can be downloaded to a corresponding computing / processing device from a computer-readable storage medium, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0033] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++) and procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit (including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA)) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit so as to perform aspects of the present invention.
[0034] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0035] These computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture, the article of manufacture including instructions for implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0036] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0037] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions noted in the boxes may occur out of the order noted in the figures. For example, depending on the functionality involved, two boxes shown in succession may actually be completed as one step, executed simultaneously, substantially simultaneously, in partial or all of a time-overlapping manner, or the boxes may sometimes be executed in the reverse order. It will also be noted that each box of the block diagrams and / or flowchart, and combinations of boxes in the block diagrams and / or flowchart, can be implemented by a system based on dedicated hardware that performs the specified functions or acts or a combination of dedicated hardware and computer instructions.
[0038] Embodiments of the present invention may be provided to end users via a cloud computing infrastructure. Cloud computing generally refers to the provision of scalable computing resources as a service over a network. More formally, cloud computing can be defined as providing an abstraction of computing capabilities between computing resources and their underlying technical architecture (e.g., servers, storage devices, networks), such that a shared pool of configurable computing resources can be conveniently and on-demand accessed, and the configurable computing resources can be rapidly configured and released with minimal management effort or service provider interaction. Thus, cloud computing allows users to access virtual computing resources in the "cloud" (e.g., storage, data, applications, and even complete virtualized computing systems) without regard to the underlying physical systems (or the location of those systems) used to provide the computing resources.
[0039] Typically, cloud computing resources are provided to users on a pay-per-use basis, where the user is charged only for the computing resources actually used (e.g., the amount of storage space consumed by the user or the number of virtualization systems instantiated by the user). The user can access any resources residing in the cloud at any time and from anywhere on the Internet. In the context of the present invention, the user can access applications available in the cloud (e.g., via a client node such as client node 145) or related data. For example, the systems and methods disclosed herein can be executed on a computing system in the cloud and the method of performing DSD replacement and testing can be performed in the context of a cloud computing environment. In such a case, the systems and methods disclosed herein can monitor the ongoing operations of the DSD and DSA including the cloud environment, store information related to such monitoring so that the data is available to other applications for inspection and analysis. Doing so allows the user to access the information from any computing system attached to a network (e.g., the Internet) connected to the cloud.
[0040] While the foregoing is directed to embodiments of the present invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof, and the scope thereof is determined by the appended claims.
Claims
1. A system for data storage device replacement and testing, comprising: A data storage array DSA, the data storage array DSA including a plurality of data storage devices DSD in a housing, the DSA being configured to distribute redundant copies of data among the plurality of DSDs on a logical path of the DSA; An I / O server, the I / O server coupling the DSA to a client node and being configured to provide data access between the client node and the DSA; And A management server, the management server being coupled to the DSA, the management server being configured to: Detect a failed DSD in the DSA, wherein detecting a failed DSD in the DSA includes: detecting that the DSA has removed the failed DSD from the logical definition of the DSA; Store a first log including a first state of the DSA and the I / O server; Remove the failed DSD from the logical path of the DSA; Distribute redundant copies of data among the plurality of DSDs in the logical path of the DSA; Detect a replacement DSD in the housing that replaces the failed DSD; Read data in a metadata sector of the replacement DSD and erase the data in the metadata sector; Add the replacement DSD to the logical path of the DSA; Store a second log including a second state of the DSA and the I / O server; Compare the first log and the second log; and Display an indication of the state of the DSA based on the comparison.
2. The system according to claim 1, wherein the management server is further configured to: after placing the replacement DSD in the logical path of the DSA, distribute redundant copies of data among the plurality of DSDs in the logical path of the DSA.
3. The system according to claim 1, wherein The DSA is one of a clustered array and a RAID array.
4. The system according to claim 1, wherein, Detecting a failed DSD in the DSA includes putting a good DSD among the plurality of DSDs in a failed state.
5. The system according to claim 1, wherein, Detecting a replacement DSD in the housing includes displaying an indicator for removing the failed DSD on the failed DSD.
6. A computer program product for data storage device replacement and testing, the computer program product comprising: Computer-readable program code, the computer-readable program code being executable by one or more computer processors to: Detect a failed data storage device DSD in a data storage array DSA, wherein detecting a failed DSD in the DSA includes: detecting that the DSA has removed the failed DSD from the logical definition of the DSA; Store a first log including a first state of the DSA and the I / O server; Remove the failed DSD from the logical path of the DSA; Distribute redundant copies of data among the plurality of DSDs in the logical path of the DSA; Detect a replacement DSD in the housing that replaces the failed DSD; Read data in a metadata sector of the replacement DSD and erase the data in the metadata sector; Add the replacement DSD to the logical path of the DSA; Store a second log including a second state of the DSA and the I / O server; Compare the first log and the second log; and An indication of the status of the DSA is displayed based on the comparison.
7. The computer program product according to claim 6, wherein the computer-readable program code is further configured to: after placing the replacement DSD in the logical path of the DSA, distribute redundant copies of data among a plurality of DSDs in the logical path of the DSA.
8. The computer program product according to claim 6, wherein, The DSA is one of a clustered array and a RAID array.
9. The computer program product according to claim 6, wherein, Detecting a failed DSD in the DSA includes placing a good DSD among the plurality of DSDs in a failed state.
10. The computer program product according to claim 6, wherein, Detecting a replacement DSD in the enclosure includes displaying an indicator for removing the failed DSD on the failed DSD.
11. A method for data storage device replacement and testing, comprising: Detecting a failed data storage device DSD in a data storage array DSA, wherein detecting the failed DSD in the DSA includes: detecting that the DSA has removed the failed DSD from the logical definition of the DSA; Storing a first log including a first state of the DSA and an I / O server; Removing the failed DSD from the logical path of the DSA; Distributing redundant copies of data among a plurality of DSDs in the logical path of the DSA; Detecting a replacement DSD in the enclosure that replaces the failed DSD; Reading data in the metadata sector of the replacement DSD and erasing the data in the metadata sector; Adding the replacement DSD to the logical path of the DSA; Storing a second log including a second state of the DSA and an I / O server; Comparing the first log and the second log; and Displaying an indication of the status of the DSA based on the comparison.
12. The method according to claim 11, the method further comprising: After placing the replacement DSD in the logical path of the DSA, distribute redundant copies of data among a plurality of DSDs in the logical path of the DSA.
13. The method according to claim 11, wherein, The DSA is one of a clustered array and a RAID array.
14. The method according to claim 11, wherein Detecting a failed DSD in the DSA includes placing a good DSD among the plurality of DSDs in a failed state.
Citation Information
Patent Citations
Cloud storage system capable of automatically detecting and replacing failure nodes and method thereof
CN103354503A
Storage management method, equipment and computer readable medium
CN108733307A