Raid maintenance system
The RAID maintenance system addresses the long rebuild time issue by using virtual drives and differential data transfer to expedite the replacement and rebuild process, enhancing maintenance efficiency and productivity.
Patent Information
- Application Number
- JP2023216103
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-03
AI Technical Summary
The long rebuild time during RAID maintenance poses a challenge, especially with increasing storage device capacities, leading to prolonged downtime and reduced maintenance efficiency.
A RAID maintenance system comprising a disk array device and a network storage device that allows for the creation of virtual drives, enabling data copying from physical disks to virtual drives, followed by differential data transfer to a replacement storage device, thereby facilitating rapid rebuilds.
This system significantly reduces rebuild time by allowing pre-copying of data to a replacement storage device, improving maintenance productivity and enabling remote rebuilds even when the maintenance site is distant from the storage location.
Smart Images

Figure 2025099439000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments relate to a RAID maintenance system.
Background Art
[0002] A computer operates by reading and writing various data to an external storage device such as an HDD or an SSD. When the external storage device fails, reading and writing of various data may become impossible. As a method to prevent this, the storage device may be configured as a RAID (Redundant Array of Independent Disks or Redundant Array of Inexpensive Disks). There are various levels in the RAID configuration. RAID0 is striping, RAID1 is mirroring, and RAID5 is parity distribution, etc. For example, RAID1 is a technology that writes the same data to two disks to make the data redundant. Even if one of the external storage devices fails, the computer can continue to operate with the other normal external storage device, and the data can also be saved normally.
[0003] In RAID, a failed external storage device is replaced with a new one. After replacing it with a new external storage device, data is copied from a normal storage device to the new storage device to ensure redundancy again. Thus, copying the data of the existing storage device to the new storage device and restoring it to the original state is called a rebuild.
[0004] In recent years, the capacity of storage devices has been increasing, and it often takes a long time for a rebuild. In such a case, there is a problem that the waiting time for the maintenance worker until the rebuild is completed during the maintenance work becomes long.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] An embodiment of the present invention aims to provide a RAID maintenance system capable of shortening the time until the completion of a rebuild.
Means for Solving the Problems
[0007] According to the embodiment, the RAID maintenance system includes a disk array device having a plurality of storage devices, and a network storage device capable of setting a plurality of virtual drives and communicably connected to the disk array device via a first communication network. The network storage device can be communicably connected to a RAID storage device via a second communication network. The RAID storage device has a plurality of physically configured disks. The plurality of virtual drives copy the data of the plurality of physical disks respectively. In response to one of the plurality of physical disks stopping, the copy of the data to one of the plurality of virtual drives is stopped, all the data of the one virtual drive is copied to one of the plurality of storage devices, and after the copy to the one storage device, the differential data between the data of the one storage device and the data of the remaining virtual drives other than the one virtual drive among the plurality of virtual drives is copied to the one storage device, and then the one physical drive is replaced with the one storage device to implement RAID.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Mode for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the drawings are schematic or conceptual, and the relationship between the thickness and width of each part, the ratio of the sizes between parts, etc. are not necessarily the same as those in reality. Also, even when representing the same part, the dimensions and ratios may be represented differently in the drawings. In the present specification and each figure, elements similar to those described above with respect to the previously shown figures are denoted by the same reference numerals, and detailed descriptions thereof are omitted as appropriate.
[0010] (First Embodiment) FIG. 1 is a schematic block diagram exemplifying the RAID maintenance system according to the first embodiment. As shown in FIG. 1, the RAID maintenance system 100 according to the present embodiment includes a disk array device 10 and a network storage device 20. The disk array device 10 is connected to the network storage device 20 via a communication network 1.
[0011] In the RAID maintenance system 100, the disk array device 10 is installed at site A, and the network storage device 20 is installed at site B. Site A and site B may be in remote locations, or site B may be within the premises of site A or adjacent to site A. Site B may be any location where the network storage device 20 can be installed. Site A is, for example, the office of a maintenance contractor for information processing devices that operates the data center of site B. For example, the network storage device 20 is a data center owned by the maintenance contractor of site A.
[0012] The RAID maintenance system 100 can be connected to the information processing device 50 via the communication network 2. The information processing device 50 is installed at site C. Site C can be a location different from sites A and B and may be a remote location from site A. The information processing device 50 has a RAID storage device 52. In this example, the RAID storage device 52 has a RAID configuration at the RAID1 level. The RAID storage device 52 constitutes a single logical drive LD0 composed of two physical disks (D0, D1) 52a, 52b.
[0013] In the RAID maintenance system 100, the network storage device 20 includes a network RAID storage device 22. The network RAID storage device 22 has the same configuration as the RAID storage device 52 at site C. That is, the network RAID storage device 22 is a virtual logical drive VLD0 composed of two virtual drives (VLD0, VLD1) 22a, 22b. In the network storage device 20, the reason for calling it a virtual drive is that the virtual drive and the physical drive do not necessarily physically coincide. For example, in the network storage device 20, there may be one physical drive, and in the example of FIG. 1, one physical drive is assigned to eight virtual drives. In the virtual drive of the network storage device 20, a disk area is set to correspond to the physical disks 52a, 52b of the RAID storage device 52 installed at site C, and it appears to operate as two drives.
[0014] The operation of the RAID maintenance system 100 according to this embodiment will be described. FIG. 1 shows the configuration of the RAID maintenance system 100 and the state of the normal operation of the RAID maintenance system 100. As shown in FIG. 1, two virtual drives 22a and 22b configured in the network storage device 20 of the RAID maintenance system 100 are connected to the information processing device 50 via the communication network 2. As shown by the arrows in FIG. 1, the virtual drives 22a and 22b copy the data of the physical disks 52a and 52b that constitute the RAID storage device 52 of the information processing device 50, respectively. That is, the data of the physical disk 52a is copied to the virtual drive 22a, and the data of the physical disk 52b is copied to the virtual drive 22b.
[0015] The copying of data from the physical disks 52a and 52b to the virtual drives 22a and 22b may be executed sequentially, and preferably, it is executed at night or the like when the load on the information processing device 50 is small. In this way, in the RAID maintenance system 100, the data of the virtual drives 22a and 22b is kept identical to the data of the physical disks 52a and 52b, respectively.
[0016] FIGS. 2 to 4 are schematic block diagrams for explaining the operation of the RAID maintenance system according to the first embodiment. FIGS. 2 to 4 show a situation where any one of the physical disks has failed in the RAID storage device 52 managed by the information processing device 50 installed at site C. Specifically, in FIGS. 2 to 4, the physical disk (D1) 52b has failed, and the procedure for rebuilding the RAID storage device 52 is shown.
[0017] As shown in FIG. 2, in the RAID storage device 52, one of the two physical disks 52a and 52b, i.e., the physical disk 52b, stops due to a failure. Due to the stop of the physical disk 52b, the operation of the virtual drive 22b also stops. The physical disk 52a continues to operate, and the virtual drive 22a continues to copy the data of the physical disk 52a.
[0018] In response to the stop of the virtual drive 22b, the network storage device 20 transmits a notification indicating that the virtual drive 22b has stopped to the disk array device 10 via the communication network 1, as indicated by the arrow of the curve in FIG. 2.
[0019] The disk array device 10 that has received the notification indicating that the virtual drive 22b has stopped selects one storage device (D1') 12a, receives the data of the virtual drive 22b via the communication network 1, and copies the received data to the storage device 12a. The data of the virtual drive 22b is copied to the storage device 12a. The data of the virtual drive 22b up to immediately before the stop is stored in the storage device 12a, that is, the data of the physical disk 52b up to immediately before the stop is stored.
[0020] As shown in FIG. 3, the storage device 12a in which the copy of the data of the virtual drive 22b is completed is removed from the disk array device 10 and transported to site C. In the information processing device 50 at site C, the operation continues using the normally operating physical disk 52a.
[0021] As shown in FIG. 4, the RAID storage device 52 is once stopped, and the physical disk 52b is replaced with the storage device 12a in which the data of the virtual drive 22b is copied. In response to the stop of the RAID storage device 52, a comparison between the data of the virtual drive 22a and the data of the virtual drive 22b is started from the network storage device 20. Based on these comparison results, as indicated by the arrow in FIG. 4, information regarding the difference data α between the data of the virtual drive 22a and the data of the virtual drive 22b is notified to the information processing device 50 via the communication network 2.
[0022] Information about the differential data α is, for example, the data with the latest timestamp among the data copied from the virtual drive 22b. In the virtual drive 22b, data input or updated after the time of the data with the latest timestamp exists only in the physical disk 52a and the virtual drive 22a. The information processing apparatus 50 copies data having a timestamp newer than the latest timestamp of the virtual drive 22b from the physical disk 52a to the storage device 12a carried to site C and performs a rebuild.
[0023] After the completion of the rebuild, a predetermined process is performed to set the physical disk 52a and the storage device 12a, which is a new physical disk, as the RAID storage device 52.
[0024] In this way, in the RAID maintenance system 100 according to the present embodiment, the failed physical disk installed at site C can be replaced and the RAID storage device 52 can be rebuilt.
[0025] The effects of the RAID maintenance system 100 according to the present embodiment will be described. The RAID maintenance system 100 according to the present embodiment includes a network storage device 20 in which virtual drives 22a and 22b having the same data as the physical disks 52a and 52b of the RAID storage device 52 to be maintained are set. In the virtual drives 22a and 22b, when any of the corresponding physical disks 52a and 52b stops, the copying of data stops in synchronization with the stopped physical disk. Therefore, in the RAID storage device 52, the differential data α between the data at the time of stop and the data at the time of physical disk replacement after the stop can be easily obtained. The data size of the differential data α is often sufficiently smaller than the data size of the physically operating disk.
[0026] After the physical disk stops, in the disk array device 10, the data of the virtual drive corresponding to the stopped physical disk can be copied in advance to the storage device 12a for replacement.
[0027] The storage device 12a that has copied the data until the physical disk stops is replaced with the stopped physical disk, and by rebuilding the storage device 12a using the differential data α from after the stop to the present, the rebuilding time can be shortened.
[0028] The copy operation of the data from the virtual drive to the storage device 12a until the physical disk stops can be performed at site A where the disk array device 10 is installed. Therefore, the maintenance worker can perform other operations within site A to which he belongs, and the productivity regarding the maintenance work can be improved.
[0029] Since the differential data α can be acquired at site C different from site A, even if site C where the information processing device 50 to be maintained is installed is a remote place from site A, the rebuilding work can be surely performed.
[0030] (Second Embodiment) FIG. 5 is a schematic block diagram illustrating a RAID maintenance system according to the second embodiment. As shown in FIG. 5, the RAID maintenance system 200 according to the present embodiment includes a network storage device 220 different from the RAID maintenance system 100 shown in FIG. 1. In other respects, the configuration of the RAID maintenance system 200 according to the present embodiment is the same as the configuration of the RAID maintenance system 100 in FIG. 1, and the same reference numerals are given to the same components and the detailed description is appropriately omitted. In the specific example of FIG. 5, the information processing device 250 at site C has a RAID storage device 252 at the RAID5 level, but it may be the same as the RAID storage device in FIG. 1 as the RAID1 level.
[0031] As shown in FIG. 5, the network storage device 220 has a virtual drive 222a. The virtual drive 222a is a single drive, and stores the data of the plurality of physical disks 252a to 252c of the RAID storage device 252 by setting the areas in the drive.
[0032] In addition, the virtual drive 222a stores information related to RAID, such as the RAID level, in the area set in the drive. With these, the virtual drive 222a can store the data of a plurality of physical disks 252a to 252c in a single logical drive.
[0033] Specifically, the virtual drive (VLD0) 222a includes areas 222a1 and 222a2. RAID configuration information is stored in area 222a1. The RAID configuration information is, for example, the RAID level, the number of physical disks, and the stripe size. The RAID configuration information can be appropriately set arbitrarily according to the RAID level. In area 222a2, the respective data of the physical disks 252a to 252c that make up the RAID storage device 252 of the information processing device 250 are stored in association with the labels of the physical disks. Note that areas 222a1 and 222a2 only need to be set so that they can be logically identified and do not need to be physically distinguished. For example, when the data of the physical disk 252a is stored in area 222a2, the data may be stored not only within the same physical drive but also within different physical drives.
[0034] The operation of the RAID maintenance system 200 according to this embodiment will be described. FIG. 5 shows the configuration of the RAID maintenance system 200 and the state of the normal operation of the RAID maintenance system 200. As shown in FIG. 5, the virtual drive 222a of the network storage device 220 is connected to the information processing device 250 via the communication network 2. As indicated by the arrows in FIG. 5, the virtual drive 222a copies the data of the physical disks 252a to 252c that make up the RAID storage device 252 of the information processing device 250 to area 222a2 in association with the labels for identifying the physical disks 252a to 252c.
[0035] FIGS. 6 to 8 are schematic block diagrams for explaining the operation of the RAID maintenance system according to the second embodiment. Figures 6 to 8 show the procedure for rebuilding the RAID storage device 252 managed by the information processing device 250 installed at Site C when the physical disk (D2) 252c fails.
[0036] As shown in FIG. 6, in the RAID storage device 52, out of the three physical disks 252a to 252c, one physical disk 252c stops due to a failure. Due to the stop of the physical disk 252c, the copy of the data associated with the label of the physical disk 252c in the area 222a2 of the virtual drive 222a also stops. The physical disks 252a and 252b continue to operate, and the virtual drive 222a continues to copy the data of the physical disks 252a and 252b and associate it with the labels of the physical disks 252a and 252b.
[0037] In response to the stop of the copy of the data associated with the label of the physical disk 252c of the virtual drive 222a, the network storage device 220 transmits a notification indicating that the copy of the data has stopped to the disk array device 10 via the communication network 1, as shown by the arrow of the curve in FIG. 6.
[0038] The disk array device 10 that has received the notification indicating that the copy of the data has stopped selects one storage device (D2’) 12b, receives the data associated with the label of the physical disk 252c in the area 222a2 of the virtual drive 222a via the communication network 1, and copies the received data to the storage device 12b. The data associated with the label of the physical disk 252c is copied to the storage device 12b.
[0039] As shown in FIG. 7, the storage device 12b in which the copy of the data of the virtual drive 222a is completed is removed from the disk array device 10 and transported to Site C. At the information processing device 250 at Site C, the operation continues using the normally operating physical disks 252a and 252b.
[0040] As shown in FIG. 8, RAID252 is once stopped. The failed physical disk 252c is replaced with the storage device 12b in which the data of the virtual drive 222b is copied. Then, in the virtual drive 222a from the network storage device 220, the data associated with the labels of the physical disks 252a and 252b and the data associated with the label of the physical disk 252c are compared. As shown by the straight arrows in FIG. 8, based on these comparison results, information regarding the differential data β between the data associated with the labels of the physical disks 252a and 252b and the data associated with the label of the physical disk 252c is notified to the information processing device 250 via the communication network 2.
[0041] The information regarding the differential data β can be the time stamp of the data, similar to the case of the first embodiment. The differential data β consists of data with a time stamp newer than the latest time stamp among the data associated with the label of the physical disk 252c. The information processing device 250 copies the differential data β from the physical disk 252b to the transferred storage device 12b and performs a rebuild. In this example, since the RAID level is RAID5, the copy source of the differential data β can be one physical disk 252b of the RAID storage device 252. The restoration work can be performed more quickly via the communication network 2 than transferring data from the virtual drive 222a.
[0042] In this way, in the RAID maintenance system 200 according to this embodiment, the failed physical disk installed at site C can be replaced to rebuild the RAID storage device 252.
[0043] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and the equivalent scope thereof.
Explanation of Signs
[0044] 1, 2... communication network, 10... disk array device, 12, 12a, 12b... storage device, 20, 220... network storage device, 22... network RAID storage device, 22a, 22b, 222a... virtual drive, 50, 250... information processing device, 52, 252... RAID storage device, 52a, 52b, 252a~252c... physical disk, 100, 200... RAID maintenance system
Claims
1. A disk array device having a plurality of storage devices, A network storage device capable of setting a plurality of virtual drives and communicably connected to the disk array device via a first communication network, Comprising: The network storage device can be communicably connected to a RAID storage device via a second communication network, The RAID storage device has a plurality of physically configured disks configured in RAID, The plurality of virtual drives, Copy the data of the plurality of physical disks respectively, In response to one of the plurality of physical disks stopping, stop copying data to one of the plurality of virtual drives, Copy all the data of the one virtual drive to one of the plurality of storage devices, After copying to the one storage device, copy the difference data between the data of the one storage device and the data of the remaining virtual drives other than the one virtual drive among the plurality of virtual drives to the one storage device, and then replace the one physical drive with the one storage device and implement RAID. A RAID maintenance system.
2. In response to stopping copying data to the one virtual drive, the network storage device transmits a notification indicating that copying data to the one virtual drive has stopped to the disk array device via the first communication network, The disk array device starts copying data to the one storage device according to the notification. The RAID maintenance system according to claim 1.
3. A disk array device having a plurality of storage devices, A network storage device capable of setting virtual drives and communicably connected to the disk array device via a first communication network, Comprising: The network storage device can be communicably connected to a RAID storage device via a second communication network, The RAID storage device has a plurality of physically configured disks configured in RAID, The virtual drive, Copy the plurality of data stored in the plurality of physical drives respectively in association with a plurality of identification labels for identifying the plurality of physical drives respectively. In response to one of the plurality of physical drives stopping, stop copying the data associated with the identification label of the one physical drive among the plurality of data, stop copying the data to one of the plurality of virtual drives, copy all the data of the one virtual drive to one of the plurality of storage devices, After copying to the one storage device, copy the differential data between the data of the one storage device and the data of the remaining virtual drives other than the one virtual drive among the plurality of virtual drives to the one storage device, and then replace the one physical drive with the one storage device and implement RAID. A RAID maintenance system.
4. In response to stopping the copy of the data associated with the identification label of the one physical drive, the network storage device transmits, via the first communication network, a notification indicating that the copy of the data associated with the identification label of the one physical drive has been stopped to the disk array device, The disk array device starts copying data to the one storage device according to the notification. The RAID maintenance system according to claim 1.
5. The disk array device is installed at a first location separated from the location where the plurality of physical drives are installed. The RAID maintenance system according to any one of claims 1 to 4.
Citation Information
Patent Citations
Disk array device and failure handling method in disk array device
JP2020119233A