Storage device and error processing method
The storage device with multiple controllers and a control unit that dynamically manages cache area errors ensures continuous data operations by excluding affected units from allocation, addressing the challenge of cache area errors leading to system failures.
Patent Information
- Application Number
- JP2023185681
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-14
Smart Images

Figure 2025074687000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a storage device and an error processing method, and is suitable for application to a storage device, for example, a technology that prevents the entire controller from being blocked even if an error occurs in the cache area of the storage device. [Background technology]
[0002] Conventionally, a storage device is equipped with multiple controllers each having a memory and a cache area, and performs data read and write operations with a host via a host I / F (Interface). In the storage device, data is duplicated in the memory and cache area of each controller, and when an error occurs in any memory or cache area, the controller on the side where the error occurred is blocked, and data read and write operations are continued using the other controller (for example, see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2004-199420 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the storage device disclosed in Patent Document 1, when an error occurs in the memory or cache area as described above, the controller on the side where the error occurs cannot be used to continue reading and writing data via the host I / F. Therefore, the storage device as a whole had to be operated in a state where the above-mentioned data duplication was lost.
[0005] The present invention has been made in consideration of the above points, and aims to propose a storage device and an error handling method that, even if an error occurs in a cache area, can minimize the blocking of the controller on the side where the error occurred. [Means for solving the problem]
[0006] In order to solve such problems, the present invention provides a system that includes a plurality of controllers that control data read and write operations with at least one host computer, and the plurality of controllers include a cache area to which a plurality of management units are assigned in which data can be temporarily stored in accordance with the data read and write operations, and a control unit that controls the data read and write operations, while, when an error occurs, determines whether the error occurred within the cache area, and if it is determined that the error occurred within the cache area, excludes a specific management unit among the plurality of management units that includes the location of the error from the allocation targets in the cache area, and controls the data read and write operations using the remaining management units among the plurality of management units.
[0007] In addition, in the present invention, an error processing method in a storage device equipped with multiple controllers each having a control unit that controls data read and write operations with at least one host computer includes a management unit allocation step in which the control unit allocates to the cache area of the multiple controllers multiple management units in which the data can be temporarily stored in accordance with the data read and write operations, a judgment step in which, when an error occurs, the control unit judges whether the error occurred within the cache area, and, if it is determined that the error occurred within the cache area, a control step in which the control unit excludes a specific management unit among the multiple management units that includes the error location from the allocation targets in the cache area, and controls the data read and write operations using the remaining management units among the multiple management units. Effect of the Invention
[0008] According to the present invention, even if an error occurs in a cache area, it is possible to prevent as much as possible the controller on the side where the error occurred from being blocked. [Brief description of the drawings]
[0009] [Figure 1] 1 is a system configuration diagram mainly showing an example of the configuration of a storage device according to an embodiment of the present invention. [Diagram 2] 2 is a diagram illustrating an example of a connection configuration between a CPU and a memory illustrated in FIG. 1; [Diagram 3] FIG. 2 is a diagram illustrating an example of memory settings illustrated in FIG. [Figure 4] 4 is a diagram showing an example of supplementary information regarding settings of the memory shown in FIG. 3. [Diagram 5] FIG. 13 is a diagram illustrating an example of the configuration of a hardware failure management table; [Figure 6] FIG. 13 is a diagram illustrating an example of the configuration of a replacement target management table. [Figure 7] FIG. 2 is a diagram illustrating a configuration of a cache directory. [Figure 8] FIG. 13 is a diagram illustrating an example of the configuration of a reverse lookup table. [Figure 9] FIG. 13 is a diagram illustrating an example of a free list. [Figure 10] FIG. 13 is a diagram illustrating an example of an exclusion list. [Figure 11] 11 is a flowchart illustrating an example of a procedure for error processing in the storage device according to the present embodiment. [Figure 12] 12 is a flowchart showing an example of a specific procedure for the segment removal process shown in FIG. 11 . [Figure 13] FIG. 11 is a diagram showing an example of a procedure for replacement recovery processing. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0011] 1 is a system configuration diagram mainly showing an example of the configuration of a storage device 200 according to this embodiment. The storage device 200 is connected to a host computer (hereinafter abbreviated as "host") 100, and performs data read and write operations between the host 100 and the storage device 200. This storage device 200 is made up of at least one computer, and executes an error processing method described below by operating a control program 204A on the computer.
[0012] The storage device 200 comprises a plurality of controllers, for example, controller 200A and controller 200B. The storage device 200 may be equipped with more than two controllers. The controller 200A has duplicated functions with the controller 200B, and has a configuration substantially similar to that of the controller 200B. Therefore, unless it is particularly necessary in relation to the configuration of the controller 200B, a description of the configuration of the controller 200B will be omitted.
[0013] The controller 200A comprises a host I / F 201A, a CPU (Central Processing Unit) 202A, a memory 203A, a non-volatile memory 207A, a drive I / F 209A, and a drive device 210A.
[0014] The host I / F 201A is connected to the CPU 202A by a signal line of a standard such as PCI express. The host I / F 201A is an interface with the host 100, and is connected to the host 100 by a fiber channel or Ethernet. The host I / F 201A exchanges data with the host 100 by a protocol such as iSCSI (Internet Small Computer System Interface). The host I / F 201A receives read and write requests from the host 100, interprets the contents of the request, and passes it to the CPU 202A.
[0015] The CPU 202A controls data read / write operations and write operations between the cache area 206A and the drive device 210A via the drive I / F 209A in response to a request from the host 100 via the host I / F 201A. Specifically, the CPU 202A performs data read / write operations between the memory 203A using, for example, Channels "0" to "4". Furthermore, the CPU 202A controls data read / write operations between the drive devices 210A and 210B via the drive I / F 209A by the control program 204A while referring to the configuration information 208A of the non-volatile memory 207A and accessing the memory 203A (the cache management area 205A, the cache area 206A, etc.).
[0016] The CPU 202A includes a memory setting unit (corresponding to "memory setting" in the figure) 202A1. This memory setting unit 202A1 is a register for managing the memory space of the memory 203A. The memory setting unit 202A1 stores setting information (see FIG. 3 and FIG. 4 described later) related to memory setting, which will be described later. The setting information related to memory setting is stored in a non-volatile manner in the non-volatile memory 207A as a part of the configuration information 208A, and is read from the non-volatile memory 207A and stored in the memory setting unit 202A1. For example, when viewed from the CPU 202A, this setting information related to memory setting indicates information regarding which start address and which end address each memory space in the memory 203A (for example, memories of ranks "0" and "1", cache area 206A, etc.) continues in the address space of the memory 203A.
[0017] The CPU 202A is connected to the CPU 202B of the other controller 200B via a specified signal line. The CPU 202A may store a request or the like received from the host 100 via the host I / F 201A in the cache area 206B of the other controller 200B. In such a case, the CPU 202A is capable of issuing an instruction to the CPU 202B of the other controller 200B via the signal line to acquire and transfer the request or the like from the cache area 206B and store it in the cache area 206A.
[0018] CPU 202A constantly monitors whether an error has occurred, and when an error occurs, outputs an error number to identify the error, the location where the error occurred (if it is within memory 203A, the start address to end address indicating the location where the error occurred), error information regarding whether an uncorrectable error has occurred, and the number of times a correctable error has occurred.
[0019] The memory 203A is, for example, a dual inline memory module (DIMM) standard memory with a dual rank (Rank "0", "1") configuration. The memory 203A has a control program 204A, a cache management area 205A, and a cache area 206A. The control program 204A controls data read and write operations with the host 100 via the host I / F 201A under the control of the CPU 202A.
[0020] The cache area 206A is a storage area to which a plurality of segments are allocated as an example of a plurality of management units in which data can be temporarily stored in response to data read / write operations between the control program 204A and the host 100. A segment number is assigned to each segment in the cache area 206A. In the cache area 206A, data can be stored in segments that are the subject of allocation under the control of the control program 204A, and data is not stored in specific segments that are excluded from the subjects of allocation.
[0021] Control program 204A is an example of a control unit, and controls data read and write operations under the control of CPU 202A. If an error occurs, it determines whether the error occurred within cache area 206A, and if it determines that the error occurred within cache area 206A, it excludes a specific segment among the multiple segments that includes the location of the error from the allocation targets in cache area 206A, and controls data read and write operations using the remaining segments among the multiple segments.
[0022] The control program 204A determines whether the error is correctable, and if the error is not correctable, excludes the particular segment from being allocated.
[0023] The cache management area 205A is a storage area for storing management information for the cache area 206A. The management information includes the segment numbers of the multiple segments in the cache area 206A and the start addresses to end addresses of each of the multiple segments.
[0024] The cache management area 205A manages a cache directory, a reverse lookup table, a free list, and an exclusion list, which will be described later. The cache directory manages the correspondence between a plurality of segments (storage units) of each logical volume (hereinafter simply referred to as "volume") that stores data, and the segments (management units) of the cache area in the cache area 206A. A segment number is assigned to each segment of each volume.
[0025] The reverse lookup table 205C is a table for deriving, from the segment number of a segment in the cache area 206A that stores certain data, the volume number of the volume that stores the data and the segment number within the volume.
[0026] The usage exclusion list manages identification information (segment numbers within the cache area) indicating the specific management units excluded from allocation targets. The free list indicates the segment numbers of unused segments within the cache area 206A. The cache directory, reverse lookup table, free list, and usage exclusion list will be described in detail later.
[0027] The non-volatile memory 207A stores configuration information 208A in a non-volatile manner. This configuration information 208A includes setting information related to the settings of the memory described above, a hardware failure management table 211, and a replacement target management table 212, the details of which will be described later. The hardware failure management table 211 is a table for managing the location of an error in the cache area 206A. The replacement target management table 212 is a table for managing a cache, including the cache area 206A having a specific segment excluded from the allocation target, as a replacement target. The details of the configuration information 208A will be described later.
[0028] The drive device 210A is at least one drive device capable of storing data in a non-volatile manner, for example, an SSD (Solid State Drive). In this embodiment, for example, a RAID (Redundant Array of Independent Disks) is configured using a plurality of drive devices 210A. The drive device 210A may also be, for example, a magnetic disk device. In this embodiment, at least one volume is configured using a plurality of drive devices 210A. Data as a target for read / write operations between the host 100 and the drive device 210A can be stored in each segment of the volume.
[0029] The drive I / F 209A is connected to the CPU 202A by a signal line of a standard such as PCI-Express (Peripheral Component Interconnect-Express). The drive I / F 209A is an interface with the drive device 210A and the drive device 210B, and controls data read / write operations using a protocol such as Serial Attached SCSI (Small Computer System Interface).
[0030] In the storage device 200 described above, when the CPU 202A receives a data write request from the host 100 via the host I / F 201A, the data is temporarily stored in the cache area 206A and also stored in the cache area 206B of the other controller 200B via the specified signal line, thereby duplicating and managing the data.
[0031] When duplication is completed in this manner, CPU 202A notifies host 100 via host I / F 201A of the completion of the write operation to cache area 206A. Furthermore, CPU 202A stores the data temporarily stored in cache area 206A in a volume constituted by drive devices 210A and 210B as described below, and when the storage is completed, notifies host 100 via host I / F 201A of the completion of the write operation to the volume.
[0032] On the other hand, when the CPU 202A receives a data read request from the host 100 via the host I / F 201A, the CPU 202A reads the data from the volume and temporarily stores it in the cache area 206A. The CPU 202A provides the data thus temporarily stored in the cache area 206A to the host 100 via the host I / F 201A.
[0033] Fig. 2 is a diagram showing an example of a connection configuration between the CPU 202A and the memory 203A shown in Fig. 1. As described above, the memory 203A is, for example, a memory conforming to the DIMM (Dual Inline Memory Module) standard with a dual rank (Ranks "0" and "1") configuration, and the CPU 202A is connected to the memory 203A of Rank 0 and the memory 203A of Rank 1 in each of the four channels, and controls the data read / write operation between the CPU 202A and the memory 203A for each channel.
[0034] Fig. 3 is a diagram showing an example of the settings of memory 203A shown in Fig. 1. The setting information regarding the settings of memory 203A is a part of configuration information 208A described above. The item "Channel selection bit" has a setting value of "6-7". The item "Rank selection bit" has a setting value of "8".
[0035] Fig. 4 is a diagram showing an example of supplementary information regarding the settings of memory 203A shown in Fig. 3. The physical addresses are in the range of 0 to 40, for example. Of these, physical addresses "6-7" manage the above-mentioned "00" of Channel "0", "01" of Channel "1", "10" of Channel "2", and "11" of Channel "3". When physical address "8" is "0", it indicates rank "0", and when physical address "8" is "1", it indicates rank "1".
[0036] 5 is a diagram showing an example of the configuration of the hardware failure management table 211. As described above, the hardware failure management table 211 is a part of the configuration information 208A. The hardware failure management table 211 manages, for each error that occurs in the memory 203A, an error number, an address where the error occurred, information on whether or not an uncorrectable error has occurred, and the number of times a correctable error has occurred.
[0037] When an error occurs in the memory 203A, the CPU 202A registers in the hardware failure management table 211 the error number, the address where the error occurred, error information regarding whether or not an uncorrectable error has occurred, and the number of times the correctable error has occurred.
[0038] 6 is a diagram showing an example of the configuration of the replacement target management table 212. As described above, the replacement target management table 212 is a part of the configuration information 208A. The replacement target management table 212 shows a list of items that have become replacement targets due to the occurrence of an error (hereinafter also referred to as a "replacement target list"). The replacement target management table 212 manages, for example, that the rank "1" of the channel "0" of the DIMM (memory 203A) standard is the item to be replaced, and that the serial number of the item is "0x12345678". The item to be replaced is written in this replacement target management table 212 by the CPU 202A every time an uncorrectable error occurs.
[0039] 7 is a diagram showing the configuration of the cache directory 213. The cache directory 213 is stored in the above-mentioned cache management area 205A. The cache directory 213 manages volume numbers for identifying each volume, segment numbers within the volumes, segment numbers within the cache area 206A, and attributes. The cache directory 213 is, for example, a table for looking up attributes of segment numbers within the cache area 206A from volume numbers and segment numbers within the volumes. Note that the cache directory 213 may be configured to manage addresses corresponding to each of the segment numbers.
[0040] The attributes are, for example, "Clean" and "Dirty." The attribute "Clean" means that the data stored in the segment of the volume with the segment number of the volume number matches the data stored in the segment of the segment number in the cache area 206A. On the other hand, the attribute "Dirty" means that the data stored in the segment of the volume with the segment number of the volume number does not match the data stored in the segment of the segment number in the cache area 206A.
[0041] That is, the attribute "Clean" indicates a case where the data stored in the segment of the segment number in the cache area 206A matches the data written in the drive device 210A, and it is acceptable to lose the data in the cache area 206A. On the other hand, the attribute "Dirty" indicates a case where the data stored in the segment of the segment number in the cache area 206A does not match the data written in the drive device 210A, and the data in the cache area 206A must not be lost. Note that when there is such a mismatch, a data write request from the host 100 is received but is not reflected in the volume, but between the multiple controllers 200A and 200B, the data has been duplicated in the same way as when there is a match. Also, the attribute "-" indicates that the data is not stored in any segment in the cache area 206A.
[0042] For example, it can be seen that the data stored in the segment with segment number "3" in cache area 206A corresponding to segment number "0" of volume with volume number "0" has the attribute "Clean", and the data stored in the segment with the segment number in cache area 206A matches the data written to drive device 210A. Note that in Fig. 7, not all entries are held as text data, but may have, for example, a tree structure or a hash structure.
[0043] 8 is a diagram showing an example of the configuration of the reverse lookup table 205C. The reverse lookup table 205C is stored in the above-mentioned cache management area 205A. As described above, the reverse lookup table 205C is a table for deriving the volume number of the volume storing certain data and the segment number within the volume from the segment number of the segment within the cache area 206A storing the data. In other words, by referring to the reverse lookup table 205C, reverse lookup of the above-mentioned cache directory 213 is possible.
[0044] The reverse lookup table 205C manages segment numbers, volume numbers, and segment numbers within the volumes within the cache area 206A. The segment numbers, volume numbers, and segment numbers within the volumes within the cache area 206A are the same as those already explained in FIG. 8, so explanations will be omitted.
[0045] 9 is a diagram showing an example of the free list 205D. The free list 205D is stored in the cache management area 205A described above, and is updated by the control program 204A under the control of the CPU 202A. As described above, the free list 205D indicates the segment numbers of unused segments in the cache area 206A. In the example shown in the figure, it can be seen that at least segment numbers "0" and "5" are unused.
[0046] 10 is a diagram showing an example of the usage exclusion list 205E. The usage exclusion list 205E is stored in the above-mentioned cache management area 205A, and is updated by the control program 204A under the control of the CPU 202A. In the illustrated example, the segment with the segment number "2" in the cache area 206A is the target of usage exclusion.
[0047] The storage device 200 according to the present embodiment has the above configuration, and next, an example of an error processing method in the storage device 200 will be described. The error processing method is an error processing method in the storage device 200 equipped with a plurality of controllers 200A, 200B each having a CPU 202A that controls data read / write operations with at least one host computer 100, and includes a management unit allocation step in which the control program 204A causes the CPU 202A to allocate a plurality of segments in which data can be temporarily stored in accordance with data read / write operations to the cache area 206A of the plurality of controllers 200A, 200B, a determination step in which, when an error occurs, the control program 204A causes the CPU 202A to determine whether the error occurred within the cache area 206A, and, when it is determined that the error occurred within the cache area 206A, a control step in which the control program 204A causes the CPU 202A to exclude a specific segment including the error occurrence location from the allocation targets in the cache area 206A, and causes the CPU 202A to control data read / write operations using the remaining segments among the plurality of segments.
[0048] 11 is a flow chart showing an example of the procedure of error processing of the storage device 200 according to this embodiment. The error processing is executed by the control program 204A under the control of the CPU 202A, but in the following explanation, for the sake of simplicity, it is explained as being executed by the control program 204A. In the following explanation, the controller 200A will be mainly explained unless the controller 200B is particularly related.
[0049] First, as described above, the free list 205D manages the segment numbers of segments in which no data is stored in the cache management area 205A, and the control program 204A uses each segment in the cache management area 205A while referring to the free list 205D.
[0050] In the storage device 200, the control program 204A of the controller 200A constantly detects whether there are any errors, and when an error occurs, it determines whether the error is within the cache area 206A, and if the error occurs within the cache area 206A, it obtains the error number, whether the error is correctable, and the error address corresponding to the location where the error occurred.
[0051] In step S10, when an error occurs, the control program 204A obtains the address where the error occurred from the hardware failure management table 211. In step S20, the control program 204A judges whether the error is correctable or not. If the control program 204A judges that the error is not correctable, it executes step S30, whereas if the control program 204A judges that the error is correctable, it executes step S50, which will be described later.
[0052] In step S30, the control program 204A judges whether or not continued use is possible. Specifically, the control program 204A judges that continued use is possible if the error is correctable and the number of errors occurring in the segment including the address where the error occurred is equal to or less than a predetermined threshold.
[0053] If the control program 204A determines that continuous use is possible, it executes step S40. Conversely, if the control program 204A determines that continuous use is not possible, it executes step S50, which will be described later.
[0054] In step S40, the control program 204A records the error number indicating the occurrence of the error in the hardware failure management table 211, and ends the error processing.
[0055] On the other hand, in step S50, the control program 204A judges whether the error occurred within the cache area 206A. If it is not judged that the error occurred within the cache area 206A, the control program 204A executes a controller blocking process. In this controller blocking process, the control program 204A blocks the controller in which the error occurred (here, the controller 200A is exemplified) (step S60).
[0056] On the other hand, if it is determined that the error occurred within the cache area 206A, the control program 204A determines whether the error occurs locally (step S70). Specifically, in a situation where errors occur at multiple locations in the cache area 206A, the control program 204A determines whether the errors in the cache area 206A are concentrated within a specific area.
[0057] If the control program 204A determines that there is no locality, it executes the controller blocking process described above, whereas if it determines that there is locality, it refers to the usage exclusion list 205E and identifies the segment number in the cache area 206A (step S80).
[0058] The control program 204A executes a segment exclusion process (step S90). Specifically, the control program 204A adds the segment number of the segment to the usage exclusion list 205E (see FIG. 10), which is a part of the configuration information 208A. Note that, as described above, when the segment number of the segment is added to the usage exclusion list 205E, the control program 204A no longer uses the segment with that segment number. Details of this segment exclusion process will be described later.
[0059] The control program 204A executes an error occurrence memory identification process (step S100) to identify the memory in which the error occurs. Specifically, as described above, the control program 204A constantly acquires error information, and identifies the memory in which the error occurs (here, the memory 203A is illustrated) based on the error information.
[0060] The control program 204A registers the memory 203A in question in the replacement target management table 212 (see FIG. 6) (step S110).
[0061] Fig. 12 is a flow chart showing an example of a specific procedure of the segment removal process shown in Fig. 11. In step S21, the control program 204A judges whether or not the segment is in use. Specifically, the control program 204A judges whether or not the segment is in use by referring to the free list 205D. If the segment is not in use, the control program 204A executes step S27 described later.
[0062] On the other hand, if the control program 204A determines that the cache area 206A is in use, it refers to the reverse lookup table 205C and identifies the volume number of the volume that is using the cache area 206A where the error occurred and the segment number within the volume (step S22).
[0063] The control program 204A refers to the cache directory 213 and checks the attribute corresponding to the volume number (step S23). If the attribute is "Clean", the control program 204A executes step S25, which will be described later, whereas if the attribute is "Dirty", the control program 204A executes step S24.
[0064] In this step S24, the control program 204A requests the other controller 200B to reflect the data of the segment of the relevant volume. Next, in step S25, the control program 204A changes the relevant entry in the cache directory 213 of its own controller 200A to "-" (not cached).
[0065] Next, in step S26, the control program 204A deletes the reference from the reverse lookup table 205C. Furthermore, the control program 204A adds the segment number in the corresponding cache area 206A to the usage exclusion list 205E.
[0066] 13 is a diagram showing an example of a procedure for the replacement recovery process. In step S31, the control program 204A refers to the serial number of the memory 203A registered as the replacement target list in the replacement target management table 212.
[0067] The control program 204A judges whether or not it matches the serial number in the replacement target management table 212. If the control program 204A judges that it does not match (for example, when the memory 203A is replaced), it deletes from the error occurrence address list the address group corresponding to the replaced memory 203A (step S33).
[0068] On the other hand, if the control program 204A determines that they match (for example, if the memory 203A has not been replaced), it executes step S34. In step S34, the control program 204A extracts a group of addresses that correspond to the memory 203A from the error occurrence address list.
[0069] Next, in step S35, the control program 204A calculates the segment number in the cache area 206A that corresponds to each address. In step S36, the control program 204A adds the calculated segment number to the usage exclusion list 205E.
[0070] The storage device 200 according to this embodiment includes a plurality of controllers 200A, 200B that control data read and write operations with at least one host computer. The plurality of controllers 200A, 200B include a cache area 206A to which a plurality of segments (management units) are assigned in which data can be temporarily stored depending on the data read and write operations, and a control unit (CPU 202A, control program 204A) that controls the data read and write operations, while determining, when an error occurs, whether the error occurred within the cache area 206A or not, and if it is determined that the data originated within the cache area 206A, excluding a specific segment among the plurality of segments including the location of the data origin from the allocation targets in the cache area 206A, and controlling the data read and write operations using the remaining segments among the plurality of segments.
[0071] With this configuration, even if an error occurs in the cache area 206A, it is possible to prevent the controller on the side where the error occurred from being blocked as much as possible. Also, even if an error occurs in the cache area 206A, the controller on the side where the error occurred can continue to operate, so that the storage device 200 can continue data read / write operations by the multiple controllers 200A and 200B as much as possible. Therefore, when an error occurs in the other controller and the other controller is blocked, it is possible to reduce the frequency of system downtime for the entire storage device 200.
[0072] In this embodiment, the control program 204A, under the control of the CPU 202A, determines whether an error that has occurred in the cache area 206A is correctable or not, and if the error is not correctable, excludes a specific segment from the allocation target. In this way, it is possible to prevent the controller including the cache area 206A in which a correctable error has occurred from being blocked, and to ensure stable operation of the storage device 200 as a whole.
[0073] In this embodiment, the multiple controllers 200A, 200B are provided with a usage exclusion list 2005E that manages segment numbers indicating specific segments that have been excluded from allocation targets. In this way, it is possible to perform control so that the control program 204A does not erroneously use segments of the cache area 206A having segment numbers managed in the usage exclusion list 205E.
[0074] In this embodiment, the multiple controllers 200A, 200B include a cache directory 213 that manages the correspondence between multiple segments of a volume that stores data and multiple segments in the cache area 206A that stores the data. In this way, the control program 204A can easily grasp the correspondence between the segment of the volume in which the data is stored and the segment of the cache area 206A in which the data is stored.
[0075] In this embodiment, the multiple controllers 200A, 200B are provided with a hardware failure management table 211 that manages the location of data occurrence in the cache area 206A. In this way, by referencing the hardware failure management table 211, it is possible to easily identify the hardware to be replaced.
[0076] In this embodiment, the multiple controllers 200A, 200B are provided with a replacement target management table 212 that manages, as a replacement target, a cache including a cache area 206A having a specific segment excluded from allocation targets. In this way, by referring to the hardware failure management table 211, it is possible to easily recognize the cache to be replaced.
[0077] The present invention is not limited to the above-described embodiment, and includes various modified examples and equivalent configurations within the spirit of the appended claims. For example, the above-described embodiment has been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to those having all of the described configurations. In addition, each element described in parallel in this embodiment may be in a form in which at least one of the elements is connected in series to the other elements. [Industrial Applicability]
[0078] The present invention can be applied to a storage device relating to a technique for preventing the entire controller from being blocked even when an error occurs in the cache area of the storage device. [Explanation of symbols]
[0079] 100... host computer, 200A, 200B... controller, 201A, 201B... host I / F, 202A, 202B... CPU, 203A, 203B... memory, 204A, 204B... control program, 205A, 205B... cache management area, 206A, 206B... cache area, 207A, 207B... non-volatile memory, 207A, 207B... configuration information, 211... hardware failure management table, 212... replacement target management table, 213... cache directory
Claims
1. A plurality of controllers for controlling data read / write operations between at least one host computer, The plurality of controllers include a cache area to which a plurality of management units are assigned in which the data can be temporarily stored in response to a read / write operation of the data; a control unit which controls the read and write operations of the data, and when an error occurs, determines whether the location of the error is within the cache area, and when it is determined that the location of the error is within the cache area, excludes a specific management unit among the plurality of management units that includes the location of the error from targets for allocation in the cache area, and controls the read and write operations of the data using the remaining management units among the plurality of management units; A storage device comprising:
2. The control unit is determining whether the error is correctable or not, and excluding the specific management unit from the allocation target if the error is not correctable; 2. The storage device according to claim 1.
3. The plurality of controllers include A usage exclusion list for managing identification information indicating the specific management unit excluded from allocation targets is provided.
3. The storage device according to claim 2.
4. The plurality of controllers include a cache directory that manages the correspondence between a plurality of storage units of a volume that stores the data and the plurality of management units in the cache area that stores the data; 2. The storage device according to claim 1.
5. The plurality of controllers include a hardware failure management table for managing the location of the failure in the cache area; 2. The storage device according to claim 1.
6. The plurality of controllers include a replacement target management table for managing a cache including the cache area having the specific management unit excluded from the allocation target as a replacement target; 2. The storage device according to claim 1.
7. 1. An error processing method in a storage device equipped with a plurality of controllers each having a control unit for controlling data read / write operations between at least one host computer, comprising: a management unit allocation step of allocating a plurality of management units in which the data can be temporarily stored in accordance with a read / write operation of the data to the cache areas of the plurality of controllers by the control unit; a determination step in which, when an error occurs, the control unit determines whether the error has occurred within the cache area; a control step in which, when it is determined that the occurrence location is within the cache area, the control unit excludes a specific management unit among the plurality of management units that includes the occurrence location from allocation targets in the cache area, and controls the read / write operation of the data using the remaining management units among the plurality of management units; 13. An error processing method comprising the steps of:
8. The control unit, in the management unit allocation step, determining whether the error is correctable or not, and excluding the specific management unit from the allocation target if the error is not correctable; 8. The method of claim 7, wherein the error processing is performed by the first processor.
9. The plurality of controllers include a usage exclusion list for managing identification information indicating the specific management unit excluded from allocation targets; In the control step, the control unit Controlling the data read / write operation so as not to use the specific management unit corresponding to the identification information managed in the usage exclusion list 9. The method of claim 8, wherein the error is processed by a processor.
10. The plurality of controllers include a cache directory for managing a correspondence relationship between a plurality of storage units of a volume for storing the data and the plurality of management units in the cache area for storing the data; In the control step, the control unit The cache directory is referenced, and a predetermined storage unit in the volume corresponding to a predetermined management unit in the cache area is identified.
8. The method of claim 7, wherein the error processing is performed by the first processor.
Citation Information
Patent Citations
Computer system, magnetic disk device, and method for controlling disk cache
JP2004199420A