Memory system

The memory system addresses the challenge of chip failures in distributed data frames by using a controller that performs parallel error correction across multiple non-volatile memories, ensuring efficient data restoration and low latency.

JP7682741B2Active Publication Date: 2025-05-26KIOXIA CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021144690
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-06
Publication Date
2025-05-26
Estimated Expiration
2041-09-06

AI Technical Summary

Technical Problem

Existing memory systems face challenges in efficiently handling chip failures within distributed data frames across multiple non-volatile memories, leading to potential data loss and increased latency due to repeated error correction attempts.

Method used

The proposed memory system employs a controller that distributes data frames across multiple non-volatile memories, including a second parity for data restoration, and executes error correction in parallel using a combination of ECC decoders and XOR restoration circuits, eliminating the need to identify failed chips.

Benefits of technology

This approach enables effective data error correction that can accommodate chip failures while maintaining low latency, reducing the risk of data loss and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682741000001
    Figure 0007682741000001
  • Figure 0007682741000002
    Figure 0007682741000002
  • Figure 0007682741000003
    Figure 0007682741000003
Patent Text Reader

Abstract

To provide a memory system capable of achieving data error correction which can cope with a chip failure while maintaining low latency.SOLUTION: According to an embodiment, a memory system comprises a plurality of non-volatile memories and a controller. The controller can communicate with a host and controls the plurality of non-volatile memories. The controller dispersedly writes data received from the host and a data frame including a first parity for detecting and correcting data errors into N (N is a natural number of 2 or above) first non-volatile memories among the plurality of non-volatile memories. The controller writes a second parity for restoring data on one first non-volatile memory in the data frame dispersedly written into the N first non-volatile memories from data on N-1 first non-volatile memories other than one first non-volatile memory into a second non-volatile memory other than the N first non-volatile memories in the plurality of non-volatile memories.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to a memory system.

Background Art

[0002] In recent years, as a new layer in the memory hierarchy to bridge the performance gap between main memory (primary memory) and storage (secondary memory), SCM (Storage Class Memory) has attracted attention.

[0003] In an SCM module (a memory system having an SCM and a controller that controls the SCM), in order to reduce latency, a data frame may be distributed and arranged across a plurality of memory chips. A data frame is composed of data received from a host and, for example, ECC (error-correcting code) for detecting and correcting errors in the data. A data frame is also referred to as an ECC frame or the like.

[0004] If a chip failure occurs in any one of the plurality of memory chips in which the data frame is distributed, it exceeds the error correction ability of the ECC significantly. In this case, the data frame in which a part of the data is distributed on the failed chip will result in the loss of all data. Therefore, as a countermeasure against this chip failure, for example, assume that XOR parity for restoring the data on the failed chip from the data on other memory chips is further added and the XOR parity is arranged in a memory chip different from the plurality of memory chips in which the data frame is distributed.

[0005] However, even if XOR parity is added, if the failed chip cannot be identified, assuming the failed chip, the repetition of (1) data restoration by XOR parity and (2) detection and correction of data errors by ECC in the order of (1) → (2) may occur up to the number of memory chips at worst. The possibility of such a situation occurring is fatal in SCM characterized by low latency. [Prior art documents] [Patent documents]

[0006] [Patent Document 1] U.S. Patent No. 10901840 Summary of the Invention [Problem to be solved by the invention]

[0007] One embodiment of the present invention provides a memory system that provides data error correction that can accommodate chip failures while maintaining low latency. [Means for solving the problem]

[0008] According to an embodiment, a memory system includes a plurality of non-volatile memories and a controller. The controller is capable of communicating with a host and controls the plurality of non-volatile memories. The controller performs a read / write operation on data received from the host or the data received from the host. Data on which compression processing has been performed data Or data to which data necessary for management has been added to the data received from the host and a data frame including a first parity for detecting and correcting an error in the data, and the data frame is distributed and written to N (N is a natural number equal to or greater than 2) first nonvolatile memories among the plurality of nonvolatile memories. See A second parity for restoring data in one first nonvolatile memory among a data frame written in a distributed manner among the N first nonvolatile memories from data in N-1 first nonvolatile memories other than the one first nonvolatile memory is written to a second nonvolatile memory among the multiple nonvolatile memories other than the N first nonvolatile memories. The controller generates a first data frame from N pieces of data read from each of N first non-volatile memories, and N-1 pieces of data read from each of N-1 first non-volatile memories among the N first non-volatile memories, and N-1 pieces of data and a second parity read from the second non-volatile memory, and generates N different second data frames that can be generated from the data on one first non-volatile memory other than the N-1 first non-volatile memories restored thereby, and at least a part of the decoding process for detecting and correcting data errors by the first parity for the first data frame or the second data frame is executed in parallel. [Brief description of the drawings]

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Embodiments for Carrying Out the Invention

[0010] Hereinafter, embodiments will be described with reference to the drawings. (First Embodiment) First, the first embodiment will be described.

[0011] FIG. 1 is a diagram showing a configuration example of a memory system 1 according to the first embodiment. FIG. 1 also shows a configuration example of an information processing system including the memory system 1 and a host 2 connected to the memory system 1. The host 2 is an information processing device such as a server or a personal computer.

[0012] As shown in FIG. 1, the memory system 1 includes a controller 11 and a non-volatile memory (memory chip) 12.

[0013] The controller 11 is a device that controls writing of data to the non-volatile memory 12 and reading of data from the non-volatile memory 12 in response to commands from the host 2. The controller 11 is configured, for example, as a SoC (system-on-a-chip).

[0014] The non-volatile memory 12 is, for example, an SCM. The SCM includes a phase change memory (PCM), a magnetoresistive RAM (MRAM), a resistive RAM (ReRAM), a ferroelectric RAM (FeRAM), and the like. That is, here, an example in which the memory system 1 is realized as an SCM module is shown.

[0015] The controller 11 has an error correction circuit 100. In the non-volatile memory 12, data errors occur with a certain probability. The error correction circuit 100 is a device that detects and corrects errors in the read data, for example, when reading data from the non-volatile memory 12. The memory system 1 according to the first embodiment realizes error correction of data that can handle chip failures while maintaining low latency of the memory system 1 for data of data frames distributed and arranged in a plurality of non-volatile memories 12. This will be described in detail below.

[0016] Here, first, with reference to FIG. 2, a method of arranging data frames in the non-volatile memory (SCM) 12 in the memory system 1 will be described.

[0017] As described above, in the non-volatile memory 12, data errors occur with a certain probability. The controller 11 writes data to the non-volatile memory 12 in units of data frames (ECC frames) in which error correction bits (ECC parity) are added to the data body (data). When reading data from the non-volatile memory 12, the controller 11 detects an error using the error correction bit, and if an error is detected, executes correction of the detected error.

[0018] Generally, write requests from host 2 are made in units such as 64B (bytes) or 128B. According to this unit, the controller 11 calculates error correction bits for detecting and correcting errors for every 64B or 128B of data, for example.

[0019] On the other hand, the minimum access unit of the non-volatile memory (memory chip) 12 may be smaller than the access request size from host 2, such as 8B or 16B. In this case, as shown in Fig. 2(A), the controller 11 can write data of the access request size from host 2 (more specifically, a data frame including the data body and error correction bits) into one non-volatile memory 12 by performing multiple memory accesses. However, this method is not suitable for a system that requires low latency performance.

[0020] Therefore, as shown in Fig. 2(B), it is generally performed to shorten the latency by the controller 11 accessing a plurality of non-volatile memories 12 in parallel. That is, a method of dispersing and arranging data in a plurality of non-volatile memories 12 has become mainstream. Also in the memory system 1 of the first embodiment, it is assumed that the controller 11 writes the data received from host 2 (more specifically, a data frame including the data body and error correction bits) in a dispersed manner to a plurality of non-volatile memories 12. The data body may be the data received from host 2, or may be data obtained by performing a predetermined process such as compression processing on the data received from host 2. Also, data necessary for management in the memory system 1 may be added. That is, the data body may be data obtained from the data received from host 2, or data including data based on the data received from host 2, or data for which writing to a plurality of non-volatile memories 12 has become necessary by receiving a write request from host 2.

[0021] In FIG. 2(B), an example is shown in which the data body (data) and the error correction bit (ECC parity) are separately arranged in different non-volatile memories 12. However, they may be mixed on the same non-volatile memory 12. FIG. 3 shows an example in which the data body and the error correction bit are mixed and arranged on the same non-volatile memory 12.

[0022] FIG. 3(A) shows an example in which, in addition to the data body, the error correction bits are also distributed and arranged in a plurality of non-volatile memories 12. On the other hand, FIG. 3(B) shows an example in which the error correction bits are arranged in the surplus area after arranging the end portion of the data body, so that the data body and the error correction bits are mixed in the same non-volatile memory 12. Since the controller 11 manages the storage positions of the data body and the error correction bits, the arrangement of the data body and the error correction bits in the plurality of non-volatile memories 12 is not limited to the examples shown in FIG. 2(B) and FIGS. 3((A), (B)), and can be performed in various ways.

[0023] Subsequently, FIGS. 4 and 5 show comparative examples to explain the problems when a data frame is distributed and arranged in a plurality of non-volatile memories.

[0024] FIG. 4 shows a state in which a data frame including a data body (User Data) and an error correction bit (ECC Parity) is distributed and arranged in 10 non-volatile memories respectively connected to 10 channels #1 to #10. Channels #1 to #10 include communication lines (memory buses) for the controller to communicate with the non-volatile memories. One or more non-volatile memories are connected to each of channels #1 to #10.

[0025] When the controller receives a read command from a host, for example, it executes reading of data from the non-volatile memory based on the logical address specified by the read command. Here, an example is given of the case of reading a data frame including data corresponding to address X (addr X) from 10 non-volatile memories. Also, let the error correction ability of the ECC decoder that corrects error bits in the data frame using the error correction bits included in the data frame be T. The error correction ability T is generally about several bits to several tens of bits.

[0026] Therefore, even if there are bit errors in the data frame (input data frame) read from 10 non-volatile memories, if the error bits in the entire data frame are equal to or less than the error correction ability T of the ECC decoder (when the number of error bits ≤ T), a data frame (output data frame) without bit errors in which the error bits are corrected can be obtained.

[0027] On the other hand, in FIG. 5, similar to FIG. 4, it is an example of the case of reading a data frame including data corresponding to address X (addr X) from 10 non-volatile memories, but in the case where one of the 10 non-volatile memories has a chip failure. The failed chip is a non-volatile memory connected to channel #6.

[0028] When a chip failure occurs, the part corresponding to the failed chip in the data frame is completely lost, so all of that part becomes error bits, and the error bits in the entire data frame greatly exceed the error correction ability T of the ECC decoder. Therefore, correction using error correction bits is no longer possible.

[0029] If not only the data frame containing the data corresponding to address X but also all data frames are distributed and arranged in ten non-volatile memories including the faulty chip, all data in the SCM module will be lost. The risk of data loss increases as the number of non-volatile memories in which the data frames are distributed increases.

[0030] Based on this comparative example, next, with reference to FIG. 6, a method for distributing and arranging data in the memory system 1 of the first embodiment will be described.

[0031] In the memory system 1 of the first embodiment, in addition to distributing and arranging a data frame including error correction bits in a plurality of non-volatile memories 12, the bit XOR (XOR parity) of the entire data frame is further arranged in a non-volatile memory 12 different from the plurality of non-volatile memories 12 in which the data frame is distributed. Here, in addition to the ten channels #1 to #10 to which the ten non-volatile memories 12 for distributing and arranging the data frame are connected, a new channel #11 to which the non-volatile memory 12 for arranging the XOR parity is connected is prepared. Hereinafter, the channel to which the non-volatile memory 12 for arranging the XOR parity is connected may be referred to as an XOR channel. Also, writing the XOR parity to the non-volatile memory 12 connected to the XOR channel may be referred to as writing the XOR parity to the XOR channel.

[0032] When the controller 11 writes a data frame to the non-volatile memory 12, it writes the XOR parity of the data frame to the XOR channel. On the other hand, when the data frame is read from the non-volatile memory 12, if a chip failure occurs in any of the non-volatile memories 12, the controller 11 (error correction circuit 100) restores the data on the failed chip from the data on the non-volatile memory 12 where the data frames of the non-failed chips are distributed and the data on the non-volatile memory 12 where the XOR parity is arranged. Here, the data (data frame) on the non-volatile memory 12 connected to channels #1 to #5, #7 to #10 and the data (XOR parity) on the non-volatile memory 12 connected to channel #11 are used to restore the data frame on the non-volatile memory 12 which is the failed chip and connected to channel #6. That is, in the memory system 1 of the first embodiment, the error correction circuit 100 of the controller 11 has an XOR restoration circuit 102 in addition to the ECC decoder 101. When the non-volatile memory 12 connected to the XOR channel (here, channel #11) has a chip failure, since the disappearance of the data frame does not occur, subsequent processing including detection and correction of error bits can be continued.

[0033] There may be bit errors in the data frame in which the data on the failed chip is restored using the XOR parity. If the number of error bits in the entire data frame is equal to or less than the error correction ability T of the ECC decoder 101 (when the number of error bits ≤ T), the error correction circuit 100 can obtain a data frame (output data frame) without bit errors in which the error bits are corrected.

[0034] Note that the restoration of the missing part in the data frame using the XOR parity can be applied not only when the entire non-volatile memory 12 fails but also when only a limited area in the non-volatile memory 12 fails. That is, the chip failure here includes not only the failure of the entire memory chip but also the failure of only a limited area in the memory chip.

[0035] Incidentally, Fig. 6 shows an example in which the non-volatile memory 12 connected to channel #6 caused a chip failure. However, when the failed chip cannot be identified, as shown in Fig. 7, the error correction circuit 100 needs to repeat the loop of (1) XOR restoration by the XOR restoration circuit 102 and (2) ECC decoding attempt by the ECC decoder 101, assuming the failed chip, until the ECC decoding is successful. That is, while changing the non-volatile memory 12 to be XOR-restored, which is indicated by the code a1, it is necessary to repeat the above loop of (1) → (2) until the ECC decoding is successful. When the data frame is distributed and arranged in 10 non-volatile memories 12, in the worst case, it may be repeated 10 times.

[0036] Fig. 8 is a flowchart showing the error correction procedure when assuming a failed chip one by one from among a plurality of non-volatile memories in which the data frame is distributed and arranged. Here, the channels to which the data frame is written are set to 1 to 10, and the channel to which the XOR parity is written is set to 11.

[0037] The error correction circuit 100 reads data from the SCM12 of all channels (S101). The error correction circuit 100 generates a data frame "Frame#0" from the data of channels 1 to 10 excluding the XOR channel 11 (S102). The error correction circuit 100 decodes the generated "Frame#0" by ECC (S103). When the ECC decoding of this "Frame#0" is successful (S104: Yes), the error correction circuit 100 ends the error correction process, assuming that the error correction was successful.

[0038] If the ECC decoding of “Frame#0” fails (S104: No), the error correction circuit 100 first sets the data of channel 1 as the data to be restored (S105), and generates a data frame “Frame#N” by restoring the data of the restoration target channel N (initially 1) with the bit XOR of other channels (S106). The error correction circuit 100 performs ECC decoding on the generated “Frame#N” (S107). When the ECC decoding of this “Frame#N” is successful (S108: Yes), the error correction circuit 100 ends the error correction process assuming that the error correction was successful.

[0039] If the ECC decoding of “Frame#N” fails (S108: No), the error correction circuit 100 determines whether the data to be restored was the data of channel 10 (S109). If it was not the data of channel 10 (S109: No), the error correction circuit 100 advances the channel to be restored by one (S110) and repeats the processes of S106 to S108. On the other hand, if it was the data of channel 10 (S109: Yes), the error correction circuit 100 ends the error correction process assuming that the error correction failed.

[0040] In this way, when assuming a failed chip one by one from among a plurality of non-volatile memories in which data frames are dispersedly arranged, the loop of (1) XOR restoration by the XOR restoration circuit 102 and (2) ECC decoding trial by the ECC decoder 101, i.e., (1) → (2), is repeated until ECC decoding is successful. Due to the occurrence of this loop, there is a possibility that the latency during data reading increases.

[0041] Therefore, in order to eliminate the above loop, the memory system 1 of the first embodiment causes the error correction circuit 100 to try all ECC decodings of data frames that may be correctable in parallel without identifying the failed chip. FIG. 9 shows a configuration example of the error correction circuit 100 of the memory system 1 of the first embodiment.

[0042] As shown in FIG. 9, the error correction circuit 100 of the memory system 1 according to the first embodiment has, for example, ten XOR restoration circuits 102 for XOR-restoring the data of each channel when dispersing and writing data frames to channels #1 to #10. The error correction circuit 100 also has a total of eleven ECC decoders 101, including one ECC decoder 101 for ECC-decoding the data frame “Frame#0” generated from the data of channels #1 to #10 and ten ECC decoders 101 for ECC-decoding the data frame “Frame#N” in which the data of any one of the channels is restored by the XOR restoration circuit 102.

[0043] The error correction circuit 100 having the ten XOR restoration circuits 102 and the eleven ECC decoders 101 performs error correction on eleven data frames “Frame#0 to #10” in parallel. If there is a data frame in which error correction has succeeded even in one of them, the error correction circuit 100 outputs that data frame.

[0044] FIG. 10 is a flowchart showing the error correction procedure by the error correction circuit 100 of the memory system 1 according to the first embodiment.

[0045] The error correction circuit 100 reads data from the SCMs 12 of all channels (S201). The error correction circuit 100 generates the data frame “Frame#0” from the data of channels 1 to 10 excluding the XOR channel 11 (S202). Further, the error correction circuit 100 generates all the data frames “Frame#N” in which the data of channel N is restored by the bit XOR of the other channels among the data of channels 1 to 10 (S203).

[0046] The error correction circuit 100 performs ECC decoding on the data frames “Frame#0” to “Frame#10” in parallel (S204). The error correction circuit 100 determines whether there is a data frame for which error correction has succeeded (S205). If there is a data frame for which error correction has succeeded (S205: Yes), the error correction circuit 100 outputs any one of the data frames for which the error correction has succeeded (S206), and ends the error correction process assuming that the error correction has succeeded. On the other hand, if there is no data frame for which error correction has succeeded (S205: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed.

[0047] As a result, in the memory system 1 of the first embodiment, the above loop assuming a failed chip is eliminated, so that error correction of data that can cope with chip failures can be realized while maintaining low latency.

[0048] Incidentally, as shown in FIG. 9, for example, an error correction circuit 100 having ten XOR restoration circuits 102 and eleven ECC decoders 101 has an increased circuit scale. When the circuit scale of the error correction circuit 100 increases, for example, the chip area of the controller 11 realized as an SoC increases. The increase in the chip area of the controller 11 causes an increase in the cost of the memory system 1.

[0049] Therefore, the error correction circuit 100 of the memory system 1 of the first embodiment is further configured to suppress an increase in the circuit scale. Next, this point will be described in detail.

[0050] There are various encoding methods used for ECC decoding. In the memory system 1 of the first embodiment, the ECC decoder 101 of the error correction circuit 100 applies a BCH (Bose - Chaudhuri - Hocquenghem) code. With reference to FIG. 11, an outline of the decoding of the BCH code will be described.

[0051] The decoding of BCH codes is performed in four steps: (1) syndrome calculation, (2) error location polynomial calculation, (3) error bit position calculation, and (4) error correction, as shown in Figure 11.

[0052] (1) Syndrome calculation calculates the syndrome from the data frame and outputs a plurality of syndrome values. If all the syndrome values are 0, it is determined that there are no error bits. If there is a non-zero syndrome value among the syndrome values, it is determined that there are error bits. Also, the number of error bits can be determined from the number of non-zero syndrome values. Syndrome values can be easily obtained by bit shift and XOR multiplication, and the circuit scale and latency tend to be small.

[0053] (2) Error location polynomial calculation obtains the error location polynomial from the syndrome values obtained by syndrome calculation. As a specific algorithm, the Berlekamp-Massey method is often used. This process needs to perform repetitive processing for the number of error bits. Since the amount of calculation per time is small, the circuit scale is small, but due to the occurrence of repetitive processing, the latency tends to be large.

[0054] (3) Error bit position calculation obtains the actual error bit position by calculating the roots of the error location polynomial. As a specific algorithm, the Chien search method is often used. This process needs to perform an exhaustive calculation for all bits, so the amount of calculation becomes extremely large. In order to perform it with low latency, it is necessary to increase the number of bits processed at a time (increase the parallelism), and the circuit scale tends to increase.

[0055] (4) Error correction corrects the error at the obtained error bit position. Generally, it is only necessary to invert the bit at the error position, so both the circuit scale and latency are negligibly small.

[0056] Therefore, when organizing the tendencies of the circuit scale and latency of each step of BCH code decoding, it is as follows.

[0057] (1) Syndrome calculation … Circuit scale: small, Latency: small (2) Error location polynomial calculation … Circuit scale: small, Latency: large (3) Error bit location calculation … Circuit scale: large, Latency: small (4) Error correction … Circuit scale: small, Latency: small These are just tendencies and vary depending on the design parameters. However, if the configuration is such that decoding is performed with low latency, the circuit scale of (3) error bit location calculation may reach nearly 90% of the entire circuit scale of BCH code decoding.

[0058] Considering the tendencies of the circuit scale and latency of each step of this BCH code decoding, the error correction circuit 100 of the memory system 1 of the first embodiment is configured to suppress an increase in the circuit scale. Specifically, with respect to one configuration example of the error correction circuit 100 in FIG. 9, improvements are made to the portion of the ECC decoder 101 indicated by the code b1 in FIG. 12.

[0059] FIG. 13 is a diagram showing one configuration example of the error correction circuit 100 in which improvements are made to the portion of the ECC decoder 101.

[0060] As shown in FIG. 13, the ECC decoder 101 has a configuration that executes (1) syndrome calculation and (2) error location polynomial calculation, which tend to have a small circuit scale among the four steps of BCH code decoding, in parallel. For (3) error bit location calculation, which tends to have a large circuit scale, and (4) error correction, which follows (3), they are executed sequentially. By deliberately improving (3) error bit location calculation, which tends to have a large circuit scale, to be executed sequentially, an increase in the circuit scale of the error correction circuit 100 is suppressed. Since (3) error bit location calculation and (4) error correction tend to have a small latency, the impact of executing them sequentially is limited. On the other hand, the parallel execution of (2) error location polynomial calculation, which tends to have a large latency, is maintained. Hereinafter, the operation of the ECC decoder 101 with such improvements will be described.

[0061] The ECC decoder 101 inputs all data frames that may have correction success in parallel (c1). The ECC decoder 101 calculates syndrome values in parallel for all data frames (c2). Also, the ECC decoder 101 calculates the number of error bits for all syndrome values (c3). If a valid number of error bits cannot be obtained, correction fails here. If a data frame with a syndrome value of 0 is found, it is output without performing error correction and the process ends (c4).

[0062] If a data frame with a syndrome value of 0 is not found while a valid number of error bits is obtained, the ECC decoder 101 calculates the error location polynomial in parallel for all syndrome values (c5). Then, the ECC decoder 101 sorts the error location polynomials in ascending order of the number of error bits (c6). The ECC decoder 101 calculates the error bit position using the error location polynomial with the smallest number of error bits (c7). If there is no contradiction in the calculated error bit position, the ECC decoder 101 performs error correction based on the obtained error bit position (c8), outputs the data, and ends.

[0063] On the other hand, if the error bit position is uncomputable or there is a contradiction in the calculated error bit position, the ECC decoder 101 re - executes the calculation of the error bit position using the error location polynomial with the next smallest number of error bits (c7). If there are no uncalculated error location polynomials, correction fails.

[0064] FIG. 14 is a flowchart showing the error correction procedure by the error correction circuit 100 with improvements made to the part of the ECC decoder 101.

[0065] The error correction circuit 100 reads data from the SCM12 of all channels (S301). The error correction circuit 100 generates a data frame "Frame#0" from the data of channels 1 to 10 excluding the XOR channel 11 (S302). Also, the error correction circuit 100 generates all the data frames "Frame#N" in which the data of channel N among the data of channels 1 to 10 is restored by bit XOR of the other channels (S303).

[0066] The error correction circuit 100 calculates the syndromes "Synd#0" to "Synd#10" of the data frames "Frame#0" to "Frame#10" in parallel (S304). The error correction circuit 100 determines whether any of the syndromes "Synd#0" to "Synd#10" is 0 (S305). If any of them is 0 (S305: Yes), the error correction circuit 100 outputs the data frame "Frame#N" whose syndrome is 0 (S306).

[0067] If none of the syndromes "Synd#0" to "Synd#10" is 0 (S305: No), the error correction circuit 100 calculates the number of error bits "t#0" to "t#10" in parallel from the syndromes "Synd#0" to "Synd#10" (S307). The error correction circuit 100 determines whether any of the calculations of the number of error bits "t#0" to "t#10" has been successful (S308). If all the calculations have failed (S308: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed.

[0068] If the calculation of any of the error bit numbers "t#0" to "t#10" is successful (S308: Yes), the error correction circuit 100 calculates the error location polynomials "Poly#0" to "Poly#10" in parallel from the syndromes "Synd#0" to "Synd#10" (S309). The error correction circuit 100 determines whether the calculation of any of the error location polynomials "Poly#0" to "Poly#10" is successful (S310). If any of the calculations fails (S310: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed.

[0069] If the calculation of any of the error location polynomials "Poly#0" to "Poly#10" is successful (S310: Yes), the error correction circuit 100 sorts the successfully calculated error location polynomial "Poly#N" in ascending order of the error bit number "t#N" (S311). The error correction circuit 100 calculates the error bit position "ErrVec#N" from the error location polynomial "Poly#N" with the smallest error bit number "t#N" that has not been calculated (S312).

[0070] The error correction circuit 100 determines whether the error bit position "ErrVec#N" is within the range of the data frame (S313). If it is not within the range of the data frame (S313: No), the error correction circuit 100 determines whether there is an error location polynomial "Poly#N" for which the error bit position has not been calculated (S314). If not (S314: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed. If it exists (S314: Yes), the error bit position "ErrVec#N" is recalculated from that error location polynomial "Poly#N" (S312).

[0071] On the other hand, when the error bit position “ErrVec#N” is within the range of the data frame (S313: Yes), the error correction circuit 100 corrects the data at the bit position corresponding to the error bit position “ErrVec#N” for the data frame “Frame#N” (S315). Then, the error correction circuit 100 outputs the corrected data frame “Frame#N” (S316), and assuming that the error correction was successful, ends the error correction process.

[0072] FIG. 15 is a first diagram showing an estimate of the worst-case latency required for error correction. (A) shows a case where, assuming a failed chip, the XOR calculation [XOR] and the (1) syndrome calculation [SYND], (2) error location polynomial calculation [ELP], and (3) error bit position calculation [CS] of the BCH code decoding are sequentially and cyclically executed. On the other hand, (B) shows a case where error correction is attempted in parallel for all data frames for which error correction may succeed without identifying the failed chip, and all of the XOR calculation [XOR] and the (1) syndrome calculation [SYND], (2) error location polynomial calculation [ELP], and (3) error bit position calculation [CS] of the BCH code decoding are executed in parallel.

[0073] For example, assuming that the latency for each calculation is XOR calculation = 1 cycle, syndrome calculation = 5 cycles, error location polynomial calculation = 25 cycles, and error bit position calculation = 5 cycles, the worst-case latency of (A) is (1 + 5 + 25 + 5) × 11 = 396 cycles. On the other hand, the latency in the case of (B) is always 1 + 5 + 25 + 5 = 36 cycles.

[0074] FIG. 16 is a second diagram showing an estimate of the worst-case latency required for error correction. (A) reproduces Fig. 15(A). On the other hand, (C) shows the case where error correction is tried in parallel for all data frames where error correction may succeed without identifying faulty chips, and the XOR calculation [XOR], the (1) syndrome calculation [SYND] and the (2) error location polynomial calculation [ELP] of the BCH code decoding are executed in parallel.

[0075] Similar to the above, for example, assuming the latency for each calculation is XOR calculation = 1 cycle, syndrome calculation = 5 cycles, error location polynomial calculation = 25 cycles, and error bit position calculation = 5 cycles, the worst-case latency of (C) is 1 + 5 + 25 + 5×11 = 86 cycles.

[0076] Fig. 17 is a diagram showing the latency and circuit scale in the case of Fig. 15(A), the latency and circuit scale diagram in the case of Fig. 15(B), and the latency and circuit scale in the case of Fig. 16(C). Taking the latency and circuit scale in the case of Fig. 15(A) as 1, an example showing the latency and circuit scale in the case of Fig. 15(B) and the latency and circuit scale in the case of Fig. 16(C) as a ratio to the latency and circuit scale in the case of Fig. 15(A) is shown.

[0077] As shown in Fig. 17, the latency of (B) is 0.09 with respect to 1 of (A), which is significantly improved in terms of performance. On the contrary, the circuit scale significantly increases to 11.0 with respect to 1 of (A). In contrast, for (C), the latency is 0.22 and the circuit scale is 1.99. Thus, (C) has a good balance between latency and circuit scale and is practical.

[0078] Thus, in the memory system 1 of the first embodiment, the error correction circuit 100 parallelizes only syndrome calculation and error location polynomial calculation in consideration of the circuit scale and latency trends at each step of BCH code decoding. As a result, the memory system 1 of the first embodiment can perform error correction of data that can handle chip failures while maintaining low latency without incurring cost increases due to an increase in circuit scale.

[0079] (Second Embodiment) Next, the second embodiment will be described. An example in which the memory system of the second embodiment is also realized as an SCM module is shown. The same reference numerals are used for the same configurations as those of the first embodiment, and the descriptions thereof are omitted.

[0080] FIG. 18 is a diagram for explaining the difference between the memory system 1 of the second embodiment and the memory system 1 of the first embodiment.

[0081] The case of restoring a data frame using the XOR parity of the XOR channel (for example, channel #11) and performing ECC decoding is a limited case. Therefore, performing ECC decoding for all data frames that may have errors every time is wasteful, for example, from the viewpoints of power consumption and heat generation.

[0082] Therefore, in the memory system 1 of the second embodiment, the error correction circuit 100 first performs ECC decoding (d1) on normal data frames that do not perform XOR restoration, and only when the ECC decoding of this data frame fails, it performs ECC decoding in parallel for all of the XOR-restored data frames (d2).

[0083] In the memory system 1 of the second embodiment, when a chip failure occurs, ECC decoding is performed in two stages: (1) normal data frames → (2) XOR-restored data frames. Therefore, compared with the memory system 1 of the first embodiment, the latency at the time of the chip failure decreases, but the power consumption and heat generation during normal times when no chip failure has occurred can be significantly reduced.

[0084] FIG. 19 is a flowchart showing the error correction procedure by the error correction circuit 100 of the memory system 1 according to the second embodiment.

[0085] The error correction circuit 100 reads data from the SCM12 of all channels (S401). The error correction circuit 100 generates a data frame "Frame#0" from the data of channels 1 to 10 excluding the XOR channel 11 (S402).

[0086] The error correction circuit 100 performs ECC decoding on the data frame "Frame#0" (S403). Then, the error correction circuit 100 determines whether the error correction of this data frame "Frame#0" is successful (S404). If the error correction is successful (S404: Yes), the error correction circuit 100 outputs the data frame "Frame#0" for which the error correction is successful (S405), and ends the error correction process assuming that the error correction is successful.

[0087] On the other hand, if the error correction of the data frame "Frame#0" fails (S404: No), the error correction circuit 100 generates all data frames "Frame#N" obtained by restoring the data of channel N with the bit XOR of the other channels among the data of channels 1 to 10 (S406). Then, the error correction circuit 100 performs ECC decoding on the data frames "Frame#0" to "Frame#10" in parallel (S407). The error correction circuit 100 determines whether there is a data frame for which the error correction is successful (S408). If there is a data frame for which the error correction is successful (S408: Yes), the error correction circuit 100 outputs any one of the data frames for which the error correction is successful (S409), and ends the error correction process assuming that the error correction is successful. On the other hand, if there is no data frame for which the error correction is successful (S408: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed.

[0088] In this way, the memory system 1 of the second embodiment can realize error correction of data that can handle chip failures while maintaining low latency without incurring cost increases due to an increase in circuit scale. In addition, power consumption, heat generation, etc. can be significantly reduced.

[0089] (Third Embodiment) Next, the third embodiment will be described. An example in which the memory system of the third embodiment is also realized as an SCM module is shown. The same reference numerals are used for the same configurations as in the first and second embodiments, and the descriptions thereof are omitted.

[0090] Generally, the power consumption of data reading from the non-volatile memory 12 is greater than the power consumption of calculations in the controller 11 in the memory system 1. Also, as mentioned in the second embodiment, the case of restoring a data frame using the XOR parity of the XOR channel (for example, channel #11) and performing ECC decoding is a limited case. In other words, the XOR parity read from the non-volatile memory 12 connected to the XOR channel is likely not to be used. That is, the reading of the XOR parity from the non-volatile memory 12 is likely to be wasted in many cases.

[0091] Therefore, in the memory system 1 of the third embodiment, the error correction circuit 100 first performs ECC decoding on the data frame read from the non-volatile memory 12 connected to a channel other than the XOR channel without reading the XOR parity. Then, the error correction circuit 100 reads the XOR parity and executes XOR restoration and ECC decoding in parallel only when the ECC decoding of this data frame fails.

[0092] FIG. 20 is a flowchart showing the error correction procedure by the error correction circuit 100 of the memory system 1 of the third embodiment.

[0093] The error correction circuit 100 reads data from the SCM12 of all channels except the XOR channel (S501), and generates a frame "Frame#0" (S502). Then, the error correction circuit 100 performs ECC decoding on the data frame "Frame#0" (S503).

[0094] The error correction circuit 100 determines whether the error correction of the data frame "Frame#0" is successful (S504). If the error correction is successful (S504: Yes), the error correction circuit 100 outputs the data frame "Frame#0" for which the error correction is successful (S505), and ends the error correction process assuming that the error correction is successful.

[0095] If the error correction of the data frame "Frame#0" fails (S504: No), the error correction circuit 100 reads data (XOR parity) from the SCM12 of the XOR channel (S506). The error correction circuit 100 generates all data frames "Frame#N" in which the data of channel N among the data of channels 1 to 10 is restored by bit XOR of the other channels (S507). Then, the error correction circuit 100 performs ECC decoding on the data frames "Frame#0" to "Frame#10" in parallel (S508).

[0096] The error correction circuit 100 determines whether there is a data frame for which the error correction is successful (S509). If there is a data frame for which the error correction is successful (S509: Yes), the error correction circuit 100 outputs any one of the data frames for which the error correction is successful (S510), and ends the error correction process assuming that the error correction is successful. On the other hand, if there is no data frame for which the error correction is successful (S509: No), the error correction circuit 100 ends the error correction process assuming that the error correction has failed.

[0097] In this way, the memory system 1 of the third embodiment can further reduce the power consumption during normal operation when no chip failure has occurred.

[0098] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are included in the invention described in the claims and the equivalent scope thereof.

Description of Reference Numerals

[0099] 1... Memory system, 2... Host, 11... Controller, 12... Non-volatile memory (SCM), 100... Correction circuit, 101... ECC decoder, 102... XOR restoration circuit.

Claims

1. A plurality of non-volatile memories, a controller that can communicate with a host and controls the plurality of non-volatile memories, comprising: the controller writes a data frame including data received from the host or data obtained by performing compression processing on the data received from the host or data with management-necessary data added to the data received from the host, and a first parity for detecting and correcting errors in the data, to N (N is a natural number of 2 or more) first non-volatile memories among the plurality of non-volatile memories in a distributed manner, writes a second parity for restoring data on one of the first non-volatile memories among the data frames written in a distributed manner to the N first non-volatile memories, to a second non-volatile memory other than the N first non-volatile memories among the plurality of non-volatile memories, from the data on the N - 1 first non-volatile memories other than the one first non-volatile memory, the controller generates a first data frame from N pieces of data read from each of the N first non-volatile memories, generates N different second data frames that can be generated from N - 1 pieces of data read from each of N - 1 first non-volatile memories among the N first non-volatile memories, the N - 1 pieces of data, and data on one of the first non-volatile memories other than the N - 1 first non-volatile memories restored by the N - 1 pieces of data and the second parity read from the second non-volatile memory, and executes at least part of a decoding process for detecting and correcting data errors by the first parity, in parallel, for the first data frame or the second data frame, a memory system.

2. The memory system according to claim 1, wherein the controller executes the decoding process for the second data frame when the decoding process for the first data frame fails.

3. The memory system according to claim 2, wherein the controller generates the second data frame when the decoding process for the first data frame fails.

4. The memory system according to claim 3, wherein the controller reads the second parity from the second non-volatile memory when the decoding process for the first data frame fails.

5. The memory system according to any one of claims 1 to 4, wherein the first parity is a BCH (Bose-Chaudhuri-Hocquenghem) code.

6. The memory system according to claim 5, wherein the controller executes in parallel the calculation of a syndrome value and the calculation of an error location polynomial for all data frames to be decoded, and sequentially executes the calculation of an error bit position using the error location polynomial.

7. The memory system according to claim 6, wherein the controller executes the calculation of an error bit position using the error location polynomial in ascending order of the number of error bits of error bit number information obtained from the syndrome value.

8. The memory system according to claim 7, wherein when the calculation of an error bit position using the error location polynomial in a certain data frame is successful, the controller omits the calculation of an error bit position using the error location polynomial for a data frame having a larger number of error bits than the certain data frame.

9. The memory system according to any one of claims 1 to 8, wherein the non-volatile memory is a phase change memory (PCM), a magnetoresistive random access memory (MRAM), a resistive random access memory (ReRAM), or a ferroelectric random access memory (FeRAM).

10. A plurality of non-volatile memories, a controller communicable with a host and controlling the plurality of non-volatile memories, comprising: The controller: performs compression processing on first data received from the host, or writes fourth data including second data with management-required data added thereto and first parity for detecting and correcting an error of the second data, in a distributed manner to first to nth (n is a natural number of 2 or more) non-volatile memories among the plurality of non-volatile memories; writes second parity for restoring data stored in one first non-volatile memory among the first to nth non-volatile memories, from data stored in (n - 1) non-volatile memories other than the one first non-volatile memory, to an (n + 1)th non-volatile memory among the plurality of non-volatile memories; The controller: generates first to nth data frames Among the first to nth data frames, the kth (where k is an integer from 1 to n) data frame includes (n - 1) data read from each of the (n - 1) non-volatile memories excluding the kth non-volatile memory among the first to nth non-volatile memories, and data restored based on the (n - 1) data and the read second parity. The controller generates an (n + 1)th data frame based on n data read from each of the first to nth non-volatile memories. The controller executes decoding processing for each of the first to (n + 1)th data frames. The first parity is a BCH (Bose-Chaudhuri-Hocquenghem) code. The controller executes in parallel the calculation of the syndrome value and the calculation of the error-location polynomial in the decoding processing for each of the first to (n + 1)th data frames, and sequentially executes the calculation of the error-bit position using the error-location polynomial for each of the first to (n + 1)th data frames. Memory system.

11. The memory system according to claim 10, wherein the controller executes the decoding processing for the first to nth data frames when the decoding processing for the (n + 1)th data frame fails.

12. The memory system according to claim 10, wherein the controller does not generate the first to nth data frames when the decoding processing for the (n + 1)th data frame is successful.

Citation Information

Patent Citations

  • Nonvolatile memory and memory device

    JP2012027991A

  • Improved error correction in solid-state disks

    JP2012514274A

  • Controller and storage apparatus

    JP2021086406A

  • Error correction decoding with redundancy data

    US10901840B2

  • Storage system and method for controlling storage

    WO2014170984A1