Memory Controller Error Checking Process Using Internal Memory Device Codes
Through the memory controller, the ECC code is recalculated and the internal SEC code is used to solve the problem of multi-bit error recovery in the prior art, the accurate detection and recovery of errors is achieved, and the error verification capability of the memory system is improved.
Patent Information
- Application Number
- CN201811129501.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-29
- Filing Date
- 2018-09-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2038-09-27
AI Technical Summary
In the prior art, when handling multi-bit errors, the memory device itself cannot correct it, and the limited information of the onboard ECC memory device cannot accurately identify the memory device where the error occurs, resulting in difficulty in recovering errors.
Read data and ECC code from the DIMM through the memory controller, recalculate the ECC code, and use the internal SEC code of the memory device to detect and recover errors to identify the error.
It realizes effective detection and recovery of multiple-bit errors, can accurately identify the memory device where the error occurs, and improves the error checksum recovery capability of the memory system.
Smart Images

Figure CN109582494B_ABST
Abstract
Description
Technical Field
[0001] The field of the present invention generally relates to computer science and, more particularly, to a memory controller error checking process using internal memory device codes. Background Art
[0002] In many computer systems, a related problem is the system memory (also known as "main memory"). Here, as understood in the art, a computing system operates by executing program code stored in the system memory and reading data / writing data on which the program code operates from / to the system memory. Thus, during the operation of a computing system, many program code reads and data reads as well as many data writes heavily utilize the system memory. Therefore, finding ways to improve system memory access performance is a motivation for computing system engineers. Brief Description of the Drawings
[0003] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the accompanying drawings, in which:
[0004] Figure 1 Shows x4 DIMM and x8 DIMM;
[0005] Figure 2 Shows another set of x4 and x8 DIMM;
[0006] Figure 3a 、 3b And 3c relate to an error checking process;
[0007] Figure 4 Shows SEC code;
[0008] Figure 5 Shows an error checking process;
[0009] Figure 6 Shows a computing system. Detailed Description of the Invention
[0010] As is known in the art, dual in-line memory modules (DIMMs) are used to implement system memory (also known as main memory) in various types of computing systems (such as servers, personal computer towers, etc.). In a typical scenario, one or more DIMMs are correspondingly inserted into one or more memory channel connectors deployed on a motherboard.
[0011] Figure 1 Shows two different types of prior art DIMMs. The first DIMM 110 is called an "X4 DIMM". The second DIMM 120 is called an "X8 DIMM". As in Figure 1As observed, each type of DIMM 110, 120 includes eight memory chips 1 through 8 on each DIMM side. "X4" indicates that each memory device has 4 bits and "X8" indicates that each memory device has 8 bits. Accordingly, the X4 DIMM 110 has a 32-bit data bus on each side of the DIMM 110 and the X8 DIMM has a 64-bit data bus on each side of the DIMM 120 (the different DIMM sides are separated by a dashed line running through the middle of each DIMM 110, 120 depicted in Figure 1 ).
[0012] Here, the memory channel into which the DIMM cards 110, 120 are to be inserted has a 64-bit data bus. This memory channel is typically designed to have a number of individually addressable memory ranks of memory that can be inserted into the memory channel. Accordingly, the entire x4 DIMM 110 appears as a single 64-bit rank. In contrast, each side of the x8 DIMM 102 appears as a 64-bit rank.
[0013] Both DIMM 110, 120 also include on-board ECC chips. Here, in various implementations, 8 bits of the 64-bit data bus of the memory channel are also reserved for error correction coding (ECC) information. As is known in the art, the memory controller writes a 64-bit data word to a specific address in each of the memory chips 1 through 8. That is, in the case of the X4 DIMM, the memory controller writes 4 unique bits into each of the sixteen memory chips (8 chips on two DIMM sides), while in the case of the X8 DIMM, the memory controller writes 8 unique bits into each of the eight memory chips (a set of eight memory chips on a particular side of the DIMM).
[0014] In a typical system memory or main memory application, data units are accessed from the system memory as cache lines consisting of multiple 64-bit words. For example, in a computing system designed to access 512-bit cache lines, a cache line is read from or written to a particular DIMM with eight read / write bursts from / to the DIMM, in which case each burst cycle includes a 64-bit data word and 8 bits of ECC. Accordingly, each cache line corresponds to 512 bits of data (64 x 8 = 512) and 64 bytes of ECC (8 x 8 = 64). Here, during the cache line write process, the memory controller calculates a 64-bit EEC value based on the 512 bits of the write data and stores an 8-bit slice of the ECC into the DIMM during each burst cycle.
[0015] Here, the ECC information can be regarded as specific code that is a function of specific bit patterns of 512-bit cache line data. During any burst cycle of a cache line write, during the write process at the same memory address as the 64-bit data word, 8 bits of the ECC information are stored together with the 64-bit data word in the ECC memory on the DIMM. Thus, a complete write operation to the DIMM for a single burst cycle includes not only writing the 64-bit data word to the DIMM, but also writing 8 bits of the ECC information for the word to the DIMM. In various implementations, the physical address applied to the system memory reserves three low-order bits for each burst of each cache line and the high-order bits specify the system memory address of the cache line.
[0016] If the memory controller receives a subsequent read request for a previously written cache line, the memory controller will issue a burst read request for the cache line to the DIMM, which causes the DIMM card to supply not only the 64-bit data word for each burst cycle, but also an 8-bit slice of the ECC information stored together with the data word. Upon receiving the 512-bit cache line (after 8 burst cycles), the memory controller recalculates the ECC information from the 512-bit cache line and compares it with the 64-bit ECC information accumulated from the 8 read bursts read from the DIMM. If the newly calculated ECC information matches the 64-bit ECC information read from the DIMM, the cache line read from the DIMM is considered error-free and is forwarded to any one of the units of the computing system that requested the word (such as a CPU core, GPU, network interface, etc.).
[0017] If the newly calculated ECC information does not match the ECC information read from the DIMM, the memory controller will identify that there is some kind of corruption in the cache line or the ECC information read from the DIMM. However, the 512-bit cache line and the 64-bit ECC information can be processed to "recover" the lost information so that the correct 512-bit cache line and the corresponding correct ECC information can be reconstructed by the memory controller. Thus, even if there is corruption in the information read from the DIMM, the corruption can be fully recovered and the correct 512-bit cache line can be forwarded to any one of the units of the computing system that requested it.
[0018] Regarding the aforementioned burst, in the case of a double data rate (DDR) memory channel (such as a DDR memory channel having characteristics defined by JEDEC specifications (e.g., DDR4)), there are two such read / write cycles per clock cycle (one read / write cycle on the rising edge of the clock and one read / write cycle on the falling edge of the clock). As described last above, to generate a 512-bit cache line, a prior art x4 DIMM performs a 512-bit burst in 8 cycles (16 memory devices × 4 bits per memory device × 8 cycles per burst = 512 bits per burst). In contrast, to generate a 512-bit cache line, a prior art x8 DIMM performs a 512-bit burst in 8 cycles (8 memory devices × 8 bits per memory device × 8 cycles per burst = 512 bits per burst).
[0019] JEDEC is an industry standards group that promulgates engineering specifications for memory channels and devices connected to them (such as DIMMs and memory controllers). (The most recent JEDEC memory channels have generally been referred to as "DDR" memory channels because they are designed for clocking data on both the rising and falling edges of the clock signal as described above).
[0020] Figure 1 Prior art DIMM cards 110, 120 indicate that two ECC memory devices are deployed on DIMM cards 110, 120. However, future JEDEC DDR memory channels are intended to change the way DIMM cards are accessed. In particular, future JEDEC DDR memory channels will emphasize longer bursts with a smaller bit width. Figure 2 shows a logical view of newer x4 and x8 DIMMs 210, 220. As can be seen in Figure 2 both the x4 DIMM 210 and the x8 DIMM 220 can be considered to support 32-bit transfers instead of 64-bit transfers (the bit width is logically halved). However, although the bit width is logically halved, the number of cycles per burst increases from 8 cycles to 16 cycles. Accordingly, a 512-bit cache line for the new x4 DIMM 210 is implemented as 8 memory devices × 4 bits per memory device × 16 cycles per burst. For the new x8 DIMM 220, a 512-bit cache line is implemented as 4 memory devices × 8 bits per memory device × 16 cycles per burst.
[0021] Changes in the logical view of the DIMM include corresponding changes in the ECC information for x4 DIMMs from every 8 bits transferred to every 4 bits transferred. That is, for each transfer of an x4 DIMM, the ECC information (such as bit width) has been cut in half. The reduction in the ECC bit width corresponds to less protection for the ECC information stored on the on-board ECC memory devices of the x4 DIMM. Additionally, even though the new x8 DIMMs can store 8-bit ECC information per burst, in various embodiments, the x8 DIMMs will use the x4 DIMM ECC code which will correspondingly result in less error protection coverage. The error coverage in both the new and old DDR technologies for x8 DIMMs is less than that for x4 DIMMs because the data that needs to be corrected when an x8 DIMM fails is twice that when an x4 DIMM fails.
[0022] However, it is believed that the relaxation of the on-board memory ECC information allows acceptable performance because each memory chip on the newer DIMMs will include its own internal ECC function 230 (for ease of drawing Figure 2 , only one internal memory ECC function 230 is explicitly labeled).
[0023] Accordingly, the memory chips themselves will be able to recover corrupted data read from their own internal memory arrays so as to give correct data by the memory chips even if their own internal memory arrays give corrupted data. In Figure 1 the case of the prior art DIMMs 110, 120, data corruption originating from the memory chip arrays can only be corrected by the memory controller using the ECC information stored in the on-board ECC memory chips of the DIMM.
[0024] However, in various embodiments, the ECC function 230 integrated on the memory die is a single error correction (SEC) code that cannot recover from multiple bit errors. For example, in one implementation, the memory cell array of each memory chip is divided into regions of, for example, 128 bits, and the memory chip maintains a unique internal SEC code for each different 128-bit region. Here, each SEC code can recover from a single bit error in the 128-bit region, but cannot recover from an error in two or more bits in the 128-bit region (during the nominal write process to the memory chip, the memory chip calculates a new SEC value internally for the specific 128-bit region to which the write data is being written).
[0025] Newer DIMM cards 210, 220 thus present error recovery challenges in the case where a single memory device generates more than one error during a memory read. Here, the memory device will not be able to correct a multi-bit error (since it uses SEC code), and the on-board ECC memory device protection may not be able to correct the error due to its reduced information content.
[0026] Accordingly, in the case of a multi-bit error from a single memory device that the memory device itself cannot correct, the processing of the memory controller that reads the data and the on-board ECC information will be able to detect the presence of the error but not identify where the error is. That is, in the case of a multi-bit error from a particular memory device that cannot be corrected by the memory device, the limited ECC information from the on-board ECC memory device of the DIMM can identify that an error has occurred but not identify which memory device has generated the multi-bit error.
[0027] Figure 3a 、 3b And 3c regarding the read error recovery process that can be performed by the memory controller 301 using error correction information resident on the DIMM 320 that is used not only to identify which memory device on the DIMM 320 has generated a multi-bit error but also to recover lost data (replacing corrupted data with correct data). For ease of discussion, the following example will concern an x8 DIMM 220, 320 having four memory devices (not counting ECC) per burst transfer of read data.
[0028] Before the read operation, referring to Figure 3a , assume that the memory controller 301 includes ECC generation logic circuitry for generating an ECC code P that is written to the on-board ECC memory 321 of the DIMM for each memory word stored on the DIMM 320. Here, in the case of an x8 DIMM 320, each memory word is composed of four eight-bit data components D 0 、D 1 、D 2 and D 3 (one eight-bit data component per memory chip). In one embodiment, the ECC code P generated by the memory controller 301 can be expressed as:
[0029] P = S 0 (D 0 ) + S 1 (D 1 ) + S 2 (D 3 ) + S 4 (D 4 ) Equation 1
[0030] Here: 1) D x is the first, second, third, or fourth component of the cache line D being written, where D1 corresponds to information written to the first memory device in a burst requiring a full cache line write, D2 corresponds to information written to the second memory device in a burst requiring a full cache line write, D3 corresponds to information written to the third memory device in a burst requiring a full cache line write, and D4 corresponds to information written to the fourth memory device in a burst requiring a full cache line write; 2) S x (D x ) is applied to D as a function of x x 3) The operation "+" corresponds to a bitwise XOR. Here, it should be apparent that x can be any of 0, 1, 2, or 3 for an x8 DIMM with 4 memory devices per pass (the value of x represents a specific one of the four memory devices).
[0031] As in Figure 3a As observed in , the read 1 of the DIMM card 320 is performed in multiple bursts to achieve a complete transfer of the cache line for each burst cycle as usual with a data word from each memory device, which is read simultaneously with a slice of the ECC code P attached to the data word by the memory controller. In the case of an x8 DIMM 320 with four memory devices per transfer, the read cache line is D = D0, D1, D2, D3, where D0 is the information read from the first memory device on the DIMM 320 within multiple bursts, D1 is the information read from the second memory device on the DIMM 320 within multiple bursts, D2 is the information read from the third memory device on the DIMM 320 within multiple bursts, and D3 is the information read from the fourth memory device on the DIMM 320 within multiple bursts. In addition, the ECC code P for cache line D is read from the ECC memory device 321 on the DIMM 320 within multiple bursts. (128b from each x8 device in DDR5).
[0032] After reading one data word D and its code P from DIMM 320, the error checking logic circuit 302 of the memory controller recalculates the ECC code word P' from the cache line D just read from the DIMM and compares it with the ECC code word P just read from the DIMM. Here, the first logic circuit 303 of the error checking logic circuit 302 is designed to execute, for example, the formula of Equation 1 above. If P' = P, there is no data corruption and the cache line D is forwarded to the requester of the data. However, if P' ≠ P, there is data corruption that needs to be resolved by the error checking logic circuit 302. Here, again assume that the detected error is a multi-bit error from one of the memory devices on DIMM 320 (since the memory device will be able to correct any single-bit error internally).
[0033] In one embodiment, the ECC codes P / P' cannot by themselves determine which memory generated the corruption. That is, it is not possible to determine from an analysis of just P and P' which of D0, D1, D2, or D3 contains the multi-bit error.
[0034] Accordingly, referring to Figure 3b , the memory controller reads the SEC codes generated by the internal SEC function of each memory device on DIMM 320 from each memory device. Recall from above that in an embodiment, each memory device stores an SEC code for each 128-bit data region of the internal storage cell array of the memory device. In the case of an x8 DIMM, in an embodiment, each SEC code is 8 bits. When reading data from any memory device, the memory device internally recalculates the SEC code for the target read data for the 128 regions that are the constituents of the target read data. The memory device compares the recalculated SEC code for the 128-bit region with the stored SEC code for the 128-bit region. If the codes match, the memory device will infer that there is no error and will not attempt any correction (it will simply forward the read data).
[0035] Thus, to reiterate, if after the ECC check Figure 3a based on the initial read of 1 and P, the memory controller 301 determines that one of the memory devices must have generated a multi-bit error (P ≠ P), then referring to Figure 3b , the next stage of the recovery process requires reading the SEC codes SEC0, SEC1, SEC2, and SEC3, which are stored in the respective memory devices within the respective 128-bit regions corresponding to D0, D1, D2, and D3 on DIMM 320 (each 128-bit region on the memory chip for which the SEC value is calculated is calculated within 16 bursts of 8-bit write data of the memory device that wrote the cache line).
[0036] Here, in various embodiments, up to half of the multiple-bit errors that may be generated by the memory device cannot be corrected by the internal SEC code of the memory device. However, such errors can still be detected by the internal SEC code of the memory device (when the ECC is the SEC code, all multiple-bit errors cannot be corrected by the memory device ECC, but approximately half are detectable and the others are miss-corrections). That is, the internal SEC function of the memory device can detect the presence of an error using the SEC code but cannot correct the error. Here, if "SEC" is the SEC code initially stored for the 128-bit region and the code recalculated for a read operation from the 128-bit region is "SEC'", the memory device will be able to detect at least that SEC ≠ SEC'.
[0037] Accordingly, in one embodiment, using the reads 3 of the initially stored SEC0, SEC1, SEC2, and SEC3 codes, the memory controller 301 (which also has the logic circuit 304 for performing the internal SEC function of the memory device) can recalculate 4 SEC0', SEC1', SEC2', and SEC3' from the received data segments D0, D1, D2, and D3 (SEC0' can be calculated by applying the SEC function to D0, SEC1' can be calculated by applying the SEC function to D1, and so on). Here, in a scenario where only one of the memory devices has generated a multiple-bit error, only one of the SEC code pairs will not match. For example, if SEC0 ≠ SEC0', the error checking logic circuit 303 of the memory controller will identify that memory device 0 has generated a multiple-bit error, and if SEC1 ≠ SEC1', the error checking logic circuit 302 of the memory controller will identify that memory device 1 has generated a multiple-bit error, and so on.
[0038] When the faulty memory device is known, P and the custom scrambling p and its inverse application to the data applied to the failed memory device can be used to recover the corrupted data. That is, if memory device 0 generates an error, then:
[0039] P' = S 0 (D 0 + e) + S 1 (D 1 ) + S 2 (D 3 ) + S 4 (D 4 ) Equation 2
[0040] Here D 0is the correct data for memory device 0 and e is the corruption value that causes the corrupted data component received for memory device 0 when XORed bitwise with the correct data D 0 In other words, if the corrupted data received from memory device 0 is , then = D 0 XOR e. In the case where the scrambling function S 0 is distributed, Equation 2 can be expressed as:
[0041] P’ = S 0 (D 0 ) + S 0 (e) + S 1 (D 1 ) + S 2 (D 3 ) + S 4 (D 4 ) Equation 3
[0042] Taking the difference between P (the value read from the ECC memory on DIMM 320) and the value of P’ (the ECC value recomputed by circuit 303 during Figure 3a process 2) yields:
[0043] P’ – P = S 0 (e) Equation 4.
[0044] That is, taking the difference between P’ and P yields the corrupted value e scrambled according to the custom scrambling reserved for memory device 0. Therefore, applying the inverse of the custom scrambling to the difference between P’ and P yields e (i.e., ). Based on e, D can be recovered by performing XOR on 0 and e. That is, (the inverse of XOR is XOR).
[0045] The reader should understand that the use of memory device 0 as the memory device that generates the corrupted data is merely an example and that the same process can be used for any of the other memory devices if any of the other memory devices is considered to be the memory device that provides the corruption result. The only difference between the different memory devices in the recovery process is the custom inverse scrambling reserved for each memory device. The method proposed above can also correct single-bit errors caused by reasons other than internal memory device errors (e.g., bit flips on the physical memory channel).
[0046] It is also necessary to point out that, for example, as a form of error correction acceleration, if the memory device detects an internal error that it cannot correct, a certain form of metadata can be set in Figure 3a the initial read 1 to notify the memory controller 301 that the memory device has detected an uncorrectable error (for example, a specific message can be posted on the CA bus of the memory channel). In the case where the memory controller 301 recognizes that the memory device just read from it has detected an error that it cannot correct, the memory controller 301 can immediately issue a read 3 of the SEC information without having to wait for the error checking logic circuit 302 to determine that P’≠ P.
[0047] The previous example described a process in which the memory controller 301 can identify the faulty memory device (memory 0 in this example), which re-executes 4 the SEC algorithm of the internal memory on the data it receives from each memory device and compares it with the internal SEC code information generated by each memory device for its corresponding data using its internal SEC function. As elaborated above, in approximately half of the multi-bit error cases, these two codes (SEC and SEC’) will be different, which leads to directly identifying the faulty memory device.
[0048] However, the other “half” of the multi-bit error cases results in the two SEC codes SEC’ and SEC matching (both at the memory controller 301 and in the memory device). That is, even if there are multi-bit errors in the read data of the memory device, the SEC code is the same for both the correct data and the corrupted data (it is part of the limitation of the SEC code, which is only used to solve single-bit errors). However, the SEC code generation algorithm (H) and the data returned by the memory device can still be processed to isolate the one containing the multi-bit error in the memory device.
[0049] Here, as is known in the art, the generation of an error correction code (such as the internal SEC code of a memory device) includes a bit-by-bit multiplication of the data to be encoded (for example, the data in a 128-bit region in the memory) with a binary value of 1 or 0 having 1 or 0 at each intersection of the rows or columns of the matrix H.
[0050] Here, for example, the bit pattern of each row of the matrix H indicates the presence (1) or absence (0) of the coefficients of the parity check equation (thus, each row represents a different parity check equation in the encoding process), and in this case, there is a unique parity check equation for each bit in the specific code to be generated for the data. Therefore, if an 8-bit code is to be generated for the data, the H matrix has 8 rows. Additionally, in a shared H matrix structure, the matrix H has a number of columns equal to the number of bits of the data to be encoded and the number of bits of the code that the encoder is to generate for the data.
[0051] Thus, for example, if the SEC code is to generate an 8-bit SEC code for 128 bits of data, the H matrix has 8 rows (one for each parity check equation or each bit in the generated SEC code) and 134 columns (128 columns for each bit of the data to be encoded and 8 bits for each bit of the SEC code to be generated for the data). After mathematically applying the matrix H to the encoded data (e.g., in a matrix multiplication operation), an 8-bit SEC code is generated. An example of such an H matrix is provided in Figure 4 In
[0052] In cases where the SEC code cannot even identify the presence of multiple-bit errors, since the SEC code read from the storage array of the memory device matches the corresponding SEC code generated by the memory controller, for Figure 4 the SEC matrix H, if the faulty memory device generates only 2-bit or 3-bit errors, the memory controller can still recover the lost data by performing the following operation:
[0053] Equation 5
[0054] Here: 1) H’ is the aforementioned SEC matrix H but with columns (e.g., 8) for the removed SEC code bits; 2) e is the error term that, when XORed with the correct data of the memory device generating the error, produces the corrupted data received for that memory device; 3) is the inverse self-permutation for memory device m; 4) is the self-permutation for memory device n; and 4) d is the number of memory devices. Performing the operation of Equation 5 will result in a minimum value (e.g., 0) for the position of the resulting matrix, in which case m = n = the faulty memory device (the position of the minimum value identifies the faulty memory device).
[0055] When the faulty memory device is known, the corrupted data received for that memory device can be corrected by applying the same method described above with respect to Equation 4. Thus, multiple-bit errors from a single memory device can be removed even if the SEC code generated from the read data and the SEC code stored with the read data do not match each other.
[0056] In cases of even more errors (such as more than three errors from the same memory device using Figure 4 the SEC code or errors from more than one memory device), in various embodiments, such errors will more likely reflect serious hardware problems rather than nominal DRAM bit flips and occasional / expected soft errors within the system.
[0057] Accordingly, referring to Figure 3c, the error checking circuit 302 of the memory controller further includes built-in self-test (BIST) logic circuitry 305. Here, the memory controller 301 enters the BIST mode if for the data word D P’≠P, then for all memory devices SEC’ = SEC, and the execution of the recovery process described above with respect to Equation 5 does not produce the desired minimum value at any location in the resulting array. In the case of entering the BIST mode, the memory controller 301 writes a known data pattern into the memory devices of the DIMM 320 and then reads them back from the DIMM 320.
[0058] Here, in the case of a severe hardware failure of any one or more memory devices (or a larger memory channel / system), the result of the BIST operation should disclose which memory device(s) is / are not operating reliably.
[0059] Although the above examples mainly rely on the use of x8 DIMMs, the reader should understand that the teachings of these examples can be easily applied to x4 DIMMs or other DIMMs composed of memory devices with bit widths different from four or eight bits. Additionally, although as described above the above examples discuss a single data word D, this single data word D can be one of many (e.g., 16).
[0060] In the above description of either of the memory controllers 301, the error checking circuit 302 or its components can be implemented using logic circuitry deployed on a semiconductor chip. The logic circuitry can include dedicated custom hardwired logic circuitry, programmable logic circuitry (such as field programmable gate array (FPGA) logic circuitry, programmable logic array (PLA) logic circuitry, etc.), or logic circuitry that executes program code (such as embedded processor logic circuitry, embedded controller logic circuitry), or any combination thereof.
[0061] Figure 5 illustrates the method described above. As observed in Figure 5 , the method includes reading 501 data and ECC code from the DIMM. The data includes data components provided by the corresponding memory devices on the DIMM. The ECC code is provided by the corresponding memory devices on the DIMM. The method also includes recalculating 502 a second version of the ECC code based on the data components. The recalculation includes applying different data scrambling to different components in the data components. The method also includes identifying 503 that the ECC code and the second version of the ECC code do not match. The method also includes receiving 504 the corresponding ECC code from the memory device to correct corruption in the data. The corresponding ECC code was originally generated within the memory device.
[0062] Figure 6 A basic model showing a basic computing system that can represent any of the above servers. As observed in Figure 6 , the basic computing system 600 may include a central processing unit 601 (which may include, for example, multiple general-purpose processing cores 615_1 through 615_X), a main memory controller 617 deployed on a multi-core processor or application processor, a system memory 602, a display 603 (such as a touch screen, tablet), a local wired point-to-point link (such as USB) interface 604, various network I / O functions 605 (such as an Ethernet interface and / or a cellular modem subsystem), a wireless local area network (such as WiFi) interface 606, a wireless point-to-point link (such as Bluetooth) interface 607, and a global positioning system interface 608, various sensors 609_1 through 609_Y, one or more cameras 610, a battery 611, a power management control unit 612, a speaker and microphone 613, and an audio encoder / decoder 614.
[0063] The application processor or multi-core processor 650 may include one or more general-purpose processing cores 615, one or more graphics processing units 616, a memory management function 617 (such as a memory controller), and an I / O control function 618 within its CPU 601. The general-purpose processing core 615 typically executes the operating system and application software of the computing system. The graphics processing unit 616 typically executes graphics-intensive functions to generate, for example, graphic information presented on the display 603. The memory control function 617 interacts with the system memory 602 to write data to / read data from the system memory 602. The power management control unit 612 typically controls the power consumption of the system 600.
[0064] The memory control function 617 (memory controller) may include error-checking logic circuitry that, as discussed in detail above, may use internal error-checking codes of the memory devices from which it reads / writes to recover lost data.
[0065] Each of the touch screen display 603, communication interfaces 604 - 607, GPS interface 608, sensors 609, (one or more) cameras 610, and speaker / microphone codecs 613, 614 can be considered various forms of I / O (input and / or output) related to the entire computing system, which, in appropriate cases, also includes integrated peripherals (such as one or more cameras 610). Depending on the implementation, each of these I / O components may be integrated on the application processor / multi-core processor 650 or may be located off the die of the application processor / multi-core processor 650 or outside the package of the application processor / multi-core processor 650.
[0066] The computing system may also include a system memory (also referred to as main memory) having multiple levels. For example, a first (faster) system memory level may be implemented using DRAM and a second (slower) system memory may be implemented using emerging non-volatile memory such as non-volatile memory whose storage cells are composed of chalcogenides, resistive random access memory (RRAM), ferroelectric random access memory (FeFRAM), etc. Emerging non-volatile memory technologies have faster access times than traditional FLASH and can thus be used in the role of system memory rather than just being classified as mass storage devices.
[0067] Software and / or firmware executed on a general-purpose CPU core of the processor (or other functional block having an instruction execution pipeline for executing program code) may perform any of the above functions.
[0068] Embodiments of the present invention may include the various processes described above. These processes may be embodied in machine-executable instructions. The instructions may be used to cause a general-purpose or special-purpose processor to execute certain processes. Alternatively, these processes may be executed by any combination of dedicated hardware components that include hardwired logic for executing the processes or programmed computer components and custom hardware components.
[0069] The elements of the present invention may also be provided as a machine-readable medium for storing machine-executable instructions. The machine-readable medium may include, but is not limited to, floppy disks, optical disks, CD-ROMs, and magneto-optical disks, flash memory, ROM, RAM, EPROM, EEPROM, magnetic or optical cards, propagation media, or other types of media / machine-readable media suitable for storing electronic instructions. For example, the present invention may be downloaded as a computer program, and the computer program may be transmitted from a remote computer (such as a server) to a requesting computer (such as a client) in the form of a data signal embedded in a carrier wave or other propagation medium via a communication link (such as a modem or network connection).
[0070] In the foregoing specification, the invention has been described with reference to its specific exemplary embodiments. However, it will be apparent that various modifications and changes can be made thereto without departing from the broader spirit and scope of the invention as set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. An apparatus, comprising: a memory controller for receiving data from a memory device, the memory controller including error checking logic circuitry for receiving an error checking code from the memory device, the error checking code being generated within the memory device based on the data, the error checking logic circuitry including circuitry for generating a second version of the error checking code based on the data received from the memory device and comparing the received error checking code with the second version of the error checking code to determine whether the data received from the memory controller is corrupted, wherein the error checking logic circuitry further includes an ECC code generation circuit for generating ECC based on the data received from the memory device and other components of data received from other memory devices on the same DIMM as the memory device, and wherein the memory controller receives the error checking code in response to the ECC and corresponding ECCs read from other components of the DIMM for the data and mismatched data.
2. The apparatus according to claim 1, wherein the memory device is deployed on an x4 DIMM.
3. The apparatus according to claim 1, wherein the memory device is deployed on an x8 DIMM.
4. The apparatus according to claim 1, wherein the error checking code is a SEC code.
5. The apparatus according to claim 1, wherein the data is received in a plurality of data read bursts.
6. The apparatus according to claim 1, wherein if the corresponding second version of the error checking code for each memory device matches each corresponding error checking code received from the memory device and other memory devices, the error checking logic circuitry applies the inverse of a first scrambling customized for a first memory device to a second scrambling customized for a second memory device with the corrupted received data.
7. The apparatus according to claim 1, wherein the error checking logic circuitry is used to apply different data scramblings.
8. The apparatus according to claim 1, wherein the error checking logic circuitry is used to apply a BIST sequence in the case where the corrupted data cannot be corrected using the error checking code.
9. A computing system, comprising: a plurality of processing cores: a main memory including DIMMs; A memory controller coupled to a DIMM, the memory controller being configured to receive data from a memory device, the memory controller including error checking logic circuitry configured to receive an error checking code from the memory device, the error checking code being generated within the memory device based on the data, the error checking logic circuitry including circuitry configured to generate a second version of the error checking code based on the data received from the memory device and to compare the received error checking code with the second version of the error checking code to determine whether the data received from the memory controller is corrupted, wherein the error checking logic circuitry further includes an error checking code ECC generation circuitry configured to generate an ECC based on the data received from the memory device and other components of data received from other memory devices on the same DIMM as the memory device, and wherein the memory controller receives the error checking code in response to the ECC and corresponding ECCs read from other components of the DIMM for the data and mismatches of the data.
10. The computing system according to claim 9, wherein the memory device is deployed on an x4 DIMM.
11. The computing system according to claim 9, wherein the memory device is deployed on an x8 DIMM.
12. The computing system according to claim 9, wherein the error checking code is a SEC code.
13. The computing system according to claim 9, wherein the data is received in a plurality of read data read bursts.
14. A method, comprising: reading data and an error checking code ECC from a DIMM, the data including data components provided by corresponding memory devices on the DIMM, the ECC being provided by the corresponding memory devices on the DIMM; recalculating a second version of the ECC based on the data components, the recalculation including applying different data scrambling to different components of the data components; identifying a mismatch between the ECC and the second version of the ECC; receiving a corresponding ECC from the memory device to correct corruption in the data, the corresponding ECC being initially generated with respect to the data within the memory device, receiving an error checking code from the memory device, the error checking code being generated within the memory device based on the data; generating a second version of the error checking code by applying different data scrambling from the memory device; and comparing the received error checking code with the second version of the error checking code, and if the two codes do not match, concluding that the data is corrupted.
15. The method according to claim 14, wherein the DIMM is an x4 DIMM.
16. The method according to claim 14, wherein the DIMM is an x8 DIMM.
17. The method according to claim 14, wherein the error checking code is a SEC code.
18. A computer-readable medium having instructions stored thereon that, when executed, cause a computing device to perform the method according to any one of claims 14 to 17.
19. An apparatus comprising means for performing the steps of the method according to any one of claims 14 to 17.
20. A computer program product having instructions which, when executed, cause a computing device to perform the method according to any one of claims 14 to 17.
Citation Information
Patent Citations
Controller and Method for Interfacing Between a Host Controller in a Host and a Flash Memory Device
US20110041039A1