Memory Rank Design of Memory Channels Optimized for Graph Applications
PUMA with X4 memory chips and tailored ECC encoding addresses inefficiencies in conventional architectures for graph applications, ensuring consistent ECC coverage and reduced overhead.
Patent Information
- Application Number
- JP2020192354
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-27
- Filing Date
- 2020-11-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-19
AI Technical Summary
Conventional memory architectures are inefficient for graph-related applications that do not follow spatial and temporal locality, leading to suboptimal ECC overhead and varying data patterns that affect ECC coverage.
Implementing a Programmable Unified Memory Architecture (PUMA) with X4 memory chips, utilizing a rank configuration that reduces ECC overhead to 50% by distributing ECC bits within the same chip, and employing ECC encoding structures tailored to specific error patterns of memory chip manufacturers.
Achieves consistent ECC coverage and reduced overhead for all data patterns, optimizing memory access for graph applications by adapting to different error patterns across memory chip manufacturers.
Smart Images

Figure 0007697757000001 
Figure 0007697757000002 
Figure 0007697757000003
Abstract
Description
Technical Field
[0001] [Statement of Government Rights] This invention was made with government support under contract No. HR0011-17-3-0004 awarded by DARPA. The government has certain rights in this invention.
[0002] The field of the invention generally relates to computational science, and more particularly, to memory rank design for memory channels optimized for graph applications.
Background Art
[0003] A computer system with a Harvard architecture executes program code in a processing core that fetches instructions and data from memory to "feed" the executable code. However, different types of programs will execute better if the architecture of the basic memory resources is optimized with reference to how the program accesses these memory resources.
Brief Description of the Drawings
[0004] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings.
[0005]
Figure 1
[0006]
Figure 2a
Figure 2b
[0007]
Figure 3
[0008]
Figure 4a
Figure 4b
[0009]
Figure 5
[0010]
Figure 6
[0011] FIG. 1 shows a conventional computer main memory system architecture 100. As can be seen in FIG. 1, the main memory system 100 includes a memory channel whose data bus 101 is 72 bits (72b) wide. Conventionally, 64 bits (64b) of the 72-bit (72b) bus are used for raw data (raw data part of the data bus), and 8 bits (8b) are used for error correction code (ECC) protection (ECC part of the data bus). Data access has conventionally been implemented in 8-cycle bursts. That is, the standard mechanism for accessing memory is to read or write data as 8 transfers of 72b (similarly in this case, 64b of data and 8b of ECC). Therefore, raw data is generally accessed in 64-byte bursts (64b / transfer X 8 transfers = 64 bytes (64B)).
[0012] Here, in recent years, the central processing unit (CPU) cores of computers generally cache data in units of 64 bytes (64B). One unit of 64B is generally referred to as a cache line. Therefore, the conventional main memory access in 64B bursts (64b / transfer X 8 transfers) as described above corresponds to the access of one CPU cache line.
[0013] Here, conventional software applications generally operate on data that has "spatial and temporal locality," meaning that data items physically stored close to each other within the main memory generally operate within the same time frame. Therefore, even when accessing large 64B chunks of data from the main memory, the system does not access overly many memory data for each access (generally, most of the data in a cache line is processed by the CPU core within the same time frame).
[0014] Unfortunately, some specific applications, such as graph-related applications, do not follow the spatial and temporal locality paradigm. Here, such applications tend to require smaller data units whose storage locations are distributed across the main memory within the same time frame. Therefore, a new architecture called the Programmable Unified Memory Architecture (PUMA) improves main memory access (or any memory access) to 8-byte (8B) chunks of raw data instead of 64B chunks of raw data. Graph-related applications can be executed by a Graphics Processing Unit (GPU). Thus, at least some of the anticipated applications of PUMA include a computer system having at least one GPU.
[0015] In the case of the PUMA approach, as shown in Figure 2a, main memory access is implemented as a burst of 8 transfers, each transfer containing 8b of raw data for a total of 64b (8B) per burst. However, referring to commonly manufactured memory chips, there is a problem with the ECC part of the PUMA access model. Specifically, individual memory chips are available in "X4" and "X8" versions. The X4 memory chip has a 4-bit data bus. The X8 memory chip has an 8-bit data bus. Thus, for example, the conventional main memory data bus of Figure 1 could be implemented with 9 X8 memory chips (8 chips for raw data and 1 chip for ECC), or 18 X4 memory chips (16 chips for raw data and 2 chips for ECC).
[0016] Generally, system designers strive to keep the amount of ECC overhead low. That is, for the same ECC encoding algorithm, a smaller amount of memory chip resources dedicated to storing ECC information is preferred over a larger amount of memory chip resources. Figure 2b shows the possible ranks of memory chips for a PUMA implementation using X8 memory chips. As is known in the art, a rank of memory chips is a set of memory chips that can support (or are the target of) burst accesses. As seen in Figure 2b, the ECC overhead is 100%. That is, there is one X8 memory chip 202_1 for storing all the raw data of the memory channel and another X8 memory chip 202_2 for storing the ECC information. Here, the X8 memory chip 202_1 by itself can handle all the raw data traffic of a PUMA access transaction (8 bits / transfer x 8 transfers). Therefore, in an X8 memory chip implementation, it is necessary to use the entire second chip 202_2 to store the ECC information.
[0017] One PUMA approach provides for compressing data into a smaller footprint such that a second chip 202_2 need not be used. Here, the ECC bits are stored in the first chip 202_1 in the remaining space that exists after compressing the raw data to less than 64b. However, while this approach is effective for some data patterns, it will not be effective for all data patterns. Therefore, the second memory chip 202_2 will still be at least necessary for those data patterns that cannot be compressed into a smaller footprint. Additionally, those data patterns that can be compressed into a smaller footprint tend to receive less ECC coverage than those data patterns that cannot be compressed (in the case of compression, fewer ECC bits are “packed” into the modest space opened up by the payload due to compression).
[0018] Therefore, as seen in FIG. 3, a better solution is to implement a rank of memory chips in a PUMA channel with X4 memory chips. As seen in FIG. 3, in an X4 memory chip, two memory chips 302_1, 302_2 can be used to store raw data and one X4 memory chip can be used to store ECC information. In this case, the ECC overhead is dramatically reduced to 50%. Here, a nominal PUMA memory burst access includes 8 transfers, each transfer including 8b of raw data and 4b of ECC information. Therefore, for each burst transaction, there is 64b of raw data and 32b of ECC information. Thus, regardless of the pattern of the raw data, the amount of ECC protection is the same and, furthermore, the amount of ECC protection is appropriate. Some exemplary ECC striping approaches are described in more detail below with respect to FIGS. 4a and 4b.
[0019] As is known in the art, the Joint Electron Device Engineering Council (JEDEC) has published memory channel interface specifications for compatibility by computer and other electronic device manufacturers. JEDEC emphasizes a memory access technology called Dual Data Rate (DDR), in which data transfer occurs on both the rising and falling edges of the transfer clock. The nomenclature accepted in the JEDEC specifications is to number them in the order in which they are released (e.g., DDR3, DDR4, DDR5, etc.). The most recent JEDEC DDR specifications correspond to DDR4 and DDR5.
[0020] According to the first embodiment, the rank of FIG. 3 is implemented with DDR4 X4 memory chips, and according to the second embodiment, the rank of FIG. 3 is implemented with DDR5 X4 memory chips. As is known in the art, both DDR4 and DDR5 are targeted at conventional computer systems that access main memory in bursts of 64B CPU cache lines. Therefore, DDR4 specifies 8 transfers per burst of nominally 64b raw data (= 64B total per burst), while DDR5 specifies 16 transfers per burst of nominally 32b raw data (also = 64B total per burst).
[0021] Thus, an X4 memory chip designed to comply with the DDR4 standard supports a nominal 8-cycle burst, while an X4 memory chip designed to comply with the DDR5 standard supports a nominal 16-cycle burst. However, importantly, the DDR5 standard also supports a burst "chop" mode in which the burst is executed in 8 cycles instead of 16 cycles.
[0022] As described above, a rank is a group of memory chips that are accessed together to support a memory access burst over a single memory channel. Therefore, the memory solution of FIG. 3 shows a single rank of memory chips for the PUMA channel. In various embodiments, a single rank is composed of X4 DDR4 memory chips, which, as described above, operate according to a burst length of nominally eight transfers. In alternative embodiments, a single rank is composed of X4 DDR5 memory chips. However, in order to conform to the PUMA architecture with a burst length of eight transfers, the DDR5 memory chips operate the burst in chop mode rather than their nominally 16 transfers per burst.
[0023] The rank of FIG. 3 can be implemented, for example, in a dual inline memory module (DIMM) that connects to the memory channel wiring of a computer motherboard. In the case of a single rank DIMM, only one instance of the rank of FIG. 3 is implemented in the DIMM. In the case of a dual rank DIMM, two instances of the rank of FIG. 3 are implemented in the DIMM. Here, typically, when multiple ranks are implemented in the same DIMM, the data bus wires of both ranks are logically connected together (e.g., DQ_0 of rank_0 is connected to DQ_0 of rank_1, DQ_1 of rank_0 is connected to DQ_1 of rank_1, etc.).
[0024] Control signals (not shown in FIG. 3) are also logically connected together in most cases, except for the chip select (CS) control signal, which is used, for example, to establish which rank of the DIMM is targeted by the host for a burst transaction. Other memory modules are also possible using various numbers of ranks per module (such as a stacked memory chip memory module). Memory modules with more than two ranks per module (e.g., three ranks per module, four ranks per module, etc.) are also possible.
[0025] Figures 4a and 4b show two different approaches for striping the ECC information of the third memory chip 302_3 in the memory chip rank for the PUMA implementation of FIG. 3. Here, in FIGS. 4a and 4b, chip 402_1 corresponds to the first X4 memory chip that stores the 4b "left half" of an 8b raw data transfer, chip 402_2 corresponds to the second X4 memory chip that stores the 4b "right half" of an 8b raw data transfer, and chip 402_3 corresponds to the third X4 memory chip that stores the ECC information. Here, row 1 corresponds to the 8b raw data and 4b ECC information transferred during 8 transfers within the first burst, row 2 corresponds to the 8b raw data and 4b ECC information transferred during 8 transfers in the second burst, and so on.
[0026] Generally speaking, an ECC algorithm generates ECC bits that perform numerically intensive calculations on the data being protected. And the ECC information is stored using the raw data. Then, when the data is read back, both the raw data and the stored ECC information are retrieved. The ECC calculation is performed again on the just-received raw data. If the newly calculated ECC information matches the ECC information stored with the raw data, the just-read raw data is understood to be free of data corruption.
[0027] However, if the newly calculated ECC information does not match the ECC information stored with the raw data, data corruption is understood to exist in the raw data and / or the ECC information. However, if the amount of actual data corruption is below a certain threshold, the corrupted bits can be recovered. If the amount of corruption is at or above the threshold, the error cannot be corrected, but at least the presence of the error is known, and an error flag can be set accordingly.
[0028] Generally, an ECC algorithm decomposes both the raw data to be protected and the ECC information generated from the raw data into symbols. A symbol is either a bit within the raw data or a group of ECC information that acts as a unit within the ECC algorithm. Generally speaking, error recovery processing can recover all raw data and ECC symbols as long as the total number of damaged raw data symbols and ECC symbols is below a certain threshold. The threshold number of damaged symbols depends on the ECC algorithm used and the ratio of ECC information to raw data information (generally, the higher the ratio, the higher the threshold of acceptable damaged symbols).
[0029] Interestingly, different memory chip manufacturers will exhibit different patterns of data corruption. That is, for example, a first memory chip manufacturer will exhibit repeated errors via a sequence of burst transfers on the same data pin but not across multiple data pins (e.g., data pin D0 always gets corrupted in repeated transfers, but data pins D1, D2, and D3 do not get corrupted in these same transfers). In contrast, a second memory chip manufacturer will exhibit errors across multiple data pins in the same burst transfer, but other transfers in the burst do not get corrupted across all data pins (e.g., data pins D0, D1, and D2 get corrupted in one transfer of the burst, but all other transfers in the burst do not get corrupted across each data pin from D0 to D3). These differences in error patterns seen across manufacturers are due, for example, to differences in the design and / or manufacturing process of each manufacturer's chips.
[0030] The different ECC encoding approaches of FIGS. 4a and 4b define symbols in different ways based on the idea that one type of the above memory chip manufacturer will cause fewer damaged symbols with one of the ECC encoding approaches, while the other type of the above manufacturer will cause fewer damaged symbols with the other ECC encoding approach.
[0031] Specifically, the ECC encoding approach of FIG. 4a defines symbols in the "length direction" (symbol data passes through the same data pins of the memory chip), whereby the manufacturer indicates an error according to the above-mentioned first pattern (for example, the error appears on one data pin during a burst and does not appear across other data pins). As can be seen in FIG. 4a, there are 12 8b symbols running vertically along the columns of FIG. 4a. Eight of the symbols are raw data symbols while four of the symbols are ECC symbols. In this case, for example, if only one data pin of one rank of three memory chips indicates an error during a burst, only one symbol is affected and all errors can be recovered.
[0032] In contrast, the ECC encoding approach of FIG. 4b defines symbols in the "horizontal direction" (symbol data flows across multiple data pins), whereby the manufacturer indicates an error according to the above-mentioned second pattern (for example, the error appears across data pins during a specific transfer of a burst and does not appear across other transfers of the burst). Again, there are eight raw data symbols and four ECC symbols. In this case, if multiple data pins of one memory chip out of one rank of memory chips indicate an error only during one transfer of a burst, only one symbol is affected and all errors can be recovered.
[0033] According to any of these ECC encoding structures, Reed-Solomon ECC encoding can be easily derived, and it is considered that errors can be recovered when up to two symbols are damaged according to the rank structures of FIGS. 3, 4a, and 4b. In addition, such an algorithm will be able to specifically identify up to four specific symbols that are damaged.
[0034] FIG. 5 shows an embodiment of a memory controller 501 designed to interface with at least one memory rank designed according to the embodiment of FIG. 3 and to generate ECC information according to the embodiments of FIGS. 4a and 4b. As can be seen in FIG. 5, the memory controller 501 includes either or both of DDR4 and DDR5 memory channel interfaces 502, and an ECC generation logic circuit 503 for preparing ECC information according to both structures seen in FIGS. 4a and 4b. In addition, the memory controller 501 includes a configuration register space 504 that indicates to the memory controller which type of ECC structure is applied (the one in FIG. 4a or the one in FIG. 4b).
[0035] Here, if it is known that the memory controller 501 is coupled to a rank of memory chips that exhibit one type of error pattern, the memory controller is configured to apply a suitable ECC encoding structure that minimizes damaged symbols with reference to the type of error pattern (e.g., using low-level software / firmware of the computer system of the memory controller). Depending on the implementation, such a configuration can be created for each channel (e.g., such that the ECC encoding structures of the various channels can be optimized even when the various channels are coupled to respective ranks having memory chips that exhibit different error types of the error pattern).
[0036] In yet other implementations, a portion of the entire memory space controlled by the memory controller is allocated to the GPU, and the memory controller 501 accesses this memory space according to the PUMA architecture and the corresponding rank structure of FIG. 3 (for example, the memory space allocated to the GPU includes the PUMA memory channels). At the same time, another portion of the entire memory space is allocated to one or more CPU processing cores, and the memory controller 501 accesses this other portion according to conventional CPU cache line access burst processing (for example, 8 transfers of 64 bits of raw data per burst between ranks, 16 transfers of 32 bits of raw data per burst between ranks, etc.). For example, conventional DDR4 and DDR5 memory channels can be used. Thus, for example, one set of memory channel I / O from the memory controller 501 can be used to implement the first portion of the memory space, while another set of memory channel I / O from the memory controller 501 can be used to implement the second portion of the memory space.
[0037] The memory controller 501 generally includes logic circuitry for performing any / all of the communications with the memory chips as described above.
[0038] FIG. 6 provides an exemplary diagram of a computing system 600 (e.g., a smartphone, a tablet computer, a laptop computer, a desktop computer, a server computer, etc.). As observed in FIG. 6, the basic computing system 600 may include a central processing unit 601 (which may include, for example, a plurality of general-purpose processing cores 615_1 to 615_X), a main memory controller 617 disposed on a multi-core processor or an application processor, a system memory 602, a display 603 (e.g., a touch screen, a flat panel), a local wired point-to-point link (e.g., a USB interface 604), various network I / O functions 605 (such as an Ethernet® interface and / or a cellular modem subsystem, etc.), a wireless local area network (e.g., WiFi®) interface 606, a wireless point-to-point link (e.g., Bluetooth®) interface 607, a global positioning system interface 608, various sensors 609_1 to 609_Y, one or more cameras 610, a battery 611, a power management control unit 612, a speaker and a microphone 613, and an audio coder / decoder 614.
[0039] The application processor or multi-core processor 650 may include one or more general-purpose processing cores 615, one or more graphics processing units 616, a memory management function 617 (e.g., a memory controller), and an I / O control function 618 within its CPU 601. The general-purpose processing cores 615 typically execute the operating system and application software of the computing system. The graphics processing unit 616 typically executes graphics-intensive functions to generate, for example, graphic information provided on the display 603. The memory control function 617 interfaces with the system memory 602 to read and write data to and from the system memory 602. The power management control unit 612 generally controls the power consumption of the system 600.
[0040] Each of the touch screen display 603, communication interfaces 604-607, GPS interface 608, sensors 609, (multiple) cameras 610, and speaker / microphone CODECs 613, 614 can all similarly be shown as various forms of I / O (input and / or output) to the overall computing system, including integrated peripheral devices (e.g., one or more cameras 610) in appropriate locations. Depending on the implementation, various ones of these I / O components can be integrated on the application processor / multi-core processor 650 or can be placed outside the die or outside the package of the application processor / multi-core processor 650.
[0041] The computing system also includes non-volatile storage 620, which can be a mass storage component of the system. Here, for example, the mass storage can be composed of one or more SSDs configured with FLASH (registered trademark) memory chips, and its multi-bit memory cells are programmed with different storage densities depending on SSD capacity utilization, as described in detail above.
[0042] Embodiments of the present invention include various processes as described above. The above processes may be embodied in machine-executable instructions. The above instructions can be used to cause a general-purpose or special-purpose processor to execute certain processes. Alternatively, the above processes may be executed by special or custom hardware components incorporating wiring logic circuits or programmable logic circuits (e.g., FPGA, PLD) for executing the processes, or may be executed by any combination of programmed computer components and custom hardware components.
[0043] The elements of the present invention may be provided as a machine-readable medium storing machine-executable instructions. Examples of the machine-readable medium include, but are not limited to, floppy disks, optical disks, CD-ROMs, magneto-optical disks, flash memories, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, propagation media, and other types of media or machine-readable media suitable for storing electronic instructions. For example, the present invention may be downloaded as a computer program in the form of a data signal embodied in a carrier wave from a remote computer (e.g., a server) to a requesting computer (e.g., a client), or transmitted by other propagation media via a communication link (e.g., a modem or network connection).
[0044] In the foregoing specification, the present invention has been described with reference to specific exemplary embodiments thereof. However, it will be apparent that various modifications and changes can be made without departing from the broader spirit and scope of the invention as set forth in the appended claims. Accordingly, the specification and drawings are to be interpreted in an illustrative rather than a limiting sense.
[0045] According to the present specification, the configurations described in each of the following items are also disclosed. [Item 1] An apparatus comprising memory chips of a rank connected to a memory channel, the memory channel being characterized as having eight transfers of 8-bit raw data per burst access, the rank of the memory chips including first, second, and third X4 memory chips, the X4 memory chips being compliant with JEDEC dual data rate (DDR) memory interface specifications, the first and second X4 memory chips being connected to an 8-bit raw data portion of a data bus of the memory channel, and the third X4 memory chip being connected to an error correction coding (ECC) information portion of the data bus of the memory channel. [Item 2] The apparatus according to item 1, wherein the X4 memory chip is an X4 DDR4 memory chip. [Item 3] The rank of the memory chip is the device according to item 2, which is arranged on the dual in-line memory module. [Item 4] The X4 memory chip is the X4 DDR5 memory chip, and the device is according to item 1. [Item 5] The eight data transfers per burst access are executed in chop mode, and the device is according to item 4. [Item 6] The rank of the memory chip is the device according to item 4, which is arranged on the dual in-line memory module. [Item 7] The ECC information part of the data bus of the memory channel transfers the entire symbol of the ECC information across the data pins of the third X4 memory chip, and the device is according to item 1. [Item 8] The ECC information part of the data bus of the memory channel transfers the entire symbol of the ECC information through a single data pin of the third X4 memory chip, and the device is according to item 1. [Item 9] A plurality of CPU processing cores, A graphics processing unit, A memory controller, Comprising, a memory channel is issued from the memory controller, the memory channel is characterized by having eight transfers of 8-bit raw data per burst access, a memory module is coupled to the memory channel, the memory module has a rank of memory chips, The rank of the memory chips includes first, second, and third X4 memory chips, the X4 memory chips conform to the JEDEC dual data rate (DDR) memory interface specification, the first and second X4 memory chips are connected to the 8-bit raw data part of the data bus of the memory channel, and the third X4 memory chip is connected to the error correction coding (ECC) information part of the data bus of the memory channel, a computing system. [Item 10] The apparatus according to item 9, wherein the X4 memory chip is an X4 DDR4 memory chip. [Item 11] The apparatus according to item 10, wherein the memory module is a dual in-line memory module. [Item 12] The apparatus according to item 9, wherein the X4 memory chip is an X4 DDR5 memory chip. [Item 13] The apparatus according to item 12, wherein the eight data transfers per burst access are executed in chop mode. [Item 14] The apparatus according to item 13, wherein the memory module is a dual in-line memory module. [Item 15] An interface for communicating with a memory chip compliant with JEDEC dual data rate (DDR) memory interface specifications, A logic circuit, Accessing a memory chip compliant with JEDEC DDR interface specifications by eight transfers of 8-bit raw data and 4-bit ECC information per burst access, The logic circuit that calculates ECC information from the raw data of the burst access using a symbol orientation based on the manufacturer of the memory chip, A memory controller comprising: [Item 16] The memory controller according to item 15, wherein the first ECC symbol orientation includes orienting the symbols in the length direction such that a single symbol is transferred via any single data pin of the memory chip. [Item 17] The memory controller according to item 16, wherein the second ECC symbol orientation includes orienting the symbols in the horizontal direction such that a single symbol is transferred via any four data pins of the memory chip. [Item 18] The memory controller according to claim 15, wherein the eight transfers are executed in chop mode. [Item 19] The memory controller according to item 15, comprising a configuration register space for establishing a specific one of a plurality of available symbol orientations. [Item 20] The memory controller according to item 15, comprising another interface for communicating with another memory chip according to a burst process for transferring a CPU cache line.
Claims
1. A rank of memory chips connected to a memory channel, the memory channel being characterized as having eight transfers of 8-bit raw data per burst access, the rank of memory chips including a first X4 memory chip, a second X4 memory chip, and a third X4 memory chip, the first X4 memory chip, the second X4 memory chip, and the third X4 memory chip conforming to the JEDEC dual data rate (DDR) memory interface specification, the first X4 memory chip and the second X4 memory chips being connected to an 8-bit raw data portion of the data bus of the memory channel, the third X4 memory chip being connected to an error correction coded information portion (ECC information portion) of the data bus of the memory channel, the ECC information portion of the data bus of the memory channel transferring an entire symbol of ECC information via a single data pin of the third X4 memory chip device.
2. The apparatus according to claim 1, wherein the eight data transfers per burst access comply with a programmable unified memory architecture (PUMA).
3. The apparatus according to claim 1 or 2, wherein the memory channel has eight transfers of 8-bit raw data and 4-bit ECC information per burst access.
4. The apparatus according to any one of claims 1 to 3, wherein the first X4 memory chip, the second X4 memory chip, and the third X4 memory chip are X4 DDR4 memory chips.
5. The apparatus according to claim 4, wherein the rank of memory chips is disposed on a dual in-line memory module.
6. The apparatus according to any one of claims 1 to 3, wherein the first X4 memory chip, the second X4 memory chip, and the third X4 memory chip are X4 DDR5 memory chips.
7. The apparatus according to claim 6, wherein the eight data transfers per burst access are performed in a chop mode.
8. The apparatus according to claim 6 or 7, wherein the rank of memory chips is disposed on a dual in-line memory module.
9. An interface for communicating with a memory chip conforming to the JEDEC dual data rate (DDR) memory interface specification, a logic circuit, Access a memory chip compliant with the JEDEC DDR interface specification by means of eight transfers of 8-bit raw data and 4-bit ECC information per burst access, The logic circuit that calculates ECC information from the raw data of the burst access using the symbol orientation based on the manufacturer of the memory chip, Comprising A memory controller.
10. The memory controller according to claim 9, wherein the first ECC symbol orientation includes orienting the symbols in the length direction such that a single symbol is transferred via any single data pin of the memory chip.
11. The memory controller according to claim 9, wherein the second ECC symbol orientation includes orienting the symbols in the horizontal direction such that a single symbol is transferred via any four data pins of the memory chip.
12. The memory controller according to any one of claims 9 to 11, wherein the eight transfers are performed in chop mode.
13. The memory controller according to any one of claims 9 to 12, wherein the memory controller includes a configuration register space for establishing a specific one of a plurality of available symbol orientations.
14. The memory controller according to any one of claims 9 to 13, wherein the memory controller includes another interface for communicating with another memory chip according to a burst process for transferring CPU cache lines.
15. A computing system comprising the device according to any one of claims 1 to 8.
16. A computing system comprising the memory controller according to any one of claims 9 to 14.
Citation Information
Patent Citations
Method and apparatus for enabling shared bus interrupt joint signaling in a multi-rank memory subsystem
JP2010501098A
Memory transaction burst operation for supporting time-multiplexed error correction coding and memory component
JP2011243206A
Semiconductor device, memory access control method, and semiconductor device system
JP2016118950A
System and method for providing error correction and detection in a memory system
US20090049365A1
Memory system components for split channel architecture
US20140325105A1