Host-level error detection and error correction
Host-level error detection and correction in memory systems address uncorrected errors by dividing data requests into portions and using parity fetches to enhance reliability and reduce maintenance costs.
Patent Information
- Application Number
- JP2024572470
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-16
- Filing Date
- 2023-06-15
- Publication Date
- 2025-07-08
AI Technical Summary
Memory structures face reliability issues due to uncorrected errors in data storage, leading to field replacement unit events and increased maintenance costs, despite on-die error correction codes.
Implement host-level error detection and correction by dividing data fetch and write requests into portions, using error correction codes to detect and correct defects in memory systems, and reconstructing data using parity fetches and unused check bits.
Enhances system reliability by reducing the likelihood of field replacement unit events and system reboots, thereby lowering maintenance costs.
Smart Images

Figure 2025521235000001_ABST
Abstract
Description
Background Art
[0001] In a memory structure, data stored in a memory bank of a memory layer may include uncorrected errors due to corruption of the stored data. To detect and correct these errors, some memory structures implement an on-die error correction code that generates check bits for the stored data. When data is read from the memory bank, these memory structures use the generated check bits to detect and correct errors in the data. However, even with such on-die error correction, when the data is read out to the processing system, defects that are not limited to the set number of bits of the data still occur. Such non-limited defects can reduce the reliability of the processing system because they can cause the processing system to malfunction and result in a field replacement unit event. These field replacement unit events require replacement of one or more components of the processing system and thus increase the cost of maintaining the processing system.
[0002] The present disclosure may be better understood by reference to the accompanying drawings, and its various features and advantages may become apparent to those skilled in the art. The use of the same reference numerals in different drawings indicates similar or identical items.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
[0004] The techniques and systems described herein relate to providing host-level error detection and correction in a processing system. A processor uses the techniques disclosed herein to perform error detection on each fetch return in a set of fetch returns from a memory based on one or more check bits provided by the memory. The set of fetch returns collectively corresponds, for example, to cache lines read from the memory. In response to detecting an error (e.g., referred to as a bad return) in any fetch return, the processor reconstructs data for bad returns from other fetch returns in the set and for additional parity fetches provided by the memory. Thereby, the processor improves error correction of the fetched data and thus improves the overall reliability of the processing system.
[0005] To illustrate the present technology and system, a processing system includes one or more processing devices (e.g., CPU, GPU) communicatively coupled to a memory configured to implement one or more error correction codes (ECCs), and one or more portions of data (e.g., stored cache lines) stored in the memory are associated with one or more check bits. A processing device of the processing system generates a set of fetch portions that identify data read from the memory and provides them to the memory. In response to receiving the set of fetch portions, the memory reads the data identified by the fetch portions and checks whether one or more defects (e.g., driver defect, bank defect) exist in the read data, and is configured to correct one or more detected defects. The memory uses one or more check bits associated with the read data to check and correct defects in the read data. Next, the memory transmits the corrected data and one or more unused check bits associated with the read data (e.g., check bits not used for defect detection or correction) to the processing device as a set of fetch returns. Further, the memory transmits an additional parity fetch including data obtained by performing one or more operations on the fetch returns, such as on the data included in the fetch returns. In response to receiving the fetch returns, the processing device is configured to detect defects in the received fetch returns using the unused check bits. After detecting a defect in the received fetch returns, the processing device erases the fetch returns and reconstructs the fetch returns using the data and parity fetches in one or more other received fetch returns. In this way, any detected defects isolated in a single fetch return are corrected, and the reliability of the processing system is improved by improving the detection and correction of such defects. Accordingly, the likelihood of in-field replacement unit events, termination of an application using the memory, system reboot, or any combination thereof occurring in the system is reduced, which helps to reduce the maintenance cost of the system.
[0006] In addition, the processing device is configured to generate a set of write portions that identify data to be written to the memory. The set of write portions corresponds to, for example, cache lines written to the memory. Also, the processing device is configured to generate a write parity that includes data obtained by performing one or more operations on the data identified by the write portions based on the set of write portions. Next, the processing device generates respective check values for each of the write portions and the write parity based on an ECC implemented by the processing device. After generating the check values, the write portions, the check values, and the write parity are transmitted to and stored in the memory. Thereby, the reliability of the data written to the memory is improved. For example, in response to detecting one or more errors in the data written to the memory, the data is reconstructed using other data stored in the memory and the data of the write parity. Therefore, the likelihood of a field replaceable unit event, termination of an application using the memory, system reboot, or any combination thereof occurring is reduced as the reliability of the data is improved.
[0007] FIG. 1 is a block diagram of a processing system 100 for high memory bandwidth host error correction according to some embodiments. The processing system 100 includes or has access to a memory 106 or other storage component implemented using a non-transitory computer-readable storage medium such as, for example, dynamic random access memory (DRAM), static random access memory (SRAM), non-volatile RAM, or any combination thereof. In an embodiment, the memory 106 includes a three-dimensional (3D) stacked synchronous dynamic random access memory (SDRAM) having one or more memory banks, memory sub-banks, or both, each having one or more memory layers. The memory 106 includes an interface, such as a high bandwidth memory (HBM) interface, a second generation high bandwidth memory (HBM2) interface, a third generation high bandwidth memory (HBM3) interface, etc., configured to communicatively couple the memory 106 to one or more portions of the processing system 100. According to an embodiment, the memory 106 is external to a processing unit implemented within the processing system 100, and in other embodiments, the memory 106 is internal to a processing unit implemented within the processing system 100. Also, the processing system 100 includes a bus 112 for supporting communication between entities implemented within the processing system 100, such as a central processing unit (CPU) 102 and a graphics processing unit (GPU) 114. Some embodiments of the processing system 100 include other buses, bridges, switches, routers, etc., which are not shown in FIG. 1 for clarity.
[0008] The techniques described herein are utilized in various embodiments with any of a variety of parallel processors (e.g., vector processors, CPUs, GPUs, general-purpose GPUs (GPGPUs), non-scalar processors, high parallel processors, artificial intelligence (AI) processors, inference engines, machine learning processors, other multi-threaded processing units, etc.), scalar processors, serial processors, or any combination thereof. FIG. 1 shows an example of a parallel processor, particularly a GPU 114, according to some embodiments. The GPU 114 renders an image shown on the display 120. For example, the GPU 114 renders an object to generate pixel values provided to the display 120, and the display 120 uses the pixel values to display an image representing the rendered object. The GPU 114 implements a plurality of processor cores 116-1 to 116-N that execute instructions simultaneously or in parallel. According to an embodiment, one or more of the processor cores 116 operate as single instruction multiple data (SIMD) units that execute the same operation on different data sets. In the exemplary embodiment shown in FIG. 1, three cores (116-1, 116-2, 116-N) representing N cores are shown, but the number of processor cores 116 implemented within the GPU 114 is a matter of design choice. Thus, in other embodiments, the GPU 114 can include any number of cores 116. Some embodiments of the GPU 114 are used for general-purpose computing. The GPU 114 executes instructions such as program code 108 stored in the memory 106, and the GPU 114 stores information such as the results of the executed instructions in the memory 106.
[0009] The processing system 100 also includes a CPU 102 that is connected to the bus 112 and thus communicates with the GPU 114 and the memory 106 via the bus 112. The CPU 102 implements a plurality of processor cores 104-1 to 104-N that execute instructions simultaneously or in parallel. In an embodiment, one or more of the processor cores 104 operate as SIMD units that perform the same operations on different data sets. In the exemplary embodiment shown in FIG. 1, three cores (104-1, 104-2, 104-M) representing M cores are shown, but the number of processor cores 104 implemented within the CPU 102 is a matter of design choice. Thus, in other embodiments, the CPU 102 can include any number of cores 104. In some embodiments, the CPU 102 and the GPU 114 have an equal number of cores 104, 116, but in other embodiments, the CPU 102 and the GPU 114 have different numbers of cores 104, 116. The processor core 104 executes instructions such as program code 110 stored in the memory 106, and the CPU 102 stores information such as the results of the executed instructions in the memory 106. Also, the CPU 102 can initiate graphics processing by issuing a draw call to the GPU 114. In an embodiment, the CPU 102 implements a plurality of processor cores (not shown in FIG. 1 for clarity) that independently execute instructions simultaneously or in parallel.
[0010] In an embodiment, the CPU 102, the GPU 114, or both are configured to generate a fetch request (e.g., a cache line fetch) and send it to the memory 106. For example, the CPU 102, the GPU 114, or both are configured to generate a load operation or a read operation that requests data (e.g., a cache line) from the memory 106. The fetch request identifies and requests data that is necessary for, supports, or is useful for one or more instructions executed by the CPU 102, the GPU 114, or both. That is, the fetch request identifies and requests data for a read operation or a load operation. In response to receiving the fetch request, the memory 106 is configured to read the data identified by the fetch request and provide it to the CPU 102, the GPU 114, or both. According to an embodiment, the CPU 102, the GPU 114, or both are configured to divide the fetch request into two or more fetch portions, each containing data representing at least a part of the fetch request. For example, the CPU 102, the GPU 114, or both are configured to divide the fetch request into two or more fetch portions, each representing an individual part of the requested data identified by the fetch request. In an embodiment, each of the resulting fetch portions is of equal size. As an example, the CPU 102, the GPU 114, or both are configured to divide a 128-bit fetch request (i.e., a fetch request that identifies 128 bits of requested data) into four fetch portions, each representing an individual 32-bit part of the fetch request. In response to receiving the fetch portion, the memory 106 is configured to read the corresponding fetch return containing the data identified by the fetch portion and provide it to the CPU 102, the GPU 114, or both.
[0011] In an embodiment, the CPU 102, the GPU 114, or both are configured to generate a write request (e.g., a cache line write request) and send it to the memory 106. The write request identifies data to be written from one or more caches included in the CPU 102, the GPU 114, or both, or communicatively coupled thereto, to one or more portions of the memory 106. The memory 106 is configured to write the data identified by the write request to one or more portions of the memory 106 in response to receiving the write request. In an embodiment, the CPU 102, the GPU 114, or both are configured to divide the write request into four write portions, each representing an individual portion of the data identified by the write request. The memory 106 is configured to write the data identified by the write portion to one or more portions of the memory 106 in response to receiving the write portion.
[0012] According to an embodiment, the CPU 102, the GPU 114, or both are configured to generate a parity request in response to generating one or more fetch portions from a fetch request. That is, the CPU 102, the GPU 114, or both generate a parity request when the fetch request is divided into one or more fetch portions. The parity request identifies each fetch portion and one or more operations to be performed. For example, the parity request identifies one or more operations to be performed on the data identified by each fetch portion. The memory 106 is configured to generate a parity fetch and return it to the CPU 102, the GPU 114, or both in response to receiving the parity request. The parity fetch includes data obtained by performing the operations identified by the parity request on the data identified by one or more fetch requests. As an example, the memory 106 receives a parity request that identifies an XOR operation and four fetch portions, each indicating a part of the data (e.g., D0, D1, D2, D3), and in response thereto, the data from each fetch return (e.g.,
Number
Number
[0013] . The memory 106 is configured to store the data of the write parity in one or more portions of the memory 106 in response to receiving the write parity.
[0013] According to an embodiment, the memory 106 is configured to implement one or more error correction codes (ECCs) for data stored in the memory 106. For example, the memory 106 is configured to implement one or more ECCs such that one or more check bits are generated for data to be stored in response to receiving a request (e.g., a write request) to store data in the memory 106. In response to generating the check bits, the memory 106 stores the memory and the check bits in at least a portion of the memory 106. Such ECCs include, for example, on-die ECC, single error correction (SEC) codes, symbol-based codes (e.g., Reed-Solomon codes), cyclic redundancy check (CRC) codes, horizontal redundancy check (LRC) codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes, or any combination thereof. As an example, the memory 106 is configured to generate 32 check bits for every 128 bits (e.g., cache line) stored in the memory 106 and store the check bits in at least a portion of the memory 106. In response to receiving one or more fetch portions, the memory 106 reads out the data specified by the received fetch portion and is configured to implement an ECC such that one or more defects in the read data are detected and corrected. According to some embodiments, the memory 106 is configured to detect and correct defects limited to a predetermined number of bits in the read data. That is, the defects are limited to each respective portion of the read data, each containing a predetermined number of bits. For example, the memory 106 is configured to detect and correct defects limited to a 16-bit portion of the read data. The memory 106 is configured to detect and correct defects, for example, by comparing one or more check bits associated with the read data with one or more values (e.g., threshold values, values, character strings), performing one or more operations on the check bits associated with the data, or both.Memory 106 is configured to generate a fetch return using the corrected data and send the fetch return to CPU 102, GPU 114, or both, in response to detecting one or more defects, correcting them, or both. In an embodiment, in response to the memory 106 detecting and correcting a defect limited to a predetermined number of bits in the read data, one or more check bits associated with the read data are left unused. That is, only a part of the check bits associated with the read data is used so that one or more unused check bits remain to detect and correct errors in the read data, in response to the memory 106 detecting and correcting a defect limited to a predetermined number of bits. As an example, in response to the memory 106 detecting and correcting a defect in the read data limited to a 16-bit portion of the read data, only 16 of the 32 check bits associated with the read data are necessary to detect and correct the defect in this 16-bit portion. Thus, only 16 check bits are used and 16 unused check bits remain.
[0014] According to an embodiment, the memory 106 is configured to send one or more unused check bits to the CPU 102, the GPU 114, or both, along with one or more fetch returns. For example, the memory 106 is configured to send a set of fetch returns including the read request data and one or more unused check bits to the CPU 102, the GPU 114, or both in response to receiving a set of fetch portions that identify the request data (e.g., cache line). In some embodiments, the memory 106 is configured to add one or more unused check bits associated with the data in the fetch return to each fetch return, while in other embodiments, the memory 106 sends the unused check bits separately from the fetch return. For example, the memory 106 is configured to add the unused check bits associated with the read data in the fetch return to the fetch return such that each fetch return includes the read data identified by the respective fetch portion and one or more unused check bits. The unused check bits include, for example, metadata associated with the fetch return indicating, for example, the error state of the fetch portion. As an example, the memory 106 is configured to provide a set of fetch returns, each including 32-bit read data identified by the respective fetch portion and 16 unused check bits (e.g., check bits not used for detecting or correcting errors in the data of the fetch return), in response to receiving a set of 32-bit fetch portions (i.e., fetch portions that identify 32-bit request data) corresponding to a fetch request (e.g., requested cache line). The CPU 102, the GPU 114, or both are configured to implement one or more ECCs to detect one or more defects in the received fetch return in response to receiving the fetch return. Such ECCs include, for example, CRC codes (e.g., CRC-9, CRC-12, CRC-16), symbol-based codes, LRC codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes, or any combination thereof, etc.For example, the CPU 102, GPU 114, or both are configured to implement a CRC-12 code for each fetch return. In an embodiment, the CPU 102, GPU 114, or both are configured to detect one or more defects in a fetch return using one or more unused check bits associated with the fetch return (e.g., included in the fetch return, returned with the fetch return, or both) according to one or more ECCs implemented by the CPU 102, GPU 114, or both. For this purpose, the CPU 102, GPU 114, or both are configured to, for example, compare one or more unused check bits associated with the fetch return to one or more values (e.g., threshold values, values, strings), perform one or more operations on one or more unused check bits associated with the fetch return according to the implemented ECC, or both, to detect a defect in the fetch return. As an example, the CPU 102, GPU 114, or both are configured to perform one or more operations on one or more unused check bits associated with the fetch return according to a CRC-12 implemented by the CPU 102, GPU 114, or both to detect one or more defects in the fetch return.
[0015] The CPU 102, GPU 114, or both are configured to correct a detected defect in response to detecting one or more defects in a fetch return. The CPU 102, GPU 114, or both are configured to correct a detected defect, for example, by erasing a fetch return including one or more defects and reconstructing the erased fetch return. That is, the CPU 102, GPU 114, or both are configured to erase a fetch return in response to detecting one or more defects in the fetch return. The CPU 102, GPU 114, or both are configured to reconstruct the erased fetch return from other received fetch returns and parity fetches after erasing the fetch return. For example, the CPU 102, GPU 114, or both are configured to reconstruct the erased fetch return by performing an XOR operation on the data and parity fetches within other received fetch returns. In this way, any defect isolated to a single fetch return (e.g., word line driver defect, bank defect) is corrected. By correcting defects in this way, the defect detection coverage and correction of the system are improved, and the reliability of the system is enhanced. Accordingly, the likelihood of a field replaceable unit event, termination of an application using the memory, system reboot, or any combination thereof occurring in the system is reduced.
[0016] In an embodiment, the CPU 102, the GPU 114, or both are configured to implement one or more ECCs to generate one or more check values for each write portion and write parity generated by the CPU 102, the GPU 114, or both. Such ECCs include, for example, CRC codes (e.g., CRC-9, CRC-12, CRC-16), symbol-based codes, LRC codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes, or any combination thereof. For example, the CPU 102, the GPU 114, or both are configured to implement a CRC-12 code for each write portion and write parity such that a 12-bit check value is generated for each write portion and write parity associated with the write portion. According to an embodiment, each check value generated for each write portion is included in the write portion when the write portion is transmitted to the memory 106, and each check value for each write parity is included in the write parity when the write parity is transmitted to the memory 106. Thus, the memory 106 is configured to store the data or write parity and check values specified in the write portion in one or more portions of the memory 106 in response to receiving the write portion. Thereby, the reliability of the data written to the memory is improved. For example, in response to detecting one or more errors in the data written to the memory, the data is reconstructed using other data and write parity stored in the memory, improving the reliability of the system.
[0017] Next, referring to FIG. 2, a block diagram of a processing system 200 for host-level error detection and correction is shown. The processing system 200 includes a CPU 102, a GPU 114, or both, similar to or the same as the cache 236 and communicatively coupled to a memory 206 similar to or the same as the memory 106. The cache 236 includes, for example, one or more private caches, shared caches, or both, included in or communicatively coupled to the processing device 230. According to an embodiment, the processing system 200 is disposed within a single package in which at least the memory 206 and the processing device 230 are disposed on the same substrate (not shown for clarity) of the package. In an embodiment, the processing device 230 and the memory 206 are communicatively coupled by a silicon interposer disposed on the substrate.
[0018] Memory 206 includes one or more stacked memory layers 224, each of which includes one or more memory banks, memory sub - banks, or both. For example, memory 206 includes a 3D stacked SDRAM having one or more memory layers 224, each of which includes one or more memory banks. The exemplary embodiment shown in FIG. 2 shows that memory 206 has eight memory layers (224 - 1, 224 - 2, 224 - 3, 224 - 4, 224 - 5, 224 - 6, 224 - 7) representing N memory layers, but in other embodiments, memory 206 can have any number of memory layers 224. According to an embodiment, each memory layer 224 includes two channels, e.g., two 64 - bit channels, configured to enable access, modification, and deletion of data stored in the memory banks, memory sub - banks, or both of the memory layer 224. As an example, each channel of the memory layer 224 is configured to enable access, modification, and deletion of data within a separate memory bank or memory sub - bank of the memory layer 224. In some embodiments, memory 206 is configured to operate each channel as two or more virtual channels such that each virtual channel of the channel enables access, modification, and deletion of data within an individual portion of the memory bank or memory sub - bank associated with the channel. For example, memory 206 is configured to operate a 64 - bit channel as two 32 - bit virtual channels, each configured to enable access, modification, and deletion of data within an individual portion of the memory bank or memory sub - bank associated with the 64 - bit channel.
[0019] In an embodiment, the memory layer 224 includes one or more silicon memory dies stacked on top of each other. That is, each memory layer 224 includes a silicon memory die, and the dies of each memory layer 224 are arranged to form a stack of memory dies. According to an embodiment, one or more through silicon vias (TSVs) (not shown for clarity) penetrate one or more memory layers 224 of the stack so as to communicatively couple the one or more memory layers 224. To control the data stored in the memory layer 224 of the memory 206, the memory 206 includes a logic layer 226 including hardware and software configured to access, modify, and delete data in memory banks and memory sub-banks of the memory layer 224. For example, the logic layer 226 includes one or more memory controllers each configured to access, modify, and delete data in memory banks and memory sub-banks of one or more memory layers 224. In an embodiment, the logic layer 226 is communicatively coupled to each memory layer 224, for example, by one or more TSVs.
[0020] According to an embodiment, the logic layer 226 is configured to store data in one or more memory banks, memory sub-banks, or both of one or more memory layers 224. For example, the logic layer 226 includes one or more memory controllers, each associated with a memory bank or a memory sub-bank and configured to store data in the associated memory bank or sub-bank. In an embodiment, the logic layer 226 is configured to implement one or more ECCs (e.g., CRC codes (e.g., CRC-9, CRC-12, CRC-16), symbol-based codes, LRC codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes) for data written to one or more data banks or data sub-banks of the memory 206. The logic layer 226 is configured to implement ECC such that one or more check bits are generated for a predetermined amount of data written to one or more memory banks or memory sub-banks. For example, the logic layer 226 is configured to implement on-die ECC such that 32 check bits are generated for every 256 bits of stored data. As another example, the logic layer 226 is configured to implement on-die ECC such that one check bit is generated for every byte of stored data. The logic layer 226 is configured to store the check bits in the same memory bank or memory sub-bank assigned to hold, for example, the data being written (i.e., the data used to generate the check bits) in response to generating one or more check bits.
[0021] In an embodiment, the processing device 230 is configured to generate one or more fetch requests (e.g., cache line requests) that request data stored in one or more memory layers 224 of the memory 206. For example, the processing device 230 generates a fetch request that requests 128-bit data from the memory 206. The processing device 230 is further configured to split the fetch request into a set of fetch portions, each including one or more fetch portions that each identify at least a part of the data requested by the fetch request. For example, the processing device 230 is configured to split a 128-bit fetch request (i.e., a fetch request that identifies 128-bit requested data) into a set of four 32-bit fetch portions, each identifying a separate 32-bit portion of the data requested by the 128-bit fetch request. According to an embodiment, the processing device 230 is configured to split the fetch request into one or more fetch portions based on the channels of the memory 206. For example, the processing device 230 is configured to split a 128-bit fetch request into two 64-bit fetch portions in response to the memory 206 having two 64-bit channels per memory layer 224. As another example, the processing device 230 is configured to split a 128-bit fetch request into four 32-bit fetch portions in response to the memory 206 having two 64-bit channels operating in a pseudo-channel mode (e.g., having two 32-bit pseudo-channels per channel). After splitting the fetch request into one or more fetch portions, the processing device 230 is configured to send the fetch portions to the logical layer 226 of the memory 206. According to an embodiment, the processing device 230 is configured to generate a parity request that identifies each fetch portion and one or more operations to be performed on the data identified by the fetch portion in response to splitting the fetch request into one or more fetch portions. For example, the parity request identifies four fetch portions and an XOR operation to be performed on the data identified by each fetch portion. After generating the parity request, the processing device 230 is configured to send the parity request to the logical layer 226 along with the fetch portions.
[0022] In response to receiving the fetch portion and the parity request, the logic layer 226 is configured to read the data specified in the fetch portion from one or more memory layers 224 and determine whether one or more defects exist in the read data. The logic layer 226 checks the read data according to the ECC implemented by the logic layer 226 in order to determine whether one or more defects exist in the read data. For example, the logic layer 226 checks whether a defect exists in the data read from the data sub-bank of the memory layer 224 according to the on-die ECC implemented by the logic layer 226. In an embodiment, the logic layer 226 is configured to determine whether one or more defects exist in the read data by, for example, comparing one or more check bits associated with the read data with one or more values (e.g., a threshold value, a value, a character string), performing one or more operations on the check bits associated with the read data according to the implemented ECC, or both. For example, the logic layer 226 is configured to compare 16 check bits associated with the read data with one or more values to determine the existence of one or more defects in the read data. According to an embodiment, the logic layer 226 is configured to detect and correct defects limited to a predetermined number of bits in the read data. That is, the defects are limited to each part of the read data, each containing a predetermined number of bits. For example, the logic layer 226 is configured to detect and correct defects limited to a 16-bit portion of the read data. In response to determining that one or more defects exist in the read data, the logic layer 226 is configured to correct one or more of the detected errors based on one or more check bits associated with the requested data, the data stored in one or more memory layers 224 of the memory 206, one or more parity bits, or any combination thereof. According to an embodiment, when detecting and correcting defects limited to a predetermined number of bits in the read data, the logic layer 226 is configured to use only a part of the check bits associated with the read data so that one or more unused check bits remain.For example, the logic layer 226 is configured to detect and correct defects in the read data limited to 16 bits by using 16 out of a total of 32 check bits so that 16 unused check bits remain.
[0023] After the correction of the requested data, the logic layer 226 is configured to generate a set of fetch returns each including one or more fetch returns each containing the corrected read data associated with each received fetch portion (e.g., the corrected read data specified by each fetch portion). For example, in response to receiving four 32-bit fetch portions (i.e., fetch portions each specifying 32-bit requested data), the logic layer 226 is configured to generate a set of four 32-bit fetch returns each containing the corrected data specified by each 32-bit fetch portion. After generating the fetch returns, the logic layer 226 is configured to generate a parity fetch based on the received parity request. As an example, the logic layer 226 is configured to perform one or more operations specified by the received parity request on the corrected data within each generated fetch return. For example, in response to receiving a parity request specifying an XOR operation and generating four fetch returns each containing a part of the read data (e.g., D0, D1, D2, and D3), the read data within the generated fetch returns (e.g.,
Number
[0024] The processing device 230 includes an error correction engine 234 that includes hardware and software configured to implement one or more ECCs (e.g., CRC codes (e.g., CRC-9, CRC-12, CRC-16), symbol-based codes, LRC codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes) for fetch returns received by the processing device 230 and write requests (e.g., write portions) generated by the processing device 230. The error correction engine 234 is configured to check, based on the implemented ECC, whether one or more defects exist in each fetch return in response to the processing device 230 receiving one or more fetch returns and one or more parity fetches. For example, the error correction engine 234 is configured to check whether one or more defects exist in each fetch return based on the CRC-12 code. In an embodiment, the error correction engine 234 is configured to use one or more unused check bits of the fetch return to determine whether the fetch return includes one or more defects (e.g., uncorrected errors). For example, the error correction engine 234 compares one or more unused check bits of the fetch return with one or more values (e.g., threshold values, values, character strings), performs one or more operations on the unused check bits of the fetch return, or both, to determine whether the fetch return includes one or more defects. The error correction engine 234 is configured to erase and reconstruct the fetch return in response to the fetch return having one or more defects (e.g., uncorrected errors). The error correction engine 234 is configured to reconstruct the fetch return based on one or more other fetch returns, parity fetches, or both. For example, the error correction engine 234 is configured to reconstruct the fetch return based on the data in each other fetch return returned with the fetch return and the data in the parity fetch returned with the fetch return.By reconstructing the fetch return using parity fetch, the detection coverage and correction of system defects are improved, and the reliability of the system is enhanced. Therefore, the likelihood of in-field replacement unit events, termination of memory-using applications, system reboots, or any combination thereof occurring in the system is reduced. After reconstructing the data in the return fetch, each other fetch return returned together with the reconstructed return fetch and fetch return is stored in the cache 236. For example, the return fetch is provided to a data fabric (not shown for clarity) communicatively coupled to the processing device 230 and the cache 236.
[0025] According to an embodiment, the processing device 230 is configured to generate one or more write requests that identify data to be written to one or more portions of the memory 206. For example, the processing device 230 is configured to generate a write request that identifies data in the cache 236 to be written to one or more memory layers 224 of the memory 206. In an embodiment, the processing device 230 is configured to split a write request into one or more write portions, each of which identifies an individual portion of the data identified by the write request. For example, the processing device 230 is configured to split a write request into one or more write portions based on the size of the channels of the memory 206. As an example, the processing device 230 is configured to split a 128-bit write request into two 64-bit write requests in response to each memory layer 224 having two 64-bit channels. As another example, the processing device 230 is configured to split a 128-bit write request into four 32-bit write requests in response to each memory layer 224 having two channels operating in a pseudo-channel mode (e.g., each channel having two pseudo-channels). In an embodiment, the processing device 230 is configured to generate a write parity that includes data obtained by performing one or more operations on the data identified by the one or more write portions in response to generating the one or more write portions. For example, the write parity includes data obtained by performing an XOR operation on the data identified by four write portions split from a write request.
[0026] According to an embodiment, the error correction engine 234 is configured to generate one or more check values for each piece of data specified by each write portion and each write parity based on the ECC implemented by the error correction engine 234. For example, the error correction engine 234 is configured to generate a 12-bit check value for each piece of data specified by each write portion and write parity based on the CRC-12 code implemented by the error correction engine 234. After the generation of the check value, each check value is added to its respective write portion or write parity. That is, each write portion includes data specifying the data to be written to the memory 206 and the check value, and each write parity includes data obtained by performing one or more operations on the data to be written and the check value. The write portion and the write parity are transmitted to the logical layer 226 of the memory 206 in response to adding the check value to the write portion and the write parity. In response to receiving one or more write portions and write parities, the logical layer 226 is configured to store the data specified by the write portion, the data of the write parity, and the check value in one or more memory layers 224. For example, the logical layer 226 is configured to store the data specified by the write portion and the check value of the write portion in a memory sub-bank of the memory layer 224. By storing the write portion together with the check value and the write parity, the reliability of the data written to the memory is improved. For example, in response to detecting one or more errors in the data written to the memory, the data is reconstructed using other data and write parities stored in the memory, thereby reducing the likelihood of an in-field replacement unit event, termination of an application using the memory, system reboot, or any combination thereof occurring.
[0027] Next, referring to FIG. 3, a block diagram of a processing system 300 for host-level error correction and detection across one or more memory emulation channels is shown. The processing system 300 includes a CPU 102, a GPU 114, or a processing device similar to or the same as the processing device 230 communicatively coupled to a memory layer 324 similar to or the same as the memory layers 224 of the memory 106, 206. In an embodiment, the memory layer 324 is any one of a plurality of layers in a 3D stacked SDRAM and includes one or more memory banks, memory sub-banks, or both, each configured to store data. According to an embodiment, the memory layer 324 includes one or more channels 338 each configured to enable access to data within one or more memory banks, memory sub-banks, or both of the memory layer 324. In some embodiments, each channel 338 is configured to enable access to data within one or more individual memory banks of the memory layer 324. For example, in the exemplary embodiment shown in FIG. 3, the memory layer 324 includes channel 0 338-1 configured to enable access to a first memory bank (not shown for clarity) and channel 1 338-2 configured to enable access to a second memory bank (not shown for clarity). Each channel 338 has a width representing the maximum amount of data that can be simultaneously read from or written to the memory bank associated with the channel 338. For example, each channel 338 has a width of 64 bits.
[0028] In an embodiment, the memory layer 324 is configured to operate in a pseudo-channel mode. While in the pseudo-channel mode, each channel 338 of the memory layer 324 operates as two or more individual pseudo-channels 340. For example, in the exemplary embodiment shown in FIG. 3, while in the pseudo-channel mode, channel 0 338-1 operates as pseudo-channel 0 340-1 and pseudo-channel 1 340-2, and channel 1 338-2 operates as pseudo-channel 2 340-3 and pseudo-channel 3 340-4. Each pseudo-channel 340 is configured to enable access to data within at least a portion of the memory bank, memory sub-bank, or both, associated with its respective channel 338. For example, in the exemplary embodiment shown in FIG. 3, pseudo-channel 0 340-1 is configured to enable access to a first memory sub-bank 342-1, which is a memory sub-bank of the memory bank associated with channel 0 338-1, and pseudo-channel 1 340-2 is configured to enable access to a second different memory sub-bank 342-2, which is a second memory sub-bank of the memory bank associated with channel 0 338-1. Each pseudo-channel 340 has a width representing the maximum amount of data that can be simultaneously read from, or written to, the memory bank, memory sub-bank, or both, associated with the pseudo-channel 340. According to an embodiment, each pseudo-channel has a width equal to half the width of its associated channel. For example, in the exemplary embodiment shown in FIG. 3, channel 0 338-1 has a width of 64 bits, whereby the widths of pseudo-channel 0 340-1 and pseudo-channel 1 340-2 are 32 bits.
[0029] In an embodiment, the processing device 330 is communicatively coupled to the memory layer 324 via a channel 338. For example, the processing device 330 is coupled to the memory layer 324 via a channel 338 through a logic layer similar to or the same as the logic layer 226, including one or more memory controllers configured to control access to, modification of, and deletion of data stored in the memory layer 324. The processing device 330 is configured to generate one or more fetch requests (e.g., cache line fetches) each identifying data read from the memory layer 324 and transmit the fetch requests to the memory layer 324 via one or more channels 338 associated with the memory layer 324. According to an embodiment, the processing device 330 is configured to divide a fetch request into one or more fetch portions based on the width of the channel 338. For example, the processing device 330 is configured to divide a 128-bit fetch request into two 64-bit fetch portions based on a 64-bit channel width. In an embodiment, the processing device 330 is configured to divide a fetch request into a set of one or more fetch portions based on the width of a virtual channel 340 while the memory layer 324 is operating in virtual channel mode. For example, the processing device 330 is configured to divide a 128-bit fetch request (i.e., a fetch request requesting 128 bits of data) into a set of four 32-bit fetch portions based on each virtual channel 340 having a width of 32 bits. After dividing a fetch request into two or more fetch portions, the processing device 330 is configured to provide the fetch portions to a logic layer that controls access to the memory layer 324 via one or more channels 338, virtual channels 340, or both. That is, the processing device 330 uses one or more channels 338, virtual channels 340, or both to provide the fetch portions to the logic layer.For example, in the exemplary embodiment shown in FIG. 3, the processing device 330 splits a 128-bit fetch request (i.e., a fetch request that requests 128 bits of data) into a set of four 32-bit fetch portions, and is configured to transmit two of the fetch portions to the logical layer using the pseudo-channel 0 340-1 and to transmit another two of the fetch portions to the logical layer using the pseudo-channel 1 340-2.
[0030] The memory layer 324 is configured to read out data specified by the fetch portion from one or more memory banks, memory sub-banks, or both, associated with the channel 338, the pseudo-channel 340, or both, in response to receiving one or more fetch portions via one or more channels 338, pseudo-channels 340, or both. For example, in the exemplary embodiment shown in FIG. 3, the memory layer 324 is configured to read out data specified by the fetch portion from the memory sub-bank 342-3 (e.g., via the logic layer) in response to receiving the fetch portion via the pseudo-channel 2 340-3. According to an embodiment, the memory layer 324 is configured to check whether there are defects in the data specified by one or more received fetch portions and read out from one or more memory banks, memory sub-banks, or both. That is, the memory layer 324 is configured to check whether there are one or more defects (e.g., uncorrected errors) in the read data. For example, the memory layer 324 is configured to check the read data based on one or more ECCs implemented by the memory layer 324 (e.g., using the logic layer). In an embodiment, the memory layer 324 is configured to check the read data by comparing one or more check bits associated with the read data with one or more values (e.g., threshold values, values, character strings) (e.g., using the logic layer), performing one or more operations on the read data, or both. According to an embodiment, the check bits associated with the read data are stored in a memory band or a memory sub-bank that stores the data. In response to determining that there are one or more defects in the read data, the memory layer 324 is configured to correct the read data (e.g., using the logic layer) using, for example, other data stored in one or more memory banks, data stored in one or more memory sub-banks, one or more parity bits, or both.In an embodiment, the memory layer 324 is configured to send the read data as one or more fetch returns to the processing device 330 using one or more channels 338, pseudo-channels 340, or both. According to an embodiment, the memory layer 324 is configured to send the fetch return on the same channels and pseudo-channels that received the fetch part that identifies the data included in the fetch return. For example, in the exemplary embodiment shown in FIG. 3, the memory layer 324 is configured to read out the data identified by the fetch part as a fetch return and send it on the pseudo-channel 1 340-2 in response to receiving the fetch part that identifies the data to be read on the pseudo-channel 1 340-2. In an embodiment, the memory layer 324 is configured to send one or more unused check bits (e.g., check bits not used to detect or correct defects in the read data) to the processing device 330 along with the fetch return using one or more channels 338, pseudo-channels 340, or both.
[0031] According to an embodiment, the processing device 330 is configured to generate one or more parity requests each associated with one or more fetch requests, fetch portions, or both. For example, the processing device 330 is configured to generate parity requests associated with one or more fetch portions split from the same fetch request. Each parity request identifies, for example, one or more fetch portions and one or more operations to be performed on the data identified by the fetch portions. In an embodiment, the processing device 330 is configured to send the parity requests to the memory layer 324 using one or more channels 338, pseudo-channels 340, or both (e.g., via the logical layer). For example, the processing device 330 is configured to send the parity requests using the same pseudo-channel 340 through which the fetch portion associated with the parity request is sent to the memory layer 324. In response to receiving the parity requests, the memory layer 324 is configured to generate and return a parity fetch. The memory layer 324 is configured to perform one or more identified operations (e.g., using the logical layer) on the read data identified by the parity request to generate the parity fetch. For example, the memory layer 324 is configured to perform an XOR operation on the read data identified by four fetch portions identified by the parity request. In some embodiments, the memory layer 324 is configured to check the data for defects using one or more ECCs and, after correcting the defects, perform the operations identified by the parity request on the data identified by the parity request. In this way, the memory layer 324 is configured to generate the parity fetch such that the parity fetch includes the data obtained by performing one or more operations (e.g., using the logical layer) on the data identified by the parity request. For example, the memory layer 324 is configured to generate a parity fetch that includes the data obtained by performing an XOR operation on the data identified by four fetch portions associated with the parity fetch based on the parity request.The memory layer 324 is configured to send the parity fetch to the processing device 330 via one or more channels 338, pseudo-channels 340, or both (e.g., using the logic layer) in response to generating the parity fetch. For example, the memory layer 324 is configured to send the parity fetch to the processing device 330 using the pseudo-channel 340 that was used to send an associated parity request (e.g., the parity request used to generate the parity fetch) from the processing device 330 to the memory layer 324.
[0032] The processing device 330 is configured to check whether one or more defects (e.g., uncorrected errors) exist in the received fetch return in response to receiving one or more fetch returns on one or more channels or pseudo-channels. The processing device 330 includes an error correction engine 334 similar to or the same as the error correction engine 234, which includes hardware and software configured to check whether one or more defects exist in the fetch return by checking whether one or more defects exist in the fetch return based on one or more ECCs (e.g., CRC codes (e.g., CRC-9, CRC-12, CRC-16), symbol-based codes, LRC codes, checksum codes, parity check codes, Hamming codes, binary convolutional codes). According to an embodiment, the error correction engine 334 is configured to check whether one or more defects exist in the one or more fetch returns using unused check bits returned with the one or more fetch returns. For example, the error correction engine 334 is configured to check whether one or more defects exist in the data of the fetch return based on the CRC code and the unused check bits in response to the processing device 330 receiving the fetch return and the one or more check bits. In an embodiment, the error correction engine 334 is configured to check whether one or more defects exist in the fetch return by comparing the one or more unused check bits with one or more values (e.g., threshold values, values, character strings), performing one or more operations, or both. For example, the error correction engine 334 is configured to check whether one or more defects exist in the fetch return by performing one or more operations on the unused check bits according to the implemented ECC.
[0033] The error correction engine 334 is configured to erase and reconstruct a fetch return in response to determining that one or more defects exist in the fetch return. In an embodiment, the error correction engine 334 is configured to reconstruct the fetch return using data within one or more other fetch returns, one or more parity fetches, or both. For example, the error correction engine 334 is configured to reconstruct the fetch return using one or more associated fetch returns (e.g., fetch returns based on the same fetch request, associated fetch portions, or both) and parity fetches associated with the fetch return (e.g., parity fetches based on the same fetch request, associated fetch portions, or both). After reconstruction of the fetch return, the reconstructed fetch return and any associated fetch returns (e.g., fetch returns based on the same fetch request, associated fetch portions, or both) are sent to a cache similar to or the same as cache 236 included in or communicatively coupled to processing device 330. In an embodiment, the reconstructed fetch return and any associated fetch returns are sent to a data fabric communicatively coupled to processing device 330 and the cache.
[0034] According to an embodiment, the processing device 330 is configured to generate one or more write requests that identify data stored in the memory layer 324 within one or more caches communicatively coupled to the processing device 330. In an embodiment, the processing device 330 is configured to divide a write request into one or more write portions based on the width of the channel 338. For example, the processing device 330 is configured to divide a 128-bit write request into two 64-bit write portions based on a 64-bit channel width. The processing device 330 is configured to divide a write request into one or more write portions based on the width of the virtual channel 340 while the memory layer 324 is operating in virtual channel mode. For example, the processing device 330 is configured to divide a 128-bit write request into four 32-bit write portions based on each virtual channel 340 having a width of 32 bits. In an embodiment, the error correction engine 334 is configured to generate a check value for each write request according to one or more implemented ECCs. That is, the error correction engine 334 generates a check value for the data specified in the write portion according to one or more ECCs. For example, the error correction engine 334 is configured to generate a check value using a CRC-9 code. In response to generating a check value for a write portion, the check value is added to the write portion such that the write portion includes the data identifying the data to be written to the memory layer 324 (e.g., data stored in one or more caches communicatively coupled to the processing device 330) and the check value generated for the write portion. After generating a write portion for each write portion of a write request (e.g., all write portions divided from the write request), the processing device is configured to transmit the write request to a logical layer communicatively coupled to the memory layer 324 using one or more of the channels 338, virtual channels 340, or both. The logical layer is configured to store the data specified in the write request and the check value of the write request in one or more memory banks associated with the channel 338 or virtual channel 340 on which the write request was received.For example, the logic layer is configured to store the data specified by the write portion and the check value of the write portion in the memory sub-bank 342-4 in response to receiving the write portion on the pseudo-channel 3340-4.
[0035] Next, referring to FIG. 4, a flowchart of an exemplary process 400 for host-level error correction during a read operation is shown. Process 400 includes a processing device 430 similar to or the same as CPU 102, GPU 114, processing devices 230, 330, or any combination thereof, configured to generate a fetch request that identifies requested data. That is, processing device 430 generates a fetch request that identifies data read from a memory 406 similar to or the same as memories 106, 206. Processing device 430 is further configured to divide the fetch request into one or more fetch portions (e.g., 405-1, 405-2, 405-3, 405-4) each identifying an individual portion of the data identified by the fetch request. For example, processing device 430 is configured to divide a 128-bit fetch request (i.e., a fetch request that requests 128 bits of data) into a set of four 32-bit fetch portions 405. In the exemplary embodiment shown in FIG. 4, processing device 430 divides the fetch request into a set of four fetch portions (405-1, 405-2, 405-3, 405-4), but in other embodiments, processing device 430 may be configured to divide the fetch request into any number of fetch portions. In response to generating one or more fetch portions 405, processing device 430 is configured to transmit a first number of fetch portions 405 to memory 406 using a first virtual channel 440-1 and transmit a second number of fetch portions to memory 406 using a second virtual channel 440-2. The exemplary embodiment shown in FIG. 4 shows that two fetch portions (fetch portion 0 405-1 and fetch portion 1 405-2) are transmitted on virtual channel 0 440-1 and two fetch portions (fetch portion 2 402-3 and fetch portion 3 402-4) are transmitted on virtual channel 1 440-2, but in other embodiments, processing device 430 can transmit any number of fetch portions 405 on virtual channel 0 440-1 and any number of fetch portions 405 on virtual channel 1 440-2.
[0036] According to an embodiment, the processing device 430 is configured to generate a parity request 410 that identifies each fetch portion 405 associated with a fetch request and one or more operations to be performed on the data identified by the fetch portion, in response to splitting the fetch request into two or more fetch portions. For example, the parity request 410 identifies fetch portion 0 405-1, fetch portion 1 405-2, fetch portion 2 405-3, and fetch portion 3 405-4, and identifies the XOR operation to be performed on the data identified by fetch portion 0 405-1, fetch portion 1 405-2, fetch portion 2 405-3, and fetch portion 3 405-4 (e.g., the data of fetch portion 1
Number
Number
Number
[0037] In response to receiving one or more fetch portions 405, memory 406 is configured to read data specified by the fetch portion 405 from one or more memory banks, memory sub-banks, or both, associated with the pseudo-channel 440 that received the fetch portion. For example, in response to receiving fetch portion 0 405-1 on pseudo-channel 1 440-2, memory 406 is configured to read data specified by fetch portion 405-1 from one or more memory banks, memory sub-banks, or both, associated with pseudo-channel 1 440-2. According to an embodiment, memory 406 is configured to check whether one or more defects (e.g., uncorrected errors) exist in the read data in response to reading the data specified by the fetch portion 405. Memory 406 is configured to compare one or more check bits associated with the read data to one or more values (e.g., threshold values, values, strings), perform one or more operations on the check bits, or both, based on one or more ECCs (e.g., CRC codes) implemented by memory 406 to check whether one or more defects exist in the read data. In an embodiment, the check bits associated with the read data are stored in the same memory bank, memory sub-bank, or both as the read data. Memory 406 is configured to correct the defect by using other data stored in one or more memory banks, memory sub-banks, or both of memory 406, using one or more parity bits, or both, in response to detecting a defect in the read data by memory 406. According to an embodiment, memory 406 is configured to check whether one or more defects limited to a predetermined number of bits exist in the read data. That is, the defects are limited to each part of the read data, each of which includes a predetermined number of bits. Memory 406 is configured to use a plurality of check bits based on the predetermined number of bits to check whether one or more defects limited to the predetermined number of bits exist in the read data.For example, the memory 406 is configured to use, for example, 16 check bits to check whether there are defects limited to 16 bits in the read data. Therefore, depending on whether one or more defects limited to a predetermined number of bits are present in the read data, one or more of the check bits associated with the read data remain unused (for example, they are not used to check whether there are defects in the read data). That is, the memory 406 is configured to check whether one or more defects are present using a plurality of check bits such that one or more unused check bits remain, depending on whether one or more defects limited to a predetermined number of bits are present in the read data.
[0038] According to an embodiment, the memory 406 is configured to generate one or more fetch returns 415 for each received fetch portion 405. Each fetch return 415 includes the read data specified (i.e., requested) in the respective fetch portion 405. In an embodiment, the memory 406 is configured to transmit a plurality of fetch returns 415 on a first virtual channel 440-1 and transmit a second plurality of fetch returns 415 on a second virtual channel 440-2. For example, the memory 406 is configured to transmit each fetch return 415 on the same virtual channel 440 on which the associated fetch portion 405 (for example, the fetch portion that identifies the data included in the fetch return) is received. As an example, the exemplary embodiment shown in FIG. 4 shows a fetch return 0 415-1 and a fetch return 1 415-2 transmitted on the virtual channel 1 440-2 on which the fetch portion 0 405-1 and the fetch portion 405-2 are received, and a fetch return 2 415-3 and a fetch return 3 415-4 transmitted on the virtual channel 0 440-1 on which the fetch portion 2 405-3 and the fetch portion 3 405-4 are received.
[0039] Memory 406 is configured to generate a parity fetch 420 in response to generating one or more fetch returns 415. The parity fetch 420 includes data obtained by performing one or more operations specified by a parity request 410 on the read data included in the fetch return 415. For example, in the exemplary embodiment shown in FIG. 4, fetch return 0 415-1 includes a first set of read data (D0), fetch return 1 415-2 includes a second set of read data (D1), fetch return 2 415-3 includes a third set of read data (D2), and fetch return 3 415-4 includes a fourth set of read data (D3). In the exemplary embodiment, memory 406 is configured to generate parity fetch 420 by performing one or more operations specified by parity request 410, such as an XOR operation, on D0, D1, D2, and D3. That is, parity fetch 420 is the read data of fetch return 415 (e.g.,
Number
[0040] The processing device 430 is configured to check, according to one or more implemented ECCs, whether there are defects (e.g., uncorrected errors) in each fetch return 415 in response to receiving a fetch return 415, a parity fetch 420, and a check bit 425. For example, the processing device 430 is configured to check whether there are defects in each fetch return 415 by using one or more unused check bits 425 associated with the fetch return 415 (e.g., the unused check bits 425 returned with the fetch return 415, included in the fetch return 415, or both). The processing device 430 includes an error correction engine similar to or the same as the error correction engines 234 and 334, which is configured to check whether there are defects in each fetch return 415 by using one or more unused check bits 425. In an embodiment, the processing device 430 is configured to determine the presence of one or more errors in the fetch return 415 by comparing one or more of the unused check bits 425 with one or more values (e.g., a threshold value, a value, a character string), performing one or more operations on one or more of the check bits 425, or both. For example, the processing device checks whether there are one or more defects in the fetch return 415 by performing one or more operations on one or more of the check bits 425 according to the implemented ECC (e.g., CRC-12). The processing device 430 is configured to erase and reconstruct the fetch return 415 in response to detecting one or more defects in the fetch return 415. According to an embodiment, the processing device 430 is configured to reconstruct the fetch return 415 by using data from one or more other fetch returns, the parity fetch 420, or both. For example, in the exemplary embodiment shown in FIG. 4, the processing device 430 is configured to reconstruct the fetch return 2 415-3 by using data from the fetch return 0 415-1, the fetch return 1 415-2, the fetch return 3 415-4, and the parity fetch 420. In this way, any defects separated in a single fetch return are corrected.Correcting the defects in this way helps to improve the system reliability and reduce the likelihood of in-field replacement unit events, termination of applications using memory, system reboot, or any combination thereof.
[0041] Next, referring to FIG. 5, a flowchart of an exemplary process 500 for correcting one or more defects in a fetch return is shown. In an embodiment, the process 500 includes a CPU 102, a GPU 114, a processing device 230, 330, 430, or a processing device similar to or the same as any combination thereof, configured to perform a defect detection operation 502 on one or more received fetch returns 515 similar to or the same as the fetch return 415. The exemplary embodiment shown in FIG. 5 shows four fetch returns (fetch return 0 515-1, fetch return 1 515-2, fetch return 2 515-3, fetch return 3 515-4) received by the processing device, but in other embodiments, any number of fetch returns 515 may be received. According to an embodiment, each fetch return 515 includes at least a portion of a data read from a memory similar to or the same as the memories 106, 206, 306, 406. For example, in the exemplary embodiment of FIG. 5, fetch return 0 515-1 includes a first portion (D0) of the read data, fetch return 1 515-2 includes a second portion (D1) of the read data, fetch return 2 515-3 includes a third portion (D2) of the read data, and fetch return 3 515-4 includes a fourth portion (D3) of the read data. The defect detection operation 502 includes the processing device checking whether one or more defects exist in the data (e.g., D0, D1, D2, D3) of each fetch return 515 based on an ECC implemented by the processing device and one or more unused check bits received from the memory (e.g., check bits not used by the memory for checking whether one or more defects exist in D0, D1, D2, or D3). For example, the processing device checks whether one or more defects exist in D0, D1, D2, and D3 based on a CRC-12 code and one or more unused check bits.
[0042] In response to the processing device detecting a defect in one or more fetch returns 515 (e.g., a defect in the data of one or more fetch returns), the processing device performs a fetch return deletion operation 504 on the fetch return including the defect (e.g., a defective fetch). For example, in the exemplary embodiment of FIG. 5, in response to the processing device detecting a defect in D1, the processing device performs a fetch return deletion operation 504 on fetch return 1 515-2. The fetch return deletion operation 504 includes deleting the fetch return including one or more defects (e.g., defective fetches). In response to the processing device deleting one or more fetch returns (e.g., defective fetches), the processing device performs a fetch return reconstruction operation 506 including reconstructing one or more deleted fetch returns 515 based on one or more other fetch returns 515 associated with the deleted fetch returns (e.g., fetch returns 515 returned together with the deleted fetch return, fetch returns 515 in the same set as the deleted fetch, fetch returns 515 corresponding to the same fetch request), parity fetches 520 similar to or the same as the parity fetches 420 associated with the deleted fetch returns (e.g., returned together with the deleted fetch return and corresponding to the same fetch request), or both, to generate reconstructed fetch returns 530. For example, in the exemplary embodiment of FIG. 5, the fetch reconstruction operation 506 includes reconstructing fetch return 1 515-1 based on fetch return 0 515-1, fetch return 2 515-3, fetch return 3 515-4, and parity fetch 520. According to an embodiment, the parity fetch 520 includes data obtained by the memory performing one or more operations on the data of the fetch return before the data of the fetch return is received by the processing device. For example, in the exemplary embodiment of FIG. 5, the parity fetch 520 includes data (e.g.,
Number
[0043] Next, referring to FIG. 6, a flowchart of an exemplary method 600 for host-level error detection during a read operation is shown. In step 605 of method 600, a processing device similar to or the same as CPU 102, GPU 114, processing devices 230, 330, 430, or any combination thereof is configured to generate a fetch request that requests to read data from a memory similar to or the same as memories 106, 206, 406. According to an embodiment, the memory includes a 3D stacked SDRAM having two or more stacked memory layers each having two or more channels. In an embodiment, the processing device divides the fetch request into one or more fetch portions each identifying an individual portion of the data requested by the fetch request. For example, the processing device divides a 128-bit fetch request (i.e., a fetch request identifying 128 bits of requested data) into four 32-bit fetch portions each identifying a 32-bit portion of the data requested by the 128-bit fetch request. Further, in step 605, the processing device generates a parity request based on the fetch portions divided from the fetch request. The parity request indicates the data identified in each fetch portion and one or more operations performed on the identified data. For example, the parity request indicates the data identified in each fetch portion and an XOR operation. In response to generating the fetch portions and the parity request, the processing device transmits the fetch portions and the parity request to the memory using one or more channels, pseudo-channels, or both associated with the memory.
[0044] In step 610, in response to receiving one or more fetch portions, the memory reads data specified by the fetch portion from one or more memory banks, memory sub-banks, or both of the memory. For example, the memory reads data specified by the fetch portion from a memory sub-bank associated with the pseudo-channel that received the fetch portion. In response to reading the data specified by the fetch portion, the memory checks, based on one or more implemented ECCs, whether one or more defects (e.g., uncorrected errors) exist in the read data. For example, the memory detects one or more defects in the read data based on on-die ECC implemented by the memory. According to an embodiment, the memory detects one or more defects in the read data, for example, by comparing one or more check bits associated with the read data with one or more values (e.g., threshold values, values, character strings), performing one or more operations on one or more check bits according to one or more implemented ECCs, or both. In response to detecting one or more defects in the read data, the memory corrects the defects by using data stored in one or more memory banks, sub-banks, or both of the memory, using one or more parity bits stored in the memory, or both. In an embodiment, the memory is configured to detect one or more defects limited to a predetermined number of bits in the read data. For example, the memory is configured to detect defects limited to 16 bits in the read data. According to an embodiment, to detect one or more defects limited to a predetermined number of bits in the read data, the memory is configured to use a plurality of check bits based on the predetermined number of bits such that one or more unused check bits remain. For example, to determine one or more defects limited to 16 bits in the read data, the memory is configured to use 16 check bits such that 16 unused check bits remain.
[0045] After detecting and correcting any errors, the memory generates one or more fetch returns. Each fetch return includes the read data specified in the respective received fetch portion. That is, the memory generates a respective fetch return for each received fetch portion. Further, in step 610, the memory generates a parity fetch based on the received parity request and the fetch returns. The parity fetch includes data obtained by performing the operation specified by the parity request on the read data of the generated fetch returns. For example, the parity fetch includes data obtained by performing an XOR operation on the read data of each fetch return. After generating the fetch returns and the parity fetch, the memory transmits the fetch returns, one or more unused check bits (e.g., check bits not used for detecting or correcting errors in the read data), and the parity fetch via one or more channels, pseudo-channels, or both of the memory, and these are received at the processing device.
[0046] In step 615, the processing device determines whether one or more of the fetch returns contain a defect (e.g., an uncorrected error). In an embodiment, the processing device compares one or more received unused check bits with one or more values (e.g., a threshold value, a value, a character string), performs one or more operations on one or more unused check bits based on an ECC implemented by the processing device, or both, to determine whether one or more of the fetch returns contain a defect. As an example, the processing device determines whether there is a defect in the fetch return by performing one or more operations on one or more unused check bits according to the CRC-9 code. In response to there being no fetch return containing a defect, the system moves to step 630, and the processing device transmits the fetch return to a data fabric coupled to one or more caches similar to or the same as the processing device and cache 236. In response to one or more fetch returns containing a defect, the system proceeds to step 620. In step 620, the processing device erases the fetch return containing the defect. In step 625, the processing device reconstructs the erased fetch return. In an embodiment, the processing device reconstructs the erased fetch return using data in one or more other fetch returns and parity fetches. That is, the processing device reconstructs the erased fetch return based on the read data included in the other fetch returns and the data in the parity fetch. In this way, any defect (e.g., a driver defect, a bank defect) separated into a single fetch return is corrected, and the reliability of the system is improved. After reconstructing the fetch return containing the defect, the system proceeds to step 630, and the fetch return (e.g., including the reconstructed fetch return) is transmitted to a data fabric communicatively coupled to the processing device and one or more caches.
[0047] Next, referring to FIG. 7, a flowchart of an exemplary process 700 for host-level error correction for a write operation is shown. Process 700 includes a processing device 730 similar to or the same as CPU 102, GPU 114, processing devices 230, 330, 430, or any combination thereof, configured to generate a write request that identifies data to be written to one or more portions of a memory 706 similar to or the same as memories 106, 206, 406. The processing device 730 is further configured to divide the write request into one or more write portions (e.g., 705-1, 705-2, 705-3, 705-4), each identifying an individual portion of the data to be written as specified by the write request. For example, the processing device 730 is configured to divide a 128-bit write request into four 32-bit write portions 705. In the exemplary embodiment shown in FIG. 6, the processing device 730 divides the write request into four write portions (705-1, 705-2, 705-3, 705-4), but in other embodiments, the processing device 730 may be configured to divide the write request into any number of write portions. In response to generating one or more write portions 705, the processing device 730 generates a check value for each write portion 705 and configures the generated check value to be added to the write portion 705 such that the write portion 705 includes the data identifying the data to be written and the check value associated with the data to be written. In an embodiment, the processing device 730 is configured to generate the check value for the write portion 705 based on one or more ECCs implemented by the processing device 730, such as a CRC-12 code. After adding each respective check value to each write portion 705, the processing device 730 is configured to transmit a first number of write portions 705 to the memory 706 using a first virtual channel 740-1 and transmit a second number of write portions to the memory 706 using a second virtual channel 740-2.The exemplary embodiment shown in FIG. 7 shows that two write portions (write portion 0 705-1 and write portion 1 705-2) are transmitted on pseudo-channel 0 740-1 and two write portions (write portion 2 602-3 and write portion 3 602-4) are transmitted on pseudo-channel 1 740-2. However, in other embodiments, the processing device 730 can transmit any number of write portions 705 on pseudo-channel 0 740-1 and any number of write portions 705 on pseudo-channel 1 740-2.
[0048] According to an embodiment, the processing device 730 is further configured to generate a write parity 710 based on the data to be written specified in each write portion 705. The write parity 710 includes, for example, data obtained by performing one or more operations on the data to be written specified in each write portion 705. For example, the processing device 730 is based on the data to be written specified in write portion 0 705-1 (W0), write portion 1 705-2 (W1), write portion 2 705-3 (W2), and write portion 3 705-3 (W3), W0, W1, W2, and W3 (for example,
Number
[0049] As disclosed herein, in some embodiments, a method, in response to receiving check bits and a plurality of fetch returns at a processing device, determines, based on the check bits, whether any of the plurality of fetch returns includes a defect; reconstructs a fetch return based on a parity fetch associated with the plurality of fetch returns and one or more other fetch returns of the plurality of fetch returns in response to determining that the fetch return includes a defect; and transmits the reconstructed fetch to one or more caches communicatively coupled to the processing device. In one aspect, the method includes transmitting a plurality of fetch portions, each identifying requested data, to a memory; and transmitting a parity request identifying the plurality of fetch portions and one or more operations to the memory. In another aspect, determining that any of the plurality of fetch returns includes a defect includes performing an operation on the check bits according to an error correction code implemented by the processing device.
[0050] In one aspect, the error correction code includes a cyclic redundancy check code. In another aspect, the method includes checking, in a memory communicatively coupled to the processing device, whether one or more defects exist in requested data identified by the plurality of fetch returns based on an error correction code implemented by the memory. In yet another aspect, the check bits include check bits not used to check in the memory whether one or more defects exist in requested data identified by the plurality of fetch returns. In another aspect, the error correction code implemented by the memory includes an on-die error correction code. In yet another aspect, a first number of fetch returns are received at the processing device on a first virtual channel of the memory, and a second number of fetch returns are received at the processing device on a second virtual channel of the memory. In yet another aspect, the memory includes a three-dimensional stacked synchronous dynamic random access memory.
[0051] In some embodiments, the system includes a memory having two or more virtual channels, and a processing device coupled to the memory by the two or more virtual channels, wherein in response to receiving check bits and a plurality of fetch returns from the memory, based on the check bits, it determines whether any of the plurality of fetch returns includes a defect, and in response to determining that a fetch return includes a defect, it reconstructs the fetch return based on a parity fetch associated with the plurality of fetch returns and one or more other fetch returns of the plurality of fetch returns, and transmits the reconstructed fetch to a data fabric communicatively coupled to the processing device. In one aspect, the processing device is further configured to transmit to the memory a plurality of fetch portions each specifying requested data, and to transmit to the memory a parity request specifying the plurality of fetch portions and one or more operations.
[0052] In one aspect, the memory is further configured to check whether one or more defects exist in the requested data specified by the plurality of fetch returns based on an error correction code implemented by the memory. In another aspect, the memory is further configured to generate check bits based on an error correction code implemented by the memory. In yet another aspect, the error correction code implemented by the memory includes an on-die error correction code. In yet another aspect, a first number of fetch returns are received at the processing device on a first virtual channel of the memory, and a second number of fetch returns are received at the processing device on a second virtual channel of the memory. In one aspect, the memory includes a three-dimensional stacked synchronous dynamic random access memory.
[0053] In some embodiments, the method comprises, in a processing device, generating a check value for each of a plurality of write portions based on an error correction code implemented by the processing device, wherein each of the plurality of write portions identifies a respective portion of data to be written to the memory, determining a write parity based on a respective portion of data to be written to the memory and an operation for each write portion of the plurality of write portions, and transmitting the write portion, the check value, and the write parity to the memory. In one aspect, the memory comprises a three-dimensional stacked synchronous dynamic random access memory. In another aspect, a first number of write portions are transmitted to the memory on a first virtual channel of the memory and a second number of write portions are transmitted to the memory on a second virtual channel of the memory. In yet another aspect, the error correction code comprises a cyclic redundancy check code.
[0054] In some embodiments, the above-described apparatus and techniques are implemented in a system that includes one or more integrated circuit (IC) devices (also referred to as integrated circuit packages or microchips), such as a processing system for host error correction for the read and write operations described above with reference to FIGS. 1-6. Electronic design automation (EDA) and computer aided design (CAD) software tools can be used for the design and manufacture of these IC devices. These design tools are typically represented as one or more software programs. The one or more software programs operate a computer system to operate on code representing the circuits of the one or more IC devices to perform at least a portion of the process for designing or adapting a manufacturing system for fabricating the circuits. The code can include instructions, data, or a combination of instructions and data. Software instructions representing design tools or manufacturing tools are typically stored on a computer-readable storage medium accessible to a computing system. Similarly, code representing one or more stages of the design or manufacture of an IC device is stored on and accessed from the same or a different computer-readable storage medium.
[0055] A computer-readable storage medium includes any non-transitory storage medium or combination of non-transitory storage media that is accessible by a computer system during use to provide instructions and / or data to the computer system. Such storage media include, but are not limited to, optical media (e.g., compact discs (CDs), digital versatile discs (DVDs), Blu-ray (registered trademark) discs), magnetic media (e.g., floppy (registered trademark) discs, magnetic tapes, magnetic hard drives), volatile memory (e.g., random access memory (RAM) or cache), non-volatile memory (e.g., read-only memory (ROM) or flash memory), or microelectromechanical systems (MEMS)-based storage media. A computer-readable storage medium (e.g., system RAM or ROM) may be incorporated within a computing system, a computer-readable storage medium (e.g., a magnetic hard drive) may be fixedly attached to a computing system, a computer-readable storage medium (e.g., an optical disc or a universal serial bus (USB)-based flash memory) may be removably attached to a computing system, or a computer-readable storage medium (e.g., network-accessible storage (NAS)) may be coupled to a computer system via a wired or wireless network.
[0056] In some embodiments, certain aspects of the above-described techniques are implemented by one or more processors of a processing system executing software. The software includes one or more sets of executable instructions stored on a non-transitory computer-readable storage medium or otherwise tangibly embodied. The software may include instructions and certain data, and when the instructions and certain data are executed by one or more processors, the one or more processors are operated to execute one or more aspects of the above-described techniques. The non-transitory computer-readable storage medium may include, for example, magnetic or optical disk storage devices, solid state storage devices such as flash memory, cache, random access memory (RAM), or other non-volatile memory device(s), etc. The executable instructions stored on the non-transitory computer-readable storage medium may be implemented in source code, assembly language code, object code, or other instruction formats interpretable or otherwise executable by one or more processors.
[0057] In addition to the above, it should be noted that not all activities or elements described in the general description are required, some activities or parts of a particular device may not be required, one or more additional activities may be performed, and one or more additional elements may be included. Further, the order in which activities are listed is not necessarily the order in which they are performed. Also, the concepts have been described with reference to particular embodiments. However, those skilled in the art will understand that various changes and modifications can be made without departing from the scope of the invention as set forth in the claims. Accordingly, the specification and drawings are to be considered in an illustrative rather than a limiting sense, and all such modifications are intended to be included within the scope of the invention.
[0058] The preposition "or" used in the context of "at least one of A, B, or C" is used herein to mean "inclusive disjunction". That is, in the above and similar contexts, "or" is used to mean "at least one of, or any combination of". For example, "at least one of A, B, and C" is used to mean "at least one of A, B, C, or any combination thereof".
[0059] Advantages, other advantages and solutions to problems have been described above with respect to specific embodiments. However, advantages, solutions to problems, and features that may give rise to or manifest any advantages, benefits or solutions are not to be construed as important, essential or indispensable features of any or all of the claims. Furthermore, the disclosed invention can be modified and implemented in different but similar ways that will be apparent to those skilled in the art having the benefit of the teachings herein, so the specific embodiments described above are merely illustrative. There is no limitation on the details of the construction or design shown herein other than as set forth in the appended claims. Therefore, it is clear that the specific embodiments described above may be changed or modified, and all such variations are considered to be within the scope of the disclosed invention. Accordingly, the protection sought herein is set forth in the appended claims.
Claims
Claim 1 A method comprising: in a processing device, in response to receiving check bits and a plurality of fetch returns, determining, based on the check bits, whether any of the plurality of fetch returns includes a defect; in response to determining that a fetch return includes a defect, reconstructing the fetch return based on a parity fetch associated with the plurality of fetch returns and one or more other fetch returns of the plurality of fetch returns; transmitting the reconstructed fetch return to one or more caches communicatively coupled to the processing device. The method. Claim 2 further comprising transmitting to a memory a plurality of fetch portions each specifying requested data; transmitting to the memory a parity request specifying the plurality of fetch portions and one or more operations. The method of claim 1. Claim 3 Determining that any of the plurality of fetch returns includes a defect comprises performing an operation on the check bits according to an error correction code implemented by the processing device. The method of claim 1. Claim 4 The error correction code includes a cyclic redundancy check code. The method of claim 3. Claim 5 further comprising, in a memory communicatively coupled to the processing device, checking, based on an error correction code implemented by the memory, whether one or more defects exist in requested data specified in the plurality of fetch returns. The method of claim 1. Claim 6 The check bits include check bits not used to check whether one or more defects exist in requested data specified in the plurality of fetch returns in the memory. The method of claim 5. Claim 7 The error correction code implemented by the memory includes an on-die error correction code. The method of claim 5. Claim 8 A first number of fetch returns are received in the processing device on a first virtual channel of a memory, and a second number of fetch returns are received in the processing device on a second virtual channel of the memory. The method of claim 1. Claim 9 The memory includes a three-dimensional stacked synchronous dynamic random access memory. The method of claim 8.
10. A system comprising: a memory having two or more virtual channels; and a processing device coupled to the memory by the two or more virtual channels, wherein the processing device is configured to: receive check bits and a plurality of fetch returns from the memory, and based on the check bits, determine whether any of the plurality of fetch returns includes an error; in response to determining that a fetch return includes an error, reconstruct the fetch return based on a parity fetch associated with the plurality of fetch returns and one or more other fetch returns of the plurality of fetch returns; and transmit the reconstructed fetch return to a data fabric communicatively coupled to the processing device. The system is configured to perform the above operations. System.
11. The processing device is further configured to: transmit to the memory a plurality of fetch portions each specifying requested data; and transmit to the memory a parity request specifying the plurality of fetch portions and one or more operations. The system of claim 10 is configured to perform the above operations. System according to claim 10.
12. The memory is configured to check whether one or more errors exist in the requested data specified in the plurality of fetch returns based on an error correction code implemented by the memory. System according to claim 10. The system of claim 10.
13. The memory is configured to generate the check bits based on an error correction code implemented by the memory. System according to claim 10. The system of claim 10.
14. The error correction code implemented by the memory includes an on-die error correction code. System according to claim 13.
15. A first number of fetch returns are received at the processing device on a first virtual channel of the memory, and a second number of fetch returns are received at the processing device on a second virtual channel of the memory. System according to claim 10.
16. The memory includes a three-dimensional stacked synchronous dynamic random access memory. System according to claim 10.
17. A method comprising: In a processing device, generating a check value for each of a plurality of write portions based on an error correction code implemented by the processing device, wherein each of the plurality of write portions identifies a respective portion of data written to a memory determining a write parity based on each respective portion of data written to the memory for each of the plurality of write portions and an operation including transmitting the write portion, the check value, and the write parity to the memory A method
18. The memory includes a three-dimensional stacked synchronous dynamic random access memory The method of claim 17
19. A first number of write portions are transmitted to the memory on a first pseudo-channel of the memory, and a second number of write portions are transmitted to the memory on a second pseudo-channel of the memory The method of claim 17
20. The error correction code includes a cyclic redundancy check code The method of claim 17
Citation Information
Patent Citations
Soft decoder for generalized product codes
US20170279468A1
Error correction in a cache memory
US7302619B1