Storage-class memory, data processing method, and processor system
By adopting a two-level memory error correction mechanism and memory particle division in storage-level memory, the reliability and access delay problems of SCM when compatible with DDR protocol are solved, and high reliability and low latency storage performance are achieved, improving universality.
Patent Information
- Application Number
- CN202411165652.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-13
- Filing Date
- 2022-01-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-01-30
AI Technical Summary
When existing storage-level memory (SCM) is compatible with DDR protocol, its access latency is uncertain and has low versatility, making it difficult to improve reliability.
A two-level memory error correction mechanism is adopted, and the memory particles are divided into the first and second categories, which are used to store data and run/re-error error correction codes respectively. Combined with the memory bit width requirements of the DDR protocol, high reliability and low access delay are achieved.
High reliability and low access latency are achieved in storage-level memory, while improving compatibility with different processors, meeting the requirements of the DDR protocol.
Smart Images

Figure CN119248182B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202111513788.2 and the application title "Method for Transmitting Data and Storage Device", which was filed with the National Intellectual Property Administration on December 13, 2021. The entire content thereof is incorporated herein by reference.
[0002] This application is a divisional application. The application number of the original application is 202210114650.3, and the original application date is January 30, 2022. The entire content of the original application is incorporated herein by reference. Technical Field
[0003] This application relates to the field of computers, and in particular, to a storage-class memory, a data processing method, and a processor system. Background Art
[0004] Storage class memory (SCM) is a new type of storage medium, and its storage capacity, access speed, and cost are between those of main memory (such as dynamic random access memory (DRAM)) and hard disks (such as NAND flash). For example, compared with DRAM, SCM has the ability of persistence, and data will not be lost after power failure, and the storage capacity is larger. Compared with NAND flash, SCM has a faster access speed.
[0005] Currently, SCM realizes memory error correction based on a built-in controller. Since the storage medium in the controller may cache data, the access latency of SCM is uncertain, and it cannot be compatible with the Double Data Rate (DDR) protocol, resulting in low generality of SCM. Therefore, how to improve the reliability of SCM and reduce the access latency of SCM as much as possible while SCM is compatible with the DDR protocol is an urgent problem to be solved currently. Summary of the Invention
[0006] This application provides a storage-class memory, a data processing method, and a processor system, thereby improving the reliability of SCM and reducing the access latency of SCM as much as possible while SCM is compatible with the DDR protocol.
[0007] In a first aspect, a storage-class memory is provided. The multiple memory dies included in the storage-class memory are divided into at least one group; each group in the at least one group includes a first type of memory die and a second type of memory die; the second type of memory die is used to store running error correction codes, and the first type of memory die is used to store data and redundant error correction codes. The running error correction codes are used to perform primary memory error correction on the data stored in the first type of memory die in the same group. The redundant error correction codes are used to perform secondary memory error correction on the data stored in the first type of memory die when the primary memory error correction fails.
[0008] In this way, while ensuring the high reliability of the data stored in the storage-class memory through a two-level memory error correction mechanism, the access latency is made as low as possible. Compared with the controller built in the SCM using a more complex error correction algorithm for memory error correction, which results in a longer access latency, in the embodiment of the present application, through the two-level memory error correction mechanism, when the error rate of the storage-class memory is relatively low, primary memory error correction can effectively shorten the access latency while ensuring the high reliability of the data stored in the storage-class memory; when the error rate of the storage-class memory is relatively high, secondary memory error correction can also ensure the high reliability of the data stored in the storage-class memory, and the overall access latency is shortened in combination with primary memory error correction. In addition, by arranging the memory dies of the storage-class memory, the bit width of the storage-class memory meets the memory bit width indicated by the Double Data Rate (DDR) protocol, so as to implement the compatibility of the storage-class memory with the DDR protocol, enabling the storage-class memory to be connected to more types of processors, and improving the versatility of the storage-class memory.
[0009] Among them, the number of the first type of memory die and the number of the second type of memory die in each group are determined according to the bit width of the first type of memory die, the bit width of the second type of memory die, and the memory bit width indicated by the DDR protocol.
[0010] Exemplarily, the memory bit width indicated by the DDR protocol is 80 bits (bit), the bit widths of both the first type of memory die and the second type of memory die are 8 bits, and the 10 memory dies included in the storage-class memory are divided into 2 groups, and each group includes 4 first type of memory dies and 1 second type of memory die; the sum of the bit widths of the 4 first type of memory dies and 1 second type of memory die included in each group is 40 bits, and the unit access data volume per bit is 16 bytes (byte) or 32 bytes.
[0011] In a possible implementation manner, the first type of memory die includes M first type of units and N second type of units. The first type of units are used to store data. The second type of units are used to store redundant error correction codes. Thus, the memory controller or firmware (FW) can perform online switching or upgrade of the number of the first type of units and the number of the second type of units according to the usage scenario, so as to improve the flexibility of adapting to the reliability requirements of different systems.
[0012] In another possible implementation, the unit access data volume of the first type of unit and the unit access data volume of the second type of unit are both the unit access data volume of one bit in the bit width of the first type of memory die. Thus, the unit access data volume of the storage-class memory meets the cache line length of the processor connected thereto, thereby enhancing the versatility of the storage-class memory.
[0013] In another possible implementation, the first type of memory die further includes a third type of unit. The third type of unit is used to implement the function of bad block management for the units of the first type of memory die. Thus, the space of the storage-class memory is saved, and the storage-class memory is more easily compatible with memories that meet the DDR protocol (such as DRAM), enhancing the versatility of the storage-class memory.
[0014] In a second aspect, a data processing method is provided. The method is executed by a controller, and the controller is connected to the storage-class memory and the processor in the first aspect or any possible storage-class memory in the first aspect; the method includes: the controller obtains the data stored in the first type of memory die in the same group in the storage-class memory for primary memory error correction; when the primary memory error correction fails, the controller obtains each first type of memory die to perform secondary memory error correction on the data stored in the first type of memory die.
[0015] Exemplarily, the controller performs primary memory error correction on the data stored in the first type of memory die in the same group in the storage-class memory, including: the controller uses Hamming code or block code to perform primary memory error correction on the data stored in the first type of memory die in the same group in the storage-class memory. The controller performs primary memory error correction on a single bit in the range of 128 bytes to 512 bytes of the data stored in the first type of memory die in the same group in the storage-class memory.
[0016] Exemplarily, the controller performs secondary memory error correction on the data stored in the first type of memory die, including: the controller uses block code or low-density parity-check code to perform secondary memory error correction on the data stored in the first type of memory die. The controller performs secondary memory error correction on a hundred bits in the range of 2048 bytes to 4096 bytes of the data stored in the first type of memory die.
[0017] In a possible implementation, the method further includes: the controller converts the instructions of the processor into the instructions of the storage-class memory, and converts the instructions of the storage-class memory into the instructions of the processor. Thus, it is convenient for the processor to perform read and write operations on the storage-class memory.
[0018] In a third aspect, a processor system is provided, which includes a controller and the above-mentioned storage-class memory and processor according to any possible implementation of the first aspect. The controller is respectively connected to the processor and at least one storage-class memory, and the controller is configured to execute the operation steps of the method according to any possible implementation of the second aspect or the second aspect.
[0019] Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. Description of the Drawings
[0020] Figure 1 It is a schematic structural diagram of a storage-class memory provided by an embodiment of the present application;
[0021] Figure 2 It is a schematic physical partitioning diagram of a storage-class memory provided by an embodiment of the present application;
[0022] Figure 3 It is a schematic structural diagram of a first type of memory die provided by an embodiment of the present application;
[0023] Figure 4 It is a schematic diagram of generating a running error correction code and a double error correction code provided by an embodiment of the present application;
[0024] Figure 5 It is a schematic structural diagram of another storage-class memory provided by an embodiment of the present application;
[0025] Figure 6 It is a schematic structural diagram of yet another storage-class memory provided by an embodiment of the present application;
[0026] Figure 7 It is a schematic structural diagram of a processor system provided by an embodiment of the present application;
[0027] Figure 8 It is a schematic structural diagram of another processor system provided by an embodiment of the present application;
[0028] Figure 9 It is a schematic diagram of a data processing method provided by an embodiment of the present application. Detailed Embodiments
[0029] In digital circuits, the smallest data unit is a bit, which is also the smallest data unit of memory (or main memory). The value of a single bit is either "0" or "1", and eight consecutive bits form a byte. In machine language, a byte represents a letter or a number. Interference from electric fields, magnetic fields, or even cosmic rays can cause the value of a single bit stored in memory to change. If the value of a single bit in a byte that is crucial for system operation changes, it may lead to system errors, resulting in downtime or other failures.
[0030] Error Correcting Code (ECC) is a technology that can detect and correct errors. Error-correcting code memory (ECC memory) is a type of memory that applies ECC technology, i.e., memory that can detect and correct errors. ECC memory is widely used in servers and graphics workstations to improve the stability and reliability of computer operation.
[0031] The embodiments of this application provide a storage-class memory, especially a storage-class memory with high reliability and low latency. That is, by using a two-level memory error correction mechanism, the high reliability of the data stored in the storage-class memory is ensured while minimizing the access latency. In addition, by arranging the memory dies of the storage-class memory, the bit width of the storage-class memory meets the memory bit width specified by the Double Data Rate (DDR) protocol, thereby enabling the storage-class memory to be compatible with the DDR protocol and allowing it to connect to more types of processors, improving the versatility of the storage-class memory.
[0032] In the embodiments of this application, the memory die can refer to Phase Change Memory (PCM). Phase Change Memory is a storage device that stores data by utilizing the difference in conductivity exhibited when a special material (such as chalcogenide) transforms between the crystalline state and the amorphous state. The memory bit width refers to the amount of data that the memory can transfer at one time. The larger the bit width, the greater the amount of data that can be transferred in one operation. The memory bit width can also be referred to as the data bit width or simply the bit width.
[0033] The multiple memory dies included in the storage-class memory are divided into at least one group. Each group in the at least one group includes a first type of memory die and a second type of memory die. It can be understood that the at least one group refers to at least one channel of DDR. The number of the first type of memory dies and the number of the second type of memory dies in each group are determined based on the bit width of the first type of memory die, the bit width of the second type of memory die, and the memory bit width specified by the DDR protocol, so that the bit width of the storage-class memory meets the memory bit width specified by the DDR protocol.
[0034] The second type of memory particles is used to store running error correction codes. The running error correction codes are used for performing primary memory error correction on the data stored in the first type of memory particles in the same group. It should be understood that the running error correction codes can be generated based on the data stored in the first type of memory particles belonging to the same group. Primary memory error correction refers to performing memory error correction on the data stored in the first type of memory particles belonging to the same group. For example, the running error correction codes stored in the second type of memory particles included in Group 1 are generated based on the data stored in the first type of memory particles included in Group 1. The running error correction codes stored in the second type of memory particles included in Group 1 are used for performing primary memory error correction on the data stored in the first type of memory particles included in Group 1.
[0035] The first type of memory particles is used to store data and heavy error correction codes. The heavy error correction codes are used for performing secondary memory error correction on the data stored in the first type of memory particles when the primary memory error correction fails. It should be understood that the heavy error correction codes can be generated based on the data stored in the first type of memory particles. Secondary memory error correction refers to performing memory error correction on the data stored in each first type of memory particle. For example, the heavy error correction codes stored in the first type of memory particles included in Group 1 are generated based on the data stored in the same first type of memory particle. The heavy error correction codes stored in the first type of memory particles included in Group 1 are used for performing secondary memory error correction on the data stored in the same first type of memory particle.
[0036] In each embodiment of the present application, the running error correction codes can also be referred to as first-level correction codes, and the heavy error correction codes can also be referred to as second-level correction codes.
[0037] In some other embodiments, the first type of memory particles is also used to implement the function of bad block management for the first type of memory particles. Compared with setting memory particles for implementing bad block management on the storage-class memory, setting the medium for implementing bad block management within the first type of memory particles saves the space of the storage-class memory, makes the storage-class memory more easily compatible with memories that meet the DDR protocol (such as DRAM), and improves the versatility of the storage-class memory.
[0038] The storage-class memory may also include other chips, such as a clock chip, etc.
[0039] The storage-class memory provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0040] Taking the memory bit width of the storage-class memory compatible with DDR5 as 80 bits (bit) as an example. Figure 1A structural schematic diagram of a storage-class memory provided by an embodiment of this application. The storage-class memory 100 includes 10 memory dies 110, and the bit width of each memory die 110 is 8 bits. The 10 memory dies 110 are divided into 2 groups, each group includes 5 memory dies 110, and the sum of the bit widths of the 5 memory dies 110 included in each group is 40 bits. The sum of the bit widths of the 5 memory dies 110 included in the 2 groups is 80 bits, so the storage-class memory 100 is compatible with the memory bit width of DDR5. Understandably, the 2 groups obtained by dividing the 10 memory dies 110 refer to the 2 sub-channels of DDR5, which is compatible with the dual-channel standard of DDR5. The bit width of each sub-channel is 40 bits.
[0041] Among them, each group contains 4 first-type memory dies 111 and 1 second-type memory die 112. The unit access data volume of each bit in the bit width of the memory die 110 is 16 bytes or 32 bytes, that is, after a certain time delay for each access to the storage-class memory, 16B or 32B of data can be obtained. If 4 first-type memory dies 111 are accessed at one time, 64B of data can be read or written. If 8 first-type memory dies 111 are accessed at one time, 128B can be read or written. 64B or 128B is equal to the cache line length of the processor, thereby improving the versatility of the storage-class memory.
[0042] The second-type memory die 112 is used to store the running error correction code. The first-type memory die 111 is used to store data and the heavy error correction code. The length of the running error correction code word is determined by the first-type memory dies 111 and the second-type memory die 112 in the group. For example, if 4 first-type memory dies 111 are accessed at one time and 64B of data can be read or written, the length of the running error correction code word is 64B + 16B; if 8 first-type memory dies 111 are accessed at one time and 128B of data can be read or written, the length of the running error correction code word is 128B + 32B. The heavy error correction code can set the length of the heavy error correction code word according to the reliability requirements.
[0043] Understandably, as Figure 2 shown, the physical division of the storage-class memory 100 from large to small can be divided into rank 120, memory die (chip) 110, bank 130, and storage cell 140. The rank 120 includes all the memory dies 110 installed on the storage-class memory 100. Each memory die 110 includes multiple banks 130. Each bank 130 includes storage cells 140. Each storage cell 140 is determined by row and column.
[0044] In some embodiments, the medium for the first type of memory die 111 to store the heavy error correction code can be some of the multiple cells 130 included in the first type of memory die 111. For example, the first type of memory die 111 includes M first type of cells 131 and N second type of cells 132. The first type of cells 131 are used to store data. The second type of cells 132 are used to store the heavy error correction code. In practical applications, the number M of the first type of cells 131 and the number N of the second type of cells 132 can be flexibly configured according to the reliability requirements. If the reliability requirement is low and the amount of data for secondary memory error correction is small, fewer second type of cells 132 and more first type of cells 131 can be configured; if the reliability requirement is high and the amount of data for secondary memory error correction is large, more second type of cells 132 can be configured. The memory controller or firmware (FW) can perform online switching or upgrade of the number M of the first type of cells 131 and the number N of the second type of cells 132 according to the usage scenario, so as to improve the flexibility in adapting to different system reliability requirements.
[0045] Exemplarily, as Figure 3 shown, the first type of memory die 111 includes 16 independently operable cells 130. The 1st cell 130 to the 14th cell 130 serve as the first type of cells 131 and are used to store data. The 15th cell 130 serves as the second type of cell 132 and is used for secondary memory error correction. Optionally, the 16th cell 130 in the first type of memory die 111 serves as the third type of cell 133 and is used to implement the function of bad block management for the cells 130 of the first type of memory die 111. When a certain cell among the 1st cell 130 to the 15th cell 130 fails, the 16th cell 130 can be used to replace the failed cell 130, thereby preventing the storage-class memory 100 from failing. For example, when some addresses of a certain cell among the 1st cell 130 to the 15th cell 130 are damaged, the controller remaps the corresponding 16B address to the reserved 16th cell 130, and the bad block management function can be achieved.
[0046] One access to a cell 130 can read or write 16B or 32B of data. Assuming that one access to a cell 130 can read or write 16B of data, the length of the heavy error correction codeword can be 224B (16 * 14) + 16B. Since the heavy error correction codeword is relatively long, the error correction ability of secondary memory error correction is also relatively strong. This is only one of the ways to construct the heavy error correction codeword. In some other embodiments, the data stored in some of the first type of cells 131 among the multiple first type of cells 131 can also be used to construct the heavy error correction codeword. During secondary memory error correction, the cells 130 included in the first type of memory die 111 are cyclically checked for errors and corrected.
[0047] Exemplarily, as Figure 4As shown, the data of the primary memory error correction and the data of the secondary memory error correction can form a two-dimensional matrix. The x dimension represents four first-class memory particles 111 and one second-class memory particle 112, and the y dimension represents the address of the minimum access unit 16B of the memory particle. The four first-class memory particles 111 store data and heavy error correction codes, and the one second-class memory particle 112 stores running error correction codes. Each row constitutes a 64B+16B running error correction code codeword, and the entire two-dimensional matrix constitutes a 512B+208B heavy error correction code codeword. When the running error correction code decoding of any row fails, the system reads out the two-dimensional matrix data for secondary memory error correction, that is, performs secondary memory error correction on the data of each column. The system can perform horizontal decoding and vertical decoding in parallel, effectively reducing the error correction rate of the system.
[0048] In some embodiments, the data of the first type memory cell 111 in each column of the two-dimensional matrix may be data of any eight cells 130 from the first cell 130 to the fourteenth cell 130, and the eight cells 130 may be continuous. For example, the data of the first type memory cell 111 in the first column may be data from the first cell 130 to the eighth cell 130, and the heavy error correction code stored in the first type memory cell 111 in the first column is generated based on the data from the first cell 130 to the eighth cell 130. The heavy error correction code stored in the first type memory cell 111 in the first column performs memory error correction on the data from the first cell 130 to the eighth cell 130.
[0049] For another example, the data of the first type memory cell 111 in the second column may be the data of the third unit 130 to the tenth unit 130, and the heavy error correction code stored in the first type memory cell 111 in the second column is generated based on the data of the third unit 130 to the tenth unit 130. The heavy error correction code stored in the first type memory cell 111 in the second column performs memory error correction on the data of the third unit 130 to the tenth unit 130.
[0050] In addition, the first type memory cell 111 is selected according to the reliability requirement for generating the data of the heavy error correction code, that is, if the reliability requirement is high, each column in the two-dimensional matrix includes more data of the cells 130; if the reliability requirement is low, each column in the two-dimensional matrix includes less data of the cells 130. The second type cells 132 in the first type memory cell 111 can store the heavy error correction code generated by different first type cells 131, for example, the second type cells 132 in the first type memory cell 111 in the first column can store the heavy error correction code generated according to the data of the first cell 130 to the eighth cell 130, and the heavy error correction code generated according to the data of the third cell 130 to the tenth cell 130. This allows the controller to perform memory error correction on the data of different first type cells 131.
[0051] The error correction algorithms used for primary memory error correction and secondary memory error correction provided in the embodiments of this application include any one of parity check code, Hamming code, block code, and error correction code. The error correction algorithm used for primary memory error correction and the error correction algorithm used for secondary memory error correction can be the same or different, without limitation.
[0052] In one example, primary memory error correction can reuse the error correction method of DDR5. For example, primary memory error correction can adopt Hamming code or block code. Since the error correction method of DDR5 has low latency and its error correction ability can cover most error scenarios of the storage-class memory, when the error rate of the storage-class memory is relatively low, it can effectively reduce the access latency under the condition of ensuring the high reliability of the data stored in the storage-class memory.
[0053] When the primary memory error correction fails to correct errors or corrects errors unsuccessfully, the re-error correction code within a single or multiple first type of memory grains 111 can be used for secondary memory error correction. For example, secondary memory error correction can adopt block code or low-density parity check code. When the error rate of the storage-class memory is relatively high, secondary memory error correction can effectively reduce the error rate of the storage-class memory, ensure the high reliability of the data stored in the storage-class memory, and extend the service life of the storage-class memory.
[0054] The above example shows that the primary memory error correction is set between the first type of memory grains 111, that is, the data of the first type of memory grains 111 belonging to the same group (channel) is used to construct the running error correction code word, and the secondary memory error correction is set within the first type of memory grains 111, that is, the data of the first type of memory grains 111 is used to construct the re-error correction code word. In practical applications, other setting methods can also be adopted for the two-level memory error correction of the storage-class memory.
[0055] The above embodiments are described by taking the bit width of the memory grain as 8 bits as an example. In another embodiment, the bit width of the memory grains used in the storage-class memory can also be 4 bits.
[0056] Exemplarily, as Figure 5 shown, the storage-class memory 500 includes 20 memory grains 510, and the bit width of each memory grain 510 is 4 bits. The 20 memory grains 510 are divided into 4 groups, each group includes 5 memory grains 510, and the sum of the bit widths of the 5 memory grains 510 included in each group is 20 bits. The sum of the bit widths of the 20 memory grains 510 included in the 4 groups is 80 bits, so the storage-class memory 500 is compatible with the memory bit width of DDR5. Each group contains 4 first type of memory grains 511 and 1 second type of memory grain 512.
[0057] Optionally, as Figure 6As shown, the 20 memory particles 510 included in the storage-class memory 500 can also be divided into 2 groups, each group includes 10 memory particles 510, and the sum of the bit widths of the 10 memory particles 510 included in each group is 40 bits. The sum of the bit widths of the 20 memory particles 510 included in the 2 groups is 80 bits, and the storage-class memory 500 is compatible with the memory bit width of DDR5. Each group includes 9 first-class memory particles 511 and 1 second-class memory particle 512.
[0058] For a detailed explanation of the first-type memory cell 511 and the one second-type memory cell 512 , reference may be made to the above description of the first-type memory cell 111 and the one second-type memory cell 112 .
[0059] When the storage-class memory is used as the main memory, the storage-class memory can be installed on the computer device using a dual-inline-memory-module (DIMM) interface, so that the processor of the computer device can read and write the storage-class memory. The computer device can be an independent server or a computing device in a computing cluster.
[0060] A controller is also required between the storage-level memory and the processor, and the controller is used to implement functions such as instruction conversion, memory error correction, bad block management, and address mapping. Compared with the solution in which the built-in controller of the storage-level memory implements memory error correction, and the DRAM cache data connected to the controller causes uncertainty in the access delay of the SCM, the storage-level memory provided in the embodiment of the present application does not require a built-in controller and DRAM, and uses the memory controller of the processor connected to the storage-level memory to control the storage-level memory, or connects other controllers between the processor and the storage-level memory to control the storage-level memory and ensure deterministic latency in accessing the storage-level memory, which is compatible with the DDR protocol.
[0061] Figure 7 A schematic diagram of a processor system provided in an embodiment of the present application. The processor system 700 includes a controller 710, a storage-level memory 720, and a processor 730. The processor system 700 may be a computer device, a server, or a computing device in a computing cluster.
[0062] The processor 730 is the arithmetic and control core of the processor system 700. The processor 730 can be a very large scale integrated circuit. An operating system 731 and other software programs are installed in the processor 730, enabling the processor 730 to access the storage-class memory 720 and various Peripheral Component Interconnect Express (PCIe) devices. The processor 730 includes one or more processor cores 732. The processor cores 732 in the processor 730 are, for example, a Central Processing Unit (CPU) or other application specific integrated circuits (ASICs). The processor 730 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In practical applications, the processor system 700 can also include multiple processors.
[0063] In the embodiments of the present application, the processor 730 is used to write data to the storage-class memory 720 or read data from the storage-class memory 720.
[0064] The storage-class memory 720 can be the main memory of the processor system 700. The storage-class memory 720 is usually used to store various running software in the operating system 731, input and output data, and information exchanged with the external memory, etc. In order to improve the access speed of the processor 730, the storage-class memory 720 needs to have the advantage of fast access speed. The processor 730 can access the storage-class memory 720 at high speed through the controller 710, and perform read and write operations on any storage unit in the storage-class memory 720.
[0065] The controller 710 is a bus circuit controller that internally controls the storage-class memory 720 in the processor system 700 and is used to manage and plan the data transmission from the storage-class memory 720 to the processor core 732. Through the controller 710, data can be exchanged between the storage-class memory 720 and the processor core 732.
[0066] The controller 710 can be integrated into the processor 730, built into the northbridge, or be an independent memory controller chip. The controller 710 can be a separate chip and be connected to the processor core 732 via the system bus. For example, the difference from the processor system shown above Figure 7 is that, as shown in Figure 8 , the controller 710 can be an external controller between the storage-class memory 720 and the processor 730. This external controller can be a centralized controller that can be connected to multiple storage-class memories 720, thereby expanding the storage capacity of the main memory of the processor system and making the storage-class memory more compatible with the DDR protocol and reducing the cost of the controller.
[0067] The embodiments of the present application do not limit the specific location and form of existence of the memory controller. In practical applications, the controller 710 can control the necessary logic to write data to the storage-class memory 720 or read data from the storage-class memory 720. The controller 710 can be a memory controller in a processor system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, an embedded processor, etc.
[0068] The processor system 700 further includes various input / output (I / O) devices 740. The I / O device 740 refers to the hardware for data transmission and can also be understood as the device docked with the I / O interface. Common I / O devices include network cards, printers, keyboards, mice, etc. All external memories can also be used as I / O devices, such as hard disks, floppy disks, optical discs, etc.
[0069] The processor 730, the storage-class memory 720, the controller 710, and the I / O device 740 are connected via a bus 750. The bus 750 may include a path for transmitting information between the above components (such as the processor 730, the storage-class memory 720). In addition to the data bus, the bus 750 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, all kinds of buses are labeled as bus 750 in the figure. The bus 750 may be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. For example, the processor 730 may access these I / O devices 740 through the PCIe bus. The processor 730 is connected to the storage-class memory 720 via a double data rate (DDR) bus. Here, different storage-class memories 720 may communicate with the processor 730 using different data buses. Therefore, the DDR bus may also be replaced by other types of data buses. The embodiments of the present application do not limit the type of the bus.
[0070] The processor system 700 further includes a DPU 760, which may be connected to the processor 730 via a PCIe bus. The DPU 760 offloads the artificial intelligence, storage, and other related applications running on other chips in the processor system 700 (such as the processor 730), improves the data processing performance of the processor system 700, and reduces the load of the processor system 700. The processor system 700 may also include other dedicated processors. A dedicated processor is a processor for a specific application or field. For example, a graphics processing unit (GPU) for processing image data or a DSP for processing signals.
[0071] It should be noted that the processor system 700 may be referred to as a host. Figure 7 This is only a schematic diagram. The processor system 700 may further include other components, such as a hard disk, an optical drive, a power supply, a chassis, a cooling system, and other input / output controllers and interfaces, which are not drawn in Figure 7 The embodiments of the present application do not limit the number of processors, storage-class memories, and controllers included in the processor system 700.
[0072] Next, Figure 9Schematic diagram of a data processing method provided by an embodiment of this application. Taking Figure 7 the structure shown as an example for illustration. Assume that the processor 730 writes to or reads data from the storage-class memory 720. The data processing method provided by the embodiment of this application includes the following steps.
[0073] In some embodiments, when the processor 730 writes data to the storage-class memory 720, S910 to S930 are executed.
[0074] S910. The controller 710 performs instruction conversion, that is, the controller 710 converts the write instruction of the processor 730 into a write instruction recognizable by the storage-class memory 720.
[0075] S920. The controller 710 generates a running error correction code and a redundant error correction code according to the data to be written.
[0076] The controller 710 determines the position of the first type of memory particles in the storage-class memory 720 where the data to be written is to be written. For example, the controller 710 performs address mapping, that is, converts the logical block address into a physical block address. The logical block address (LBA) describes the virtual address of the data block on the storage device, and is generally used in auxiliary memory devices such as hard disks. The LBA can refer to the address of a certain data block or the data block pointed to by a certain address. The physical block address (PBA) describes the physical address of the data block on the storage device. There is a one-to-one mapping relationship between the LBA and the PBA, and this mapping relationship is usually stored in the main memory. In the embodiment of the application, the mapping relationship between the LBA and the PBA can be stored in the storage-class memory. When the controller 710 determines to write the data to be written into the storage-class memory 720, it first determines the logical block address, and then determines the physical block address of the storage-class memory 720 according to the logical block address and the mapping relationship between the LBA and the PBA, and writes the data to be written into the storage-class memory 720 according to this physical block address.
[0077] Furthermore, the controller 710 generates a running error correction code based on the data written to the same group of the first type of memory particles in the data to be written by using an error correction algorithm (such as: Hamming code or block code), and generates a redundant error correction code based on the data written to the same first type of memory particles by using an error correction algorithm (such as: block code or low-density parity-check code).
[0078] S930. The controller 710 writes the running error correction code, the redundant error correction code, and the data to be written into the storage-class memory 720.
[0079] The controller 710 writes the data to be written to the determined positions of the first type of memory particles, and writes the running error correction code to the second type of memory particles. The heavy error correction code is written to the second type of cells of the first type of memory particles.
[0080] In some other embodiments, when the processor 730 reads data from the storage-class memory 720, S940 to S970 are executed.
[0081] S940. The controller 710 performs instruction conversion, that is, the controller 710 converts the read instruction of the processor 730 into a read instruction recognizable by the storage-class memory 720.
[0082] S950. The controller 710 reads data, the running error correction code, and the heavy error correction code from the storage-class memory 720.
[0083] The controller 710 determines the storage position of the data in the storage-class memory 720, that is, first determines the logical block address, and then determines the physical block address of the storage-class memory 720 according to the logical block address and the mapping relationship between the LBA and the PBA, and reads the data from the storage-class memory 720 according to the physical block address.
[0084] S960. The controller 710 performs first-level memory error correction on the read data, that is, uses the running error correction code to verify the read data.
[0085] In some embodiments, the controller 710 determines the storage position of the data, reads the data and the running error correction code according to the storage position. The controller 710 generates a temporary running error correction code using the read data based on an error correction algorithm (such as: Hamming code or block code), compares the temporary running error correction code with the running error correction code. If the temporary running error correction code is the same as the running error correction code, it means the read data is correct; if the temporary running error correction code is different from the running error correction code, it means the read data has an error. If the read data does not match the stored code, the read data can be decrypted using the parity bit to determine which bit is in error, and then correct the in-error bit.
[0086] Optionally, if the first-level memory error correction fails, S970 is executed. The controller 710 performs second-level memory error correction, that is, uses the heavy error correction code to verify the data of the first type of memory particles. For example, the controller 710 uses block code or low-density parity-check code to perform second-level memory error correction on the data stored in the first type of memory particles using the heavy error correction code.
[0087] During the data processing, the controller 710 can continuously scan the data using an algorithm (such as: Hamming code) to check and correct the memory errors of each bit.
[0088] The embodiments of the present application do not limit the method for determining the position where the controller 710 writes data to the storage-class memory 720 and the position when reading data from the storage-class memory 720, and the prior art can be referred to. Moreover, with regard to the generation method of the run error correction code and the double error correction code, and the verification method for the read data using the run error correction code and the double error correction code, the prior art can also be reused.
[0089] It can be understood that, in order to implement the functions in the above embodiments, the computing device includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application scenario and design constraint conditions of the technical solution.
[0090] The method steps in this embodiment can be implemented in the way of hardware, or can be implemented by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), register, hard disk, removable hard disk, CD-ROM or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor, so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in the computing device. Of course, the processor and the storage medium can also exist as discrete components in a network device or a terminal device.
[0091] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid state drive (SSD).
[0092] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A storage-class memory, characterized in that, Applied to a processor system, the processor system includes the storage-class memory and a controller, the storage-class memory includes a plurality of memory dies, and the plurality of memory dies are divided into at least one group; the first group in the at least one group includes a first type of memory die and a second type of memory die; The second type of memory die is used to store a running error correction code, the second type of memory die is compatible with the DDR protocol, and by using the running error correction code, the controller performs primary memory error correction on the data stored in the first type of memory die in the first group based on the DDR protocol by multiplexing the error correction method of DDR5; The first type of memory die is used to store data and a heavy error correction code, and the heavy error correction code is used to perform secondary memory error correction on the data stored in the first type of memory die when the controller fails to perform primary memory error correction on the data stored in the first type of memory die with the error correction code, and the secondary memory error correction ability is stronger than the primary memory error correction ability.
2. The storage-class memory according to claim 1, wherein The first type of memory die includes M first type of units and N second type of units, the first type of units are used to store data, and the second type of units are used to store the heavy error correction code.
3. The storage-class memory according to claim 2, wherein The unit access data amount of the first type of units and the unit access data amount of the second type of units are both the unit access data amount of one bit in the bit width of the first type of memory die.
4. The storage-class memory according to any one of claims 1-3, characterized in that, The first type of memory die further includes a third type of units, and the third type of units are used to implement the function of bad block management for the units of the first type of memory die.
5. The storage-class memory according to any one of claims 1 to 3, characterized in that, The bit width of the storage-class memory meets the memory bit width indicated by the double data rate DDR protocol.
6. The storage-class memory according to claim 5, wherein The number of the first type of memory dies and the number of the second type of memory dies in each group are determined according to the bit width of the first type of memory die, the bit width of the second type of memory die, and the memory bit width indicated by the DDR protocol.
7. The storage-class memory according to claim 6, characterized in that, The memory bit width indicated by the DDR protocol is 80 bits, the bit widths of both the first type of memory die and the second type of memory die are 8 bits, and the 10 memory dies included in the storage-class memory are divided into 2 groups, and each group includes 4 first type of memory dies and 1 second type of memory die; the sum of the bit widths of the 4 first type of memory dies and 1 second type of memory die included in each group is 40 bits, and the unit access data amount of each bit is 16 bytes or 32 bytes.
8. A data processing method, characterized in that, The method is executed by a controller, the controller is connected to at least one storage-class memory, the storage-class memory includes a plurality of memory dies, the plurality of memory dies are divided into at least one group, the first group in the at least one group includes a first type of memory die and a second type of memory die, and the second type of memory die is compatible with the DDR protocol, and the method includes: The controller performs primary memory error correction on the data stored in the first type of memory die in the first group of the storage-class memory based on the DDR protocol by multiplexing the error correction method of DDR5; When the first-level memory error correction fails, the controller acquires the data of the first type of memory particles, and performs second-level memory error correction on the data stored in the first type of memory particles, where the second-level memory error correction ability is stronger than the first-level memory error correction ability.
9. The method according to claim 8, wherein The controller performing first-level memory error correction on the data stored in the first type of memory particles includes: The controller uses Hamming code or block code to perform first-level memory error correction on the data stored in the first type of memory particles in the storage-level memory; The controller performing second-level memory error correction on the data stored in the first type of memory particles includes: The controller uses block code or low-density parity-check code to perform second-level memory error correction on the data stored in the first type of memory particles.
10. The method according to claim 8 or 9, characterized in that, The controller performing first-level memory error correction on the data stored in the first type of memory particles includes: The controller performs memory error correction on a single bit in 128 bytes to 512 bytes of the data stored in the first type of memory particles; The controller performing second-level memory error correction on the data stored in the first type of memory particles includes: The controller performs second-level memory error correction on a hundred bits in 2048 bytes to 4096 bytes of the data stored in the first type of memory particles.
11. The method according to claim 8 or 9, characterized in that, The method further includes: The controller converts the instructions of the processor into the instructions of the storage-level memory, and converts the instructions of the storage-level memory into the instructions of the processor.
12. A processor system, characterized in that, The processor system includes a controller and at least one storage-level memory and a processor as described in any one of claims 1-7. The controller is respectively connected to the processor and at least one of the storage-level memories, and the controller is configured to execute the operation steps of the method described in any one of claims 8-11 above.
Citation Information
Patent Citations
Dual data protection in storage devices
US11170869B1