IRDIMM storage access method and system
By detecting weak and failed memory units in the DRAM chip and mapping their addresses to reserved memory units, the memory reliability problems caused by the differences in memory units in the DRAM chip are solved, and the reliability and availability of memory are improved.
Patent Information
- Application Number
- CN202510765306.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
In the prior art, there are weak memory units and failed memory units in the DRAM chip, which fail to effectively identify and process, resulting in reduced memory reliability, which may lead to inability to read and write correctly and errors.
By detecting the storage unit situation in the DRAM chip, establishing an address mapping table, mapping the addresses of weak storage units and failed storage units to the reserved storage units, increasing the refresh frequency, and issuing an early warning signal when there is insufficient reserved space to ensure that the storage units with strong data retention capabilities replace weak or failed units.
It improves memory reliability, reduces errors on the application side, extends the service life of the DRAM chip, and enhances the usability of the system.
Smart Images

Figure CN120277012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of memory technologies, and particularly to a storage access method and system. Background Art
[0002] The number of storage units in a DRAM chip is in the tens of billions. Among the tens of billions of storage units, there may be defective storage units. In the prior art for defective storage units, for example, in the related art with the publication number CN118051444A, usually the bad row address is replaced so that when accessing the bad row address, the actual accessed address is a valid row address, thereby ensuring that the memory can be read and written normally.
[0003] However, in the related art, only defective storage units and how to replace defective storage units are concerned, and it has not been found that among the valid storage units, the data retention capabilities of each valid unit are different and the differences are very large. The JEDEC standard usually requires the data retention capability to be 64 mS or more, but the DRAM native design margin is calculated in "multiples". The Mean of the normal distribution is mostly in the "second level" (more than ten times the margin, "low-risk strong storage units"), and only a very small number of storage units have relatively poor data retention capabilities (referred to as "discrete units" or "high-risk weak storage units"). Discrete units have two characteristics. The first is that the data retention capability is low, and the second is that they are relatively easy to age (the data retention capability ages and decreases).
[0004] Since the number of these weak storage units is relatively small compared to strong storage units, they are easily overlooked. However, if these weak storage units continue to be used, the memory reliability decreases, which may cause the memory to be unable to be read and written correctly, thereby generating errors at the application end. Summary of the Invention
[0005] The main object of the present invention is to propose an iRDIMM storage access method and system, aiming to improve the memory reliability and reduce errors generated at the application end.
[0006] To achieve the above object, the present invention proposes an iRDIMM storage access method. The iRDIMM has at least one DRAM chip, and the DRAM chip includes a plurality of storage units. The iRDIMM storage access method includes: Detecting the storage conditions of the plurality of storage units; When it is detected according to the storage conditions that there are weak storage units among the plurality of storage units, writing the weak storage unit address into the mapped address of the address mapping table, and writing the reserved storage unit address into the mapping address of the address mapping table, so that when accessing the weak storage unit address, according to the weak storage unit address and the address query table, the reserved storage unit address is accessed; wherein, The address query table includes a weak storage unit index table and a weak storage unit mapping table. The weak storage unit index table is used to query the addresses of weak storage units; the address mapping table is used to store the correspondence between the addresses of weak storage units and the mapped addresses; The storage area indicated by the reserved storage unit address is located in the reserved space of the DRAM chip.
[0007] In one embodiment, after the step of detecting the storage conditions of a plurality of the storage units, the iRDIMM storage access method further includes: When it is detected that there are weak storage units among a plurality of the storage units according to the storage conditions, detecting the reserved space of the DRAM chip; When the reserved space of the DRAM chip is less than a preset threshold, increasing the refresh frequency of the weak storage units to a preset refresh frequency, and refreshing the weak storage units at the preset refresh frequency.
[0008] In one embodiment, after the step of detecting the storage conditions of a plurality of the storage units, the iRDIMM storage access method further includes: When it is detected that there are weak storage units among a plurality of the storage units according to the storage conditions, detecting the reserved space of the DRAM chip; When the reserved space of the DRAM chip is less than a preset threshold, comparing the data retention capabilities of the weak storage units; Writing the addresses of the weak storage units with weaker data retention capabilities into the mapped addresses of the address mapping table, and writing the reserved storage unit addresses into the mapping addresses of the address mapping table, so that when accessing the addresses of the weak storage units, according to the addresses of the weak storage units and the address query table, accessing the reserved storage unit addresses corresponding to the addresses of the weak storage units with weaker data retention capabilities; And, increasing the refresh frequency of the weak storage units with stronger data retention capabilities.
[0009] In one embodiment, after the step of detecting the storage conditions of a plurality of the storage units, the iRDIMM storage access method further includes: When it is detected that there are failed storage units among a plurality of the storage units according to the storage conditions, writing the addresses of the failed storage units into the mapped addresses of the address mapping table, and writing the reserved storage unit addresses into the mapping addresses of the address mapping table, so that when accessing the addresses of the failed storage units, according to the addresses of the failed storage units and the address query table, accessing the reserved storage unit addresses corresponding to the failed storage units.
[0010] In one embodiment, after the step of detecting the storage conditions of a plurality of the storage units, the iRDIMM storage access method further includes: When it is detected that there are weak memory cells and failed memory cells in multiple said memory cells according to the storage situation, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip does not meet the remapping requirements of the weak memory cells and the failed memory cells, write the address of the failed memory cell into the mapped address of the address mapping table, and write the address of the reserved memory cell into the mapping address of the address mapping table, so that when accessing the address of the failed memory cell, according to the address of the failed memory cell and the address query table, access the address of the reserved memory cell corresponding to the failed memory cell; And, increase the refresh frequency of the weak memory cells to a preset refresh frequency, and refresh the weak memory cells at the preset refresh frequency.
[0011] In one embodiment, after the step of detecting the storage situation of multiple said memory cells, the iRDIMM memory access method further includes: When it is detected that there are weak memory cells and failed memory cells in multiple said memory cells according to the storage situation, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip does not meet the remapping requirements of the failed memory cells, and it is determined according to the address query table that there is a mapped weak memory cell address, cancel the correspondence between the weak memory cell address and the mapping address, so as to release the weak memory cell address from the address query table; And, increase the refresh frequency of the weak memory cells to a preset refresh frequency, and refresh the weak memory cells at the preset refresh frequency.
[0012] In one embodiment, after the step of detecting the storage situation of multiple said memory cells, the iRDIMM memory access method further includes: When it is detected that there are weak memory cells and failed memory cells in multiple said memory cells according to the storage situation, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip meets the remapping requirements of the weak memory cells and the failed memory cells, write the weak memory cell address and the address of the failed memory cell into the mapped address of the address mapping table respectively, and write the corresponding reserved memory cell address into the mapping address of the address mapping table, so that when accessing the weak memory cell address, according to the weak memory cell address and the address query table, access the address of the reserved memory cell corresponding to the weak memory cell; And / or, when accessing the address of the failed memory cell, according to the address of the failed memory cell and the address query table, access the address of the reserved memory cell corresponding to the failed memory cell.
[0013] In one embodiment, the iRDIMM memory access method further includes: When it is detected that the storage units in the reserved space are mapped by failed storage units and the reserved space is less than a preset threshold, a warning signal is issued.
[0014] In one embodiment, the storage conditions include data retention ability and read / write data accuracy; The specific process of detecting the storage conditions of multiple storage units includes: Detecting the data retention time, the charge amount of the storage unit, and the leakage rate of the storage unit of the storage unit to determine the data retention ability; Obtaining the data written to the storage unit and the data read from the storage unit, and comparing the data written to the storage unit with the data read from the storage unit to determine the read / write data accuracy according to the comparison result.
[0015] The present invention also provides an iRDIMM memory access system, including: An iRDIMM, including at least one DRAM chip, a non-volatile memory, and an iRCD; wherein, The DRAM chip has a storage space and a reserved space. The storage space has a plurality of storage units. The reserved space is set at a position with stronger data preservation ability of the DRAM chip. The reserved space has a plurality of reserved storage units; The non-volatile memory is used to store error information of failed storage units and / or information of weak storage units; The iRCD is used to store an address query table; A system warning module; A processor, configured to detect the storage conditions of multiple storage units, write the addresses of weak storage units and failed storage units to the mapped addresses of the address mapping table, write the addresses of reserved storage units to the mapped addresses of the address mapping table, and access the storage unit addresses, weak storage unit addresses, and failed storage unit addresses; the processor is further configured to control the system warning module to issue a warning signal when it is detected that the storage units in the reserved space are mapped by failed storage units and the reserved space is less than a preset threshold.
[0016] The present invention detects the storage conditions of multiple said storage units, and when it is detected according to the storage conditions that there are weak storage units among the multiple said storage units, writes the weak storage unit address to the mapped address of the address mapping table, and writes the reserved storage unit address to the mapping address of the address mapping table, so as to, when accessing the weak storage unit address, access the reserved storage unit address according to the weak storage unit address and the address query table. By replacing the weak storage units, the present invention enables the actual access of the strong storage unit address in the reserved space when accessing the weak storage units, which is beneficial to improving the memory reliability and reducing errors generated at the application side. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of an embodiment of the iRDIMM storage access method of the present invention; Figure 2 It is a schematic flowchart of another embodiment of the iRDIMM storage access method of the present invention; Figure 3 It is a schematic flowchart of yet another embodiment of the iRDIMM storage access method of the present invention; Figure 4 It is a schematic flowchart of still another embodiment of the iRDIMM storage access method of the present invention; Figure 5 It is a schematic flowchart of a further embodiment of the iRDIMM storage access method of the present invention; Figure 6 It is a schematic flowchart of yet another embodiment of the iRDIMM storage access method of the present invention; Figure 7 It is a schematic flowchart of an additional embodiment of the iRDIMM storage access method of the present invention; Figure 8 It is a schematic diagram of the functional modules of an embodiment of the iRDIMM storage access system of the present invention; Figure 9 It is a schematic diagram of the composition of the memory inspection program of an embodiment of the iRDIMM storage access system of the present invention; Figure 10 It is a schematic diagram of the address pass-through / replacement process of an embodiment of the iRDIMM storage access system of the present invention; Figure 11Schematic diagram of the weak storage unit row address mapping process in the iRDIMM storage access system of the present invention without a failed storage unit row address; Figure 12 Schematic diagram of the weak storage unit row address mapping process in the iRDIMM storage access system of the present invention with a failed storage unit row address; Figure 13 Schematic diagram of the weak storage unit block address mapping process in the iRDIMM storage access system of the present invention without a failed storage unit block address; Figure 14 Schematic diagram of the weak storage unit block address mapping process in the iRDIMM storage access system of the present invention with a failed row address.
[0019] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0021] It should be noted that if there are directional indications (such as up, down, left, right, front, back,...) involved in the embodiments of the present invention, the directional indications are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0022] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of the technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of the technical solutions does not exist and is not within the protection scope required by the present invention.
[0023] As used herein, the term "and / or" is merely a description of the relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.
[0024] Computer blue screens and mobile phone lags. Such "quality problems" have never been truly solved since the products were launched. With the emergence of new applications such as autonomous driving, medical intelligence, and artificial intelligence... etc., there is an increasing "zero tolerance" for "errors". Therefore, the quality requirements for electronic products have been raised from "reliable" to "dependable". "Dependable" means that we can entrust our health, wealth, and even lives to the relevant systems. DRAM (Dynamic Random Access Memory) is an indispensable component in all computing systems and is also the single factor that has the greatest impact on the quality of electronic products. According to data from many global leading enterprises, product returns (RMA) due to DRAM errors account for 25% - 75% of the overall proportion. What's more troublesome is that DRAM errors are often a multi-factor coupling phenomenon. For example, when running a certain algorithm in a specific environment (combined with temperature, voltage, algorithm, process, cause and effect... etc.), it is difficult to reproduce, diagnose, and repair. DRAM chips are the most difficult to design, manufacture, and test... and are also... the most prone to aging. Moreover, the failure error rate during the life cycle of DRAM chips follows the "bathtub curve" rule, that is, in the initial stage of product application (usually 3 to 9 months), it is the peak period of chip aging. During this period, the internal structure of the chip is constantly changing, and its performance and parameters fluctuate greatly. If the performance or parameters drop below the application requirements, functional problems, that is, "errors", will occur and require repair.
[0025] Therefore, the present invention proposes a storage access method for iRDIMM (Intelligent Registered Dual In-line Memory Module).
[0026] Referring to Figure 1 and Figures 7 to 10 , in an embodiment of the present invention, the iRDIMM storage access method includes: Step S100, detecting the storage conditions of multiple said storage units; The iRDIMM has at least one DRAM (Dynamic Random Access Memory) chip, and the DRAM chip includes a plurality of memory cells. The number of memory cells in a DRAM chip is in the tens of billions. The data retention capabilities of each memory cell are different, and the differences are extremely large. The JEDEC (Joint Electron Device Engineering Council) standard usually requires the data retention capability to be 64 mS or above, but the native design margin of DRAM is calculated in "multiples". Among the multiple memory cells, there are usually low-risk and strong memory cells, and there may also be high-risk and weak memory cells and failed memory cells. Among them, the low-risk and strong memory cells, also known as discrete memory cells, are memory cells whose data retention capabilities are within the main range of the Mean. Usually, the Mean of the normal distribution is mostly in the "second level" (more than ten times the margin, which is in line with the product design and production process, and the chance of problems in this part during the entire product life cycle is very low).
[0027] The high-risk and weak memory cells are memory cells whose data retention capabilities are "lower than" the 99.9% range of the normal distribution, significantly deviating from the product design and being the weaker (defective) part of the production process. Such cells are relatively easy to age. Because of the low margin, their data retention capabilities are relatively poor. The discrete memory cells may have relatively low data retention capabilities or be relatively easy to age (the data retention capabilities age and decrease). According to a large amount of experimental data, there are extremely few discrete cells whose data retention capabilities are close to the 64 mS "lifeline" of JEDEC. And the discrete memory cells that deviate from the normal distribution of the chip may have aged. If these discrete memory cells are still used to store data, errors may occur because their data retention capabilities are weak and lower than the minimum requirements of the application side.
[0028] The failed memory cells may experience storage errors due to reasons such as differences in the memory manufacturing process of the storage chip, hardware damage, or software operation errors, resulting in the failure of the memory cells.
[0029] Among them, the storage unit can be in units of rows, that is, one storage unit represents a row of storage space. The storage unit can also be in units of blocks, that is, one storage unit can be composed of multiple rows of storage space to form a storage block. The storage conditions include data retention ability and read / write data accuracy; the data retention ability can be determined by detecting the data retention time of the storage unit, the charge amount of the storage unit, and the leakage rate of the storage unit; it can be understood that a storage unit is usually provided with a transistor and a capacitor. Among them, the capacitor is used to store charge, and whether the data is at the "high logic" level (i.e., logic "1") or at the "low logic" level (i.e., logic "0") depends on whether the capacitor has stored charge, that is, whether the voltage across the capacitor is high or low. When the capacitor is charged, the voltage across the capacitor is high, representing data 1; when the capacitor is discharged, the voltage across the capacitor is low, representing data 0. Due to the leakage current caused by the junction of the switching transistor, such as a MOS transistor, the initial charge amount stored in the capacitor can be greatly reduced or completely disappear, that is, the capacitor will gradually leak electricity. Therefore, without replenishing the stored charge, the data stored in the capacitor may be lost. DRAM needs to be refreshed regularly to maintain the correctness of the data. The refresh operation is to prevent data loss by periodically rewriting data into the capacitor. If DRAM is not refreshed for a long time, the data will be lost, resulting in unreliable stored information.
[0030] The data retention ability can be represented by the length of the data retention time, the amount of charge in the storage unit, the leakage rate of the storage unit, etc. Among them, the retention time of the data depends on the time for which the accumulated charge remains in the capacitor. The more charge in the capacitor of the storage unit, the longer the time for the charge amount to disappear. At the same time, the lower the leakage rate, the longer the time for the charge to remain in the capacitor. Therefore, the strength of the data retention ability can be determined by detecting the data retention time of the storage unit, the charge amount of the storage unit, and the leakage rate of the storage unit.
[0031] The read / write data accuracy can be determined through the steps of Write-Read-Compare. Specifically, it can be to obtain the data written to the storage unit and the data read from the storage unit, and compare the data written to the storage unit with the data read from the storage unit, so as to determine the read / write data accuracy according to the comparison result. Specifically, if the read / write data is inconsistent, the storage unit has failed and cannot be used as an effective storage area for data; if the read / write data is consistent, it means that the storage unit is effective and can continue to be used as an effective storage area for data.
[0032] In practical applications, when determining data retention ability and the accuracy of reading and writing data, the actual operations of system software, built-in self-check tools of the system, externally attached professional memory detection programs, or any behavior of reading and writing to memory can be used to confirm that "the functions and performance of the memory are working properly". If a memory error occurs, error information of the storage unit is collected, including but not limited to the physical address of the faulty memory (Channel, Sub-channel, Rank, Bank Group, Bank, Row, Col, etc.), DQ (bit), time, temperature, CPU process vectors, etc. Modern high-performance memories may be affected by internal and external factors of the system and result in errors, which are related to transmission, logic, memory banks, etc. Some are random errors, some are reproducible, and some are not reproducible. Therefore, for the collected error information, it is first stored to form a database, and then the data is repeatedly analyzed and compared until a conclusion is drawn for each error. The resources for analysis include but are not limited to product design architecture, original manufacturer data, big data comparison, artificial intelligence summary, etc.
[0033] In summary, this embodiment can determine whether there are weak storage units by detecting the storage conditions of multiple said storage units, specifically by scanning the memory space of the memory and checking the data retention ability of iRDIMM, and can determine whether there are failed storage units by checking whether there are storage units with inconsistent reading and writing.
[0034] Step S200: When it is detected according to the storage condition that there are weak storage units among multiple said storage units, write the weak storage unit address into the mapped address of the address mapping table, and write the reserved storage unit address into the mapping address of the address mapping table, so that when accessing the weak storage unit address, the reserved storage unit address can be accessed according to the weak storage unit address and the address query table; where The address query table includes a weak storage unit index table and a weak storage unit mapping table. The weak storage unit index table is used to query the weak storage unit address; the address mapping table is used to store the correspondence between the weak storage unit address and the mapping address; The storage area indicated by the reserved storage unit address is located in the reserved space of the DRAM chip.
[0035] In this embodiment, the reserved space can be calculated based on the probabilities of weak storage units and failed storage units in the DRAM chip and the total number of row addresses in the DRAM chip. The probabilities of weak storage units and failed storage units can be obtained through experimental tests or based on empirical values. The iRDIMM includes multiple memory dies, that is, multiple DRAM chips. In each DRAM chip, a reserved space can be set, and each reserved space includes at least one row of reserved storage units. The weak storage unit index table includes an index number and the address of the weak storage unit. The address mapping table includes a mapped address column and a mapping address column. The mapped address column is written with the weak storage unit address, and the mapping address column is written with the reserved storage unit address. When a weak storage unit is detected among multiple storage units, the address of the weak storage unit, that is, the weak storage unit address, is written into the mapped address column, and the address of a reserved storage unit, that is, the reserved storage unit address, is written into the mapping address column. Thus, the reserved storage unit address is replaced with the reserved storage unit address, and data is written into the reserved storage unit address. When accessing the weak storage unit address, the weak storage unit address and the address query table can be used to index the reserved storage unit address that has a mapping relationship with the weak storage unit address. Among them, the iRDIMM can also be provided with non-volatile memories, such as SPD (Serial Presence Detect) and iRCD (Intelligent Registering Clock Driver). Among them, the non-volatile memory is used to store the error information of the failed storage unit and the information of the weak storage unit. The BIOS (Basic Input Output System) is used to store the memory inspection program for detecting whether there are failed storage units and high-risk weak storage units in the storage unit. When the system is working, the CPU of the system will execute the above memory inspection program. The iRCD is used to store the added address query table inside the chip. When storing the address query table, it receives the list / address mapping relationship of high-risk weak storage units sent by the system firmware and performs configuration. Without the awareness of the service, it performs pass-through or replacement operations on the commands / addresses initiated by the service. Specifically, when the command / address initiated by the service is not the address of the failed storage unit and the weak storage unit address, the address is passed through. When the command / address initiated by the service is the failed storage address and the weak storage unit address, the address replacement is performed according to the address query table.
[0036] In actual application, after the system is powered on, the system firmware applied to the memory module, i.e., iRDIMM, can read the information stored in the non-volatile memory, such as the information stored in the non-volatile memory, the list of failed storage unit addresses, the list of high-risk weak storage units, and the list of low-risk storage units, through the management channel. (Before the iRDIMM manufacturer ships the product or before the Server goes online, a test device can be provided to initialize the iRDIMM to find out the "list of high-risk weak storage units" of the chip and the "list of low-risk storage units" located at the main position of the normal distribution. The number of low-risk storage units only needs to be sufficient for the address mapping of the weak storage units. The reserved space can be set at the position where the data storage capacity of the low-risk storage units is the strongest. For example, when the position of the low-risk storage unit with the strongest data storage capacity is at the edge of the chip, the reserved space is set at the edge of the chip (by the edge, it means at the highest, lowest, or their nearby positions). Then, the addresses of the weak storage units are written into the mapped address column of the address mapping table, and the low-risk storage units are written into the mapping address column of the address mapping table, and the function of "replacing the weak with the strong" can be completed). The system firmware reserves memory space in the low-risk area according to the three lists (the list of failed storage unit addresses, the list of high-risk weak storage units, and the list of low-risk storage units), and feeds it back to the CPU that executes the memory inspection program. The CPU synthesizes the three lists, records them in the software database, and forms an address mapping relationship according to the reserved space, and passes it to the system firmware. The system firmware configures the basic functions of the iRCD according to the error information of the failed storage units and the information of the weak storage units stored in the non-volatile memory. The system firmware writes the list of high-risk weak storage units / address mapping relationship into the iRCD chip to form an address query table.
[0037] Alternatively, during the actual application process, error data information can also be collected by the CPU from system software and hardware channels during the system operation phase. The system software and hardware channels include application-side monitoring software and system monitoring hardware (such as SystemRAS), and all software and hardware modules that can obtain memory error data (such as memory controller, providing dedicated registers open to BMC). When the system supports, for example, during low-load or idle periods of the system, release the bus occupancy of memory resources, and use a memory inspection program to scan the memory space within a specified area to check the data storage ability of the memory module and check whether there are new addresses of failed storage units. The CPU comprehensively updates the three lists in real time, records them in the software database, and forms a new address mapping relationship and a new three lists according to the reserved space. The system firmware writes the new three lists into the non-volatile memory through the management channel. Of course, in other embodiments, the three lists can also be written into a "non-volatile storage space" somewhere in the system, such as a hard disk, SSD, Flash, EEPROM, ROM, etc. The system firmware writes the new list of high-risk weak storage units / address mapping relationship into the iRCD chip. In this way, when accessing the address of a weak storage unit, the replaced mapped address can be accessed through the address query table. By replacing the weak storage unit, the present invention enables the actual access to the address of the strong storage unit in the reserved space when accessing the weak storage unit, which is beneficial to improving memory reliability and reducing errors generated at the application side.
[0038] Referring to Figure 2 , in one embodiment, after the step S100 of detecting the storage conditions of multiple storage units, the iRDIMM storage access method further includes: Step S300: When detecting that there are weak storage units among the multiple storage units according to the storage conditions, detecting the reserved space of the DRAM chip; Step S400: When the reserved space of the DRAM chip is less than a preset threshold, increasing the refresh frequency of the weak storage unit to a preset refresh frequency and refreshing the weak storage unit at the preset refresh frequency.
[0039] It is understandable that the reserved space of the DRAM chip is limited. After the reserved space is occupied by the mapping address corresponding to the weak storage unit or the failed storage unit detected first, the data of the weak storage unit detected later cannot be stored in the storage unit of the reserved space by mapping, that is, the reserved space is not enough to map all the weak storage unit addresses to the reserved storage unit addresses of the reserved space. However, in order to prevent data loss, the data of the storage unit must be read before the data is lost to generate read information, and then the capacitor is recharged according to the read information to maintain the initial charge amount. Therefore, DRAM needs to be refreshed regularly to maintain the correctness of the data. The data storage capacity of the weak storage unit is weaker than that of the low-risk strong storage unit. If the original digital display frequency is continued to be used to refresh the storage unit, there will inevitably be a risk of data loss. To this end, this embodiment increases the refresh frequency of the weak storage unit so that the data retention capacity of the weak storage unit can be consistent or substantially consistent with the data retention capacity of the strong storage unit. Optionally, the refresh frequency of the weak storage unit can be increased to N times of the refresh frequency of the strong storage unit, such as 1 time, 1.5 times, 2 times, 2.5 times, 3 times, etc., and can be specifically set according to the data retention capability, that is, the weaker the data retention capability of the weak storage unit, the faster the refresh frequency of the weak storage unit. Specifically, the refresh frequency of the weak storage unit can be determined according to the storage capacity of the weak storage unit, and the refresh frequency of the weak storage unit is adjusted to the refresh frequency of the determined weak storage unit, so that the storage unit with a shorter data retention time is refreshed more frequently than the storage unit with a longer data retention time, so that the capacitor is recharged by frequent read and write operations to maintain the initial charge amount. Such setting makes it possible to refresh the weak storage unit address faster than the conventional refresh frequency when it cannot be replaced, thereby ensuring that the weak storage unit can be used normally without data loss. In actual application, the reserved space in iRCD is almost completely occupied by the real error unit, and the system early warning module will give a warning to the user or system administrator, prompting the system self-check and self-repair resources to be almost exhausted. When the mapping resources on the iRCD are exhausted, the AiMS system can speed up the DRAM refresh rate, such as double refresh, to extend system availability and buy maintenance time for server managers.
[0040] Reference Figure 3 In another embodiment, when the reserved space is insufficient to map all weak storage unit addresses to the reserved storage unit addresses of the reserved space, the following method can be used to process the weak storage units, namely: Step S300, when it is detected according to the storage situation that there is a weak storage unit in the plurality of storage units, detecting the reserved space of the DRAM chip; Step S500: When the reserved space in the DRAM chip is less than a preset threshold, compare the data retention capabilities of the weak storage cells; Step S600: Write the addresses of the weak storage cells with relatively weak data retention capabilities into the mapped addresses of the address mapping table, and write the addresses of the reserved storage cells into the mapping addresses of the address mapping table, so that when accessing the addresses of the weak storage cells, according to the addresses of the weak storage cells and the address query table, access the addresses of the reserved storage cells corresponding to the addresses of the weak storage cells with relatively weak data retention capabilities; In addition, increase the refresh frequency of the weak storage cells with relatively strong data retention capabilities.
[0041] It can be understood that the more frequently the storage cells are refreshed, that is, the faster the refresh frequency, the more system resources are occupied, and the performance of the system will also decline accordingly. Therefore, in the case where the reserved space is not sufficient to map all the addresses of the weak storage cells to the reserved storage cell addresses in the reserved space, the data retention capabilities of the already mapped weak storage cells and the newly detected weak storage cells can be compared, and the weak storage cells can be classified according to the high and low data retention capabilities, so as to distinguish the weak storage cells that need to be mapped to the reserved storage cell addresses in the reserved space and the weak storage cells that need to increase the refresh frequency. Thus, write the addresses of the weak storage cells with relatively weak data retention capabilities into the mapped addresses of the address mapping table, and write the addresses of the reserved storage cells into the mapping addresses of the address mapping table, while the addresses of the weak storage cells with relatively strong data retention capabilities are not replaced, but their refresh frequencies are increased. Moreover, the weaker the storage ability, the shorter the data retention time, and the faster the refresh frequency. Optionally, the refresh frequency of the weak storage cells can be increased to a preset refresh frequency, and the weak storage cells can be refreshed at the preset refresh frequency, so that the storage cells with a shorter data retention time are refreshed more frequently than the storage cells with a longer data retention time, and the capacitor can be recharged through frequent read and write operations to maintain the initial charge amount.
[0042] It should be noted that in the above embodiments, the data retention capabilities of all weak storage units are compared, that is, the data retention capabilities of both the already mapped weak storage units and the newly detected weak storage units are compared. When the data retention capability of the mapped weak storage unit is stronger than that of the newly detected weak storage unit, the correspondence between the address of the already mapped weak storage unit and the mapped address is cancelled, so as to release the address of the weak storage unit from the address lookup table, and the refresh frequency is reset according to the data retention capability of the weak storage unit to increase the refresh frequency of the weak storage unit. When the data retention capability of the mapped weak storage unit is weaker than that of the newly detected weak storage unit, the refresh frequency is directly reset according to the data retention capability of the newly detected weak storage unit to increase the refresh frequency of the weak storage unit.
[0043] Of course, in other embodiments, when the reserved space is less than the preset threshold and the reserved space can store the data of some of the newly detected weak storage units but not all of the data of the weak storage units, it may also be not to cancel the correspondence between the address of the already mapped weak storage unit and the mapped address to release the address of the weak storage unit from the address lookup table. Instead, only the data retention capabilities of the newly detected weak storage units are compared, and according to the data retention capabilities, the addresses of the weak storage units with weaker data retention capabilities are written into the mapped addresses of the address mapping table, and the addresses of the reserved storage units are written into the mapping addresses of the address mapping table, so as to access the corresponding reserved storage unit address according to the weak storage unit address and the address lookup table when accessing the weak storage unit address. At the same time, the refresh frequency of the weak storage units with stronger data retention capabilities that fail to store data in the reserved space is increased.
[0044] Refer to Figure 4 , in one embodiment, after step S100, the step of detecting the storage conditions of the multiple storage units, the iRDIMM storage access method further includes: Step S700: When it is detected according to the storage conditions that there are failed storage units among the multiple storage units, write the addresses of the failed storage units into the mapped addresses of the address mapping table, and write the addresses of the reserved storage units into the mapping addresses of the address mapping table, so as to access the corresponding reserved storage unit address according to the addresses of the failed storage units and the address lookup table when accessing the addresses of the failed storage units.
[0045] In this embodiment, the address of the failed storage unit written is the mapping table of the failed storage unit addresses in the address lookup table. The failed storage unit index table includes the index number and the address of the failed storage unit. The address mapping table includes the column of the mapped address and the column of the mapping address. The column of the mapped address is written with the address of the failed storage unit, and the column of the mapping address is written with the address of the reserved storage unit. When it is detected that there is one failed storage unit among multiple storage units, the address of the failed storage unit, that is, the address of the failed storage unit, is written into the column of the mapped address, and the address of a reserved storage unit, that is, the address of the reserved storage unit, is written into the column of the mapping address. Thus, the address of the reserved storage unit is replaced with the address of the reserved storage unit, and data is written into the address of the reserved storage unit. When it is necessary to access the address of the failed storage unit, the address of the failed storage unit and the address lookup table can be used to index the address of the reserved storage unit that has a mapping relationship with the address of the failed storage unit.
[0046] Optionally, referring to Figure 5 , after step 100, the step of detecting the storage status of multiple storage units, the iRDIMM storage access method further includes: Step S800: When it is detected that there are weak storage units and failed storage units among multiple storage units according to the storage status, detect the reserved space of the DRAM chip; Step S900: When the reserved space of the DRAM chip does not meet the remapping requirements of the weak storage units and the failed storage units, write the address of the failed storage unit into the mapped address of the address mapping table, and write the address of the reserved storage unit into the mapping address of the address mapping table, so that when accessing the address of the failed storage unit, according to the address of the failed storage unit and the address lookup table, access the address of the reserved storage unit corresponding to the failed storage unit; and, increase the refresh frequency of the weak storage unit to a preset refresh frequency, and refresh the weak storage unit at the preset refresh frequency.
[0047] It is understandable that the address access of high-risk weak storage units does not necessarily result in errors, but there are potential technical risks. Therefore, when new weak storage units appear, they are updated in real time, and the storage units with weak data retention capabilities on the chip / memory module are recorded in a list. This list is "updated in real time" and repaired according to the status of the storage body, that is: after the iRDIMM is initialized, when the system encounters a memory error and decides to use the iRCD scheme after confirmation, the address lookup table in the iRCD is "updated in real time", and the non-volatile storage body data on the system and the iRDIMM is also updated. This real-time update can be continuous until all margin resources are exhausted. When new weak storage units and failed storage units are detected, if the margin of the reserved space is sufficient, the addresses of the failed storage units and the weak storage units can be mapped to the addresses of the storage units in the reserved space at the same time. When the margin is insufficient, the address of the failed storage unit is preferentially mapped to the address of the storage unit in the reserved space. At the same time, the refresh frequency of the weak storage unit can be determined according to the storage capacity of the weak storage unit, and the refresh frequency of the weak storage unit is adjusted to the determined refresh frequency of the weak storage unit, so that the storage units with shorter data retention times are refreshed more frequently than the storage units with longer data retention times.
[0048] Optionally, referring to Figure 6 , after step 100, the step of detecting the storage conditions of multiple storage units, the iRDIMM storage access method further includes: Step S800: When weak storage units and failed storage units are detected in multiple storage units according to the storage conditions, detect the reserved space of the DRAM chip; Step S1000: When the reserved space of the DRAM chip does not meet the remapping requirements of the failed storage unit and there is a mapped weak storage unit address determined according to the address lookup table, cancel the correspondence between the weak storage unit address and the mapped address to release the weak storage unit address from the address lookup table; Write the address of the failed storage unit into the mapped address of the address mapping table, and write the address of the reserved storage unit into the mapped address of the address mapping table, so that when accessing the address of the failed storage unit, the corresponding reserved storage unit address can be accessed according to the address of the failed storage unit and the address lookup table; And, increase the refresh frequency of the weak storage unit to a preset refresh frequency, and refresh the weak storage unit at the preset refresh frequency.
[0049] In this embodiment, when the margin of the reserved space is not sufficient to map the address of the failed storage unit to the reserved storage unit address of the reserved space, it is necessary to cancel the correspondence between the weak storage unit address and the mapped address, so as to release the storage unit address of the reserved space occupied by the weak storage unit, map the address of the failed storage unit to the storage unit address of the reserved space, that is, write the address of the failed storage unit to the mapped address of the address mapping table, and write the corresponding reserved storage unit address to the mapping address of the address mapping table, so as to access the reserved storage unit address corresponding to the failed storage unit according to the address of the failed storage unit and the address query table when accessing the address of the failed storage unit. At the same time, the refresh frequency of the weak storage unit can be determined according to the storage capacity of the weak storage unit, and the refresh frequency of the weak storage unit is adjusted to the determined refresh frequency of the weak storage unit, so that the storage unit with a shorter data retention time is refreshed more frequently than the storage unit with a longer data retention time. Among them, the released weak storage unit can be the one with stronger data retention ability among the replaced weak storage units.
[0050] Refer to Figures 11 to 14 , where Figure 11 shows the mapping process of the weak storage unit row address in the case of no failed storage unit row address, Figure 12 shows the mapping process of the weak storage unit row address in the case of having a failed storage unit row address, Figure 13 shows the mapping process of the weak storage unit block address in the case of no failed storage unit row address, Figure 14 shows the mapping process of the weak storage unit block address in the case of having a failed storage unit row address. Optionally, before canceling the correspondence between the weak storage unit address and the mapped address to release the storage unit address of the reserved space occupied by the weak storage unit, the data retention capabilities of the weak storage units stored in the reserved space can be compared, and the correspondence between the weak storage unit address corresponding to the one with the strongest data retention ability and the mapped address is canceled. That is, every time it is necessary to release the storage unit address of the reserved space occupied by a weak storage unit, start releasing from the storage unit address of the reserved space occupied by the weak storage unit with the strongest data retention ability until the weak storage unit does not occupy the reserved space.
[0051] In one embodiment, the iRDIMM storage access method further includes: When it is detected that the storage unit in the reserved space is mapped by a failed storage unit and the reserved space is less than a preset threshold, a warning signal is issued.
[0052] It can be understood that due to the limited reserved space, even if the weak storage units do not occupy the reserved space, as the usage time of the iRDIMM increases, during this process, there will be more and more failed storage units. As a result, when the failed storage units occupy all the reserved space and the reserved space is less than the preset threshold, a warning message can be issued. In actual application, almost all of the reserved space in the iRCD is occupied by the truly faulty units, and the system warning module will alert the user or system administrator, indicating that the system self-check and self-repair resources are almost exhausted.
[0053] Referring to Figure 7 , in one embodiment, after the step of detecting the storage conditions of the multiple storage units, the iRDIMM storage access method further includes: Step S1100: When weak storage units and failed storage units are detected among the multiple storage units according to the storage conditions, detect the reserved space of the DRAM chip; Step S1200: When the reserved space of the DRAM chip meets the remapping requirements of the weak storage units and the failed storage units, write the weak storage unit address and the failed storage unit address into the mapped addresses of the address mapping table respectively, and write the corresponding reserved storage unit address into the mapping address of the address mapping table, so that when accessing the weak storage unit address, according to the weak storage unit address and the address query table, access the corresponding reserved storage unit address of the weak storage unit; And / or, when accessing the failed storage unit address, according to the failed storage unit address and the address query table, access the corresponding reserved storage unit address of the failed storage unit.
[0054] In this embodiment, when it is detected that both weak storage units and failed storage units exist simultaneously, if the margin of the reserved space is sufficient, the addresses of the failed storage units and the weak storage units can be mapped to the storage unit addresses of the reserved space simultaneously. Otherwise, the replacement of the address of the failed storage unit is taken as the first priority, and the replacement of the address of the weak storage unit is given priority. If there is no address of the failed storage unit, the replacement of the address of the weak storage unit is taken as the second priority and can be mapped into the reserved space. For example, if there are 5 discrete high-risk weak storage unit "row addresses" (ROWs) in a DRAM chip, 16 (a multiple of 2) ROWs are set aside in the low-risk area as the "mapping and replacement" area, and an address lookup table is established in the iRCD. The 16 row addresses (ROWs) with the worst data retention ability in the DRAM chip are used as the input of the weak storage unit index table, and the 16 reserved low-risk row addresses are used as the output of the weak storage unit index table. The overall solution is that when the iRCD receives one of the "16 worst row address lists", it will remap it to one of the low-risk row addresses through the address mapping table, and replace the "weak storage unit row address" with the "low-risk row address" (Low Risk ROW address). Once an address of a failed storage unit appears, the high-risk weak storage unit row address will release a reserved row address resource, which is occupied by the failed row address, and the weak storage unit index table is updated synchronously. After the high-risk weak storage units are replaced, the memory margin is greatly improved. The process of the memory margin decreasing when a failed storage unit address appears is as follows: The initial high-risk weak storage unit list is mapped to low-risk (Rx) replacement, for example: ROW A (78mS) is mapped to R1, 1215mS; ROW B (83mS) is mapped to R2, 1213mS; ROW C (106mS) is mapped to R3, 1211mS; ROW D (112mS) is mapped to R4, 1216mS; ROW E (128mS) is mapped to R5, 1212mS; ROW F (162mS) is mapped to R6, 1213mS; ROW G (189mS) is mapped to R7, 1211mS; ROW H (201mS) is mapped to R8, 1213mS.
[0055] To sum up, the original margin without the iRCD scheme is 78mS - 64mS = 14mS, and the margin of its data retention is (201mS+) - 64mS = 137mS+... The margin increases by more than 123mS.
[0056] When the first error occurs at the application side; the application-side monitoring system confirms that the error address is BAD-1; the weak storage unit index table is updated.
[0057] ROW A (78mS) is mapped to R1, 1215mS; ROW B (83mS) is mapped to R2, 1213mS; ROW C (106mS) is mapped to R3, 1211mS; ROW D (112mS) is mapped to R4, 1216mS; ROW E (128mS) is mapped to R5, 1212mS; ROW F (162mS) is mapped to R6, 1213mS; ROW G (189mS) is mapped to R7, 1211mS; BAD-1 (59mS) is mapped to R8, 1213mS.
[0058] It should be noted that ROW H (201mS) is replaced by BAD-1 (59mS). After the update, ROW H resumes normal operation and becomes the worst storage row address (Worst ROW address in Data Retention) of the entire chip or iRDIMM. The overall margin decreases from 137mS+ to 137mS.
[0059] When the second error occurs at the application side, the application-side monitoring system confirms that the error address is BAD-2, and the weak storage unit index table is updated.
[0060] ROW A (78mS) is mapped to R1, 1215mS; ROW B (83mS) is mapped to R2, 1213mS; ROW C (106mS) is mapped to R3, 1211mS; ROW D (112mS) is mapped to R4, 1216mS; ROW E (128mS) is mapped to R5, 1212mS; ROW F (162mS) is mapped to R6, 1213mS; BAD-2 (48mS) is mapped to R7, 1211mS; BAD-1 (59mS) is mapped to R8, 1213mS.
[0061] It should be noted that ROW G (189 mS) is replaced by BAD-2 (48 mS). After the update, ROW G is put into normal operation, and the overall margin of the worst memory row address of the entire chip or iRDIMM is reduced from 137 mS to 125 mS.
[0062] … When the seventh error occurs at the application side, the application-side monitoring system confirms that the error address is BAD-7 and updates the weak memory cell index table.
[0063] ROW A (78 mS) is mapped to R1, 1215 mS; BAD-7 (63 mS) is mapped to R2, 1213 mS; BAD-6 (19 mS) is mapped to R3, 1211 mS; BAD-5 (62 mS) is mapped to R4, 1216 mS; BAD-4 (54 mS) is mapped to R5, 1212 mS; BAD-3 (62 mS) is mapped to R6, 1213 mS; BAD-2 (48 mS) is mapped to R7, 1211 mS; BAD-1 (59 mS) is mapped to R8, 1213 mS.
[0064] ROW B (83 mS) is replaced by BAD-7 (63 mS). After the update, ROW B is put into normal operation, and the overall margin of the worst memory cell row address of the entire chip or iRDIMM is reduced from 42 mS to 19 mS.
[0065] The present invention also provides an iRDIMM memory access system. Referring to Figures 8 to 11 , the iRDIMM memory access system includes: iRDIMM, including at least one DRAM chip, non-volatile memory, and iRCD; wherein, The DRAM chip has a storage space and a reserved space. The storage space has a plurality of memory cells, and the reserved space is set at a position with stronger data preservation ability of the DRAM chip. The reserved space has a plurality of reserved memory cells; The non-volatile memory is used to store error information of failed memory cells and / or information of weak memory cells; System warning module; The iRCD is used to store an address query table; A processor, configured to detect the storage conditions of multiple said storage units, write the addresses of weak storage units and failed storage units into the mapped addresses of an address mapping table, write the addresses of reserved storage units into the mapping addresses of the address mapping table, and access the storage unit addresses, weak storage unit addresses and failed storage unit addresses; the processor is further configured to, when detecting that the storage units in the reserved space are mapped by the failed storage units and the reserved space is less than a preset threshold, control the system warning module to issue a warning signal.
[0066] In this embodiment, the non-volatile memory increases the information of the failed storage unit and / or weak storage unit, the storage of the error information, the memory inspection program and the RCD chip to increase the address query table, and reserve a part of the memory space as the address remapping area. The memory inspection program can be run by the processor in the system background without affecting the normal operation of the application system. The memory inspection program includes an error collection interface, an intercommunication interface with the system firmware, a memory self-check module, an address mapping calculation module, a software database, and a system early warning module. The address mapping calculation module collects the error information of the failed address and the address information of the high-risk weak storage unit area, calculates the address mapping relationship, and passes it to the system firmware, and the system firmware configures the non-volatile memory and iRCD. When the reserved space in iRCD is almost completely occupied by the real error unit, the processor can control the action of the system early warning module, and the system early warning module will warn the user or system administrator, prompting that the system self-check and self-repair resources are almost exhausted. When the mapping resources on the iRCD are exhausted, the AiMS system can speed up the DRAM refresh speed, such as double refresh, to extend the system availability and strive for maintenance and repair time for server managers. The address lookup table receives the high-risk weak storage unit list / address mapping relationship sent by the system firmware for configuration. Without being perceived by the business, the command / address initiated by the business is transparently transmitted or replaced. The error address of the memory storage unit after confirmation is written to the dedicated location of the non-volatile memory of the iRDIMM. A certain storage unit is reserved in each iRDIMM to form a reserved space for repair, and the reserved space can be set in the location of the DRAM chip with strong data storage capacity. In actual application, an SRAM (Static Random-Access Memory) can be added to the RCD dedicated to the iRDIMM for the base lookup table. Each row of the table has two address words, one to store the fail address (weak storage unit address or failed storage unit address), and the other to point to the reserved address, that is, the mapping address. For iRDIMMs that have no errors or do not need to be repaired, the error address = reserved address during initialization. When an error address (weak memory cell address or failed memory cell address) is confirmed, the system's underlying firmware will write the address to the iRDIMM's non-volatile memory and iRCD simultaneously via I2C or I3C. When the address received by the RCD matches any fail address in the lookup table, it will be replaced with the address pointing to the reserved space in the lookup table to complete the task of re-mapping the fail address.By cooperating the system-level firmware (Frame ware) with the newly added hardware logic of the iRDIMM, the present invention can improve the reliability and service life of the DRAM memory, form the functions of "patrolling, inspecting, and repairing" the memory, and truly achieve discoverability, traceability, confirmability, and repairability.
[0067] The above are only optional embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of the present invention under the inventive concept of the present invention, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present invention.
Claims
1. An iRDIMM memory access method, characterized in that, The iRDIMM has at least one DRAM chip, and the DRAM chip includes a plurality of memory cells; The iRDIMM memory access method includes: Detecting the storage conditions of the plurality of memory cells; When it is detected according to the storage conditions that there are weak memory cells among the plurality of memory cells, writing the weak memory cell address into the mapped address of the address mapping table, and writing the reserved memory cell address into the mapping address of the address mapping table, so that when accessing the weak memory cell address, according to the weak memory cell address and the address query table, accessing the reserved memory cell address; wherein, The address query table includes a weak memory cell index table and a weak memory cell mapping table. The weak memory cell index table is used to query the weak memory cell address; the address mapping table is used to store the correspondence between the weak memory cell address and the mapping address; The storage area indicated by the reserved memory cell address is located in the reserved space of the DRAM chip.
2. The iRDIMM memory access method according to claim 1, wherein After the step of detecting the storage conditions of the plurality of memory cells, the iRDIMM memory access method further includes: When it is detected according to the storage conditions that there are weak memory cells among the plurality of memory cells, detecting the reserved space of the DRAM chip; When the reserved space of the DRAM chip is less than a preset threshold, increasing the refresh frequency of the weak memory cell to a preset refresh frequency, and refreshing the weak memory cell at the preset refresh frequency.
3. The iRDIMM memory access method according to claim 1, wherein After the step of detecting the storage conditions of the plurality of memory cells, the iRDIMM memory access method further includes: When it is detected according to the storage conditions that there are weak memory cells among the plurality of memory cells, detecting the reserved space of the DRAM chip; When the reserved space of the DRAM chip is less than a preset threshold, comparing the data retention ability of the weak memory cells; Writing the address of the weak memory cell with a weaker data retention ability into the mapped address of the address mapping table, and writing the reserved memory cell address into the mapping address of the address mapping table, so that when accessing the weak memory cell address, according to the weak memory cell address and the address query table, accessing the reserved memory cell address corresponding to the weak memory cell address with a weaker data retention ability; And, increasing the refresh frequency of the weak memory cell with a stronger data retention ability.
4. The iRDIMM memory access method according to claim 1, characterized in that, After the step of detecting the storage conditions of the plurality of memory cells, the iRDIMM memory access method further includes: When it is detected according to the storage conditions that there are failed memory cells among the plurality of memory cells, writing the failed memory cell address into the mapped address of the address mapping table, and writing the reserved memory cell address into the mapping address of the address mapping table, so that when accessing the failed memory cell address, according to the failed memory cell address and the address query table, accessing the reserved memory cell address corresponding to the failed memory cell.
5. The iRDIMM memory access method according to claim 1, wherein After the step of detecting the storage conditions of the plurality of memory cells, the iRDIMM memory access method further includes: When it is detected according to the storage situation that there are weak storage units and failed storage units among multiple said storage units, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip does not meet the remapping requirements of the weak storage units and the failed storage units, write the address of the failed storage unit into the mapped address of the address mapping table, and write the address of the reserved storage unit into the mapping address of the address mapping table, so that when accessing the address of the failed storage unit, according to the address of the failed storage unit and the address query table, access the address of the reserved storage unit corresponding to the failed storage unit; And, increase the refresh frequency of the weak storage unit to a preset refresh frequency, and refresh the weak storage unit at the preset refresh frequency.
6. The iRDIMM memory access method according to claim 1, wherein After the step of detecting the storage situation of multiple said storage units, the iRDIMM storage access method further includes: When it is detected according to the storage situation that there are weak storage units and failed storage units among multiple said storage units, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip does not meet the remapping requirements of the failed storage unit and there is a mapped weak storage unit address determined according to the address query table, cancel the correspondence between the weak storage unit address and the mapping address, so as to release the weak storage unit address from the address query table; Write the address of the failed storage unit into the mapped address of the address mapping table, and write the address of the reserved storage unit into the mapping address of the address mapping table, so that when accessing the address of the failed storage unit, according to the address of the failed storage unit and the address query table, access the address of the reserved storage unit corresponding to the failed storage unit; And, increase the refresh frequency of the weak storage unit to a preset refresh frequency, and refresh the weak storage unit at the preset refresh frequency.
7. The iRDIMM memory access method according to claim 1, wherein After the step of detecting the storage situation of multiple said storage units, the iRDIMM storage access method further includes: When it is detected according to the storage situation that there are weak storage units and failed storage units among multiple said storage units, detect the reserved space of the DRAM chip; When the reserved space of the DRAM chip meets the remapping requirements of the weak storage units and the failed storage units, write the weak storage unit address and the address of the failed storage unit into the mapped address of the address mapping table respectively, and write the corresponding address of the reserved storage unit into the mapping address of the address mapping table, so that when accessing the weak storage unit address, according to the weak storage unit address and the address query table, access the address of the reserved storage unit corresponding to the weak storage unit; And / or, when accessing the address of the failed storage unit, according to the address of the failed storage unit and the address query table, access the address of the reserved storage unit corresponding to the failed storage unit.
8. The iRDIMM memory access method according to any one of claims 1 to 7, characterized in that, The iRDIMM storage access method further includes: When it is detected that the storage unit in the reserved space is mapped by a failed storage unit and the reserved space is less than a preset threshold, issue a warning signal.
9. The iRDIMM memory access method according to any one of claims 1 to 7, characterized in that, The storage conditions include data retention capability and read and write data accuracy; The detecting the storage conditions of the plurality of storage units specifically includes: Detecting the data retention time of the storage unit, the charge amount of the storage unit, and the leakage rate of the storage unit to determine the data retention capability; The data written to the storage unit and the data read from the storage unit are obtained, and the data written to the storage unit and the data read from the storage unit are compared to determine the accuracy of the read and write data according to the comparison result.
10. An iRDIMM memory access system, characterized in that, include: iRDIMM, comprising at least one DRAM chip, a non-volatile memory and an iRCD; wherein, The DRAM chip has a storage space and a reserved space, the storage space has a plurality of storage units, the reserved space is arranged at a position where the data storage capacity of the DRAM chip is relatively strong, and the reserved space has a plurality of reserved storage units; The non-volatile memory is used to store error information of failed storage units and / or information of weak storage units; The iRCD is used to store an address query table; System early warning module; A processor is used to detect the storage status of multiple storage units, write weak storage unit addresses and failed storage unit addresses into mapped addresses of an address mapping table, write reserved storage unit addresses into mapped addresses of an address mapping table, and access storage unit addresses, weak storage unit addresses and failed storage unit addresses; the processor is also used to control the system early warning module to send an early warning signal when it is detected that the storage unit of the reserved space is mapped by a failed storage unit and the reserved space is less than a preset threshold.
Citation Information
Patent Citations
Memory device and operating method thereof
CN106601285A
Memory access method and device
CN108959106A
Memory chip having reduced baseline refresh rate with additional refreshing for weak cells
CN109559770A
Devices, systems and methods with improved refresh address generation
US20140241093A1
Efficiently Managing Unmapped Blocks to Extend Life of Solid State Drive with Low Over-Provisioning
US20170160976A1
Cited By
Security storage chip energy consumption optimization method based on multi-objective optimization
CN121541767A
A method for optimizing the energy consumption of a secure storage chip based on multi-objective optimization
CN121541767B