Large-capacity Memory System for In-memory Computing

The scalable memory system addresses deduplication challenges by using partitioned architecture and reference counters to manage virtual and physical memory spaces efficiently, enhancing performance and reducing latency.

JP7713070B2Active Publication Date: 2025-07-24SAMSUNG ELECTRONICS CO LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
JP2024108045
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-11-04
Filing Date
2024-07-04
Publication Date
2025-07-24
Estimated Expiration
2039-08-14

AI Technical Summary

Technical Problem

Conventional deduplication systems for in-memory computing face issues such as non-linear increase in deduplication translation tables and increased read and write latency due to multiple physical reads/writes for logical operations, degrading performance.

Method used

A scalable memory system architecture with system partitions, transaction managers, command queues, and write data engine managers that maintain data coherency and consistency, using reference counters and conversion tables to manage virtual and physical memory spaces efficiently, reducing memory requirements and latency.

Benefits of technology

The system provides improved performance and scalability by maintaining data integrity and reducing memory usage, achieving high throughput and low latency even with increasing virtual memory sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713070000002
    Figure 0007713070000002
  • Figure 0007713070000003
    Figure 0007713070000003
  • Figure 0007713070000004
    Figure 0007713070000004
Patent Text Reader

Abstract

To provide a memory system for providing deduplication of user data in a physical memory space of a system.SOLUTION: A system partition 200 that a system architecture 100 has includes: a physical memory having a physical memory space; a write data engine manager 204 that includes a write command queue, receives a data write request from a transaction manager for managing a memory space, and manages at least one of data consistency and data integrity to a physical memory space on the basis of determination of a memory write conflict by using the write command queue; and a write data engine for storing data corresponding to a data write request in a physical memory space on the basis of a write command stored in a write command queue in response to the write command queue.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a memory system, and more particularly, to a memory system that provides deduplication of user data in the physical memory space of the system for user data replicated in the virtual memory space of a host system.

Background Art

[0002] Artificial Intelligence (AI), big data, and in-memory processing use increasingly large memory capacities. To meet such requirements, in-memory deduplication systems (dedupe-DRAM) have been developed. Unfortunately, conventional deduplication systems have several problems. For example, the deduplication translation table increases non-linearly as the virtual memory size increases. Furthermore, the deduplication operation usually causes some increase in read and write latency. That is, a single logical read or write requires multiple physical reads or writes, thereby degrading performance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] The present invention has been made in view of the above conventional problems, and an object of the present invention is to provide a scalable architecture that enables a large-capacity memory system for in-memory computing with improved performance.

Means for Solving the Problems

[0006] A memory system according to one aspect of the present invention made to achieve the above object includes at least one system partition. The at least one system partition includes a physical memory having a physical memory space of a first predetermined size, at least one transaction manager that receives a data write request and corresponding data from a host system and maintains data coherency and data consistency for the virtual memory space of the host system using a transaction table, at least one command queue that maintains data coherency and data consistency for the physical memory space using an outstanding bucket number and the at least one command queue, receives the data write request from the transaction manager, and transmits a write command corresponding to each of the data write requests to a selected command queue among the at least one command queue, and at least one write data engine manager corresponding to each of the at least one command queue. When the data corresponding to the data write request is not replicated in the virtual memory space, the write data engine manager stores the data in an overflow memory area to respond to the write command stored in the corresponding command queue, or when the data corresponding to the data write request is replicated in the virtual memory space, the write data engine manager increments a reference counter for the data. The memory system is characterized by comprising the above components.

[0007] The virtual memory space includes a second predetermined size that is equal to or greater than the first predetermined size of the physical memory space. The memory system includes a translation table that includes a first predetermined number of bits corresponding to the second predetermined size of the virtual memory space, a second predetermined number of bits corresponding to the first predetermined size of the physical memory, and a third predetermined number of bits for data granularity. The second predetermined number of bits and the third predetermined number of bits may be a subset of the first predetermined number of bits.

[0008] A memory system according to another aspect of the present invention made to achieve the above object includes a plurality of system partitions, and at least one of the plurality of system partitions has a physical memory having a physical memory space of a first predetermined size including a plurality of memory regions, and receives a data write request and corresponding data from a host system, and uses a transaction table to maintain data consistency and data integrity for a virtual memory space of the host system including a second predetermined size greater than the physical size of the first predetermined size, and uses a conversion table to convert the virtual memory space of the host system into a physical memory space of the at least one system partition. At least one transaction manager, at least one command queue, and uses an outstanding bucket number and the at least one command queue to maintain data consistency and data integrity for the physical memory space, receives the data write request from the transaction manager, and transmits a write command corresponding to each of the data write requests to a selected command queue among the at least one command queue. At least one write data engine manager, corresponding to each of the at least one command queue, and when the data corresponding to the data write request is not replicated in the virtual memory space, stores the data in an overflow region, or when the data is replicated in the virtual memory space, increases a reference counter for the data, and a write data engine that responds to the write command stored in the corresponding command queue, corresponding to each memory region of the physical memory, and includes a reference counter storage space for the data stored in each memory region, and a memory regional manager that controls access of the write data engine to the corresponding each memory region.

[0009] The conversion table includes a first predetermined number of bits corresponding to the second predetermined size of the virtual memory space, a second predetermined number of bits corresponding to the first predetermined size of the physical memory, and a third predetermined number of bits for data granularity, and the second predetermined number of bits and the third predetermined number of bits may be subsets of the first predetermined number of bits.

[0010] A duplicate elimination memory system according to one aspect of the present invention made to achieve the above object includes a plurality of system partitions, and at least one of the plurality of system partitions includes a physical memory having a physical memory space of a first predetermined size, receives a data write request and corresponding data from a host system, and uses a transaction table to maintain data consistency and data integrity with respect to a virtual memory space of the host system including a second predetermined size equal to or greater than the first predetermined size of the physical memory space. The virtual memory space of the host system is converted into the physical memory space of the at least one system partition using a conversion table including a first predetermined number of bits corresponding to the second predetermined size of the virtual memory space, a second predetermined number of bits corresponding to the first predetermined size of the physical memory, and a third predetermined number of bits for data granularity, wherein the second predetermined number of bits and the third predetermined number of bits are subsets of the first predetermined number of bits. At least one transaction manager, at least one command queue, which uses an outstanding bucket number and the at least one command queue to maintain data consistency and data integrity with respect to the physical memory space, receives the data request from the transaction manager, and transmits a write command corresponding to each of the data write requests to a selected command queue among the at least one command queue. At least one write data engine manager, a write data engine corresponding to each command queue, which stores the data in an overflow area when the data corresponding to the data write request is not replicated in the virtual memory space, or increases a reference counter when the data corresponding to the data write request is replicated in the virtual memory space, and responds to the write command stored in the corresponding command queue.

Advantages of the Invention

[0011] According to the present invention, a scalable architecture is provided that enables a large-capacity memory system for in-memory computing with improved performance.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

Mode for Carrying Out the Invention

[0013] In the following detailed description, various embodiments are described to provide a complete understanding of the present invention. However, the idea of the present invention can be implemented by those skilled in the art without the detailed description herein. In other examples, well-known methods, procedures, configurations, and circuits are not described so as not to obscure the embodiments of the present invention. Further, the described aspects can be embodied in an imaging device or system including, but not limited to, a smartphone, a user equipment, and / or a laptop computer, to perform low-power, 3D depth measurement.

[0014] Throughout the detailed description of this specification, the term "one embodiment" means that the specific feature, structure, or characteristic associated with the embodiment is included in at least one embodiment described in this specification. That is, what is described in various parts of the detailed description in the expression "in one embodiment" or "according to one embodiment" (or other expressions having a similar meaning) does not necessarily indicate the same embodiment. Furthermore, the specific features, structures, or characteristics can be combined in one or more embodiments in a suitable manner. In contrast, the term "exemplary" as used in this specification means "provided as an example, instance, or illustration". Some embodiments described as "exemplary" in this specification are not construed as necessarily being more preferred or advantageous than other embodiments. Furthermore, the specific features, structures, or characteristics can be combined in one or more embodiments in a suitable manner. Also, depending on the context described in this specification, singular terms include the corresponding plural forms, and plural terms include the corresponding singular forms.

[0015] The various drawings (including the configuration drawings) described in this specification are shown merely for convenience of explanation. The reference numerals are repeated in the drawings to indicate corresponding and / or similar elements.

[0016] The terms used in this specification are for the purpose of describing some embodiments, and the present invention is not limited thereto. The term "comprising", as used herein, while indicating the presence of the recited features, integers, steps, operations, elements, and / or components, does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms "first", "second", etc., used herein are used as labels preceding nouns and are not limited to a particular type of order (e.g., spatial, temporal, logical, etc.) unless explicitly defined. Further, the same reference numerals are used in two or more drawings to refer to parts, components, blocks, circuits, units, or modules having the same or similar functions. However, such use is for the sake of simplicity of the drawings and convenience of explanation, and does not limit that the configurational or structural details of such a configuration or unit are the only way to embody some of the exemplary embodiments described herein where the parts / modules commonly referred to in all embodiments are the same or common.

[0017] When an element or layer is described as being connected to another element or layer, it can be directly connected to the other element or layer or there may be intervening elements or layers. Conversely, when an element is described as being directly connected to another element or layer, there are no intervening elements or layers. The term "and / or" as used herein includes any combination of one or more of the associated listed items.

[0018] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms such as those defined in a general dictionary are to be interpreted as having the same meaning as in the context of the relevant art and are not to be interpreted in an ideal or overly formal sense unless explicitly defined herein.

[0019] As used herein, the term "module" refers to any combination of software, firmware, and / or hardware configured to provide the functions described herein in association with the module. Software is embodied as a software package, code, and / or command set or commands, and the term "hardware" as used herein, for example, includes any one or combination of hard-wired circuits, programmable circuits, state machine circuits, and / or firmware storing commands executed by a programmable circuit. A module is embodied either collectively or individually as a circuit forming part of a larger system, such as, but not limited to, an integrated circuit (IC), a system-on-chip (SoC), etc. The various configuration and / or functional blocks described herein are embodied as modules including software, firmware, and / or hardware configured to provide the functions described herein in relation to the various configuration and / or functional blocks.

[0020] The present disclosure provides a large-scale deduplication memory system architecture having a conversion table in which the size does not increase non-linearly even when the virtual memory size of the system increases. In one embodiment, the architecture of the memory system described herein includes a plurality of system partitions, and the functions within each system partition are parallelized to provide high throughput and low latency.

[0021] In one embodiment, the system partition is configured to manage the virtual memory space of the host system described later and includes one or more transaction managers that provide concurrency and coherency of the virtual memory space. Further, each transaction manager supports multiple outstanding transactions and may support multiple threads, each including multiple outstanding transactions.

[0022] The system partition includes one or more data engine managers configured to manage the physical memory space of the system partition. The data engine manager is configured to manage a plurality of parallel data engines. The plurality of data engines may execute interleaving of memory addresses for a plurality of orthogonal memory regions of the system partition to improve throughput. The term "orthogonal" as used herein with respect to the plurality of memory regions means that any data engine can access any memory region of the system partition. For example, for a write access, a memory write conflict is managed by the write data engine manager (WDEM) before the data arrives at the write data engine (WDE). After the data arrives at the WDE, the memory access conflict is eliminated. For a read access, the read data engine (RDE) can access any memory region without memory access conflict and without limitation. The memory regional manager manages memory regions associated with an application, such as different deduplicated memory regions. That is, the physical memory conflict is eliminated "orthogonally" by the WDEM.

[0023] In addition, the architecture of the memory system described in this specification provides a reference counter (RC) that can be incorporated into user data stored in the physical memory of the system to reduce the memory size (memory amount) used by the memory system. In one embodiment, the reference counter provides an indication of the number of times user data containing the reference counter is replicated in the virtual memory space of the memory system. In one embodiment, incorporating the reference counter into the user data reduces the system memory requirements by approximately 6% at the granularity of 64-byte user data. In other embodiments, incorporating the reference counter into the user data reduces the system memory requirements by approximately 12% at the granularity of 32-byte user data.

[0024] In other embodiments where the number of times user data is replicated in the virtual memory space of the memory system is greater than or equal to a predetermined number of times, the field used for the reference counter is alternated with a field that provides an index to an extended reference counter table or a specific data pattern table. In yet another embodiment, the translation table used by the memory system includes an entry having a pointer or index to a specific data pattern table when the number of times user data is replicated in the virtual memory space of the memory system is greater than or equal to a predetermined number of times, thereby reducing latency and improving throughput. By arranging a pointer or index in the translation table, user data that is frequently replicated in the virtual memory space is automatically detected when the translation table is accessed. Further, such frequently replicated user data can be more easily analyzed using the automatic detection provided by such a configuration of the translation table.

[0025] FIG. 1 is a block diagram showing an example of a scalable deduplication memory system architecture according to an embodiment of the present invention. The scalable deduplication memory system architecture (hereinafter referred to as system architecture 100) includes one or more host devices or host systems 101, a host interface 102, a front-end scheduler 103, and a plurality of system partitions (200A to 200K). The host system 101 is communicably connected to the host interface 102, and the host interface 102 is communicably connected to the front-end scheduler 103. The host interface 102 includes one or more direct memory access (DMA) devices (DMA0 to DMAH-1). The front-end scheduler 103 is communicably connected to each of the plurality of system partitions (200A to 200K). The host system 101, the host interface 102, and the front-end scheduler 103 operate in a well-known manner.

[0026] The system partition 200 comprises a write data path 201 and a read data path 202. The write data path 201 includes a transaction manager (TM) 203, a write data engine (WDE) manager (WDEM) 204, one or more memory region managers (MRM) 207, and one or more memory regions (MR) 209. In one embodiment, the write data path 201 includes a transaction manager (TM) 203. In one embodiment, there is one TM 203 and one WDEM 204 per system partition. As shown in FIG. 1, the MR 209 includes a memory controller (MEM CNTRLR) and a DIMM (dual in-line memory module). In other embodiments, it includes equivalents for the MR 209. The memory of the MR 209 may include, but is not limited to, dynamic random access memory (DRAM), static random access memory (SRAM), volatile memory, and / or non-volatile memory.

[0027] Compared with the write data path 201, the read data path 202 is relatively simple because there are no issues of data coherency or data consistency related to the read data path. As shown in FIGS. 1 and 2, the read data path 202 includes one or more MRs 209, one or more memory region managers (MRM) 207 (FIG. 1), a read data engine (RDE) manager (RDEM) 212, and one or more RDEs 213. A read request for data received from the host system 101 includes a conversion table (FIG. 4) and, in response to the read access of the data, accesses an appropriate memory region and provides the requested read data via the read data path 202.

[0028] Figure 2 is a more detailed block diagram showing an example of the system partition shown in Figure 1. The architecture of the system partition 200 is parallelized so that multiple transactions and multiple threads (and their multiple transactions) are processed in a way that provides high throughput.

[0029] The write data path 201 of the system partition 200 includes a transaction manager (TM) 203, a write data engine manager (WDEM) 204, one or more WDEs 205 (WDE0 to WDEN-1), a memory region arbiter 206, MRMs 207 (MRM0 to MRMP-1), an interconnect 208, and one or more memory regions 209 (MR0 to MRP-1). The read data path 202 of the system partition 200 includes a memory region 209, an interconnect 208, a memory region arbiter, a read data engine manager (RDEM) 212, and one or more read data engines (RDE) 213. The system partition 200 includes a translation table manager (TT MGR) 210 and a TT memory (TT MEM) 211. Each of the system partitions (200A to 200K) shown in Figure 1 is similarly configured.

[0030] Components such as the memory region arbiter 206 and the interconnect 208 operate in a well-known manner and are not described in detail here. Further, the memory region arbiter 206 is located on both the write data path 201 and the read data path 202, as shown by the dotted line in FIG. 2. However, each path may actually include individual memory region arbiters that communicate with and cooperate with other memory region arbiters. As shown in FIG. 1, the read data path 202 passes through the MRM 207, but is not shown as such in FIG. 2. This is because the read data engine manager (RDEM) 212 only needs to read the conversion table to obtain the PLID for the physical memory location, and then the RDE 213 reads the memory location indexed by the PLID. Further, the configurations shown in FIGS. 1 and 2 may be embodied as modules that include software, firmware, and / or hardware that provide the functions described herein in association with various configurations and / or functional blocks.

[0031] FIG. 3 is a block diagram showing another configuration example of the write data path of the system partition according to an embodiment of the present invention. More specifically, the components of the write data path 201 of the system partition 200 shown in FIG. 3 include the TM 203, the WDEM 204, the WDE 205, the interconnect 208, and the MR 209. Further, the write data path 201 includes a write data buffer controller 214.

[0032] Referring to FIGS. 1 to 3, TM203 includes a TM controller (TM CTRL) 215 and one or more transaction tables (TM table) 216. The transaction table 216 is used to track unprocessed write transactions occurring between the host system 101 and the memory area 209. WDEM204 includes an outstanding bucket number table (OBN table) 217, a command dispatcher 218, an overflow memory invalidate manager (OVMI manager) 219, and one or more command queues (CQ) 220 (CQ0 to CQN-1). WDE205 includes one or more WDEs (WDE0 to WDEN-1). The write data buffer controller 214 includes one or more write data buffers 221.

[0033] TM203 is configured to manage data coherency and data consistency with respect to the virtual memory space of host system 101. In one embodiment, TM203 supports a plurality of outstanding write transactions and a plurality of threads, and provides high data throughput. In one embodiment, transaction table 216 includes a list of outstanding write transactions. In other embodiments, transaction table 216 includes a list of a plurality of threads, each having a plurality of outstanding transactions. When receiving a data write request from host system 101, TM203 assigns a transaction identification number (ID) to the write data request, and the transaction ID and other metadata are entered into the selected transaction table 216. The data associated with the write data request is stored in write data buffer 221. The transaction ID and information associated with the data in write data buffer 221 are downstreamed for further processing by write data path 201. TM203 uses one or more TM tables 216 to maintain the completion of sequential memory writes posted in order to the host system / virtual memory space.

[0034] FIG. 4 is a detailed diagram showing an example of a conversion table used by TM203 and RDEM212 to track the location of data within the physical memory space of a system partition according to an embodiment of the present invention. A particular exemplary embodiment of conversion table 400 corresponds to a virtual memory space of 1TB (i.e., 2 40 bytes) and is referred to as VA_WIDTH. The deduplication granularity (data granularity) is 64 bits (i.e., 2 6 bits). The physical memory space of conversion table 400 is 256GB (i.e., 2 38bytes) and is referred to as PHY SPACE. The index (TT_IDX) of the conversion table of the system shown in FIG. 4 has 34 bits. The partitioning provided by the system partition 200 results in maintaining the index of the conversion table within 32 bits, so as not to extremely increase the size of the conversion table as the size of the user data stored in MR209 increases. In other embodiments, the virtual memory size, physical memory size, and deduplication granularity (data granularity) may be different from those shown in FIG. 4.

[0035] FIG. 5 is a diagram showing the relationship between a hash table entry of an exemplary hash table and a conversion table entry of an exemplary conversion table according to an embodiment of the present invention. Referring to FIGS. 2 and 5, in one embodiment, the conversion table 500 is stored in TT MEM 211, and the hash table 501 is stored in the HTRC MEM of MR209. The logical address 502 received from the host system 101 includes a conversion table index (TT_IDX in FIG. 4) and a granularity. The conversion table index provides an index to the physical line identifier (PLID: physical line identification) entry of the conversion table 500. The format 503 of the PLID entry includes a row (i.e., hash bucket) index (R_INDX) and a column index (COL_INDX) for the hash table 501. The content included in a specific row and a specific column of the hash table 501 is a physical line (PL: physical line) that can be specific user data having a granularity of 2 g and may be, for example, when g is 6, the granularity of the PL is 64 bytes.

[0036] In one embodiment, the hash table index (i.e., both the row index and the column index) is generated by a hash function h executed on user data. Since the hash table entry is generated by the hash function h executed on the user data C, only a part C” of the user data needs to be stored in the PL. In other words, if C represents user data with granularity 2 g and C” represents the part of the user data C that needs to be stored in the hash table, then C’ represents the part of the user data C that can be restored from C” using the hash function h.

[0037]

Number

[0038] In the PL, the space or field corresponding to C’ is used for other purposes such that the user data C stores reference counter (RC) information related to the number of times it is replicated in the virtual memory of the host system 101. In the PL, the space or field corresponding to C’ is referred to as a reconstructed field.

[0039] FIG. 6 is a flowchart of an exemplary control process provided by a transaction manager according to an embodiment of the present invention. In control process 600, at step 601, TM203 is in an idle loop waiting to receive a data write command from host system 101 or a response from WEDM204. At step 602, if a data write command is received from host system 101, the flow proceeds to step 603 where the transaction table 216 is searched. At step 604, the data write command is assigned a transaction ID and inserted into the appropriate transaction table 216. At step 605, the write data is input into the write data buffer 221. At step 606, the transaction table 216 is read to obtain old physical line identifier (PLID) information, and then the flow returns to step 601.

[0040] At step 602, if the data write command is not received, the flow proceeds to step 607 where it is determined whether a response has been received from WEDM204. If no response is received, the flow returns to step 601. If a response is received, the flow proceeds to step 608 where the transaction table 216 is updated. At step 609, the conversion table stored in TT MEM211 is updated. Then, the flow returns to step 601.

[0041] WDEM204 is configured to manage data coherency and data consistency of the physical memory space within partition 200 by managing WDE205. In one embodiment, WDEM204 uses command (CMD) dispatcher 218 and command queues 220 (i.e., CQ0~CQN-1) to maintain internal read / write transactions (read / write threads) to other WDE205s and multiple unprocessed writes. WDEM204 operates to ensure that write transactions to the same hash bucket are executed in order by transmitting write transactions to the same hash bucket to the same command queue (CQ). WDEM204 merges partial cache line writes and performs memory management for the OV MEM (overflow memory) area of MR209.

[0042] WDEM204 includes an OBN (outstanding bucket number) table 217 that is used to track hash table bucket numbers and the status of unprocessed write commands. Command dispatcher 218 assigns commands to other WDE205s by storing the write commands in the selected CQ220. WDEM204 may further include an OVM invalidation table (not shown).

[0043] FIG. 7 is a flowchart of an exemplary control process provided by WEDM according to an embodiment of the present invention. The control process 700 starts from step 701. At step 702, it is determined whether a write data command has been received from the transaction manager (TM) 203 together with the write data buffer ID and the old PLID information. If received, the flow proceeds to step 703 where the write data table is read. At step 704, if the write data read from the write data buffer 221 is represented in a specific data pattern, the flow proceeds to step 706 to transmit a response to the transaction manager. If, at step 704, the write data is not the specific data, the flow proceeds to step 705 where the outstanding command buffer (OCB) table is updated. At step 707, the opcode is added to the appropriate command queue (CQ) 220, and at step 708 the process ends. If, at step 709, a response is received from the WDE 205, the flow proceeds to step 710 where the outstanding command buffer (OCB) table is updated. At step 711, the updated status is transmitted to the transaction manager (TM) 203, and at step 708 the process ends.

[0044] Each WDE 205 receives commands (opcodes) from the corresponding CQ 220 of the WEDM 204. The WDE 205 is configured to perform duplicate elimination determination and other related calculations. The WDE 205 can operate in parallel to improve the overall throughput of the system partition 200. Each WDE 205 is coordinated with the MRM 207 of each memory area to interleave memory access.

[0045] Figure 8 is a flowchart of an exemplary control process provided by the WDE according to an embodiment of the present invention. At step 801, the WDE 205 is in an idle loop waiting to receive a command (opcode) from the WDEM 204. At step 802, it is determined whether the command compares the signature of the user data with the signatures of the data of other users in the hash bucket, or whether the command is to decrement the reference counter (RC). If the command (i.e., the opcode) is to decrement the reference counter, the flow proceeds to step 803 where the reference counter for the user data is decremented. Thereafter, the flow proceeds to step 804 where the request status is returned to the WDEM 204, and then returns to step 801.

[0046] At step 802, if the request received from the WDEM 204 is to write user data (i.e., a physical line), it proceeds to step 805 where the signature of the user data is compared with the signatures of other user data and compared with the user data already stored in the hash table. At step 806, if there is a match, the flow proceeds to step 807 where the reference counter for the user data is incremented. Thereafter, it proceeds to step 804 where the request status is reported (returned) to the WDEM 204 again. If there is no match at step 806, it proceeds to step 808. If no previous user data is found, the HTRC manager is called to store the user data in the hash table with a reference counter of "1", or if the hash bucket is full or a hash collision occurs, the OVM manager is called to store the user data in the overflow area. Thereafter, the flow proceeds to step 804 to report the request status to the WDEM 204.

[0047] MRM207 manages the memory access to MR209 from WDE205. MRM207 includes a hash table / reference counter manager (HTRC MGR) 222, an overflow memory manager (OVM MGR) 223, and a signature manager (SIG MGR) 224. Each MRM207 receives control and data information for HTRC MGR 222 (HR R / W), OVM MGR 223 (OVM R / W), and SIG MGR 224 (SIG R / W) from each of the WDE205s.

[0048] FIG. 9A is a diagram showing an example of a physical line (PL) 900 including a reconstructed field 901 having an embedded reference counter (embedded RC) according to an embodiment of the present invention. In this embodiment, since the number of times user data is replicated in the virtual memory space is small, the reconstructed field 901 includes an embedded RC having a relatively small size. For such a system situation, the embedded RC is configured as a base RC having a size of (r - 1) bits.

[0049] In other embodiments where the number of times user data is replicated in the virtual memory space of the host system exceeds the size of the embedded RC in the reconstructed field of PL900, the reconstructed field includes an index to an extended RC table entry. In another embodiment, an indicator flag, such as a selected bit of the PL, is used to indicate whether the number of times user data is replicated in the virtual memory space exceeds the size of the embedded RC.

[0050] Figure 9B is a diagram showing an example of PL910 including an indicator flag 911 according to an embodiment of the present invention. As shown in Figure 9B, the indicator flag 911 is set to indicate that the number of replications of user data is less than or equal to the size of the built-in RC in the restoration field 912. Therefore, the content of the restoration field 912 indicates the number of replications of user data in the virtual memory space of the host system 101.

[0051] Figure 9C is a diagram showing a situation where the indicator flag 911 is set to indicate that the number of replications of user data exceeds the size of the built-in RC in the restoration field 912 when the content of the restoration field includes an index to an extended RC table 913 including an extended count of the replicated user data. Alternatively, the restoration field 912 may include an index or pointer to a specific data pattern table 914 including, for example, frequently replicated data. In one embodiment, the conversion table includes a PLID that is an index or pointer to a specific data pattern table 914 for frequently replicated user data, as shown in Figure 8, and can reduce the latency associated with the user data. That is, by placing a pointer or index in the conversion table, user data frequently replicated in the virtual memory space can be automatically detected when the conversion table is accessed.

[0052] Each of the MR209 includes a memory area for metadata and a plurality of data memory areas. When the system partition 200 is configured as a deduplication memory, the metadata area includes a conversion table stored in the TT MEM211. The data memory areas of the MR209 include a hash table / reference counter (HTRC) memory (HTRC MEM), a signature memory (SG MEM), and an overflow memory (OV MEM).

[0053] Writes from other virtual addresses to the same physical memory region in MR209 are managed by MRM207. The HTRC manager 222 and the SIG manager 224 perform deduplication functions and calculate the memory location with respect to the row (alias, bucket) and column positions of the hash table. Writes from other virtual addresses to the same HTRC bucket may reach the HTRC bucket in an order different from the original order from the host. Consistency is managed by WDEM204. When write data, for example, write data A and write data B, reach the same PLID (A == B), the reference counter is simply incremented. When write data from other virtual addresses is not identical but has the same bucket number but a different column (alias, way) number, the write data is stored in the same bucket but differently in other ways. When only one entry remains in the bucket, either A or B is stored in the last entry, and the other one is stored in the overflow region. When write data A and write data B are different but have the same hash bucket and row number, the second write data is stored in the overflow region.

[0054] As will be recognized by those skilled in the art in the technical field of the present invention, the idea of the present invention described in the detailed description can be modified or corrected in a wide range of applications. Therefore, the technical scope of the present invention is not limited to the specific embodiments described above.

Description of Reference Numerals

[0055] 100 System Architecture 101 Host System 102 Host Interface 103 Front-End Scheduler 200, 200A~200K System Partitions 201 Write Data Path 202 Read Data Path 203 Transaction Manager (TM) 204 Write Data Engine Manager (WDEM) 205 Write Data Engine (WDE) 206 Memory Area Arbiter 207 Memory Area Manager (MRM) 208 Interconnect 209 Memory Area (MR) 210 Translation Table Manager (TT MGR) 211 TT Memory (TT MEM) 212 Read Data Engine Manager (RDEM) 213 Read Data Engine (RDE) 214 Write Data Buffer Controller 215 TM Controller (TM CTRL) 216 Transaction Table (TM Table) 217 OBN Table 218 Command (CMD) Dispatcher 219 OVMI Manager 220 Command Queue (CQ) 221 Write Data Buffer 222 Hash Table / Reference Counter Manager (HTRC MGR) 223 Overflow Manager (OVM MGR) 224 Signature Manager (SIG MGR) 400, 500 Translation Table 501 Hash Table 502 Logical Address 503 Format 900, 910 Physical Line (PL) 901, 912 Recovery Field 911 Indicator Flag 913 Extended RC Table 914 Specific Data Pattern Table

Claims

1. A memory system, comprising: A physical memory having a physical memory space; A write data engine manager that includes a write command queue, receives a data write request from a transaction manager configured to manage the memory space, eliminates memory write conflicts, and uses the write command queue to maintain at least one of data consistency or data integrity with respect to the physical memory space; A write data engine corresponding to the write command queue and storing data corresponding to the data write request in the physical memory space based on a write command stored in the write command queue; When the data is not replicated in the virtual memory space of the host system, the write data engine stores the data corresponding to the write command in the write command queue in an overflow memory area; The write data engine further changes a reference counter for the data based on the fact that the data corresponding to the data write request is replicated in the virtual memory space. A memory system characterized by this.

2. The memory system according to claim 1, wherein the physical memory space includes a first predetermined size, and the virtual memory space of the host system includes a second predetermined size.

3. The memory system according to claim 2, wherein the transaction manager receives the data write request and corresponding data from the host system and maintains data consistency or data integrity with respect to the virtual memory space of the host system.

4. The memory system according to claim 1, wherein the transaction manager determines a physical address corresponding to a virtual address of the data write request based on receipt of the data write request.

5. The memory system according to claim 1, wherein the transaction manager maintains at least one of data consistency after sequential memory writes or data integrity after sequential memory writes with respect to the virtual memory space of the host system.

6. The transaction manager tracks one or more data write requests received from the host system. The memory system according to claim 1, wherein the transaction manager further tracks one or more threads of write data requests. **Claim 7** The transaction manager assigns a transaction identification number (ID) to the data write request, The memory system according to claim 1, wherein the transaction manager transmits one or more write completion messages to the host system in an order corresponding to the transaction ID of the received data write request. **Claim 8** A memory system, A physical memory having a physical memory space including a plurality of memory regions, A write data engine manager that includes a write command queue, receives a data write request from a transaction manager configured to manage the memory space, eliminates memory write conflicts, and uses the write command queue to maintain at least one of data consistency or data integrity for the physical memory space. A write data engine that corresponds to the write command queue and stores data corresponding to the data write request in an overflow region when the data is not replicated in the virtual memory space of the host system in response to a write command stored in the write command queue. A memory system characterized by comprising: **Claim 9** A memory region manager that corresponds to each memory region of the physical memory, includes a storage space for a reference counter for data stored in each memory region, and controls access of the write data engine to the corresponding each memory region. Further provided, The memory system according to claim 8, wherein the write data engine changes a reference counter for the data based on the data corresponding to the data write request replicated in the virtual memory space. **Claim 10** A memory region manager that corresponds to each memory region of the physical memory, includes a storage space for a reference counter for data stored in each memory region, and controls access of the write data engine to the corresponding each memory region. Further provided, The memory system according to claim 8, wherein the physical memory space includes a first predetermined size, and the virtual memory space of the host system includes a second predetermined size.

11. The transaction manager receives the data write request and corresponding data from the host system, and maintains at least one of data consistency or data integrity with respect to the virtual memory space of the host system, The memory system according to claim 8, wherein the virtual memory space of the host system is converted into the physical memory space.

12. The memory system according to claim 8, wherein the transaction manager maintains at least one of data consistency after sequential memory writes or data integrity after sequential memory writes with respect to the virtual memory space of the host system.

13. The memory system according to claim 8, wherein the transaction manager determines a physical address corresponding to a virtual address of the data write request based on the data write request.

14. The transaction manager tracks one or more data write requests received from the host system, The transaction manager assigns one or more transaction identification numbers (IDs) to at least one of the one or more data write requests received from the host system, and transmits one or more write completion messages to the host system in an order corresponding to the one or more transaction IDs of at least one of the one or more data write requests received from the host system. The memory system according to claim 8, characterized in that.

15. A method of operating a memory system, comprising: receiving a data write request from a transaction manager configured to manage a memory space via a write data engine manager; the write data engine manager eliminates a memory write conflict and uses a write command queue to maintain at least one of data consistency or data integrity with respect to the physical memory space of the physical memory. A step of storing data corresponding to the data write request in the physical memory space based on a write command stored in the write command queue by a write data engine corresponding to the write command queue, and including, The method further includes a step of increasing a reference counter for the data based on that the data corresponding to the data write request is replicated in a virtual memory space of a host system by the write data engine based on a write command stored in the write command queue.

16. The physical memory space includes a first predetermined size, the virtual memory space of the host system includes a second predetermined size, and the method further Receiving the data write request and corresponding data from the host system by the transaction manager, and Maintaining at least one of data consistency or data integrity for the virtual memory space of the host system by the transaction manager, and Converting the virtual memory space of the host system into the physical memory space by the transaction manager. The method according to claim 15, further including the above features.

17. The method according to claim 15, further including a step of controlling, by a memory area manager associated with a corresponding memory area of the physical memory, one or more accesses by the write data engine to the memory area corresponding to the memory area manager.

18. The method according to claim 15, further including a step of tracking, by the transaction manager, one or more threads of one or more write requests.

Citation Information

Patent Citations

  • Main storage sharing type multiprocessor system

    JP2000348000A

  • Transaction-based memory systems and memory modules, and methods of operating master controller and slave controller

    JP2017045452A

  • Arithmetic processing unit, control device, information processing device and method for controlling information processing device

    JP2017146786A

  • Data collection and storage method and duplication removal module

    JP2017208096A

  • System and method for maximized dedupable memory

    JP2018120594A