Near memory processing apparatus, processor, and data processing method

By performing metadata processing operations in near memory processing devices, the problem of inefficiency of the ZFS file system in PB-level storage scenarios is solved, and more efficient metadata management and performance improvement is achieved.

CN119988319APending Publication Date: 2025-05-13SAMSUNG (CHINA) SEMICONDUCTOR CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510089018.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the PB-level large-capacity storage scenario, the ZFS file system needs to manage TB-level metadata, resulting in low IO operation efficiency, and memory access-intensive computing leads to a large amount of data interaction between the CPU and memory, affecting performance.

Method used

A near memory processing device, processor and data processing method are provided to reduce IO operations of loading metadata from the storage device by performing block tracking, deduplication table searching, and range tree traversal in the near memory processing device, and offloading memory access-intensive calculations to the near memory processing device.

Benefits of technology

By reducing IO operations and memory data interaction, the performance of the ZFS file system is significantly improved, the efficiency of metadata management is improved, and better solutions are provided for PB-level large-capacity storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988319A_ABST
    Figure CN119988319A_ABST
Patent Text Reader

Abstract

The invention discloses a near memory processing apparatus, a processor, and a data processing method. The near memory processing device includes: a memory; a computing core including: a block tracking core configured to determine an address of a first target data block from the memory in response to receiving an indirect block parsing request from the processor, and calculate a checksum of a second target data block and write the calculated checksum to the memory in response to receiving a checksum update request from the processor; the deduplication table search core is configured to execute deduplication table search processing in response to a deduplication table query request received from the processor so as to determine whether a third target data block exists in the memory or not; and a tree traversal core configured to traverse the range tree in the memory in response to receiving the range tree query request from the processor to determine a target node corresponding to the range tree query request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of memory technology, and more particularly, to a near memory processing device, a processor, and a data processing method. Background Art

[0002] The ZFS file system is a next-generation advanced open source file system with excellent features such as unlimited scalability, rich functions, and high integrity.

[0003] Currently, the surge in global data volume has stimulated the growth of storage device capacity. SSDs with a single capacity of 30TB to 60TB and single-node storage servers with PB levels have begun to appear on the market. Therefore, for file systems, managing PB-level data volumes will be a huge challenge.

[0004] Usually, for a file system, a portion of the storage space is used to store file system metadata. For PB-level user file data, TB-level metadata needs to be stored. The ZFS file system uses the ARC (Adjustable Replacement Cache) mechanism to automatically cache hot metadata in memory. In this case, the larger the allocated memory, the better the performance of the ZFS file system.

[0005] Every time the ZFS file system reads or writes a file, it accesses metadata such as the index node (dnode), deduplication table (DDT), and space map. If there is not enough memory to store the metadata, it will be obtained through IO (input and output), which is inefficient. In addition, there are many memory access-intensive calculations in metadata processing, resulting in massive amounts of data moving between the CPU and memory. Summary of the invention

[0006] The present disclosure provides a near memory processing device, a processor, and a data processing method, which can reduce the IO operation of loading metadata from the storage device, reduce a large amount of data interaction between the CPU and the memory, and improve the performance of the ZFS file system in a large-capacity storage scenario.

[0007] According to a first aspect of an embodiment of the present disclosure, a near memory processing device is provided, the near memory processing device comprising: a memory; a computing core, the computing core comprising: a block tracking core, configured to determine the address of a first target data block from the memory in response to an indirect block resolution request received from a processor, and to calculate the checksum of a second target data block and write the calculated checksum to the memory in response to a checksum update request received from the processor; a deduplication table search core, configured to perform a deduplication table search process in response to a deduplication table query request received from the processor to determine whether a third target data block exists in the memory; and a tree traversal core, configured to traverse a range tree in the memory in response to a range tree query request received from the processor to determine a target node corresponding to the range tree query request.

[0008] Optionally, the memory stores ZFS file system metadata, which includes an index node (dnode), a deduplication table (DDT), and a space map, wherein the index node metadata is linked to each data block through a multi-level indirect block in the memory, and wherein a range tree is generated in the memory based on the space map metadata.

[0009] Optionally, the multi-level indirect block includes N levels of indirect blocks, N is an integer greater than 1, the i-th level indirect block includes multiple blocks, i is an integer greater than or equal to 1 and less than N, each block in the j-th level indirect block includes a block pointer pointing to a corresponding block in the j-1-th level indirect block, including an address and a checksum of the corresponding block in the j-1-th level indirect block, j is an integer greater than 1 and less than or equal to N, and each block in the 1st level indirect block includes a block pointer pointing to a data block, including an address and a checksum of the data block.

[0010] Optionally, the block tracking core is configured to: in response to receiving an indirect block resolution request from a processor, based on the ID of the first target data block, start from the Nth level indirect block, load a block pointer pointing to a corresponding block in the next level indirect block until the first level indirect block; determine the address of the first target data block from the block pointer of the corresponding block in the first level indirect block, and return the address of the first target data block to the processor, wherein the indirect block resolution request is generated based on reading / writing the first target data block and includes the ID of the first target data block.

[0011] Optionally, the block tracking core is configured to: load a second target data block from a memory in response to receiving a checksum update request from a processor; calculate a checksum of the second target data block, and write the calculated checksum to an indirect block stored in the memory, wherein the checksum update request is generated based on writing the second target data block.

[0012] Optionally, the deduplication table search core is configured to: in response to receiving a deduplication table query request from the processor, calculate a first-level hash value based on the fingerprint of the third target data block; load the deduplication table leaf node from the memory based on the first-level hash value; calculate a second-level hash value based on the fingerprint of the third target data block; load the deduplication table entry from the memory based on the second-level hash value and the loaded deduplication table leaf node; in response to the fingerprint of the third target data block being equal to the key value of the loaded deduplication table entry, determine that the third target data block is stored in the memory, wherein the deduplication table query request is generated based on reading / writing the third target data block.

[0013] Optionally, the tree traversal core is configured to: in response to receiving a range tree query request from a processor, read a current range tree node from a memory; based on the read current range tree node, perform a binary search according to the start and end addresses of the space targeted by the range tree query request; in response to finding the range tree node corresponding to the range tree query request through the binary search, determine the found range tree node as the target node corresponding to the range tree query request; in response to not finding the range tree node corresponding to the range tree query request through the binary search, read the next range tree node from the memory, use the read next range tree node as the current range tree node, and return to the step of performing the binary search, wherein the range tree query request is generated based on the read / write data.

[0014] According to a second aspect of an embodiment of the present disclosure, a processor is provided, comprising: an index node management module, configured to: generate an indirect block resolution request in response to an instruction to read / write a first target data block, and send the indirect block resolution request to a near memory processing device; generate a checksum update request in response to an instruction to write the first target data block, and send the checksum update request to the near memory processing device; a deduplication table management module, configured to: generate a deduplication table query request in response to an instruction to read / write a second target data block, and send the deduplication table query request to the near memory processing device; a space management module, configured to: generate a range tree query request in response to an instruction to read / write a third target data block, and send the range tree query request to the near memory processing device.

[0015] Optionally, the index node management module is also configured to: receive the address of the first target data block determined in response to an indirect block resolution request from the near memory processing device, and / or receive the result of calculating the checksum of the first target data block in response to a checksum update request from the near memory processing device; the deduplication table management module is also configured to: receive the result of determining whether the second target data block exists in the near memory processing device in response to a deduplication table query request from the near memory processing device; the space management module is also configured to: receive the range tree node determined in response to a range tree query request from the near memory processing device.

[0016] According to a third aspect of an embodiment of the present disclosure, there is provided a data processing method executed in a near memory processing device, the data processing method comprising: in response to receiving an indirect block resolution request from a processor, determining an address of a first target data block from a memory, and in response to receiving a checksum update request from the processor, calculating a checksum of a second target data block and writing the calculated checksum to the memory; in response to receiving a deduplication table query request from the processor, performing a deduplication table search process to determine whether a third target data block exists in the memory; in response to receiving a range tree query request from the processor, traversing a range tree in the memory to determine a target node corresponding to the range tree query request.

[0017] Optionally, the memory stores ZFS file system metadata, the ZFS file system metadata including an index node (dnode), a deduplication table (DDT), and a space map, wherein the index node metadata is linked to each data block through a multi-level indirect block in the memory, wherein a range tree is generated in the memory according to the space map metadata.

[0018] Optionally, the multi-level indirect block includes N levels of indirect blocks, N is an integer greater than 1, the i-th level indirect block includes multiple blocks, i is an integer greater than or equal to 1 and less than N, each block in the j-th level indirect block includes a block pointer pointing to a corresponding block in the j-1-th level indirect block, including an address and a checksum of the corresponding block in the j-1-th level indirect block, j is an integer greater than 1 and less than or equal to N, and each block in the 1st level indirect block includes a block pointer pointing to a data block, including an address and a checksum of the data block.

[0019] Optionally, the step of determining the address of the first target data block includes: in response to receiving an indirect block resolution request from the processor, based on the ID of the first target data block, starting from the Nth level indirect block, loading a block pointer pointing to a corresponding block in the next level indirect block until the first level indirect block; determining the address of the first target data block from the block pointer of the corresponding block in the first level indirect block, and returning the address of the first target data block to the processor, wherein the indirect block resolution request is generated based on reading / writing the first target data block and includes the ID of the first target data block.

[0020] Optionally, the step of calculating the checksum of the second target data block includes: loading the second target data block from the memory in response to receiving a checksum update request from the processor; calculating the checksum of the second target data block, and writing the calculated checksum to an indirect block stored in the memory, wherein the checksum update request is generated based on writing the first target data block.

[0021] Optionally, the step of performing a deduplication table search process to determine whether a third target data block exists in a memory includes: in response to receiving a deduplication table query request from a processor, calculating a first-level hash value based on a fingerprint of the third target data block; loading a deduplication table leaf node from a memory based on the first-level hash value; calculating a second-level hash value based on the fingerprint of the third target data block; loading a deduplication table entry from a memory based on the second-level hash value and the loaded deduplication table leaf node; in response to the fingerprint of the third target data block being equal to the key value of the loaded deduplication table entry, determining that the third target data block is stored in the memory, wherein the deduplication table query request is generated based on reading / writing the third target data block.

[0022] Optionally, the step of traversing the range tree in the memory to determine the target node corresponding to the range tree query request includes: in response to receiving the range tree query request from the processor, reading the current range tree node from the memory; based on the read current range tree node, performing a binary search according to the start and end addresses of the space targeted by the range tree query request; in response to finding the range tree node corresponding to the range tree query request through the binary search, determining the found range tree node as the target node corresponding to the range tree query request; in response to not finding the range tree node corresponding to the range tree query request through the binary search, reading the next range tree node from the memory, using the read next range tree node as the current range tree node, and returning to the step of performing the binary search, wherein the range tree query request is generated based on the read / write data.

[0023] According to a fourth aspect of an embodiment of the present disclosure, a data processing method executed by a processor is provided, the data processing method comprising: generating an indirect block resolution request in response to an instruction to read / write a first target data block, and sending the indirect block resolution request to a near memory processing device; generating a checksum update request in response to an instruction to write the first target data block, and sending the checksum update request to the near memory processing device; generating a deduplication table query request in response to an instruction to read / write a second target data block, and sending the deduplication table query request to the near memory processing device; generating a range tree query request in response to an instruction to read / write a third target data block, and sending the range tree query request to the near memory processing device.

[0024] Optionally, the data processing method further includes: receiving an address of a first target data block determined in response to an indirect block resolution request from a near memory processing device, and / or receiving a result of calculating a checksum of the first target data block in response to a checksum update request from a near memory processing device; receiving a result of determining whether a second target data block exists in the near memory processing device in response to a deduplication table query request from a near memory processing device; and receiving a range tree node determined in response to a range tree query request from a near memory processing device.

[0025] According to a fifth aspect of an embodiment of the present disclosure, a computer-readable storage medium storing a computer program is provided. When the computer program is executed by a processor, the data processing method as described above is implemented.

[0026] According to a sixth aspect of an embodiment of the present disclosure, a host system is provided, the host system comprising: the near memory processing device as described above; the processor as described above; and a memory.

[0027] According to the near memory processing device, processor and data processing method of the embodiments of the present disclosure, by utilizing the near memory processing device (e.g., CMM-DC device) to expand the memory capacity, more metadata can be cached, the metadata cache hit rate can be improved, and the IO operations for loading metadata from the storage device can be reduced. At the same time, by offloading memory access intensive calculations to the near memory processing device, a large amount of data interaction between the CPU and the memory can be reduced. In addition, according to the near memory processing device, processor and data processing method of the embodiments of the present disclosure, a solution for improving metadata management can be provided for PB-level large-capacity storage, thereby improving the execution efficiency of the ZFS file system. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other objects, features and advantages of the present disclosure will become more apparent through the following detailed description in conjunction with the accompanying drawings, in which: Figure 1 is a diagram showing an example of an index node (dnode) access process of a ZFS file system; Figure 2 is a diagram showing an example of a deduplication process for a ZFS file system; Figure 3 is a diagram showing an example of space management of a ZFS file system; Figure 4 is a diagram showing the architecture of a host system according to an embodiment of the present disclosure; Figure 5 is a diagram showing an example of a process flow of block tracking and checksum according to an embodiment of the present disclosure; Figure 6 is a diagram showing an example of a deduplication table search process according to an embodiment of the present disclosure; Figure 7 is a diagram illustrating an example of a range tree search process according to an embodiment of the present disclosure; Figure 8 is a flowchart illustrating a data processing method executed by a processor according to an embodiment of the present disclosure; Fig. 9 is a flowchart illustrating a data processing method performed in a near memory processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] Hereinafter, various embodiments of the present disclosure are described with reference to the accompanying drawings, wherein the same reference numerals are used to represent the same or similar elements, features and structures. However, it is not intended that the present disclosure be limited to specific embodiments by the various embodiments described herein, and it is intended that: the present disclosure covers all modifications, equivalents and / or substitutes of the present disclosure, as long as they are within the scope of the attached claims and their equivalents. The terms and words used in the following specification and claims are not limited to their dictionary meanings, but are only used to enable a clear and consistent understanding of the present disclosure. Therefore, it should be apparent to those skilled in the art that the following description of various embodiments of the present disclosure is provided for illustrative purposes only, and not for the purpose of limiting the present disclosure defined by the attached claims and their equivalents.

[0030] It should be understood that the singular includes the plural unless the context clearly indicates otherwise.The terms "include", "comprising" and "having" used herein indicate the presence of disclosed functions, operations or elements, but do not exclude other functions, operations or elements.

[0031] For example, the expression "A or B" or "at least one of A and / or B" may indicate A and B, A or B. For example, the expression "A or B" or "at least one of A and / or B" may indicate (1) A, (2) B, or (3) both A and B.

[0032] In various embodiments of the present disclosure, it is intended that when a component (e.g., a first component) is referred to as being “coupled” or “connected” to another component (e.g., a second component) or being “coupled” or “connected” to another component (e.g., the second component), the component may be directly connected to the other component or may be connected through another component (e.g., a third component). In contrast, when a component (e.g., a first component) is referred to as being “directly coupled” or “directly connected” to another component (e.g., the second component) or being directly coupled to or directly connected to another component (e.g., the second component), there is no other component (e.g., the third component) between the component and the other component.

[0033] The expression "configured to" used in describing various embodiments of the present disclosure may be used interchangeably with expressions such as "suitable for", "having the ability to", "designed to", "suitable for", "manufactured to", and "capable of", for example, depending on the circumstances. The term "configured to" may not necessarily indicate being "specially designed to" in terms of hardware. On the contrary, the expression "a device configured to..." in some cases may indicate that the device and another device or part "can...". For example, the expression "a processor configured to perform A, B, and C" may indicate a dedicated processor (e.g., an embedded processor) for performing the corresponding operations or a general-purpose processor (e.g., a central processing unit CPU or an application processor (AP)) for performing the corresponding operations by executing at least one software program stored in a storage device.

[0034] The terms used herein are intended to describe certain embodiments of the present disclosure, but are not intended to limit the scope of other embodiments. Unless otherwise noted herein, all terms (including technical or scientific terms) used herein may have the same meaning as those generally understood by those skilled in the art. Generally, the terms defined in the dictionary should be considered to have the same meaning as the contextual meaning in the relevant field, and, unless clearly defined herein, should not be understood differently or understood to have an overly formal meaning. In any case, the terms defined in the present disclosure are not intended to be interpreted as excluding embodiments of the present disclosure.

[0035] The following first describes how the ZFS system handles metadata.

[0036] Every time the ZFS file system reads or writes a file, it accesses metadata such as the index node (dnode), deduplication table (DDT), and space map.

[0037] Figure 1 is a diagram showing an example of an index node (dnode) access flow of a ZFS file system.

[0038] Reference Figure 1 , similar to the inode in the Linux operating system, the ZFS file system uses the inode and the multi-level indirect blocks linked to the dnode to manage files. For a file with N (N is an integer greater than 2) levels of indirect blocks (L1 to LN), when accessing a data block, it is necessary to access each indirect block from top to bottom along the access path and parse the corresponding block pointer (block ptr) (such as Figure 1 The block tracking process is shown in the orange arrow in Figure 1). The checksum of a block is recorded in the block pointer of the block in the previous level. Therefore, whenever a data block is modified, all its related indirect blocks must be updated in a chain (such as Figure 1The verification and update process is shown by the blue arrows in the figure).

[0039] In the case where there is no indirect block in memory (i.e., the indirect block misses the memory), for a file with N levels of indirect blocks, each time a user accesses (reads / writes) a data block, N storage device IOs will occur, and N block pointer resolutions will be generated (memory access intensive calculations). In addition, when writing (or modifying) a data block, an additional N+1 write / modify checksums will be generated (memory access intensive calculations).

[0040] Figure 2 is a diagram illustrating an example of a deduplication process of a ZFS file system.

[0041] Reference Figure 2 , ZFS file system supports block-level deduplication. A fingerprint is generated for each data block. If the fingerprint hits the deduplication table (DDT), the storage device IO of the data block will not be triggered, which can greatly reduce the storage space. However, the cost of using this technology is that more memory is required to cache the DDT, otherwise the IO operation accessing the DDT will reduce the performance of the ZFS file system.

[0042] like Figure 2 As shown in Figure 1, DDT is a two-level hash table. For data with a repetition rate of 5 and a block size of 128k, 1TB of data requires 0.5GB of DDT. When DDT does not exist in memory (that is, DDT misses memory), accessing each data block will cause storage device IO and require multiple comparison calculations (memory access intensive calculations).

[0043] Figure 3 is a diagram illustrating an example of space management of a ZFS file system.

[0044] The ZFS file system divides the storage device into M segments of equal size (called metaslabs). Each metaslab has a space map, which records the space allocation and release. When a metaslab is selected as the space allocation object, a range tree of free space is generated in memory according to the space map.

[0045] For a TB-level storage device, a range tree of about 100MB is required in memory. When the space map does not exist in memory (i.e., the space map does not hit the memory), each time a data block is allocated, the space map needs to be read from the storage device, resulting in storage device IO. At the same time, each time a data block is allocated, a search operation on the range tree is required (memory access intensive calculation).

[0046] like Figures 1 to 3As shown in the figure, every time the ZFS file system reads or writes a file, it accesses metadata such as the index node (dnode), deduplication table (DDT), and space map. If there is not enough memory to store the metadata, it will be obtained through the storage device IO, which is inefficient. Therefore, by expanding the memory capacity, unnecessary IO can be reduced. In addition, in metadata processing, there are many memory access intensive calculations, which cause massive amounts of data to move between the CPU and memory. This memory access intensive calculation is simple to calculate, highly repeatable, and does not require complex logical control. If these memory access intensive calculations are transferred to a location close to the metadata, the data movement between the CPU and memory can be greatly reduced, fundamentally improving efficiency.

[0047] Table 1 shows the IO times and memory access intensive calculation times when accessing each data block in the prior art.

[0048] Table 1

[0049] Near memory processing (also known as near memory processing, PNM) is a technology that integrates memory and logic chips into advanced integrated circuit packages, which uses memory for data calculations and reduces data movement between the CPU and memory. The open standard CXL can be used in conjunction with PNM devices to facilitate the expansion of memory capacity. In tests, PNM solutions based on the CXL interface have been shown to more than double the performance of applications such as recommendation systems or in-memory databases that require high memory bandwidth.

[0050] Therefore, according to the embodiments of the present disclosure, a near memory processing device, a processor, and a data processing method are proposed. As a memory access intensive computing acceleration solution for the ZFS file system, the near memory processing device, the processor, and the data processing method can cache more metadata in the memory, reduce storage device IO, significantly reduce data transmission between the CPU and the memory, release CPU computing resources, reduce CPU pressure, and ultimately improve the performance of the ZFS file system by offloading these calculations to the near memory processing device (for example, CMM-DC (CXL Memory Module-DRAM Compute, CXL processor module-DRAM computing) device).

[0051] Figure 4 is a diagram illustrating an architecture of a host system according to an embodiment of the present disclosure.

[0052] Reference Figure 4 , the host system may include a processor 410 and a near memory processing device 420. In addition, the host system may also include a memory for storing data. Figure 4 In the embodiment, the host system adopts the ZFS file system, the processor 410 may be a CPU, and the near memory processing device 420 may be a CMM-DC device.

[0053] According to an embodiment of the present disclosure, the calculation of the metadata of the ZFS file system is divided into two modes. One is non-memory access intensive calculation, which does not have frequent memory access, and all data exchanges are between the host and the memory. The other is memory access intensive calculation, which has the following characteristics and is suitable for near memory processing devices such as CMM-DC devices: frequent data access; simple logic, repeated "read-calculate" or "read-calculate-write" mode; the final required result data volume is relatively small.

[0054] Reference Figure 4 The processor 410 may include multiple metadata management modules, for example, an index node (dnode) management module 411 , a deduplication table (DDT) management module 412 , and a space management module 413 .

[0055] The dnode management module 411 may generate an indirect block resolution request in response to an instruction to read / write the first target data block, and send the indirect block resolution request to the near memory processing device 420. Alternatively, the dnode management module 411 may generate a checksum update request in response to an instruction to write the first target data block, and send the checksum update request to the near memory processing device 420. On the other hand, the dnode management module 411 may receive from the near memory processing device 420 an address of the first target data block determined in response to the indirect block resolution request, and / or receive from the near memory processing device 420 a result of calculating a checksum of the first target data block in response to the checksum update request.

[0056] The DDT management module 412 may generate a deduplication table query request in response to an instruction to read / write the second target data block, and send the deduplication table query request to the near memory processing device 420. On the other hand, the DDT management module 412 may also receive from the near memory processing device 420 a result of determining whether the second target data block exists in the near memory processing device in response to the deduplication table query request.

[0057] The space management module 413 may generate a range tree query request in response to an instruction to read / write the third target data block, and send the range tree query request to the near memory processing device 420. On the other hand, the space management module 413 may receive a range tree node (e.g., a position and a corresponding value of the range tree node) determined in response to the range tree query request from the near memory processing device 420.

[0058] According to an embodiment of the present disclosure, each of the dnode management module 411, the DDT management module 412, and the space management module 413 may include a logic control unit, a non-memory access intensive algorithm unit, and a memory access intensive algorithm proxy unit.

[0059] For example, the dnode management module 411 may include a dnode control unit, a non-memory access intensive algorithm (NMI algorithm) unit, a parsing agent unit, and an update agent unit. Here, the parsing agent unit and the update agent unit belong to the memory access intensive algorithm agent unit. The non-memory access intensive algorithm unit can perform various non-memory access intensive calculations related to dnode / indirect blocks. The dnode control unit can control the parsing agent unit to generate an indirect block parsing request and send the indirect block parsing request to the near memory processing device 420 in response to an instruction to read / write the first target data block. In addition, the dnode control unit can control the update agent unit to generate a checksum update request and send the checksum update request to the near memory processing device 420 in response to an instruction to write the first target data block.

[0060] The DDT management module 412 may include a DDT control unit, a non-memory access intensive algorithm (NMI algorithm) unit, and a search agent unit as a memory access intensive algorithm agent unit. The non-memory access intensive algorithm unit may perform various non-memory access intensive calculations related to DDT. The DDT control unit may control the search agent unit to generate a deduplication table query request and send the deduplication table query request to the near memory processing device 420 in response to an instruction to read / write the second target data block.

[0061] The space management module 413 may include a space control unit, a non-memory access intensive algorithm (NMI algorithm) unit, and a traversal agent unit as a memory access intensive algorithm agent unit. The non-memory access intensive algorithm unit may perform various non-memory access intensive calculations related to the space graph. The space control unit may control the traversal agent unit to generate a range tree query request in response to an instruction to read / write the third target data block, and send the range tree query request to the near memory processing device 420.

[0062] Return to reference Figure 4 , the near memory processing device 420 may include a computing core 421 and a memory 422. The memory 422 may store ZFS file system metadata. As described above, the ZFS file system metadata may include an index node (dnode), a deduplication table (DDT), and a space map. The index node metadata may be linked to each data block through a multi-level indirect block in the memory. In addition, a range tree may be generated in the memory based on the space map metadata.

[0063] The computing core 421 may include a plurality of computing cores, for example, a block tracking core 4211, a deduplication table search core 4212, and a tree traversal core 4213. The block tracking core 4211 may determine the address of the first target data block from the memory 422 in response to receiving an indirect block resolution request from the processor 410, and calculate the checksum of the second target data block and write the calculated checksum to the memory 422 in response to receiving a checksum update request from the processor 410. The deduplication table search core 4212 may perform a deduplication table search process in response to receiving a deduplication table query request from the processor 410 to determine whether a third target data block exists in the memory 422. The tree traversal core 4213 may traverse the range tree in the memory 422 in response to receiving a range tree query request from the processor 410 to determine the target node corresponding to the range tree query request.

[0064] The block tracking core 4211 may include a block tracking control unit, an indirect block parsing unit, and a checksum calculation unit. The block tracking control unit is responsible for data interaction with the processor 410 and controlling the operation of the block tracking core 4211. In response to receiving an indirect block parsing request from the processor 410, the block tracking core 4211 may load a block pointer pointing to a corresponding block in the next level of indirect blocks starting from the Nth level indirect block based on the ID of the first target data block until the first level of indirect blocks. The block tracking core 4211 may determine the address of the first target data block from the block pointer of the corresponding block in the first level of indirect blocks, and return the address of the first target data block to the processor 410. According to an embodiment of the present disclosure, the indirect block parsing request is generated based on reading / writing the first target data block and includes the ID of the first target data block. The operations of loading the block pointer and determining the address of the first target data block described above may be performed by the indirect block parsing unit.

[0065] On the other hand, the block tracking core 4211 may load the second target data block from the memory in response to receiving a checksum update request from the processor. Then, the block tracking core 4211 may calculate a checksum of the second target data block and write the calculated checksum to an indirect block stored in the memory. According to an embodiment of the present disclosure, the checksum update request is generated based on writing the second target data block. The operations of loading the second target data block, calculating and writing the checksum described above may be performed by a checksum calculation unit.

[0066] Figure 5 is a diagram illustrating an example of a process flow of block tracking and checksum according to an embodiment of the present disclosure.

[0067] Reference Figure 5In response to receiving an indirect block resolution request from the resolution agent unit, the block tracking core 4211 may perform a block tracking operation. The indirect block resolution request may include a block tracking start instruction (trace start), an address of a dnode (dnode), and an ID of a first target data block (datablk id). The block tracking core 4211 (e.g., the indirect block resolution unit) may repeat the following operations based on the ID of the first target data block, starting from the Nth level (level = N) indirect block: calculate the indirect block ID; load a block pointer pointing to a corresponding block in the next level indirect block from the memory 422 based on the calculated indirect block ID; resolve the loaded block pointer; if the level (level) of the current indirect block is not 1, reduce the level by 1 (level--), and return to the step of calculating the indirect block ID; if the level of the current indirect block is 1, return the address in the resolved block pointer to the resolution agent unit.

[0068] Further references Figure 5 In response to receiving a checksum update request from the update agent unit, the block tracking core 4211 may perform a checksum calculation. The checksum update request may include a start instruction (start). The block tracking core 4211 (e.g., a checksum calculation unit) may load page data from the memory and then calculate the checksum of the second target data block. Here, the page data may include multiple data blocks, and the multiple data blocks may include at least one second target data block. In the case of multiple second target data blocks, it may be determined whether the current second target data block is the last second target data block. If the current second target data block is not the last second target data block, return to the step of calculating the checksum of the second target data block to calculate the checksum of the next second target data block. If the current second target data block is the last second target data block, the calculated checksums of all second target data blocks are written to the memory (i.e., written to the corresponding indirect blocks).

[0069] Return to reference Figure 4 The deduplication table search core 4212 may include a deduplication table search control unit, a hash calculation unit, a parsing calculation unit, and a comparison calculation unit. The deduplication table search control unit is responsible for data interaction with the processor 410 and controlling the operation of the deduplication table search core 4212.

[0070] Figure 6 is a diagram illustrating an example of a deduplication table search process according to an embodiment of the present disclosure.

[0071] Reference Figure 4 and Figure 6, the deduplication table search core 4212 may calculate the first level hash value according to the fingerprint of the third target data block in response to receiving the deduplication table query request from the processor (e.g., the search agent unit). According to an embodiment of the present disclosure, the deduplication table query request may be generated based on reading / writing the third target data block. The deduplication table query request may include a start instruction (search start) and a fingerprint of the third target data block, and the first level hash value may be calculated by the hash calculation unit. The deduplication table search core 4212 may load the deduplication table leaf node from the memory according to the first level hash value (L1 hash value). The parsing calculation unit may parse the DDT structure to load the DDT leaf node. The deduplication table search core 4212 may also calculate the second level hash value (L2 hash value) according to the fingerprint of the third target data block. Similarly, the second level hash value may be calculated by the hash calculation unit. The deduplication table search core 4212 may load the deduplication table entry from the memory based on the second level hash value and the loaded deduplication table leaf node. Similarly, the DDT entry may be loaded by the parsing calculation unit. In response to the fingerprint of the third target data block being equal to the key value of the loaded deduplication table entry, the deduplication table search core 4212 may determine that the third target data block is stored in the memory. Otherwise, the deduplication table search core 4212 may determine that the third target data block is not stored in the memory. The comparison calculation unit may compare the fingerprint of the third target data block with the key value of the loaded deduplication table entry. The deduplication table search control unit may return a result indicating whether the third target data block is stored in the memory to the search agent unit.

[0072] Return to reference Figure 4 The tree traversal core 4213 may include a traversal control unit, a comparison calculation unit, and a position calculation unit. The traversal control unit is responsible for data interaction with the processor 410 and controlling the operation of the tree traversal core 4213.

[0073] Figure 7 is a diagram illustrating an example of a range tree search flow according to an embodiment of the present disclosure.

[0074] Reference Figure 4 and Figure 7, the tree traversal core 4213 may read the current range tree node from the memory in response to receiving a range tree query request from the processor (e.g., a traversal proxy unit). According to an embodiment of the present disclosure, the range tree query request may be generated based on read / write data. The range tree query request may include a query start instruction (find start), the start and end addresses (start, end) of the query space, and the range tree address (rangetree). The tree traversal core 4213 may perform a binary search (including calculation of deviation and comparison) based on the read current range tree node and the start and end addresses of the space targeted by the range tree query request. For example, the binary search may be performed by a position calculation unit. The tree traversal core 4213 may determine the found range tree node as the target node corresponding to the range tree query request in response to finding the range tree node corresponding to the range tree query request through the binary search. On the other hand, the tree traversal core 4213 may read the next range tree node from the memory in response to not finding the range tree node corresponding to the range tree query request through the binary search, use the read next range tree node as the current range tree node, and return to the step of performing the binary search. For example, the range tree node found by the binary search may be compared with the start and end addresses of the query space included in the range tree query request by the comparison calculation unit to determine whether the found range tree node is the range tree node corresponding to the range tree query request. The traversal control unit may return the found range tree node corresponding to the range tree query request to the traversal proxy unit, for example, the position of the range tree node and its value may be returned to the traversal proxy unit.

[0075] As described above, by using near memory processing devices to expand memory capacity, more metadata can be cached, the metadata cache hit rate can be improved, and the IO operations for loading metadata from storage devices can be reduced. At the same time, by offloading memory access-intensive calculations to near memory processing devices, a large amount of data interaction between the CPU and memory can be reduced.

[0076] Figure 8 is a flowchart showing a data processing method executed by a processor according to an embodiment of the present disclosure. Here, the processor may be a reference Figures 4 to 7 Description of the processor.

[0077] Reference Figure 8 In step S801, an indirect block resolution request may be generated in response to an instruction to read / write the first target data block, and the indirect block resolution request may be sent to the near memory processing device; and a checksum update request may be generated in response to an instruction to write the first target data block, and the checksum update request may be sent to the near memory processing device.

[0078] In step S802, a deduplication table query request may be generated in response to an instruction to read / write the second target data block, and the deduplication table query request may be sent to a near memory processing device.

[0079] In step S803 , a range tree query request may be generated in response to an instruction to read / write the third target data block, and the range tree query request may be sent to the near memory processing device.

[0080] Optionally, the data processing method executed by the processor may also include the following steps: receiving the address of the first target data block determined in response to an indirect block resolution request from the near memory processing device, and / or receiving the result of calculating the checksum of the first target data block in response to a checksum update request from the near memory processing device; receiving the result of determining whether the second target data block exists in the near memory processing device in response to a deduplication table query request from the near memory processing device; receiving the range tree node determined in response to a range tree query request from the near memory processing device.

[0081] Fig. 9 is a flowchart showing a data processing method performed in a near memory processing device according to an embodiment of the present disclosure. Here, the near memory processing device may be a reference to Figures 4 to 7 A near memory processing device is described.

[0082] Reference Fig. 9 In step S901, in response to receiving an indirect block resolution request from a processor, an address of a first target data block may be determined from a memory of a near memory processing device, and in response to receiving a checksum update request from the processor, a checksum of a second target data block may be calculated and the calculated checksum may be written to the memory.

[0083] In step S902, in response to receiving a deduplication table query request from the processor, a deduplication table search process may be performed to determine whether a third target data block exists in the memory of the near memory processing device.

[0084] In step S903 , in response to receiving a range tree query request from the processor, a range tree in the memory of the near memory processing device may be traversed to determine a target node corresponding to the range tree query request.

[0085] According to an embodiment of the present disclosure, a memory of a near memory processing device stores ZFS file system metadata, and the ZFS file system metadata includes an index node (dnode), a deduplication table (DDT), and a space map. The index node metadata can be linked to each data block through a multi-level indirect block in the memory of the near memory processing device. In addition, a range tree can be generated in the memory of the near memory processing device according to the space map metadata.

[0086] As described above, the multi-level indirect block includes N-level indirect blocks, N is an integer greater than 1, and the i-th level indirect block includes multiple blocks, i is an integer greater than or equal to 1 and less than N. Each block in the j-th level indirect block includes a block pointer pointing to a corresponding block in the j-1-th level indirect block, including an address and a checksum of the corresponding block in the j-1-th level indirect block, j is an integer greater than 1 and less than or equal to N. Each block in the 1-th level indirect block includes a block pointer pointing to a data block, including an address and a checksum of the data block.

[0087] According to an embodiment of the present disclosure, step S901 may specifically include: in response to receiving an indirect block resolution request from a processor, based on the ID of the first target data block, starting from the Nth level indirect block, loading a block pointer pointing to a corresponding block in the next level indirect block until the first level indirect block; determining the address of the first target data block from the block pointer of the corresponding block in the first level indirect block, and returning the address of the first target data block to the processor, wherein the indirect block resolution request is generated based on reading / writing the first target data block and includes the ID of the first target data block.

[0088] Optionally, step S901 may also specifically include: in response to receiving a checksum update request from the processor, loading a second target data block from the memory of the near memory processing device; calculating a checksum of the second target data block, and writing the calculated checksum to an indirect block stored in the memory of the near memory processing device, wherein the checksum update request is generated based on writing the first target data block.

[0089] According to an embodiment of the present disclosure, step S902 may specifically include: in response to receiving a deduplication table query request from a processor, calculating a first-level hash value according to the fingerprint of the third target data block; loading a deduplication table leaf node from the memory of a near memory processing device according to the first-level hash value; calculating a second-level hash value according to the fingerprint of the third target data block; based on the second-level hash value and the loaded deduplication table leaf node, loading a deduplication table entry from the memory of the near memory processing device; in response to the fingerprint of the third target data block being equal to the key value of the loaded deduplication table entry, determining that the third target data block is stored in the memory of the near memory processing device, wherein the deduplication table query request is generated based on reading / writing the third target data block.

[0090] According to an embodiment of the present disclosure, step S903 may specifically include: in response to receiving a range tree query request from a processor, reading a current range tree node from a memory of a near memory processing device; based on the read current range tree node, performing a binary search according to the start and end addresses of the space targeted by the range tree query request; in response to finding the range tree node corresponding to the range tree query request through the binary search, determining the found range tree node as the target node corresponding to the range tree query request; in response to not finding the range tree node corresponding to the range tree query request through the binary search, reading the next range tree node from the memory of the near memory processing device, using the read next range tree node as the current range tree node, and returning to the step of performing the binary search, wherein the range tree query request is generated based on the read / write data.

[0091] According to the near memory processing device, processor and data processing method of the embodiments of the present disclosure, by utilizing the near memory processing device (e.g., CMM-DC device) to expand the memory capacity, more metadata can be cached, the metadata cache hit rate can be improved, and the IO operations for loading metadata from the storage device can be reduced. At the same time, by offloading memory access intensive calculations to the near memory processing device, a large amount of data interaction between the CPU and the memory can be reduced. In addition, according to the near memory processing device, processor and data processing method of the embodiments of the present disclosure, a solution for improving metadata management can be provided for PB-level large-capacity storage, thereby improving the execution efficiency of the ZFS file system.

[0092] Table 2 shows an example of comparison of the IO times of the storage device of the prior art and the present disclosure.

[0093] Table 2

[0094] In Table 2, memory size indicates the memory size of the host system (eg, memory size), and CMM size indicates the size of a near-memory processing device.

[0095] Table 3 shows an example of comparison of memory access intensive computation times of the prior art and the present disclosure.

[0096] Table 3

[0097] According to an embodiment of the present disclosure, the data processing method executed in the near memory processing device and the data processing method executed in the processor can be written as a computer program and stored on a computer-readable storage medium. When the computer program is executed by the processor, the data processing method as described above is implemented. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. In one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0098] Although the present disclosure includes specific examples, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the spirit and scope of the claims and their equivalents. The examples disclosed herein are to be considered in a descriptive sense and not for purposes of limitation. The description of features or aspects in each example will be considered applicable to similar features or aspects in other examples. Suitable results may be obtained if the described techniques are performed in a different order, and / or if the described systems, structures, devices, or circuits are combined in a different manner and / or replaced or supplemented by other components or their equivalents. Therefore, the scope of the disclosure is not limited by the detailed description, but by the claims and their equivalents, and all changes within the scope of the claims and their equivalents will be considered to be included in the present disclosure.

Claims

1. A near memory processing device, characterized in that: The near memory processing device comprises: Memory; A computing core, the computing core comprising: a block tracking core configured to, in response to receiving an indirect block resolution request from the processor, determine an address of a first target data block from the memory, and, in response to receiving a checksum update request from the processor, calculate a checksum of a second target data block and write the calculated checksum to the memory; a deduplication table search core configured to, in response to receiving a deduplication table query request from the processor, perform a deduplication table search process to determine whether a third target data block exists in the memory; The tree traversal core is configured to, in response to receiving a range tree query request from the processor, traverse the range tree in the memory to determine a target node corresponding to the range tree query request.

2. The near memory processing device according to claim 1, characterized in that The memory stores ZFS file system metadata, wherein the ZFS file system metadata includes an index node (dnode), a deduplication table (DDT), and a space map. The index node metadata is linked to each data block through multi-level indirect blocks in the memory. Therein, a range tree is generated in a memory according to the spatial metadata.

3. The near memory processing device according to claim 2, characterized in that The multi-level indirect block includes N-level indirect blocks, N is an integer greater than 1, the i-th level indirect block includes a plurality of blocks, i is an integer greater than or equal to 1 and less than N, Each block in the j-th level indirect block includes a block pointer pointing to a corresponding block in the j-1-th level indirect block, including an address and a checksum of the corresponding block in the j-1-th level indirect block, where j is an integer greater than 1 and less than or equal to N. Each block in the level 1 indirect block includes a block pointer to a data block, which includes the address and checksum of the data block.

4. The near memory processing device according to claim 3, characterized in that The block tracking core is configured as: In response to receiving an indirect block resolution request from the processor, based on the ID of the first target data block, starting from the Nth level indirect block, loading a block pointer pointing to a corresponding block in the next level indirect block until the first level indirect block; determining an address of a first target data block from a block pointer of a corresponding block in the first level indirect block, and returning the address of the first target data block to the processor, The indirect block resolution request is generated based on reading / writing the first target data block and includes the ID of the first target data block.

5. The near memory processing device according to claim 3, characterized in that The block tracking core is configured as: In response to receiving the checksum update request from the processor, loading a second target data block from the memory; calculating a checksum of the second target data block and writing the calculated checksum to an indirect block stored in the memory, The checksum update request is generated based on writing the second target data block.

6. The near memory processing device according to claim 2, characterized in that The deduplication table search core is configured as: In response to receiving a deduplication table query request from the processor, calculating a first-level hash value according to a fingerprint of the third target data block; According to the first-level hash value, load the deduplication table leaf node from the memory; Calculate a second level hash value based on the fingerprint of the third target data block; Loading a deduplication table entry from a memory based on the second level hash value and the loaded deduplication table leaf node; In response to the fingerprint of the third target data block being equal to the key value of the loaded deduplication table entry, determining that the third target data block is stored in the memory, The deduplication table query request is generated based on reading / writing the third target data block.

7. The near memory processing device according to claim 2, characterized in that The tree traversal core is configured as: In response to receiving a range tree query request from a processor, reading a current range tree node from a memory; Based on the read current range tree node, perform a binary search according to the start and end addresses of the space targeted by the range tree query request; In response to finding a range tree node corresponding to the range tree query request through a binary search, determining the found range tree node as a target node corresponding to the range tree query request; In response to not finding the range tree node corresponding to the range tree query request through the binary search, reading the next range tree node from the memory, taking the read next range tree node as the current range tree node, and returning to the step of performing the binary search, The range tree query request is generated based on reading / writing data.

8. A processor, characterized in that: The processor comprises: The index node management module is configured to: generate an indirect block resolution request in response to an instruction to read / write the first target data block, and send the indirect block resolution request to the near memory processing device; generate a checksum update request in response to an instruction to write the first target data block, and send the checksum update request to the near memory processing device; a deduplication table management module, configured to: generate a deduplication table query request in response to an instruction to read / write the second target data block, and send the deduplication table query request to the near memory processing device; The space management module is configured to: generate a range tree query request in response to an instruction to read / write the third target data block, and send the range tree query request to the near memory processing device.

9. The processor according to claim 8, wherein: The index node management module is further configured to: receive from the near memory processing device an address of the first target data block determined in response to the indirect block resolution request, and / or receive from the near memory processing device a result of calculating a checksum of the first target data block in response to the checksum update request; The deduplication table management module is further configured to: receive from the near memory processing device a result of determining whether the second target data block exists in the near memory processing device in response to the deduplication table query request; The space management module is further configured to receive, from the near memory processing device, a range tree node determined in response to the range tree query request.

10. A data processing method executed in a near memory processing device, characterized in that: The data processing method comprises: In response to receiving an indirect block resolution request from the processor, determining an address of a first target data block from the memory, and in response to receiving a checksum update request from the processor, calculating a checksum of a second target data block and writing the calculated checksum to the memory; In response to receiving a deduplication table query request from the processor, performing a deduplication table search process to determine whether a third target data block exists in the memory; In response to receiving a range tree query request from the processor, the range tree in the memory is traversed to determine a target node corresponding to the range tree query request.

11. A data processing method executed by a processor, characterized in that: The data processing method comprises: generating an indirect block resolution request in response to an instruction to read / write the first target data block, and sending the indirect block resolution request to the near memory processing device; generating a checksum update request in response to an instruction to write the first target data block, and sending the checksum update request to the near memory processing device; generating a deduplication table query request in response to an instruction to read / write a second target data block, and sending the deduplication table query request to a near memory processing device; In response to an instruction to read / write the third target data block, a range tree query request is generated and sent to the near memory processing device.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data processing method according to claim 10 or 11 is implemented.

13. A host system, characterized in that: The host system comprises: A near memory processing device as claimed in any one of claims 1 to 7; A processor as claimed in claim 8 or 9; Memory.