Systems and methods for collecting tracking data via memory device
By allocating private memory in memory devices and tracking physical memory addresses in real-time, the problem of delay and resource waste of host devices in virtual address translation is solved, and memory access efficiency and management efficiency are improved.
Patent Information
- Application Number
- CN202510081052.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-09
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, it is difficult for host devices to efficiently track and manage the physical memory addresses of memory devices, resulting in delays and waste of computing resources, especially during virtual address conversion.
By allocating private memory in memory devices, utilizing hardware solutions to track physical memory addresses in real time and store them in a private portion of volatile memory, providing a direct access mechanism for use by host devices.
It realizes the reduction of memory access latency, improves the performance of memory devices and the management efficiency of host devices, and provides accurate information on data access patterns to make optimization decisions.
Smart Images

Figure CN120371202A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority and the benefit of U.S. Provisional Application No. 63 / 624,422, filed on January 24, 2024, entitled "Mechanisms for Collecting and Filtering Host Physical Addresses (HPAs) on CXL Memory Devices for AI Inference", the entire content of which is incorporated herein by reference. Technical Field
[0003] One or more aspects of embodiments in accordance with the present disclosure relate to memory devices, and more particularly to the collection by a host device of trace data accumulated by a memory device. Background Art
[0004] Applications running on a host computing device may need to read data and write data to memory. As the amount of data read from and written to memory increases, the need for storage devices and memory, as well as for efficiently retrieving data from storage devices and memory, may also increase.
[0005] The above information disclosed in this background art section is only for enhancing the understanding of the background of the present disclosure, and thus may include information that does not form the prior art. Summary of the Invention
[0006] One or more embodiments of the present disclosure relate to a method, including: receiving, by a memory device, a command from a computing device; identifying, by the memory device, first data associated with the command; storing, by the memory device, the first data in a first portion of volatile memory of the memory device for reading by the computing device; and accessing, by the memory device, a second portion of the volatile memory, wherein the second portion of the volatile memory is configured to store a copy of second data stored in non - volatile memory in the memory device.
[0007] According to one or more embodiments, the command is for reading the second data from the non - volatile memory or writing the second data to the non - volatile memory, and the first data identifies a physical address associated with the second data.
[0008] According to one or more embodiments, the accessing of the second portion of the volatile memory is based on detecting the command.
[0009] According to one or more embodiments, the command is for the memory device to perform a calculation, and the first data includes the result of the calculation.
[0010] According to one or more embodiments, the reading performed by the computing device is based on detecting a trigger, where the trigger includes a signal generated by the memory device. The method further includes: detecting, by the memory device, a fullness of the first portion of the volatile memory; and generating, by the memory device, the signal based on the detection of the fullness of the first portion of the volatile memory.
[0011] According to one or more embodiments, the first portion of the volatile memory is mapped to a first physical address space of the computing device and is accessed by the computing device via a memory access operation.
[0012] According to one or more embodiments, the memory device is configured to send the second data to the computing device based on accessing the second portion of the volatile memory.
[0013] According to one or more embodiments, the computing device is configured to take an action based on the first data.
[0014] According to one or more embodiments, the action includes reconfiguring the second portion of the volatile memory to increase a cache hit rate.
[0015] According to one or more embodiments, reconfiguring the second portion includes modifying a cache algorithm.
[0016] One or more embodiments of the present disclosure also relate to a memory device, including: a controller; a volatile memory; and a non-volatile memory. The controller is configured to: receive a command from a computing device; identify first data associated with the command; store the first data in a first portion of the volatile memory for the computing device to read; and access a second portion of the volatile memory, where the second portion of the volatile memory is configured to store a copy of second data stored in the non-volatile memory.
[0017] These and other features, aspects, and advantages of the embodiments of the present disclosure will be more fully understood when considered in conjunction with the following detailed description, the appended claims, and the drawings. Of course, the actual scope of the invention is defined by the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The non-limiting and non-exhaustive embodiments of the present embodiment are described with reference to the following drawings, where like reference numerals refer to like parts throughout the views unless otherwise specified.
[0019] Figure 1 A block diagram of a system for collecting trace data on a memory device according to one or more embodiments is depicted;
[0020] Figure 2 depicts a block diagram of a memory device in accordance with one or more embodiments;
[0021] Figure 3 depicts a block diagram of the representation and management of a memory device by a host processor in accordance with one or more embodiments;
[0022] Figure 4 depicts a block diagram of trace data collected and stored in a private memory in accordance with one or more embodiments;
[0023] Figure 5 depicts a flowchart of a process for collecting trace data via a memory device; and
[0024] Figure 6 depicts a flowchart of a process performed by a collection and decision engine for collecting trace data and making decisions based on the collected trace data in accordance with one or more embodiments. DETAILED DESCRIPTION
[0025] Hereinafter, example embodiments will be described in more detail with reference to the accompanying drawings, in which like reference numerals always denote like elements. However, the present disclosure may be implemented in various different forms and should not be construed as limited to the embodiments shown herein. Instead, these embodiments are provided as examples so that the present disclosure will be sufficient and complete and will fully convey the aspects and features of the present disclosure to those skilled in the art. Accordingly, processes, elements, and techniques that are not necessary for those of ordinary skill in the art to fully understand the aspects and features of the present disclosure may not be described. Unless otherwise noted, like reference numerals denote like elements in the drawings and the written description, and thus, their description may not be repeated. Further, in the drawings, relative dimensions of elements, layers, and regions may be exaggerated and / or simplified for clarity.
[0026] Embodiments of the present disclosure will be described below with reference to block diagrams and flowcharts. Therefore, it should be understood that each block of the block diagrams and flowcharts can be implemented in the form of a computer program product, a complete hardware embodiment, a combination of hardware and a computer program product, and / or a device, system, computing device, computing entity, etc. that executes instructions, operations, steps, and similar terms that can be used interchangeably (e.g., executable instructions, instructions for execution, program code, etc.) on a computer-readable storage medium. For example, the retrieval, loading, and execution of code can be performed sequentially such that one instruction is retrieved, loaded, and executed at a time. In some example embodiments, the retrieval, loading, and / or execution can be performed in parallel such that multiple instructions are retrieved, loaded, and / or executed together. Thus, such embodiments can generate a specially configured machine that executes the steps or operations specified in the block diagrams and flowcharts. Therefore, the block diagrams and flowcharts support various combinations of embodiments for executing the specified instructions, operations, or steps.
[0027] An application can perform calculations on large amounts of data. As such types of calculations increase, the demand for memory may also increase. Memory expansion techniques can help alleviate this problem by providing a hierarchical memory subsystem that helps increase the memory capacity at a lower cost. The memory subsystem can include, for example, memory devices that follow the Compute Express Link (CXL) protocol. The memory devices can include volatile memory (e.g., dynamic random access memory (DRAM)) and non-volatile memory (NVM). An application running on a host device can read and write data to the memory devices via load and store commands and treat the memory devices as an extension of the main memory attached to the processor.
[0028] There may be latency involved in accessing memory devices in a hierarchical memory subsystem. The performance of a memory device can depend on, for example, how efficiently the memory device can use its volatile memory to cache data. Therefore, it may be desirable to track those accesses when an access to a physical memory address of a memory device occurs for controlling, for example, the data cached in the volatile memory. However, it may be impractical for the host to track accesses via a software solution. For example, an application accessing a memory device uses virtual addresses to read data from and write data to the memory device. To obtain the physical address associated with a virtual address, the host OS may need to convert the virtual address to a physical memory address via, for example, a page table. The conversion of a virtual address to a physical address may introduce undesirable latency and consume additional computing resources of the host when implementing a software solution.
[0029] Tracking memory access via a hardware solution (e.g., via the memory device itself) can help avoid some of the latency and cost associated with software solutions for tracking memory access. For example, a hardware solution can avoid translating virtual addresses to physical addresses as part of the tracking process. However, memory devices are typically exposed as host-managed device memory (HDM) for use and management by a host operating system (OS) in the same manner as main memory. Thus, any page of the memory device that could be used to store the tracked physical memory addresses may not be available at a given time because the OS can allocate the page to another host application.
[0030] Generally, embodiments of the present disclosure relate to collecting physical memory addresses (commonly referred to as trace data) on a memory device as a hardware solution and providing access to the trace data to a host for making management decisions. In some embodiments, a portion of the volatile memory of the memory device is allocated as private memory. In some embodiments, the host maps the private memory as a direct access (DAX) region at runtime. The memory device may include a trace collection engine configured to identify accesses by an application to physical memory addresses of the memory device during runtime. Accesses to the memory addresses may be via load and store commands. The trace collection engine may store the identified memory addresses in the private memory.
[0031] In some embodiments, a host (e.g., an application running on the host) accesses the trace data (e.g., physical addresses) in the private memory. The host may access the trace data in response to detecting a trigger. The trigger may be, for example, a signal from the memory device indicating that trace data is available for the host. The host may retrieve the trace data by accessing (e.g., directly accessing) the private memory using a load and store interface. Real-time access to the trace data provides the host with accurate information about the data access patterns of the application. The host may make decisions based on the trace data to improve the performance of the application and / or the memory device. For example, the host may make prefetching and / or other cache management decisions to reduce latency when accessing the memory device.
[0032] Although the trace data according to various embodiments is described as physical memory addresses, various embodiments are not limited thereto and may include other types of data generated by the memory device. For example, the memory device may be configured to perform various types of computations, including machine learning computations for inference and / or training. The trace data stored in the private memory may be the result of such computations for providing the host with direct real-time access to the results.
[0033] Figure 1A block diagram depicting a system for collecting trace data on a memory device according to one or more embodiments. The system includes a host computing device (referred to as the "host") 100, which is coupled to one or more endpoints, such as one or more memory devices 102a - 102c (collectively 102).
[0034] The host 100 includes, but is not limited to, a processor 105, a main memory 104, and a root complex (RC) interface 112. The processor 105 may include one or more central processing unit (CPU) cores 116, which are configured to execute computer program instructions and process data stored in a cache memory (referred to simply as "memory" or "cache") 118. The cache 118 may be dedicated to one of the CPU cores 116 or shared among the individual CPU cores. It should be understood that although CPUs are used to describe various embodiments, those skilled in the art will recognize that GPUs or other computing units may be used in place of or in addition to CPUs.
[0035] The cache 118 may be coupled to a memory controller 120, which in turn is coupled to the main memory 104. The main memory 104 may include, for example, dynamic random access memory (DRAM) that stores computer program instructions and / or other types of data similar to the memory devices 102 (collectively data). To enable the CPU cores 116 to execute instructions or retrieve data provided by the memory devices 102, the corresponding data may be loaded into the cache 118, and the CPU cores may consume the data from the cache memory (e.g., directly). If the data to be consumed is not yet in the cache 118, a cache miss may occur, and it may be necessary to query the memory devices 102 to load the data. For example, if the data to be consumed is not in the cache 118, the cache miss logic may query the data from the memory (e.g., the main memory (e.g., DRAM) 104 or the memory devices 102) based on the mapped virtual or physical address.
[0036] In some embodiments, the processor 105 (e.g., an application running on the processor) generates data access requests to the memory devices 102. One or more of the data access requests may include a virtual memory address of the location to write or read data. The processor 105 may call a memory management unit (MMU) 108 to convert the virtual address to a physical address for processing the request. Although the MMU 108 is Figure 1is depicted as part of the processor 105, but those skilled in the art should recognize that the MMU 108 can be a device / circuit separate from the processor 105. The MMU 108 can include a translation table 110 that maps virtual addresses to physical addresses. In some embodiments, the MMU 108 is stored in the main memory 104. A request sent to the memory device 102 to fulfill a data access request can include a physical address corresponding to a virtual address.
[0037] In some embodiments, the host 100 exchanges signals or messages with the memory device 102 via the RC interface 112 and the interface connections 106a - 106c (collectively 106). For example, the host 100 can send requests (e.g., load or store requests) for reading data from or writing data to the memory device 102 through the RC interface 112 and the interface connections 106. A message from the memory device 102 to the host 100, such as a response to a request from the host, can be transmitted through the interface connections 106 to the RC interface 112, which in turn transmits the response to the processor 105. The memory device 102 can also send signals to the host 100 through the RC interface 112 and the interface connections 106 that include, for example, certain types of notifications.
[0038] In some embodiments, the interface connections 106 (e.g., the connector and its protocol) include a memory expansion bus, such as, for example, Compute Express Link (CXL), but the embodiments are not limited thereto. For example, the interface connections 106 (e.g., the connector and its protocol) can also include a general interface, such as Ethernet, Universal Serial Bus (USB), etc. In some embodiments, the interface connections 106 can include (or can conform to) a cache coherent interconnect for accelerators (CCIX), a dual in-line memory module (DIMM) interface, a small computer system interface (SCSI), Fibre Channel, Serial Attached SCSI (SAS), iWARP protocol, InfiniBand protocol, 5G wireless protocol, Wi-Fi protocol, Bluetooth protocol, etc.
[0039] The RC interface 112 can be, for example, a CXL interface that is configured to implement a root complex for connecting the processor 105 and the main memory 104 to the memory device 102. The RC interface 112 can include one or more ports 114a - 114c to connect one or more memory devices 102 to the RC. In some embodiments, the MMU 108 and / or the translation table 110 can be integrated into the RC interface 112 to allow the RC interface to implement address translation.
[0040] The memory device 102 may include one or more of volatile computer-readable storage media and / or non-volatile computer-readable storage media. In some embodiments, one or more of the memory devices 102 include memory attached to a CPU or GPU, such as, for example, a CXL-attached memory device (including volatile and persistent memory devices), an RDMA-attached memory device, etc., but the embodiments are not limited thereto. A CXL-attached memory device (referred to simply as a CXL memory) may follow the CXL.mem protocol, where the host 100 may access the memory using commands such as load and store commands. In this regard, the host 100 may act as a requester, and the CXL memory may act as a subordinate.
[0041] In some embodiments, the memory device 102 is included in a memory system that allows memory tiering to deliver an appropriate cost or performance profile. In this regard, different types of storage media may be organized in a memory hierarchy or tier based on the characteristics of the storage media. The characteristic may be access latency. In some embodiments, the tier or level of the memory device increases as the access latency decreases. In some embodiments, when data is to be retrieved before querying a higher-level storage media, the application may query the storage media with the lowest tier.
[0042] In some embodiments, one or more of the memory devices 102 are memory devices of the same or different types that are aggregated into a storage pool. For example, the storage pool may include one or more CPU or GPU attached memory devices.
[0043] In some embodiments, one or more of the memory devices 102 are configured to monitor and collect certain types of data (hereinafter referred to as trace data) at runtime (e.g., during the execution of an application). The trace data may include, for example, the physical addresses of the memory device 102 requested by the host 100, the results of computations performed by the memory device 102, etc.
[0044] In some embodiments, the host 100 includes a collection and decision (C&D) engine 124 that is configured to retrieve the trace data collected by the memory device 102. The C&D engine 124 may be implemented via software, firmware, or hardware or a combination of software, firmware, and / or hardware. The software, firmware, and / or hardware may be part of (or executed by) the processor 105.
[0045] In some embodiments, the C&D engine 124 is configured to make decisions based on the retrieved trace data. Example decisions can include, for example, managing data stored in the cache 118, main memory 104, and / or memory device 102 (collectively referred to as storage media), generating prefetch instructions, configuring or reconfiguring the memory device 102, etc. For example, the C&D engine 124 can manage the data in the memory device 102 by promoting or demoting different levels of pages in and out of the memory device. In this regard, the C&D engine 124 can analyze the trace data for the physical addresses accessed by the application. The C&D engine 124 can determine, based on the trace data, that certain physical addresses are accessed more frequently than other physical addresses and should therefore be retained in the cache of the memory device 102. In other examples, the C&D engine 124 can modify the cache algorithm (e.g., cache replacement algorithm) based on the trace data.
[0046] Figure 2 A block diagram of a memory device 102 in accordance with one or more embodiments is depicted. In some embodiments, the memory device 102 includes a storage controller 200, a storage memory 202, and a non-volatile memory (NVM) 204. The storage controller 200 can be configured to exchange commands and / or data with the RC interface 112 via the interface connection 106. In this regard, the storage controller 120 can include at least one processor or processing component embedded thereon for interfacing with the host 100, the storage memory 202, and the NVM 204. The processing component can include, for example, digital circuitry (e.g., a microcontroller, a microprocessor, a digital signal processor, or a logic device (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or the like)) that is capable of executing data access instructions (e.g., via firmware and / or software) to provide access to and from data stored in the storage memory 202 or the NVM 204 in accordance with the data access instructions.
[0047] The storage memory 202 can be a high-performance memory of the memory device 102 and can include (or can be) volatile memory, such as, for example, DRAM, but the present disclosure is not limited thereto, and the storage memory 202 can be any suitable kind of high-performance volatile or non-volatile memory. Although a single storage memory 202 is depicted for simplicity, those skilled in the art should recognize that the memory device 102 can include other local memories for temporarily storing other data for the storage device.
[0048] In some embodiments, the storage memory 202 is partitioned into a first portion (hereinafter referred to as “private memory”) 206 and a second portion (hereinafter referred to as “cache memory”) 208 isolated from the first portion. For example, the storage memory 202 may have a total capacity of 128 GB. A first quantity of the total capacity (e.g., 8 GB) may be allocated or reserved as the private memory 206, and a second quantity of the total capacity (e.g., 120 GB) may be allocated or reserved as the cache memory 208.
[0049] The private memory 206 may be exposed to the processor 105 as a persistent memory residing on the memory bus. In this regard, the private memory 206 emulates a persistent memory. The host 100 (e.g., the C&D engine 124) may access (e.g., directly access) the private memory 206 through the memory interface 210 without the need for the MMU 108 to perform address translation. In some embodiments, the host 100 uses a memory access operation (e.g., a memory read operation) to access the private memory 206.
[0050] In some embodiments, the direct access to the private memory 206 is via a direct access (DAX) mechanism that may otherwise be used to directly access files stored in persistent memory. When using the DAX mechanism, a persistent memory-aware file system can recognize that a file is stored in persistent memory and map the persistent memory directly into the application's address space. The application can read file data and write file data to the persistent memory without the host OS (e.g., the memory controller 120) caching the file in the cache 118.
[0051] In some embodiments, the cache memory 208 is used to cache a copy of the data stored in the NVM 204. In this regard, the cache memory 208 may store a copy of the data stored in the NVM 204. For example, data that will be accessed by an application in the near future may be copied from the NVM 204 to the cache memory 208 to allow retrieval of the data from the cache memory rather than the NVM 204. In this regard, the trace data stored in the private memory 206 may be used to prefetch data into the cache memory 208. The host 100 may use the DAX mechanism to access the trace data to make prefetch decisions.
[0052] In some embodiments, the cache memory 208 has a lower access latency than the NVM 204. Thus, in some embodiments, accessing data from the cache memory 208 helps improve overall system performance and responsiveness.
[0053] In some embodiments, cache memory 208 and NVM 204 operate under a non-uniform memory access (NUMA) memory model and participate in a hierarchical memory subsystem. In this regard, cache memory 208 and NVM 204 can form NUMA nodes that serve as the last-level cache for host processor 105.
[0054] In some embodiments, NVM 204 persistently stores data received, for example, from host 100. NVM 204 can include, for example, one or more NAND flash memories, but the present disclosure is not limited thereto, and NVM 204 can include any suitable type of memory for persistently storing data depending on the implementation of memory device 102 (e.g., disks, tapes, optical discs, etc.).
[0055] In some embodiments, storage controller 200 is configured with a trace collection engine (hereinafter referred to as the trace collector) 212 for capturing data access requests from host 100 (e.g., by an application executed by processor 105). Trace collector 212 can be implemented via software, firmware (e.g., ASIC), hardware, or a combination of software, firmware, and / or hardware. For example, trace collector 212 can be an FPGA block. In another example, trace collector 212 can be software executed by storage controller 200.
[0056] Trace collector 212 can identify load or store requests from the host, as well as the host physical address included in the request. In some embodiments, trace collector 212 stores the identified host physical address as trace data (e.g., sequentially) in private memory 206. The trace data stored in private memory 206 can be directly accessed by host processor 105 (e.g., via the DAX mechanism) without address translation or caching into CPU cache 118.
[0057] In some embodiments, the trace collector 212 participates in filtering and / or sampling of the trace data for storing the filtered or sampled trace data in the private memory 206. For example, if the host 100 accesses the same host physical address multiple times, the trace collector 212 may store the physical address in the private memory 206 once and include additional data associated with the stored physical address to indicate the number of times the address has been accessed. In another example, if a repeating pattern of physical address accesses is detected (e.g., addresses 1, 3, 5, and 7 are accessed multiple times), the pattern may be recorded once along with a value indicating the number of times the pattern is repeated. In another example, if the access is to sequential memory addresses (increasing or decreasing in sequence), the starting physical address and the length of the sequence may be recorded. However, embodiments are not limited thereto, and the trace collector 212 may employ other data compression mechanisms to store data in the private memory 206.
[0058] In some embodiments, the storage controller 200 is configured to perform one or more computing functions. The computing functions may include, for example, computations for training or inference via a machine learning model. The computations may also involve encryption, decryption, compression, decompression, etc. In some embodiments, the results of the computing functions are stored in the private memory 206 for direct access by the host processor 105.
[0059] In some embodiments, the trace collector 212 monitors conditions for signaling the host 100 to access the collected trace data. The conditions may include, for example, the fullness of the private memory 206. For example, the trace collector 212 may signal the host 100 to access the collected trace data based on detecting that the private memory 206 has been filled to a threshold percentage of its allocated capacity (e.g., 100%, 95%, etc.). In this regard, the trace collector 212 may write a flag in a configuration register notifying the host 100 that the trace data is available. The host 100 (e.g., the C&D engine 124) may poll the configuration register periodically (e.g., on a regular or irregular basis) via, for example, the CXL.io protocol to check for the flag. The C&D engine may retrieve the trace data based on detecting the flag.
[0060] Figure 3Depicts a block diagram of the representation and management of a memory device by a host processor 105 in accordance with one or more embodiments. In some embodiments, the memory device 102 is pre-configured such that a first amount of the storage memory is set as the private memory 206, and a second amount of the storage memory is set as the cache memory 208. For example, if the total capacity of the storage memory 202 is 128 GB, 8 GB of the storage memory may be allocated or designated as the private memory 206, and the remaining 120 GB of the storage memory may be allocated or designated as the cache memory 208. In some embodiments, the allocation or reallocation of the private memory and the cache memory may be configured by the host 100.
[0061] The memory device 102 may communicate its storage capacity and storage type to the host (e.g., during the startup of the memory device 102). In some embodiments, the memory device 102 communicates a first amount of the private memory 206 as persistent memory that can be accessed via DAX (e.g., directly). The memory device 102 may also communicate a second amount of the cache memory 208 as a NUMA node. The capacity of the NVM 204 (e.g., the NUMA node capacity) may also be communicated to the host 100.
[0062] The host OS 300 may map the memory device 102 to the host physical address space 301 based on the communicated information. In some embodiments, the persistent memory (pmem) aware file system 308 is configured to identify the private memory 206 as emulated persistent memory and map the private memory to a first host managed device memory (HDM) region (hereinafter referred to as the DAX region 302). The size of the DAX region 302 may be equal to the allocated size of the private memory 206.
[0063] In some embodiments, the memory management system 310 is configured to identify the NVM 204 as a NUMA node and map the NVM 204 to a second HDM region of the host physical address space (hereinafter referred to as the NUMA region 304). The size of the NUMA region 304 may be equal to the size of the NVM 204. The host physical address space 301 may also include other types of memory, such as host DRAM 306 (which may be similar to the main memory 104).
[0064] In some embodiments, the pmem aware file system 308 allows direct access to the addresses mapped to the DAX region 302 to retrieve the trace data stored in the private memory 206. The access may be via a memory read operation of the private memory 206 via the memory interface 210. The read operation may be via a read API, such as a C programming language API, as follows:
[0065] read(cxl_dev, offset, length)
[0066] Where "cxl_dev" identifies the memory device 102, "offset" identifies the address of the private memory 206 that has been mapped to the DAX region 302, and "length" identifies the amount of data to be read. The length can correspond to all of the trace data or a subset of the trace data in the private memory 206.
[0067] In some embodiments, the memory management system 310 can dynamically allocate and deallocate memory addresses in the NUMA region 304 at runtime to an application. The application can use virtual addresses to perform data access operations. The memory management system 310 can call the MMU 108 to translate the virtual address into a physical address in the NUMA region 304 and forward the load or store command to the physical address for processing by the memory device 102. The memory device 102 can check the cache memory 208 to determine whether the requested data is in the cache memory. In the case where the requested data is not in the cache memory 208, the memory device 102 can access the NVM 204.
[0068] Figure 4 A block diagram depicting trace data 400 collected and stored in the private memory 206 according to one or more embodiments. The trace data 400 can include the type 402 of access (e.g., read or write access) received from an application on the host 100, and the physical address 404 accessed by the application. In some embodiments, the trace data includes one or more bits 406 for indicating whether the access results in a "hit" or "miss". For example, if the accessed address is in the cache memory 208, it can be identified as a "hit", and if the accessed address is not in the cache memory 208, it can be identified as a "miss".
[0069] Figure 5 A flowchart depicting a process for collecting trace data via the memory device 102 according to one or more embodiments. The process begins, and the memory device 102 (e.g., the storage controller 200) receives a first command from a computing device (e.g., the host 100). The first command can be a data calculation command, a load or store command, etc.
[0070] In operation 504, the trace collector 212 identifies first data associated with the first command. The first data can include, for example, the physical address included in the load or store command. In some embodiments, the first data includes the result of a calculation performed by the storage controller 200.
[0071] In operation 506, the trace collector 212 stores the first data in a first portion of the volatile memory of the memory device 102. For example, the first portion may be the private memory 206 of the storage memory 202 in the memory device 102. The first data may be stored sequentially in the private memory 206.
[0072] In an embodiment where the first command is a load or store command, the storage controller 200 accesses a second portion of the volatile memory of the memory device 102. For example, the second portion may be the cache memory 208.
[0073] A load or store command may be generated in response to an application sending a read or write request to a virtual memory address. The MMU 108 may translate the virtual memory address into a physical address. The physical address may be an address mapped to the NUMA region 304. The root complex interface 112 may send the load or store command including the physical address to the memory device 102 for processing. The storage controller 200 may check the cache memory 208 to determine whether the requested data exists in the cache memory. If the data exists in the cache memory 208, the request may be fulfilled based on the data in the cache memory without accessing the NVM 204.
[0074] Figure 6 A flowchart depicting a process for collecting trace data and making decisions based on the collected trace data, performed by the C&D engine 124, in accordance with one or more embodiments is shown. The process begins, and in operation 600, the C&D engine 124 determines whether a trigger has been detected. The trigger may be, for example, a flag in a configuration register that indicates that trace data is available for collection.
[0075] If a trigger is detected, then in operation 602, the C&D engine 124 retrieves the trace data from the private memory 206. In some embodiments, the C&D engine 124 retrieves the trace data by accessing the address in the DAX region 302 that is mapped to the private memory 206 and performing a read operation on the address. The read operation may allow the trace data to be retrieved (e.g., directly) from the private memory 206 via the memory interface 210 without any address translation by the MMU 108.
[0076] In operation 604, the C&D engine 124 processes the trace data to make a decision. For example, the C&D engine 124 may decide that the cache memory 208 should be reconfigured to improve the cache hit rate. In one example, the C&D engine 124 may detect that the miss rate is higher than a set threshold. In such a case, the C&D engine 124 may decide that the current cache algorithm, such as the cache replacement algorithm, prefetch algorithm, etc., should be modified to improve the cache hit rate.
[0077] In another example, the C&D engine 124 can detect patterns of physical memory addresses accessed by an application by analyzing trace data. For example, the C&D engine 124 can detect a group of addresses that are frequently accessed together. In such a case, the C&D engine 124 can decide that the data corresponding to the detected group of addresses should be retained in the cache memory 208 without being evicted.
[0078] In operation 606, the C&D engine 124 can send a message to the memory device 102 based on the decision made in operation 604. For example, the message can include a command to reconfigure the cache memory 208 (e.g., switch the cache algorithm from a least recently used (LRU) cache algorithm to a clock-based cache algorithm). In some embodiments, the message includes program instructions for execution by the storage controller 200 to effect the reconfiguration.
[0079] One or more embodiments of the present disclosure (e.g., the C&D engine 124, the storage controller 200, and / or the trace collector 212) can be implemented in one or more processors. The term processor can refer to one or more processors and / or one or more processing cores. One or more processors can be hosted in a single device or distributed across multiple devices (e.g., in a cloud system). The processor can include, for example, an application specific integrated circuit (ASIC), a general or special purpose central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a programmable logic device such as a field programmable gate array (FPGA). In a processor, as used herein, each function is performed by hardware configured (i.e., hardwired) to perform that function, or by more general hardware such as a CPU configured to execute instructions stored in a non-transitory storage medium (e.g., memory). The processor can be fabricated on a single printed circuit board (PCB) or distributed across several interconnected PCBs. The processor can include other processing circuitry; for example, the processing circuitry can include two processing circuits, an FPGA and a CPU, interconnected on a PCB.
[0080] It should be understood that although terms such as "first", "second", "third", etc. may be used herein to describe various elements, components, regions, layers, and / or sections, these elements, components, regions, layers, and / or sections should not be limited by these terms. These terms are only used to distinguish one element, component, region, layer, or section from another. Thus, a first element, component, region, layer, or section discussed herein can be termed a second element, component, region, layer, or section without departing from the spirit and scope of the inventive concept.
[0081] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the inventive concept. Further, unless explicitly stated, the embodiments described herein are not mutually exclusive. Aspects of the embodiments described herein may be combined in some implementations.
[0082] As used herein, the singular forms "a" and "an" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when the terms "comprises" and / or "comprising" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. Expressions such as "at least one of..." modify the entire list of elements when preceding the list of elements and do not modify the individual elements of the list. Further, when describing embodiments of the inventive concept, the use of "may" means "one or more embodiments of the present disclosure". Additionally, the term "exemplary" is intended to refer to an example or illustration. As used herein, the terms "use", "using", and "used" may be considered synonymous with the terms "utilize", "utilizing", and "utilized", respectively.
[0083] Although exemplary embodiments of systems and methods for collecting trace data via a memory device have been specifically described and illustrated herein, many modifications and variations will be apparent to those skilled in the art. Accordingly, it should be understood that systems and methods for collecting trace data constructed in accordance with the principles of the present disclosure may be embodied differently than specifically described herein. The present disclosure is also defined in the appended claims and their equivalents.
[0084] Systems and methods for collecting trace data via a memory device may include one or more combinations of the features set forth in the following statements.
[0085] Statement 1. A method, comprising: receiving, by a memory device, a command from a computing device; identifying, by the memory device, first data associated with the command; storing, by the memory device, the first data in a first portion of volatile memory of the memory device for reading by the computing device; and accessing, by the memory device, a second portion of the volatile memory, wherein the second portion of the volatile memory is configured to store a copy of second data stored in non-volatile memory in the memory device.
[0086] Statement 2. The method according to Statement 1, wherein the command is for reading the second data from the non-volatile memory or writing the second data to the non-volatile memory, and the first data identifies a physical address associated with the second data.
[0087] Statement 3. The method according to Statement 2, wherein access to the second portion of the volatile memory is based on detecting the command.
[0088] Statement 4. The method according to Statement 1, wherein the command is for performing a calculation by the memory device, and the first data includes a result of the calculation.
[0089] Statement 5. The method according to Statement 1, wherein the reading by the computing device is based on detecting a trigger, wherein the trigger includes a signal generated by the memory device, and the method further includes: detecting, by the memory device, a fullness of the first portion of the volatile memory; and generating, by the memory device, the signal based on the detecting of the fullness of the first portion of the volatile memory.
[0090] Statement 6. The method according to Statement 1, wherein the first portion of the volatile memory is mapped to a first physical address space of the computing device and is accessed by the computing device via a memory access operation.
[0091] Statement 7. The method according to Statement 1, wherein the memory device is configured to send the second data to the computing device based on accessing the second portion of the volatile memory.
[0092] Statement 8. The method according to Statement 1, wherein the computing device is configured to take an action based on the first data.
[0093] Statement 9. The method according to Statement 8, wherein the action includes reconfiguring the second portion of the volatile memory to increase a cache hit rate.
[0094] Statement 10. The method according to Statement 9, wherein the reconfiguring the second portion includes modifying a cache algorithm.
[0095] Statement 11. A memory device, comprising: a controller; volatile memory; and non-volatile memory, wherein the controller is configured to: receive a command from a computing device; identify first data associated with the command; store the first data in a first portion of the volatile memory for the computing device to read; and access a second portion of the volatile memory, wherein the second portion of the volatile memory is configured to store a copy of second data stored in the non-volatile memory.
[0096] Statement 12. The memory device according to Statement 11, wherein the command is for reading the second data from the non-volatile memory or writing the second data to the non-volatile memory, and the first data identifies a physical address associated with the second data.
[0097] Statement 13. The memory device according to Statement 12, wherein the controller is configured to access the second portion of the volatile memory based on detecting the command.
[0098] Statement 14. The memory device according to Statement 11, wherein the command is for the memory device to perform a calculation, and the first data includes a result of the calculation.
[0099] Statement 15. The memory device according to Statement 11, wherein the reading by the computing device is based on detecting a trigger, wherein the trigger includes a signal generated by the memory device, and wherein the controller is further configured to: detect a fullness of the first portion of the volatile memory; and generate the signal based on the detection of the fullness of the first portion of the volatile memory.
[0100] Statement 16. The memory device according to Statement 11, wherein the first portion of the volatile memory is mapped to a first physical address space of the computing device and is configured to be accessed by the computing device via a memory access operation.
[0101] Statement 17. The memory device according to Statement 11, wherein the controller is further configured to send the second data to the computing device based on accessing the second portion of the volatile memory.
[0102] Statement 18. The memory device according to Statement 11, wherein the computing device is configured to take an action based on the first data.
[0103] Statement 19. The memory device according to Statement 18, wherein the action includes reconfiguring the second part of the volatile memory to increase the cache hit rate.
[0104] Statement 20. The memory device according to Statement 19, wherein reconfiguring the second part includes modifying the cache algorithm.
Claims
1. A method performed by a memory device, comprising: Receiving, by the memory device, a command from a computing device; Identifying, by the memory device, first data associated with the command; Storing, by the memory device, the first data in a first portion of a volatile memory of the memory device for the computing device to read; And Accessing, by the memory device, a second portion of the volatile memory, wherein the second portion of the volatile memory is configured to store a copy of second data stored in a non-volatile memory in the memory device.
2. The method according to claim 1, wherein The command is for reading the second data from the non-volatile memory or writing the second data to the non-volatile memory, and the first data identifies a physical address associated with the second data.
3. The method according to claim 2, wherein, The accessing of the second portion of the volatile memory is based on detecting the command.
4. The method according to claim 1, wherein, The command is for the memory device to perform a calculation, and the first data includes a result of the calculation.
5. The method according to claim 1, wherein, The reading by the computing device is based on detecting a trigger, wherein the trigger includes a signal generated by the memory device, and the method further comprises: Detecting, by the memory device, a fullness of the first portion of the volatile memory; and Generating, by the memory device, the signal based on the detection of the fullness of the first portion of the volatile memory.
6. The method according to claim 1, wherein The first portion of the volatile memory is mapped to a first physical address space of the computing device and is accessed by the computing device via a memory access operation.
7. The method according to claim 1, wherein The memory device is configured to send the second data to the computing device based on accessing the second portion of the volatile memory.
8. The method according to claim 1, wherein The computing device is configured to take an action based on the first data.
9. The method according to claim 8, wherein The action includes reconfiguring the second portion of the volatile memory to increase a cache hit rate.
10. The method according to claim 9, wherein, The reconfiguring the second portion includes modifying a cache algorithm.
11. A memory device, comprising: A controller; A volatile memory; And A non-volatile memory, Wherein the controller is configured to: Receive a command from a computing device; Identify first data associated with the command; Store the first data in a first portion of the volatile memory for the computing device to read; and Access a second portion of the volatile memory, wherein the second portion of the volatile memory is configured to store a copy of second data stored in the non-volatile memory.
12. The memory device according to claim 11, wherein, The command is for reading the second data from the non-volatile memory or writing the second data to the non-volatile memory, and the first data identifies a physical address associated with the second data.
13. The memory device according to claim 12, wherein, The controller is configured to access the second portion of the volatile memory based on detecting the command.
14. The memory device according to claim 11, wherein, The command is for the memory device to perform a calculation, and the first data includes a result of the calculation.
15. The memory device according to claim 11, wherein, The reading performed by the computing device is based on detecting a trigger, where the trigger includes a signal generated by the memory device, and where the controller is further configured to: Detect the fullness of the first portion of the volatile memory; and Generate the signal based on the detection of the fullness of the first portion of the volatile memory.
16. The memory device according to claim 11, wherein, The first portion of the volatile memory is mapped to a first physical address space of the computing device and is configured to be accessed by the computing device via a memory access operation.
17. The memory device according to claim 11, wherein, The controller is further configured to send the second data to the computing device based on accessing the second portion of the volatile memory.
18. The memory device according to claim 11, wherein, The computing device is configured to take an action based on the first data.
19. The memory device according to claim 18, wherein, The action includes reconfiguring the second portion of the volatile memory to increase the cache hit rate.
20. The memory device according to claim 19, wherein, The reconfiguring of the second portion includes modifying the cache algorithm.