Mechanism for efficiently rinsing the memory - side cache of dirty data

JP7686747B2Active Publication Date: 2025-06-02ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023519087
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-09-24
Filing Date
2021-09-17
Publication Date
2025-06-02
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Memory-side caches interfere with efficient ordering of DRAM transactions due to cache hits and misses, leading to underutilization of spare main memory bandwidth and inefficiencies in DRAM efficiency.

Method used

Implement a memory rinse mechanism that utilizes spare DRAM bandwidth during high hit rate phases to perform read and write rinsing of cached data, preserving the original order of transactions and reducing the number of cache evictions.

Benefits of technology

This approach optimizes DRAM efficiency by utilizing spare bandwidth, maintaining clean data in the cache, and preserving the optimal transaction order, thereby reducing power consumption and improving overall memory performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000016_0000
    Figure 00000016_0000
  • Figure 00000017_0000
    Figure 00000017_0000
  • Figure 00000018_0000
    Figure 00000018_0000
Patent Text Reader

Abstract

The method includes, in response to each of a plurality of write requests received at a memory-side cache device coupled to the memory device, writing payload data specified by the write requests to the memory-side cache device; if a first bandwidth availability condition is met, performing a cache write-through by writing the payload data to the memory device; and recording an indication that the payload data written to the memory-side cache device matches the payload data written to the memory device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Modern computing systems typically rely on multiple caches within a cache hierarchy to improve memory performance. Compared to main memory, a cache is a smaller and faster memory device that stores data that is frequently accessed or expected to be used in the near future so that the data can be accessed with low latency. Such cache devices are typically placed between a processing unit that issues memory requests and a main memory device and are often implemented in static random access memory (SRAM). A memory-side cache is a dedicated cache attached to a specific memory device or memory partition that caches data written to the memory device and read from the memory device by other devices.

[0002] This disclosure is shown by way of example and not limitation in the figures of the accompanying drawings.

Brief Description of the Drawings

[0003] [Figure 1] A diagram showing a computing system according to one embodiment. [Figure 2] A diagram showing a plurality of computing nodes connected via a data fabric interconnect according to one embodiment. [Figure 3] A diagram showing components within a memory partition according to one embodiment. [Figure 4] A diagram showing an interface between a memory-side cache device and a main memory device according to one embodiment. [Figure 5] A diagram showing a process for rinsing a memory-side cache according to one embodiment.

Modes for Carrying Out the Invention

[0004] The following description includes numerous specific details, such as examples of particular systems, components, and methods, to enhance understanding of the embodiments. However, it will be apparent to those skilled in the art that at least some embodiments can be implemented without these specific details. In other examples, well-known components or methods are not described in detail or are presented in the form of a simplified block diagram to avoid unnecessarily obscuring the embodiments. Thus, the specific details described are merely illustrative. Certain implementations may differ from these exemplary details and are still considered to fall within the scope of the embodiments.

[0005] A memory-side cache attached to the main memory device (e.g., DRAM) improves performance by caching data that is frequently read from or written to the main memory device. Memory access requests directed to main memory are processed with lower latency from the memory-side cache if the requested data is available in the cache. However, the presence of a memory-side cache can interfere with the efficient ordering of DRAM transactions. Due to DRAM timing constraints, the ordering of a set of memory transactions (e.g., read and write requests) affects how quickly the transactions can be executed in DRAM. If some of the transactions in a set result in a cache hit and are therefore processed from the memory-side cache, the remaining transactions that result in a cache miss are processed from DRAM in a different order than the original order in which they arrived. Furthermore, a cache miss can also change the access order, as if the victim is dirty, the victim entry is removed from the cache and updated in DRAM. This tends to disable mechanisms aimed at reordering DRAM transactions to achieve the highest possible bandwidth.

[0006] The operation of the memory-side cache and its main memory is characterized by two phases: 1) when the memory-side cache experiences a low hit rate, and more accesses reach the main memory (e.g., while the cache is populated with a working set of data for a new workload); and 2) when the memory-side cache experiences a high hit rate, and spare main memory bandwidth is not fully utilized. Therefore, one embodiment of the memory-side cache device utilizes the spare DRAM bandwidth observed during the high hit rate phase to perform read and write rinsing of cached data. Rinsing is performed when available DRAM bandwidth is detected and on data that has been marked as dirty in the memory-side cache when data is accessed.

[0007] When a rinse is performed, the original order of memory transactions is preserved, which tends to be more efficient because transactions are sent to DRAM regardless of whether they are cache hits or cache misses. The original sequence of transactions is generally expected to be more efficient from a DRAM efficiency standpoint. Furthermore, performing a memory rinse when DRAM bandwidth is available reduces the number of memory transactions that would otherwise be executed due to the evicting of dirty data from the cache, which could interfere with the memory access order or consume memory bandwidth if DRAM usage is already high during a phase of a program with a high cache miss rate.

[0008] Figure 1 shows one embodiment of a computing system 100 in which a memory rinse mechanism is implemented. Generally, the computing system 100 is embodied as one of many different types of devices, including but not limited to laptop computers or desktop computers, mobile devices, servers, etc. The computing system 100 includes a number of components 102-108 that communicate with each other via a bus 101. In the computing system 100, each of the components 102-108 can communicate with any of the other components 102-108 directly via the bus 101 or via one or more of the other components 102-108. The components 101-108 in the computing system 100 are housed in a single physical enclosure, such as a laptop computer or desktop computer chassis, or a mobile phone casing. In an alternative embodiment, some of the components of the computing system 100 are embodied as peripheral devices, so that the entire computing system 100 is not housed in a single physical enclosure.

[0009] Furthermore, the computing system 100 includes a user interface device for receiving information from or providing information to the user. Specifically, the computing system 100 includes an input device 102 such as a keyboard, mouse, touchscreen, or other device for receiving information from the user. The computing system 100 displays information to the user via a display 105 such as a monitor, light-emitting diode (LED) display, liquid crystal display, or other output device.

[0010] The computing system 100 further includes a network adapter 107 for sending and receiving data via a wired or wireless network. The computing system 100 also includes one or more peripheral devices 108. The peripheral devices 108 may include mass storage devices, position detection devices, sensors, input devices, or other types of devices used by the computing system 100.

[0011] The computing system 100 includes one or more processing units 104, and in the case of multiple processing units 104, they can operate in parallel. The processing units 104 receive and execute instructions 109 stored in the memory subsystem 106. In one embodiment, each of the processing units 104 includes multiple computing nodes located on a common integrated circuit board. The memory subsystem 106 includes memory devices used by the computing system 100, such as random access memory (RAM) modules, read-only memory (ROM) modules, hard disks, and other non-temporary computer-readable storage media.

[0012] Some embodiments of the computing system 100 may include fewer or more components than the embodiment shown in Figure 1. For example, certain embodiments may be implemented without a display 105 or an input device 102. Other embodiments may have two or more specific components; for example, one embodiment of the computing system 100 may have multiple buses 101, a network adapter 107, a memory device 106, etc.

[0013] In one embodiment, the processing unit 104 and memory 106 within the computing system 100 are implemented as a plurality of processing units and memory partitions connected by a data interconnection fabric 250, as shown in Figure 2. The data interconnection fabric 250 connects a plurality of computing nodes, including processing units 201-203 and memory partitions 207-209. The processing units 201-203 are connected to the data interconnection 250 via one or more coherent master devices 210, and the memory partitions 207-209 are connected to the data interconnection 250 via coherent slave devices 204-206, respectively. In one embodiment, these nodes 201-219 reside within the same device package and on the same integrated circuit die. For example, all of nodes 201-209 may be implemented on a monolithic central processing unit (CPU) or graphics processing unit (GPU) die with multiple processing cores. In an alternative embodiment, some of nodes 201-209 reside on different integrated circuit dies. For example, nodes 201-206 can reside on multiple chiplets mounted on a common interposer, each chiplet having multiple (e.g., four) processing cores.

[0014] The interconnection 250 includes a plurality of interconnection links that provide transmission paths for nodes 201 to 209 to communicate with each other. In one embodiment, the interconnection fabric 250 provides a plurality of different transmission paths between any pair of origin and destination nodes, and provides different transmission paths for any given origin node to communicate with each possible destination node.

[0015] Memory requests issued by processing units 201-203, directed to any of the memory partitions 207-209, are transmitted via the interconnect 250 through one or more coherent master devices 210 and received by the coherent slave devices of the memory partitions. For example, a memory request to access data stored in memory partition 207 is received by coherent slave device 204. In one embodiment, coherent slave device 204 reorders memory transactions based on DRAM timing constraints to maximize the throughput of memory partition 207. Coherent slave device 204 then sends the reordered memory requests to the memory partitions for processing.

[0016] Figure 3 shows the components of a memory partition 207 in a computing system 100 according to one embodiment. The memory partition 207 includes a DRAM device 310 for storing data and a memory-side cache device 320 for caching the data stored in the DRAM 310. The cache controller within the memory-side cache device 320 includes logic circuit modules 321 to 324. The memory-side cache device 320 also includes a data array 326 for storing cached data and a tag array 325 for storing metadata of the cached data. Memory transactions, such as read requests or write requests generated by processing units 201 to 203, are received from the data fabric interconnect 250 by a coherent slave device 204, reordered by the coherent slave device 204 to maximize DRAM throughput, and transferred in the reordered sequence to the input / output (I / O) ports 321 of the memory-side cache device 320 in the memory partition 207.

[0017] The memory-side cache 320 includes a monitoring circuit 323 that determines one or more metrics indicating the bandwidth availability of the installed DRAM device 310 based on information from the memory interface 324. If one or more bandwidth availability metrics indicate that sufficient DRAM bandwidth is available, read and write rinses of the cached data in the memory-side cache 320 are enabled. In one embodiment, read rinses and write rinses are enabled independently based on different bandwidth availability metrics. A write rinse refers to updating the backing data in the DRAM device 310 to match the cached data in the memory-side cache 320 in response to a data write request. A read rinse refers to updating the backing data in the DRAM device 310 to match the cached data when data is read from the cache 320.

[0018] When write rinse is enabled, data is written via the cache for write requests received on I / O port 321. That is, each payload data of a write request received by the memory-side cache 320 is written to an entry in the cache data array 326 (by the cache read / write logic 322) and also to a memory location in the DRAM 310 (via the memory interface 324), regardless of whether the request resulted in a cache hit or cache miss. The cache read / write logic 322 marks the data as "clean" in the tag array 325 (for example, by deasserting the "dirty" bit in the tag associated with the written cache entry), indicating that the cached copy of the data matches the backing copy of the data in the DRAM 310.

[0019] For a set of write requests directed to DRAM 310, one or more upstream devices reorder the write requests to maximize throughput for writing data to DRAM 310. The coherent slave 204 sends the write requests to the memory-side cache device 320 according to the determined write sequence. If write rinse is enabled, the payload data for the write requests is written to DRAM 310 in the order corresponding to the determined write sequence. Since the data for each write request is written to DRAM 310 regardless of whether the write request causes a cache hit or cache miss, the data is written to DRAM 310 in the same order (i.e., the optimized write sequence) as it is received by the cache device 320.

[0020] Writing via cache 320 reduces the amount of dirty data in cache 320. As a result, fewer dirty lines are evicted when a cache miss occurs, and capacity is reallocated for the miss data added to cache 320. The evicted dirty data is read from cache 320 so that its backing data in DRAM 310 can be updated. However, since the backing data does not need to be updated, if the evicted data is clean, it does not need to be read. Therefore, the power consumption of cache 320 decreases due to the reduced number of reads from cache 320. Furthermore, there are fewer write transactions sent to DRAM 310, which could otherwise interfere with the optimal ordering of DRAM transactions.

[0021] When read rinse is enabled and a read request resulting in a cache hit is received in the memory-side cache 320, the cache read / write logic 322 reads the data requested by each read request from the data array 326 and returns it. Read rinse is enabled if spare DRAM bandwidth is available. Therefore, spare bandwidth is used to rinse dirty cache data when dirty cached data is read. The read / write logic 322 checks the tag array 325 to determine whether the requested data is marked as "dirty" in the cache. If it is marked as "dirty," the data is flushed via the memory interface 324, and the memory interface updates the corresponding backing data in DRAM 310 with the cached version of the data. The cache read / write logic then marks the cached data as "clean."

[0022] Figure 4 shows the components within a memory-side cache device 320 and a DRAM device 310 according to one embodiment. The monitoring circuit 323 and memory interface 324 in the memory-side cache device 320 are connected to the memory controller 311 in the DRAM device 310. The memory interface 324 transmits memory transactions (e.g., read requests and write requests) to the memory controller 311, and the memory controller performs the requested read and write operations in the DRAM circuit.

[0023] The monitoring circuit 323 determines one or more bandwidth availability metrics that control whether read and / or write lances are enabled based on information obtained from the memory interface 324. The monitoring circuit also includes a counter 401 that tracks the total number of write transactions, including read transactions and write lance transactions, that occur within a predetermined time window (e.g., defined by the number of cycles, milliseconds, etc.). The counter 401 is involved in read lance and cache write-through transactions by incrementing the counter value during the current period in which the transaction is being executed. Read lance and write-through transactions are enabled and executed while the counter value does not exceed a predetermined maximum number of transactions for that period. Thus, the counter 401 is used to limit the number of memory lance transactions over time. In one embodiment, the counter 401 separately tracks the number of read lance transactions and write lance transactions and compares each value to its own maximum value.

[0024] In one embodiment, the counter 401 is used to enforce a maximum number of memory lance transactions within each period of a series. In this case, the counter 401 is reset at the start of each period. Alternatively, the counter 401 tracks the number of lance transactions within the most recent time window, such that transactions longer than a predetermined elapsed time do not contribute to the count value.

[0025] The cache controller in the memory-side cache device 320 includes a memory interface 324 that interfaces with the memory controller 311 of the DRAM device 310. The memory interface 324 sends memory transactions to the memory controller 311, which executes the requested transactions in the DRAM cells. The memory controller 311 includes two queues 404 and 405 for storing incoming memory transactions. The data command queue 404 stores commands indicating the operations to be performed, and the write data buffer 405 stores the data to be written.

[0026] The cache controller of the memory-side cache device 320 uses a token arbitration mechanism to perform flow control for communication with the memory controller 311. As shown in Figure 4, each token represents an entry in the data command queue 404 or the write data buffer 405. For example, each of the six data command queue tokens represents one of the six available entries in the data command queue 404, and each of the four write data buffer tokens represents one of the four available entries in the write data buffer 405. The memory interface 324 in the cache controller issues a memory access request to the memory controller 311 if it indicates that a sufficient number of tokens are available and there is enough space free in buffers 404 and 405 to receive the request. When the request is sent, the token is consumed. When the memory controller 311 frees up space in buffers 404 and 405 (for example, when the request is completed), the memory controller 311 returns the token to the memory interface 324 of the cache controller. In one embodiment, each of the read rinse transaction and the write-through transaction uses entries in the data command queue 404 and entries in the write data buffer 405.

[0027] The monitoring circuit 323 determines one or more bandwidth availability metrics based on the available bandwidth indicator received from the memory device 310. In one embodiment in which a token arbitration mechanism is used, the available memory bandwidth corresponds to the amount of available space within the data command queue 404 and the write data buffer 405, and thus is indicated by the number of available tokens for each of the buffers 404 and 405. Accordingly, the monitoring circuit 323 determines one or more bandwidth availability metrics based on the number of available tokens for each of the command queue 404 and the write data buffer 405. The greater the number of available tokens for either of the buffers 404 and 405, the more available memory bandwidth corresponds.

[0028] The monitoring circuit 323 determines different bandwidth availability metrics for each of the read-through mechanism and the write-through mechanism. The threshold comparison logic 402 compares each metric to a different threshold. Accordingly, the read-through and write-through mechanisms can be enabled and disabled under different conditions. When the threshold comparison logic 402 determines that the available bandwidth indicated by the bandwidth availability metric of the read-through mechanism is greater than its corresponding threshold (e.g., the number of available tokens exceeds the threshold number of tokens for each of the buffers 404 and 405), the read-through mechanism is enabled. That is, when the bandwidth availability meets the condition (e.g., exceeds the threshold), the read-through mechanism and the write-through mechanism can be enabled. Similarly, when the threshold comparison logic 402 determines that the available bandwidth indicated by the bandwidth availability metric of the write-through mechanism is greater than its corresponding threshold, the write-through mechanism is enabled.

[0029] FIG. 5 shows a process 500 for rinsing cached data when memory bandwidth is available, according to one embodiment. The rinse process 500 is executed by components within the computing system 100 including the memory side cache device 320 and the DRAM device 310.

[0030] In block 501, the monitoring circuit 323 determines one or more bandwidth availability metrics for the DRAM device 310. The monitoring circuit 323 monitors the number of tokens 403 representing available entries in the data command queue 404 and the write data buffer 405, and thus can compare these metrics with one or more different thresholds to determine whether sufficient memory bandwidth is available to perform proactive read rinse or write-through transactions.

[0031] If no write request is received in the memory-side cache 320 in block 503, process 500 proceeds to block 505. If no read request is received in the memory-side cache 320 in block 505, process 500 returns to block 501. Therefore, if no memory request is received, the monitoring circuit 323 continues to monitor the availability of the memory bandwidth of the DRAM device 310.

[0032] In block 503, when a write request is received at I / O port 321 of cache device 320, the cache read / write logic 322 writes the payload data specified by the write request to an entry in data array 326. In block 509, threshold comparison logic 402 compares the number of write-through transactions performed during the current period (as indicated by counter 401) to the maximum number. If the maximum number of write-through transactions is exceeded, no write-through is performed, and the cache read / write logic 322 marks the cached data as "dirty" in its corresponding tag in tag array 325. Process 500 returns to block 501 and continues monitoring the memory bandwidth availability metric.

[0033] In block 509, if the maximum number of write-through transactions has not been exceeded, as indicated by counter 401, process 500 proceeds to block 513. In block 513, monitoring circuit 323 determines whether the available memory bandwidth is greater than the threshold for enabling write-through transactions. In particular, threshold comparison logic 402 compares the number of available tokens 403 to a threshold. If the number of tokens exceeds the threshold, write-through transactions are enabled. In one embodiment, the number of available tokens for each of buffers 404 and 405 is compared to their respective thresholds, and if both thresholds are exceeded, write-through transactions are enabled. If the available memory bandwidth is not greater than the threshold, write-through transactions are not enabled, and process 500 proceeds to block 511. In block 511, cache read / write logic 322 marks the cached data as "dirty," and process 500 returns to block 501 to continue monitoring the memory bandwidth availability metric.

[0034] In block 513, if the available memory bandwidth is greater than a threshold, write-through transactions are enabled. As provided in block 515, the memory interface 324 performs write-through of payload data by writing the payload data to the DRAM device 310. Counter 401 is incremented to count the write-through transactions for the current period. In block 517, the cache read / write logic 322 marks the cached data as "clean" (for example, by deasserting the "dirty" bit in the tag associated with the entry containing the cached data) to record an indicator that the payload data written to the memory-side cache device 320 matches the payload data written to the DRAM device 310. From block 517, process 500 returns to block 510 to continue monitoring the memory bandwidth.

[0035] Therefore, blocks 503-517 are repeated for each write request received at I / O port 321 of cache device 320 in order to perform cache write-through, provided that sufficient memory bandwidth is available and the number of write-through transactions does not exceed the maximum number. As a result, once the write-through mechanism is enabled, the payload data is written to memory device 310 in the order corresponding to the sequence in which the write requests were first received.

[0036] In block 505, when a read request is received at the I / O port 321 of the memory-side cache device 320, the cache read / write logic 322 checks the tag array 325 in block 519 to determine whether the requested data is in the data array 326. If the read request results in a cache miss, the requested data is read from the DRAM device 310. As provided in block 521, the cache read / write logic 322 updates the data array 326 and tag array 325 to contain the data, and the data is returned to complete the request. From block 521, process 500 returns to block 501 to continue monitoring memory bandwidth availability.

[0037] In block 519, if the read request results in a cache hit, the cache read / write logic 322 reads and returns the requested data from the data array 326, as provided in block 523. In block 525, the cache read / write logic 322 checks the tags of the data in the tag array 325 to determine whether the data is dirty (for example, whether the "dirty" bit of the data entry is asserted). If the data is not dirty, the cached data already matches its backing data in the DRAM device 310, and process 500 returns to block 501 to continue monitoring memory bandwidth availability without performing a read rinse.

[0038] In block 525, if the cached data is dirty, the cached data does not match its backing data in the DRAM device 310, and process 500 proceeds to block 527. In block 527, the monitoring circuit 323 determines whether the maximum number of read rinse transactions has been exceeded. The threshold comparison logic 402 compares the number of read rinse transactions counted by the counter 401 for the current period with the maximum number of read rinse transactions allowed for the current period. If the maximum number of read rinse transactions has been exceeded, no read rinse is performed, and process 500 returns to block 501.

[0039] In block 527, if the current time period does not exceed the maximum number of read rinse transactions, process 500 proceeds to block 529. In block 529, the monitoring logic determines whether the available memory bandwidth exceeds the threshold for enabling read rinse. In one embodiment, the threshold comparison logic 402 compares the number of available data command queue 404 and write data buffer 405 tokens with their respective thresholds, and if both thresholds are exceeded, read rinse is enabled. In one embodiment, the threshold for enabling read rinse is different from the threshold for enabling write-through. If read rinse is not enabled, process 500 returns to block 501.

[0040] In block 529, if the monitoring circuit 323 determines that sufficient memory bandwidth is available, the memory interface 324 performs a read rinse by updating the backing data corresponding to the read request so that the backing data matches its corresponding data in the cache 320. The memory interface 324 sends a write request with the data to the memory controller 311, and the counter 401 is incremented to count the read rinse transactions for the current period. In block 533, the cache read / write logic 322 deasserts the "dirty" bit of the cached entry to indicate that the cached data in the memory-side cache device 320 matches its backing data in the DRAM device 310. From block 533, process 500 returns to block 501.

[0041] Therefore, blocks 505-533 are repeated for each read request received at the I / O port 321 of the cache device 320 to perform a cache read rinse, provided that sufficient memory bandwidth is available and the number of read rinse transactions does not exceed the maximum number. As a result of the above write-through and read rinse mechanism, any spare memory bandwidth available during the phase in which the cache 320 experiences a high hit rate is used to maintain clean data in the cache 320, thus enabling the storage of an optimal memory write sequence and reducing the number of victims written back to the DRAM 310.

[0042] This method includes writing the payload data specified by a write request to the memory-side cache device in response to each of a plurality of write requests received in the memory-side cache device coupled to the memory device, and performing a cache write-through when a first bandwidth availability condition is met. The cache write-through is performed by writing the payload data to the memory device and recording an indicator that the payload data written to the memory-side cache device matches the payload data written to the memory device.

[0043] In this method, for each of multiple write requests, writing the payload data specified by the write request includes storing the data in an entry in the memory-side cache device. Recording the index includes deasserting the dirty bit in the tag associated with the entry.

[0044] This method further includes receiving multiple write requests in the memory-side cache according to a write sequence. The writing of payload data to the memory device is performed in the order corresponding to the write sequence.

[0045] The method further includes, in response to each of multiple read requests received by the memory-side cache device, updating the backing data to match the cached data in the memory-side cache according to an indicator that the cached data requested by the read request is different from the backing data in the memory device, and recording an indicator that the cached data in the memory-side cache matches the backing data in the memory device.

[0046] The method also includes determining a first bandwidth availability metric for the memory device. The first bandwidth availability condition is met when the first bandwidth availability metric is greater than a first bandwidth threshold.

[0047] The method further includes determining a second bandwidth availability metric for a memory device based on the amount of available space in the memory device's command queue and its write data buffer, respectively. Backing data updates are further performed if the second bandwidth availability metric is determined to be greater than a second bandwidth threshold.

[0048] The method further includes determining a first bandwidth availability metric based on the amount of available space in the command queue of the memory device and the write data buffer of the memory device.

[0049] This method further includes determining when a first bandwidth condition is met based on an indicator of available bandwidth received from a memory device.

[0050] In this method, performing a cache write-through for each write request of a group of write requests further includes incrementing a counter value for the current time period, and the cache write-through is performed when it is determined that the counter value is less than the maximum number of cache write-through transactions for the current time period.

[0051] The memory-side cache device includes cache read / write logic for writing payload data specified by a write request to the memory-side cache device in response to each of a plurality of write requests received in the memory-side cache device coupled to the memory device, and a memory interface for performing a cache write-through when a first bandwidth availability condition is met by writing the payload data to the memory device in response to each of the plurality of write requests. The cache read / write logic performs a cache write-through by recording an indicator that the payload data written to the memory-side cache device matches the payload data written to the memory device.

[0052] In the memory-side cache device, the cache read / write logic writes the payload data specified by the write request for each of multiple write requests by storing the data in an entry in the memory-side cache device, and records the metric by deasserting the dirty bit in the tag associated with the entry.

[0053] The memory-side cache device further includes input / output ports for receiving multiple write requests according to the write sequence. The memory interface writes payload data to the memory device in the order corresponding to the write sequence.

[0054] In the memory-side cache device, the memory interface performs a read rinse in response to each of multiple read requests received in the memory-side cache device. This rinse is performed by updating the backing data to match the cached data, based on an indicator that the read request causes a cache hit in the memory-side cache and that the cached data requested by the read request differs from the backing data in the memory device. The cache read / write logic performs a read rinse by recording an indicator that the cached data in the memory-side cache matches the backing data in the memory device.

[0055] Furthermore, the memory-side cache device includes monitoring circuitry for determining a first bandwidth availability metric based on the amount of available space in the memory device's command queue and its write data buffer, respectively. The first bandwidth availability condition is met when the first bandwidth availability metric is greater than a first bandwidth threshold.

[0056] In the memory-side cache device, the monitoring circuit determines a second bandwidth availability metric for the memory device based on the amount of available space in the memory device's command queue and write data buffer, respectively. The memory interface updates the backing data in response to determining that the second bandwidth availability metric is greater than the second bandwidth threshold.

[0057] The memory-side cache device includes a monitoring circuit for determining when a first bandwidth availability condition is met, based on an indicator of available bandwidth received from the memory device.

[0058] Furthermore, the memory-side cache device includes a counter for performing cache write-throughs by incrementing the counter value for the current time period for each write request in a set of multiple write requests. Cache write-throughs are performed when it is determined that the counter value is less than the maximum number of cache write-through transactions for the current period.

[0059] The computing system includes a memory device for storing backing data, and a memory-side cache device coupled to the memory device, which, in response to each of a plurality of write requests received in the memory-side cache device, writes the payload data specified by the write request to the memory-side cache device, and performs a cache write-through when a first bandwidth availability condition is met. The cache write-through is performed by writing the payload data to the memory device and recording an indicator that the payload data written to the memory-side cache device matches the payload data written to the memory device.

[0060] In a computing system, the memory-side cache device determines a first bandwidth availability metric for the memory device. The first bandwidth availability condition is met when the first bandwidth availability metric is greater than a first bandwidth threshold.

[0061] In a computing system, the memory-side cache device, in response to each of multiple read requests received by the memory-side cache device, updates the backing data to match the cached data according to an indicator that the read request causes a cache hit in the memory-side cache and that the cached data requested by the read request differs from the backing data in the memory device; records an indicator that the cached data in the memory-side cache matches the backing data in the memory device; and determines a second bandwidth availability metric for the memory device based on the amount of available space in the memory device's command queue and the memory device's write data buffer, respectively. Further updates of the backing data are performed if the second bandwidth availability metric is determined to be greater than a second bandwidth threshold.

[0062] The computing system further includes a coherent slave device that receives multiple write requests via a data fabric interconnect, determines a write sequence for the multiple write requests, and sends the multiple write requests to a memory-side cache device according to the write sequence, and the writing of payload data to the memory device is performed in the order corresponding to the write sequence.

[0063] The computing system further includes a data fabric interconnection coupled to a memory-side cache device for sending memory transactions, which include multiple write requests, from one or more processing units to the memory-side cache device.

[0064] As used herein, the term "coupled" may mean coupled directly or indirectly through one or more intervening components. Any of the signals provided through the various buses described herein may be time-division multiplexed with other signals and provided through one or more common buses. Furthermore, interconnections between circuit components and blocks may be represented as buses or single signal lines. Each bus may, alternatively, be one or more single signal lines, and each single signal line may, alternatively, be a bus.

[0065] Certain embodiments may be implemented as computer program products that may include instructions stored on a non-temporary computer-readable storage medium. These instructions may be used to program a general-purpose processor or a dedicated processor to perform the operations described. A computer-readable storage medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer) (e.g., software, processing applications). A non-temporary computer-readable storage medium may include, but is not limited to, magnetic storage media (e.g., floppy diskettes), optical storage media (e.g., CD-ROMs), magneto-optical storage media, read-only memory (ROM), random-access memory (RAM), erasable programmable memory (e.g., EPROMs and EEPROMs), flash memory, or other types of media suitable for storing electronic instructions.

[0066] In addition, some embodiments may be implemented in a distributed computing environment in which a computer-readable storage medium is stored on and / or executed by two or more computer systems. Furthermore, information transferred between computer systems may be pulled or pushed via a transmission medium connecting the computer systems.

[0067] Generally, a data structure representing a computing system 100 and / or a part thereof, mounted on a computer-readable storage medium, may be a database or other data structure that can be read by a program and used directly or indirectly to manufacture hardware including the computing system 100. For example, the data structure may be an action-level description or register-transfer-level (RTL) description of hardware functionality in a high-level design language (HDL) such as Verilog or VHDL. The description can be read by a synthesis tool that can synthesize the description to generate a netlist containing a list of gates from a synthesis library. The netlist contains a set of gates that also represent the functionality of the hardware including the computing system 100. The netlist can then be arranged and routed to generate a dataset describing the geometric shapes to be applied to a mask. The mask can then be used in various semiconductor manufacturing processes to manufacture semiconductor circuits(s) corresponding to the computing system 100. Alternatively, a database on a computer-readable storage medium may, as desired, be a netlist (with or without a synthetic library), a dataset, or Graphic Data System (GDS) II data.

[0068] Although the operations of the methods(s) described herein are shown and explained in a specific order, the order of operations of each method may be modified so that certain operations may be performed in reverse order, or so that certain operations may be performed at least partially concurrently with other operations. In another embodiment, the instructions or suboperations of individual operations may be performed intermittently and / or alternately.

[0069] In the above specification, embodiments have been described with reference to specific exemplary embodiments. However, it will be apparent that various modifications and changes can be made to them without departing from the broader scope of embodiments described in the appended claims. Accordingly, this specification and the drawings should be taken as illustrative rather than restrictive.

Claims

1. 1. A method comprising: In response to each of a plurality of write requests received at a memory-side cache device coupled to the memory device, writing payload data specified by the write request to the memory-side cache device; performing a cache write-through by writing the payload data to the memory device if a first bandwidth availability condition is met, and recording an indication that the payload data written to the memory-side cache device matches the payload data written to the memory device. method.

2. For each of the plurality of write requests: writing the payload data specified by the write request includes storing the payload data in an entry of the memory-side cache device; recording the indication includes deasserting a dirty bit in a tag associated with the entry.

10. The method of claim 1.

3. receiving the write requests at the memory-side cache according to a write sequence; writing the payload data to the memory device in an order corresponding to the write sequence; 10. The method of claim 1.

4. In response to each of a plurality of read requests received at the memory-side cache device, if the read requests cause a cache hit in the memory-side cache, in response to an indication that cached data requested by the read requests differs from backing data in the memory device: updating the backing data to match the cached data; and recording an indication that the cached data in the memory-side cache matches the backing data in the memory device.

10. The method of claim 1.

5. determining a first bandwidth availability metric for the memory device; the first bandwidth availability condition is met if the first bandwidth availability metric is greater than a first bandwidth threshold; 10. The method of claim 1.

6. determining a second bandwidth availability metric for the memory device based on the amount of available space in each of a command queue of the memory device and a write data buffer of the memory device; updating the backing data is performed in response to determining that the second bandwidth availability metric is greater than a second bandwidth threshold. The method of claim 5.

7. determining the first bandwidth availability metric based on an amount of available space in each of a command queue of the memory device and a write data buffer of the memory device. The method of claim 5.

8. determining when the first bandwidth availability condition is met based on the indication of available bandwidth received from the memory device.

10. The method of claim 1.

9. For each of the plurality of write requests: performing the cache write-through includes incrementing a counter value for a current time period; the cache write-through is performed in response to determining that the counter value is less than a maximum number of cache write-through transactions for the current time period.

10. The method of claim 1.

10. A memory-side cache device, cache read / write logic configured to, in response to each of a plurality of write requests received at a memory-side cache device coupled to the memory device, write payload data specified by the write requests to the memory-side cache device; a memory interface configured to perform a cache write-through if, in response to each of the plurality of write requests, writing the payload data to the memory device satisfies a first bandwidth availability condition; the cache read / write logic is configured to perform the cache write-through by recording an indication that the payload data written to the memory-side cache device matches the payload data written to the memory device. Memory-side cache device.

11. The cache read / write logic, for each of the plurality of write requests: writing the payload data specified by the write request by storing it in an entry in the memory-side cache device; configured to record the indication by deasserting a dirty bit in a tag associated with the entry.

11. The memory-side cache device of claim 10.

12. an input / output port configured to receive the plurality of write requests according to a write sequence, the memory interface configured to write the payload data to the memory device in an order corresponding to the write sequence; 11. The memory-side cache device of claim 10.

13. the memory interface is configured to, in response to each of a plurality of read requests received at the memory-side cache device, if the read requests cause a cache hit in the memory-side cache, perform a read rinse by, in response to an indication that cached data requested by the read requests differs from backing data in the memory device, updating the backing data to match the cached data; the cache read / write logic is configured to perform the read rinse by recording an indication that the cached data in the memory-side cache matches the backing data in the memory device.

11. The memory-side cache device of claim 10.

14. and a monitoring circuit configured to determine a first bandwidth availability metric based on an amount of available space in each of a command queue of the memory device and a write data buffer of the memory device, and wherein the first bandwidth availability condition is met when the first bandwidth availability metric is greater than a first bandwidth threshold.

11. The memory-side cache device of claim 10.

15. the monitoring circuitry is configured to determine a second bandwidth availability metric for the memory device based on an amount of available space in each of a command queue of the memory device and a write data buffer of the memory device; the memory interface is configured to update the backing data in response to determining that the second bandwidth availability metric is greater than a second bandwidth threshold.

15. The memory-side cache device of claim 14.

16. and a monitoring circuit configured to determine when the first bandwidth availability condition is met based on an indication of available bandwidth received from the memory device.

11. The memory-side cache device of claim 10.

17. For each of the plurality of write requests: a counter configured to perform the cache write-through by incrementing a counter value for a current time period, the cache write-through being performed in response to determining that the counter value is less than a maximum number of cache write-through transactions for the current time period.

11. The memory-side cache device of claim 10.

18. 1. A computing system comprising: a memory device configured to store backing data; a memory-side cache device coupled to the memory device; The memory-side cache device In response to each of a plurality of write requests received at the memory-side cache device, writing payload data specified by the write request to the memory-side cache device; performing a cache write-through by writing the payload data to the memory device if a first bandwidth availability condition is met and recording an indication that the payload data written to the memory-side cache device matches the payload data written to the memory device; configured to: Computing system.

19. The memory-side cache device determining a first bandwidth availability metric for the memory device, and configured such that the first bandwidth availability condition is met if the first bandwidth availability metric is greater than a first bandwidth threshold; 20. The computing system of claim 18.

20. The memory-side cache device In response to each of a plurality of read requests received at the memory-side cache device, if the read requests cause a cache hit in the memory-side cache, in response to an indication that cached data requested by the read requests differs from backing data in the memory device: updating the backing data to match the cached data; recording an indication that the cached data in the memory-side cache matches the backing data in the memory device; determining a second bandwidth availability metric for the memory device based on an amount of available space in each of a command queue of the memory device and a write data buffer of the memory device, wherein updating the backing data is performed in response to determining that the second bandwidth availability metric is greater than a second bandwidth threshold; configured to:

20. The computing system of claim 19.

21. further comprising a coherent slave device; The coherent slave device receiving the plurality of write requests over a data fabric interconnect; determining a write sequence for the plurality of write requests; sending the plurality of write requests to the memory-side cache device according to the write sequence, wherein writing the payload data to the memory device is performed in an order corresponding to the write sequence; configured to:

20. The computing system of claim 18.

22. a data fabric interconnect coupled to the memory-side cache device and configured to transmit memory transactions including the plurality of write requests from one or more processing units to the memory-side cache device; 20. The computing system of claim 18.