Active direct cache transfer from producer to consumer

Direct cache transfer mechanisms in CPU-accelerator systems address inefficiencies by initiating data transfer during cache evictions, using novel CMOs to ensure data is available in consumer caches, enhancing system efficiency and reducing latency in heterogeneous environments.

JP7681005B2Active Publication Date: 2025-05-21XILINX INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022514594
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-04
Filing Date
2020-06-08
Publication Date
2025-05-21
Estimated Expiration
2040-06-08

AI Technical Summary

Technical Problem

Current CPU-accelerator systems face inefficiencies in data transfer due to heterogeneous cache sizes and operating frequencies, leading to increased latency and bandwidth differences in cache coherent non-uniform memory access (CC-NUMA) systems, where producer-consumer pairs are located near or far from memory, and the pull model for direct cache transfer (DCT) is often too late to be effective.

Method used

Implementing direct cache transfer (DCT) mechanisms that initiate data transfer when updated data is evicted from the producer cache, using flush-stash and copyback-stash cache maintenance operations (CMOs) to directly move data from producer to consumer caches, regardless of consumer location, and leveraging hardware or software coherency management.

Benefits of technology

Enhances performance by ensuring data is already in the consumer's cache when needed, reducing reliance on main memory access and minimizing interference with other producer-consumer pairs, thus improving overall system efficiency in CC-NUMA systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681005000001
    Figure 0007681005000001
  • Figure 0007681005000002
    Figure 0007681005000002
  • Figure 0007681005000003
    Figure 0007681005000003
Patent Text Reader

Abstract

Embodiments herein create DCT mechanisms that initiate a DCT when updated data is being evicted from a producer cache. These DCT mechanisms are applied when a producer replaces updated content in its cache because the producer has moved to work on a different data set (e.g., a different task), moved to work on a different function, or the producer-consumer task manager (e.g., a management unit) has enforced software coherency by sending a cache maintain operation (CMO). One advantage of the DCT mechanisms is that, because the cache transfer occurs directly when the updated data is being evicted, by the time the consumer starts its task, the updated data is already in its own cache or another cache in the cache hierarchy.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Examples of this disclosure relate generally to performing direct cache transfers in a heterogeneous processing environment. [Background technology]

[0002] background Server CPU-accelerator systems such as those enabled by Computer eXpress Link (CXL), Cache Coherent Interconnect for Accelerators (CCIX), Quick Path Interconnect / Ultra Path Interconnect (QPI / UPI), Infinity Fabric, NVLink (trademark pending), and Open Coherent Accelerator Processor Interface (Open CAPI) connected SoCs are all inherently hardware cache coherent systems, i.e., the hardware maintains a universal, coherent view of the data being accessed, modified, and cached, regardless of whether the processor or accelerator is acting as a producer or consumer of the data and metadata (information about the data).

[0003] Current shared memory CPU-accelerator execution frameworks rely on either software coherency or hardware coherency for producer-consumer interactions in their systems. Over time, at least one of the CPU or accelerator acts as a producer or consumer of data or metadata as part of an application or as part of the execution of a function. The movement of that data between the caches of a producer-consumer pair can be done using either explicit actions of software coherency or implicit actions of hardware coherency.

[0004] These CPU-accelerator systems are also described as cache coherent non-uniform memory access systems (CC-NUMA). CC-NUMA results in differences in both latency and bandwidth depending on whether the CPU or accelerator access is near or far memory, and depending on where the data is cached when accessed by either the producer or consumer. In addition, producer-consumer pairs and their cached data may be located closer to each other than the data they are operating on. This may result in producer-consumer pairs having better latency and bandwidth for interactions with each other compared to their respective individual interactions with the data they are operating on. Furthermore, these CPU-accelerator systems are heterogeneous systems, and the capabilities of the CPU and accelerators may also differ in terms of their operating frequencies and caching capabilities. Summary of the Invention [Means for solving the problem]

[0005] overview Techniques for performing direct cache transfer (DCT). One example is a computing system comprising: a producer including a first processing element configured to generate processed data; a producer cache configured to store the processed data generated by the producer; a consumer including a second processing element configured to receive and process the processed data generated by the producer; and a consumer cache configured to store the processed data generated by the consumer. The producer is configured to perform the DCT to transfer the processed data from the producer cache to the consumer cache in response to receiving a stash cache maintenance operation (stash-CMO).

[0006] Another example herein is a method that includes generating processed data at a producer including a first hardware processing element, storing the processed data in a producer cache, in response to receiving a stash CMO, performing a DCT to transfer the processed data from the producer cache to a consumer cache, and processing the processed data at a consumer including a second hardware processing element after the DCT.

[0007] Another example herein is a computing system comprising: a producer including a first processing element configured to generate processed data; a producer cache configured to store the processed data generated by the producer; a consumer including a second processing element configured to receive and process the processed data generated by the producer; and a consumer cache configured to store the processed data generated by the consumer. In response to receiving a stash cache maintenance operation (stash-CMO), the producer is configured to perform a direct cache transfer (DCT) to transfer the processed data from the producer cache to the consumer cache.

[0008] Any of the above computing systems may further include at least one coherent interconnect communicatively coupling the producers to the consumers, wherein the computing system is a cache coherent non-uniform memory access (CC-NUMA) system.

[0009] Optionally, in any of the computing systems described above, the producer learns the location of the consumer before the producer completes a task that instructs the producer to generate processed data.

[0010] Optionally, in any of the computing systems described above, coherency of the computing system is maintained by a hardware element, and the stash-CMO and DCT are performed in response to a producer deciding to update a main memory of the computing system with processed data currently stored in the producer cache, and the stash-CMO is a flash-type stash-CMO.

[0011] Optionally, in any of the computing systems described above, coherency of the computing system is maintained by a hardware element, and the stash-CMO and DCT are performed in response to the producer cache issuing a capacity eviction to remove at least a portion of the processed data, and the stash-CMO is a copy-back stash-CMO.

[0012] Optionally, in any of the computing systems described above, the producer does not know the whereabouts of the consumer before the producer completes a task that instructs the producer to generate the processed data.

[0013] Optionally, in any of the computing systems described above, coherency of the computing system is maintained by a software management unit, the producer is configured to notify the software management unit when a task of generating the processed data is completed, and the software management unit is configured to instruct a home of the processed data to initiate a DCT. Further, the software management unit may optionally (A) send a first stash-CMO to the home to initiate a DCT on the producer, and / or (B) in response to receiving the first stash-CMO, the home sends a snoop to the producer including the stash-CMO instructing the producer to perform a DCT.

[0014] Another example herein is a method that includes generating processed data at a producer including a first hardware processing element, storing the processed data in a producer cache, performing a DCT to transfer the processed data from the producer cache to a consumer cache in response to receiving a stash-CMO, and after the DCT, processing the processed data at a consumer including a second hardware processing element, and storing the processed data generated by the consumer in the consumer cache.

[0015] Any of the methods described above may further include the DCT being performed using at least one coherent interconnect that communicatively couples the producers to the consumers, wherein the producers and consumers are part of a CC-NUMA system.

[0016] Optionally, in any of the methods described above, the method may further include the producer informing the producer of the location of the consumer before completing the task of instructing the producer to generate the processed data.

[0017] Optionally, in any of the above methods, coherency between the consumer and the producer may be maintained by a hardware element, and the stash-CMO and DCT are performed in response to the producer determining to update the main memory with the processed data currently stored in the producer cache, and the stash-CMO is a flush-type stash-CMO.

[0018] Optionally, in any of the methods described above, coherency between the consumer and the producer is maintained by a hardware element, and the stash-CMO and DCT are performed in response to the producer cache issuing a capacity eviction to remove at least a portion of the processed data, and the stash-CMO is a copy-back stash-CMO.

[0019] Optionally, in any of the methods described above, the producer does not know the whereabouts of the consumer before the producer completes a task that instructs the producer to generate the processed data.

[0020] Optionally, in any of the above methods, coherency between the consumer and the producer is maintained by a software management unit, the method further comprising notifying the software management unit when a task of generating the processed data is completed by the producer, and instructing a home of the processed data to initiate a DCT.

[0021] Optionally, in any of the above-described methods, instructing the home of the processed data to initiate DCT includes sending a first stash CMO from the software management unit to the home to initiate DCT on the producer.

[0022] Optionally, in any of the methods described above, in response to receiving the first stash-CMO, sending a snoop from the home to the producer, the snoop including the stash-CMO instructing the producer to perform the DCT. The method may optionally further include (A) in response to the snoop, updating the main memory to include the processed data stored in the producer cache, and / or (B) after sending the snoop, informing the consumer that the DCT is complete. [Brief description of the drawings]

[0023] [Figure 1] 1 is a block diagram of a computing system that implements a pull model to perform direct cache transfers, according to an example. [Diagram 2] 1 is a block diagram of a computing system in which a producer initiates direct cache transfer, according to an example. [Diagram 3] 1 is a flowchart illustrating a producer initiating a direct cache transfer according to an example. [Figure 4] 1 is a block diagram of a computing system in which a coherency manager initiates a direct cache transfer when a producer completes a task, according to an example. [Diagram 5] 1 is a flowchart illustrating a coherency manager initiating a direct cache transfer when a producer completes a task, according to an example. [Figure 6] 1 is a block diagram of a computing system using direct cache transfer to perform pipelining, according to an example. [Figure 7] FIG. 1 is a block diagram of a computing system implementing interaction between a CPU and computational memory, according to an example. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0024] Detailed Description Direct Cache Transfer (DCT) is a technique that moves data directly between two caches outside the control of the corresponding processing element (e.g., CPU or accelerator core). In some embodiments, DCT is a pull model, and when a consumer wants data, it contacts a tracing agent (e.g., home or home agent) to locate a copy of the data. The tracing agent then instructs the cache containing the copy to transfer the data to the cache corresponding to the consumer using DCT. The pull model works well for homogeneous systems where the processor and memory are at the same distance from each other. For example, a first processing element (e.g., consumer) can request data from the cache of a second processing element (e.g., producer) that is currently processing the data. Since a producer may always evict data when completing a task, the producer evicts the data from its cache, which means that the consumer gets whatever data is left in the cache when the task is completed. However, in a homogeneous system, it doesn't matter that the producer is evicting entries, because the producer and consumer have the same size cache. That is, the producer may have processed 100MB of data, but the consumer and producer caches may only be 2MB in size. Thus, the consumer cache can only store 2MB, even if the producer cache uses the DCT multiple times to send more than 2MB to the consumer cache.

[0025] However, as mentioned above, in CC-NUMA systems, there may be different processing elements with different sized caches. Thus, these systems can utilize smaller producer caches to send data to larger consumer caches using multiple DCTs. However, in the pull model, it is often too late to utilize the DCT because of the time it takes before the consumer can start the DCT. This is exacerbated when the consumer is slower (e.g., has a slower operating frequency) than the producer.

[0026] Instead of using a pull model, embodiments herein create DCT mechanisms that initiate the DCT when updated data is being evicted from a producer cache. These DCT mechanisms are applied when a producer replaces updated content in its cache because the producer has moved to work on a different data set (e.g., a different task), moved to work on a different function, or the producer-consumer task manager has implemented software coherency by sending a cache maintain operation (CMO). One advantage of the DCT mechanisms is that because the DCT is performed when the updated data is being evicted, by the time the consumer accesses the updated data, the data is already in the consumer's cache or another cache in the cache hierarchy. Thus, the consumer does not need to go to main memory, which can be far away in a CC-NUMA system, to retrieve the data set.

[0027] 1 is a block diagram of a computing system 100 that implements a pull model for performing DCT, according to one example. In one embodiment, computing system 100 is a CC-NUMA system that includes different types of hardware processing elements, e.g., producers 120 and consumers 135, that may have different types and sizes of caches, e.g., producer cache 125 and consumer cache 140. In this example, producers 120 and consumers 135 are located in the same domain 115 and communicatively coupled via a coherent interconnect 150. However, in other embodiments, producers 120 and consumers 135 may be in different domains coupled with multiple interconnects.

[0028] FIG. 1 illustrates a home 110 (e.g., a home agent or home node) that owns data stored in a producer cache 125), a consumer cache 140, and data stored in memories 105A and 105B (also referred to as main memory). As described below, the home 110 can transfer data from memory 105A (e.g., producer / consumer data 170) to the producer cache 125 and the consumer cache 140 so that the respective computation engines 130, 145 can process the data. In this example, the producer 120 first processes the data 170, and then the processed data is processed by the consumer 135. That is, the producer 120 generates data that is consumed by the consumer 135. As a non-limiting example, the producer 120 and the consumer 135 can represent respective integrated circuits or computational systems. The producers 120 and consumers 135 may be accelerators, central processing units (CPUs), graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), systems on chips (SoCs), etc. The computational engines 130, 145 may be processing cores or processing elements within the producers 120 and consumers 135.

[0029] Rather than storing data generated by producer 120 back in main memory 105A and 105B, computing system 100 can instead perform the DCT and transfer the data directly from producer cache 125 to consumer cache 140. FIG. 1 illustrates multiple actions 160 for performing the DCT using a pull model that relies on consumer 135 to initiate the DCT process. During action 160A, computation engine 130 of producer 120 processes a data set that is part of a task (e.g., a search task, a machine learning task, a compression / decompression task, an encryption / decryption task, etc.). For a subset of the data set that may be kept in producer cache 125, the updated (or processed) data may remain in producer cache 125. That is, this example assumes that producer cache 125 is not large enough to store the entire data set of the task.

[0030] During action 160B, the subset of the dataset that cannot be held in producer cache 125 is flushed to main memory 105A and stored as producer / consumer data 170. For example, processing the entire dataset for a task may result in 100 MB, but producer cache 125 can only store 2 MB.

[0031] During action 160C, when the compute engine 145 of the consumer 135 detects that the producer 120 has completed its task (this may be done using a flag or metadata sent by the producer 120), it retrieves the processed data by first sending a retrieve request to the home 110 (i.e., location of coherency management) of the dataset. For the subset of the dataset that is still in the producer cache 125, this data may remain in the cache 125.

[0032] During action 160D, the subset of the data set that is no longer in the producer cache 125 (ie, the data 170 stored in memory 105A) is sent from memory 105A to the consumer 135.

[0033] During action 160E, the home 110 sends a DCT snoop informing the producer 120 that a subset of the task's dataset still stored in the producer cache 125 should be sent directly to the consumer 135. During action 160F, the producer 120 performs a DCT to transfer any data associated with the task directly from the producer cache 125 to the consumer cache 140.

[0034] However, there are some drawbacks to using a pull model to perform DCT. If the producer-consumer interaction relies on software coherency, then the software coherency mechanism essentially negates the benefits of DCT initiated by the consumer 135. This is because software coherency typically involves either requiring the producer 120 to flush all updated content of its cache to memory 105 before the producer 120 indicates completion of the task (using metadata or the setting of a task completion flag) or requiring another processing unit (e.g., a producer-consumer task manager, which may be a CPU or an accelerator) to enforce software coherency by sending a CMO that, when executed, effectively accomplishes flushing all updated content in any cache to memory 105. The drawback is that by the time the consumer 135 initiates the DCT, the data has already been flushed from the producer cache 125 and any other caches (such as last level caches), so there is no opportunity to transfer data between the producer and consumer caches.

[0035] If the producer-consumer interaction relies on hardware coherency, the producer 120, after communication of completion in the metadata or setting of a task done flag, typically moves on to the next data set (e.g., the next task) or performs a different function on a different data set. These subsequent producer actions may result in cache capacity contention, as the previous updated data set is evicted from the producer cache 125 to make room for the new data set. Again, there is a similar drawback, since by the time the consumer 135 accesses the data, either the producer cache 125 or the hierarchical caching system between the producer 120 and memory 105 has already replaced the updated data set with the new working data set. Again, there is no opportunity for DCT. The ideal forwarding overlap window between when a producer 120 evicts its updated cached content and when a consumer 135 accesses those updated cached content may be further exacerbated by asymmetries when the producer cache 125 is smaller than the consumer cache 140 or when there are asymmetries in operating speeds, e.g., when the compute engine 130 of the producer 120 has a faster operating frequency than the compute engine 145 of the consumer 135.

[0036] Moreover, the pull model has system-wide drawbacks. In a CC-NUMA system where producer-consumer pairs are close to each other (e.g., in the same domain 115) and far from memory 105, the producer-consumer pairs have improved bandwidth and latency attributes relative to each other with respect to memory 105. However, system 100 can only take advantage of that proximity with a small subset of data still stored in producer cache 125. Data sets that had to be transferred from memory 105 to consumer 135, i.e., data sets that were evicted from producer cache 125, had to go to distant memory locations in memory 105 to complete the producer-consumer action. Not only did the producer-consumer action execute with lower performance as a result, but these actions further hindered other producer-consumer actions between other consumer-producer pairs (not shown in FIG. 1 ) on different data sets that were also homed by home 110. This can occur frequently in modern multi-processor and multi-accelerator cloud systems where overall computational efficiency is valued. The following embodiments do not rely on a secondary passive action of the consumer 135 to initiate a DCT (as is done in the pull model) and can increase the amount of data available for DCT between producer and consumer caches.

[0037] 2 is a block diagram of a computing system 200 in which a producer initiates a DCT, according to an example. The computing system 200 includes the same hardware components as the computing system 100 of FIG 1. In one embodiment, the home 110 and the memory 105 are remote from the producer 120 and the consumer 135. In other words, the home 110 is remote from the domain 115.

[0038] 2 also illustrates several actions 205 that are performed during the DCT. For ease of explanation, these actions 205 are illustrated in parallel with the blocks of the flowchart of FIG.

[0039] 3 is a flowchart of a method 300 for a producer to initiate a DCT, according to an example. In one embodiment, the method 300 is performed when maintaining coherency on data in the system 200 using hardware coherency. In contrast, maintaining data coherency using software coherency is described in Figures 4 and 5 below.

[0040] At block 305, the producer 120 identifies the consumer 135 for the data currently being processed by the producer 120. For example, when assigning a task to the producer 120, a software application may inform the producer 120 which processing elements in the system 200 are consumers of the data. Thus, the producer 120 knows the destination consumer 135 for the processed data from the beginning (or at least before the producer finishes the task).

[0041] At block 315, the producer 120 processes the data set for the task. As mentioned above, the task may be a search task, a machine learning task, a compression / decompression task, an encryption / decryption task, etc. Upon executing the task, the computation engine 130 of the producer 120 generates processed data.

[0042] In block 320, the computation engine 230 for the producer 120 stores the processed data in the producer cache 130. This corresponds to action 205A in FIG.

[0043] Method 200 has two alternative paths from block 320. The method can proceed to either block 325 where producer cache 125 performs a DCT to consumer cache 140 in response to a flush-stash CMO, or to block 330 where producer cache 125 performs a DCT to consumer cache 140 in response to a copyback-stash CMO during cache eviction. Both flush-stash and copyback-stash CMOs are new CMOs that can be used to initiate a DCT at producer 120 instead of consumer 135. That is, flush-stash and copyback-stash CMOs are new operation codes (op codes) that allow a producer to initiate a DCT transfer to a consumer cache.

[0044] In block 325, the computation engine 130 for the producer 120 executes a flush-stash CMO (e.g., an example of a flush-type stash-CMO) when the producer 120 plans to update the memory 105 with the latest contents of the data being processed by the producer 120. That is, when updating the memory 105, the computation engine 130 initiates a DCT with the flush-stash CMO, and the updated data being sent to the memory 105A is also sent from the producer cache 125 to the consumer cache 140 via the DCT. As indicated by action 205B, the producer cache 125 sends the updated data to the memory 105A. In parallel (or at a different time), as indicated by action 205C, the producer cache 125 sends the same updated data to the consumer cache 140.

[0045] In block 330, the compute engine 130 of the producer 120 executes a copyback-stash CMO (e.g., an example of a copyback type of stash-CMO) when the producer cache 125 issues a capacity eviction because the processed data for the task is being replaced with a new data set for the next task or function. Thus, the copyback-stash CMO is executed in response to an eviction when the system wants to remove the processed data in the producer cache 125 for a new data set. In this case, action 205B indicates that the data is evicted from the producer cache 125 and stored in the memory hierarchy in memory 105A. Action 205C represents the result of the execution of the copyback-stash CMO in which the evicted data corresponding to the task is sent to the consumer cache 140 using the DCT. Actions 205B and 205C may occur in parallel or at different times.

[0046] Regardless of whether block 325 or 330 is used, consumer 135 can optionally choose to accept or reject the stash operation based on the capacity of consumer cache 140. In any case, the correctness of the operation is maintained by action 205B, where memory has an updated copy.

[0047] In block 335, the computation engine 145 in the consumer 135 processes the data in the consumer cache 140 after receiving the completion status, as indicated by action 205E of Figure 2. The consumer 135 can receive the completion status information from a variety of different sources. For example, the completion status may be stored in the cache 140 as part of the DCT, or a separate memory completion flag may be located in memory 105A.

[0048] For any portion of the data set that was not accepted during the stash operation (either in block 325 or block 330), consumer 135 retrieves an updated copy from memory 105, as indicated by action 205D. For the subset of the data set that is held in consumer cache 140 and was previously accepted by a stash operation, consumer 135 benefits from already having the updated content in cache 140, thereby not only getting low-latency, high-bandwidth cache access, but also avoiding having to retrieve data from memory 105, which has low bandwidth and high latency.

[0049] Although the above embodiments discuss using a CMO to perform producer-initiated DCT between a consumer and a producer in the same domain, the embodiments are not limited to such. A producer-consumer pair can be located in different domains. For example, the producer may be in a first extension box and the consumer is a second extension box connected using one or more switches. The producer may know which domain the consumer is in, but may not know which processing element in that domain has been selected (or will be selected) as the consumer of the data. Instead of sending processed data directly from the producer cache to a consumer cache in a different domain, in one embodiment, the producer can use the above CMO to perform a DCT from the producer cache to a cache in another domain (e.g., a cache in a switch in the second extension box). Thus, the CMO can be used to send data to a different domain rather than to a specific consumer cache. Once in the domain, the consumer can take the data. Doing so still avoids retrieving that subset of data from main memory and avoids using the home 110 as an intermediary.

[0050] Figure 4 is a block diagram of a computing system 400 in which a coherency manager initiates a DCT when a producer completes a task, according to one example. Unlike Figure 2, in which coherency is managed by hardware, in Figure 4, coherency is managed by a software application, namely, a management unit 405 (Mgmt unit). Otherwise, Figure 4 includes the same components as shown in Figure 2. Figure 4 also includes a number of actions 415A-E, which are described in conjunction with the flowchart of Figure 5.

[0051] 5 is a flowchart of a method 500 for a coherency manager to initiate a DCT when a producer completes a task, according to one example. In one embodiment, the coherency manager is a software application that manages coherency in a heterogeneous processing system (e.g., management unit 405 of FIG. 4).

[0052] At block 505, the producer 120 stores the processed data for the task in the producer cache 125. This is indicated by action 415A in FIG. 4 where the computation engine 130 provides the processed data to the producer cache 125. At this point, the producer 120 may not know which processing element in the system 400 will be the consumer of the processed data. That is, the management unit 405 (e.g., a software coherency manager) may not yet have selected which processing element will perform the next task on the processed data. For example, the management unit 405 may wait until the producer 120 has finished (or is nearly finished) before selecting a processing element for the consumer based on the current workload or the idle time of the processing element.

[0053] At block 510, the producer 120 informs the software coherency manager (e.g., the management unit 405) when the task is completed. In one embodiment, after completing the task, the producer 120 sets a flag or updates metadata to inform the management unit 405 of the status. At this point, the management unit 405 selects a consumer 135 to process the data generated by the producer 120 based, for example, on the current workload or idle time of the processing elements in the computing system.

[0054] In block 515, the management unit instructs the home of the processed data to perform DCT. In one embodiment, the management unit 405 uses the compute engine 410 (which is separate from the compute engine 130 in the producer 120) to send the stash CMO to the home 110, as shown by action 415B in FIG. 4. The stash-CMO can be the clean invalid-stash CMO and its persistent memory variant, or the clean shared-stash CMO and its persistent memory variant. These two CMO-stashes are also new CMOs with new corresponding opcodes. Note that the management unit 405 is a logical unit and does not have to be mutually exclusive with the producer compute engine 130. In some embodiments, the producer compute engine 130 moves on to execute the next producer task, and in other embodiments, the producer compute engine 130 executes software coherency management actions upon completing its producer task.

[0055] In block 520, upon receiving the stash-CMO (e.g., clean invalid-stash or clean shared-stash), the home 110 sends a snoop to the producer 120. In one embodiment, the home 110 sends a stash-CMO snoop that is consistent with the original stash-CMO operation sent to the producer 120 by the management unit 405. This is illustrated by action 415C in FIG.

[0056] In block 525, in response to receiving the stash-CMO from the home 110, the producer 120 updates the memory 105 with the latest value of the processed data stored in the producer cache 125. This is indicated by action 415D in FIG.

[0057] At block 530, producer 120 also performs a DCT to consumer cache 140 in response to the received CMO-stash operation. That is, producer 120 performs a DCT to consumer cache 140 of the latest value of the data in producer cache 125 corresponding to the task, as indicated by action 415E in Figure 4. Consumer 135 can optionally choose to accept or reject the stash operation based on the capacity of consumer cache 140. In either case, correctness of the operation is maintained by block 535, since memory 105A has an updated copy of the data.

[0058] In block 535, the management unit 405 informs the consumer 135 that the DCT is complete. In one embodiment, the management unit 405 determines when the stash-CMO sequence (and accompanying DCT) is complete and sets a completion flag that is monitored by the consumer 135. The computation engine of the consumer 135 retrieves the updated copy from the consumer cache 140. However, the producer cache 125 may not be able to store all the processed data for the task, so some of the data may have been moved to memory 105A. Thus, the consumer 135 may still need to retrieve a subset of the processed data from memory 105, as indicated by action 415F. However, for the subset of the data set transferred to the consumer cache 140 that was previously accepted by the stash-CMO operation, the consumer 135 benefits from the updated content being stored in the cache 140 using low-latency, high-bandwidth cache accesses, avoiding retrieving that data from the low-bandwidth, high-latency memory 105.

[0059] In one embodiment, the software management unit 415 provides a table that defines a system address map (SAM) data structure to the hardware that executes the flush, copyback, or other CMO, so that the hardware statically knows the stash target ID of the CMO based on the address range the CMO is executing in. The stash target ID is derived from the Advanced Programmable Interrupt Controller (APIC) ID, CPU ID, CCIX agent ID, and may also be determined by other means, including the target ID included as part of the stashing DCT CMO.

[0060] Some non-limiting advantages of the above embodiments include improved performance by consumers of data because the DCT is performed when updated data is evicted. The content accessed by the consumer is already in its own cache or another cache in the cache hierarchy. The consumer does not need to go to main memory, which can be far away in a CC-NUMA system, to retrieve the data set.

[0061] The present invention also provides further performance advantages. In the case of CC-NUMA systems where producer-consumer pairs are close to each other and far from main memory, producer-consumer pairs have superior bandwidth and latency attributes that are exploited during DCT up to the capacity of the consumer cache's ability to hold the data set. Furthermore, producer-consumer actions did not interfere with actions performed (at the system level) by other producer-consumer pairs on different data sets that were also returned by remote nodes or home agents. Thus, each producer-consumer pair across the system achieves higher aggregate performance with minimal interference due to the above-described embodiments.

[0062] 6 is a block diagram of a computing system 600 that performs pipelining using the DCT, according to an example. Similar to FIGS. 2 and 4, the system 600 includes a memory 105 (e.g., main memory) and a home 110 that may be remote from the coherent interconnect 150 and the processing elements 605.

[0063] A processing element 605 may be a CPU, an accelerator, an FPGA, an ASIC, a system on a chip (SOC), etc. A processing element 605 may be both a producer and a consumer, i.e., a processing element 605 may consume data processed by another processing element 605 and then process that data to generate data that is consumed by another processing element 605.

[0064] 6, the processing elements 605 are arranged in a daisy chain to perform different subtasks of an overall task. For example, the processing elements 605 may be configured to perform different tasks associated with security, machine learning, or data compression applications.

[0065] A pipeline acceleration system can efficiently run across a daisy chain of producers / consumers, regardless of where the memory is homed. Figure 6 shows actions 620A-F where each processing element 605 (except processing element 605A) acts as a consumer of a data set on which an acceleration task was performed before another processing element 605 acts as a producer. The processing element 605 also serves as a producer of the next data set that the consumer receives. Using the CMO and DCT discussed above, large amounts of pipelined accelerated data can be moved from cache 610 to cache 610, using the DCT that leverages the high performance local connections between them, without relying on the location of the remote home 110 where the memory 105 is hosted.

[0066] 7 is a block diagram of a computing system that performs CPU-to-computational memory interactions, according to an example. Although the above embodiment describes using DCT between a consumer-producer pair that may be two processing elements, similar techniques may be applied to a consumer-producer pair that includes a CPU 702 (or other processing element) and a computational memory in an accelerator device 750. Unlike traditional I / O models, the memory (not shown) and processing elements (e.g., request agent (RA) 720 and slave agent (SA) 725) in the accelerator device 750 are in the same coherent domain as the CPU 702 and its cache 710. Thus, the home 715 of the host 701 ensures that data stored in the host 701 and accelerator device 750 are stored coherently, and requests for memory operations originating from the host 701 or accelerator device 750 receive the most up-to-date version of the data, regardless of whether the data is stored in the host 701 or accelerator device 750 memory.

[0067] In some applications, the bottleneck when processing data comes from moving the data to the computation unit that processes it, not the time it takes to process the data. The situation where the time required to complete a task is limited by moving the data to the computation engine is called a computation memory bottleneck. Moving the computation engine closer to the data can help alleviate this problem. With the embodiments herein, the computation memory bottleneck can be reduced by using DCT to transfer data between the CPU and the computation memory, which reduces the amount of data retrieved from the main memory.

[0068] 7, the accelerator device 750 (e.g., a compute-memory accelerator) initiates an ownership request 760, which may be performed with a flush-invalidate-stash CMO similar to the CMO discussed above to perform the DCT. Additionally, the CPU 705 may initiate a flush 770, which may also be performed with a flush-invalidate-stash CMO (or flush-stash CMO), such that a cache in the accelerator device 750 contains the updated CPU-modified content, and the accelerator device 750 may then quickly perform compute-memory actions with this updated CPU-modified content.

[0069] In the foregoing, reference has been made to embodiments of the present disclosure. However, the present disclosure is not limited to the specific described embodiments. Instead, any combination of the aforementioned features and elements, whether related to different embodiments or not, is contemplated to realize and practice the present disclosure. Furthermore, the embodiments of the present disclosure may achieve other possible solutions and / or advantages over the prior art, but whether or not a particular advantage is achieved by a given embodiment does not limit the present disclosure. Thus, the aforementioned aspects, features, embodiments, and advantages are merely illustrative and are not considered elements or limitations of the claims unless expressly recited in the claims. Similarly, references to "disclosure" should not be construed as a generalization of any inventive subject matter disclosed herein, and should not be considered elements or limitations of the claims unless expressly recited in the claims.

[0070] Aspects of the present disclosure may be embodied as a system, method, or computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be generally referred to herein as a "circuit," "module," or "system." Additionally, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.

[0071] Any combination of one or more computer readable media may be utilized. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer readable storage media would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0072] A computer-readable signal medium may include a propagated data signal in which computer-readable program code is embodied, for example in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including but not limited to electromagnetic, optical, or any suitable combination thereof. A computer-readable signal medium is not a computer-readable storage medium, but may be any computer-readable medium that can communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0073] The program code embodied on the computer readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, fiber optic cable, RF, etc., or any suitable combination of the foregoing.

[0074] Computer program code for carrying out operations for aspects of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or a connection may be made to an external computer (e.g., through the Internet using an Internet Service Provider).

[0075] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, executed via the processor of the computer or other programmable data processing apparatus, generate means for performing the functions / operations specified in the blocks of the flowchart illustrations and / or block diagrams.

[0076] These computer program instructions may be stored on a computer-readable medium that can direct a computer, other programmable data processing apparatus, or other device to function in a particular manner, and the instructions stored on the computer-readable medium produce an article of manufacture that includes instructions that implement the functions / acts specified in the flowchart and / or block diagram blocks.

[0077] Computer program instructions are loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to form a computer-implemented process, and the instructions executing on the computer or other programmable apparatus provide a process for implementing the function / operation specified in a block or blocks of the flowcharts and / or block diagrams.

[0078] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of code, comprising one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. Each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and computer instructions.

[0079] While the forgoing is directed to embodiments of the present invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof, the scope of the invention being determined by the claims appended hereto.

Claims

1. 1. A computing system comprising: a producer including a first processing element configured to generate processed data; a producer cache configured to store the processed data generated by the producers; a consumer including a second processing element configured to receive and process the processed data generated by the producer; a consumer cache configured to store the processed data generated by the consumer; 1. A computing system comprising: a producer configured to, in response to receiving a stash cache maintenance operation (stash-CMO), transfer the processed data from the producer cache to a main memory, and in parallel, perform a direct cache transfer (DCT) to transfer the processed data from the producer cache to the consumer cache.

2. 10. The computing system of claim 1, further comprising at least one coherent interconnect communicatively coupling said producers to said consumers, said computing system being a cache coherent non-uniform memory access (CC-NUMA) system.

3. The computing system of claim 1 , wherein the producer knows the location of the consumer before the producer completes a task of instructing the producer to generate the processed data.

4. 2. The computing system of claim 1, wherein coherency of the computing system is maintained by a hardware element, the stash-CMO and the DCT are executed in response to the producer determining to update a main memory in the computing system with the processed data currently stored in the producer cache, and the stash-CMO is a flash-type stash-CMO.

5. 2. The computing system of claim 1, wherein coherency of the computing system is maintained by a hardware element, the stash-CMO and the DCT are executed in response to the producer cache issuing a capacity eviction to remove at least a portion of the processed data, and the stash-CMO is a copy-back stash-CMO.

6. The computing system of claim 1 , wherein the producer does not know the whereabouts of the consumer before the producer completes a task that instructs the producer to generate the processed data.

7. 2. The computing system of claim 1, wherein coherency of the computing system is maintained by a software management unit, the producer is configured to notify the software management unit when a task of generating the processed data is completed, and the software management unit is configured to instruct a home of the processed data to begin the DCT.

8. The software management unit sends a first STASH-CMO to the home to start the DCT on the producer; 8. The computing system of claim 7, wherein in response to receiving the first stash-CMO, the home sends a snoop to the producer that includes the stash-CMO instructing the producer to execute the DCT.

9. 1. A method comprising: generating processed data at a producer including a first hardware processing element; storing the processed data in a producer cache; in response to receiving a STASH-CMO, transferring the processed data from the producer cache to a main memory and in parallel performing a DCT to transfer the processed data from the producer cache to a consumer cache; processing the processed data in a consumer including a second hardware processing element after the DCT; storing the processed data generated by the consumer in the consumer cache.

10. 10. The method of claim 9, wherein the DCT is performed using at least one coherent interconnect communicatively coupling the producers to the consumers, the producers and the consumers being part of a CC-NUMA system.

11. 10. The method of claim 9, further comprising: the producer informing the producer of the location of the consumer before the producer completes a task of instructing the producer to generate the processed data.

12. 10. The method of claim 9, wherein coherency between the consumer and the producer is maintained by a hardware element, the stash-CMO and the DCT are executed in response to the producer determining to update a main memory with the processed data currently stored in the producer cache, and the stash-CMO is a flush-type stash-CMO.

13. 10. The method of claim 9, wherein coherency between the consumer and the producer is maintained by a hardware element, the stash-CMO and the DCT are performed in response to the producer cache issuing a capacity eviction to remove at least a portion of the processed data, and the stash-CMO is a copy-back stash-CMO.

14. 10. The method of claim 9, wherein the producer does not know the whereabouts of the consumer before the producer completes a task of instructing the producer to generate the processed data.

15. Coherency between the consumer and the producer is maintained by a software management unit, the method further comprising: notifying the software management unit when the task of generating the processed data is completed by the producer; 10. The method of claim 9, further comprising instructing a home of the processed data to begin the DCT.

Citation Information

Patent Citations

  • Information processor, processor, control method of processor, control method of information processor, and cache memory

    JP2005316854A

  • Multiprocessor system

    JP2007241601A

  • Arithmetic processing unit, information processor and method for controlling arithmetic processing unit

    JP2014048986A

  • Cache configured to log addresses of high-availability data via a non-blocking channel

    US9336142B2