Method, apparatus, and system for prefetching exclusive cache coherence states of memory instructions
By obtaining the exclusive cache consistency state related to storage instructions in a multi-CPU system, the problem of delay in obtaining this state is solved, and system performance is improved and power waste is reduced.
Patent Information
- Application Number
- CN201980055916.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-08-27
- Filing Date
- 2019-08-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2039-08-26
AI Technical Summary
In multi-CPU systems, the process of obtaining exclusive cache coherence states associated with storage instructions may involve significant delays, resulting in system performance degradation and power waste.
Writing to the cache is performed by determining whether the cache line associated with the storage instruction should be allocated in the cache and obtaining the exclusive cache consistency state of the cache line if necessary.
Reduces the delay involved in obtaining cache consistency states exclusively by the stored instructions, improves system performance, and reduces the associated power waste when waiting for the stored instructions to complete.
Smart Images

Figure CN112602067B_ABST
Abstract
Description
[0001] Priority Claim
[0002] This patent application claims priority to non-provisional application Ser. No. 16 / 113,120, filed Aug. 27, 2018, titled "Methods, Apparatus, and Systems for Exclusive Cache Coherence States for Prefetch Store Instructions", which is assigned to the assignee of the present application and is hereby incorporated herein by reference in its entirety. Technical Field
[0003] Aspects of the present disclosure generally relate to store instructions and, more particularly, to exclusive cache coherence states for prefetch store instructions. Background Art
[0004] A computing device may execute memory access instructions (e.g., load instructions and store instructions) as part of normal processing operations. In a computing device having multiple central processing units (CPUs), the computing device may execute a hardware coherence protocol to ensure that any associated cache memory and system memory shared among the multiple CPUs are updated in a consistent manner in response to memory access instructions (and particularly store instructions).
[0005] In a system where multiple CPUs may share access to a particular memory location, one particular method of ensuring coherence of store instructions is a barrier instruction. A barrier instruction is an instruction that forces all stores prior to the barrier instruction to be visible to all CPUs in the computing device before proceeding with operations allowed after the barrier instruction. This ensures that CPUs working on shared memory values receive the correct updated data so that those CPUs can make progress, as CPUs working on old data will effectively waste cycles doing that work. To allow the barrier instruction to complete, a CPU working on a particular shared data segment (i.e., a particular memory location) will acquire the exclusive cache coherence state of that data.
[0006] However, in modern computing devices having many CPUs (especially in the case of a server system-on-chip (SoC) which can have dozens or more CPUs on a single SoC), due to system bus contention or other factors, the process of obtaining the exclusive cache coherence state of shared data may involve significant latency. In addition, some CPU architectures can aggregate outstanding store instructions and perform related memory transactions only on a periodic basis (i.e., update the main memory locations associated with those store instructions). Thus, if a CPU waits until a store instruction is otherwise completed to retrieve the exclusive cache coherence state, the CPU can be forced to stall for a relatively large number of cycles (and thus, any other CPU waiting for the data can also be forced to stall). This results in an undesired performance degradation of the system and wasted power, as the computing device must remain active but cannot make progress.
[0007] Accordingly, it is desirable to provide a mechanism for reducing the latency involved in obtaining the exclusive cache coherence state associated with a store instruction. SUMMARY OF THE INVENTION
[0008] A simplified summary of one or more aspects is presented below to provide a basic understanding of these aspects. This summary is not an extensive review of all contemplated aspects, and is neither intended to identify key or critical elements of all aspects nor to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description presented later.
[0009] In one aspect, a method includes determining whether a cache line associated with a store instruction should be allocated in a cache. The method further includes, if the cache line associated with the store instruction should be allocated in the cache, performing a write-ahead to the cache by obtaining the exclusive cache coherence state of the cache line associated with the store instruction. The write-ahead can be selectively enabled or disabled by software.
[0010] In another aspect, an apparatus includes a cache and a collection buffer coupled to the cache. The collection buffer is configured to store a plurality of cache lines, where each cache line of the plurality of cache lines is associated with a store instruction. The collection buffer is further configured to determine whether a first cache line associated with a first store instruction should be allocated in the cache. The collection buffer is further configured to, if the first cache line associated with the first store instruction is to be allocated in the cache, issue a write-ahead request to obtain the exclusive cache coherence state of the first cache line associated with the first store instruction.
[0011] In yet another aspect, a non-transitory computer-readable medium includes instructions that, when executed by a processor, cause the processor to determine whether a cache line associated with a store instruction should be allocated in a cache. If the cache line associated with the store instruction should be allocated in the cache, the instructions further cause the processor to perform a write-ahead of the cache by obtaining an exclusive cache coherence state for the cache line associated with the store instruction.
[0012] In yet another aspect, an apparatus includes means for storage and means for storing memory access instructions coupled to the means for storage. The means for storing memory access instructions is configured to store multiple lines, each line being associated with a store instruction. The means for storing memory access instructions is further configured to issue a write-ahead request to the means for storage to obtain an exclusive coherence state for a first line associated with a first store instruction if the first line associated with the first store instruction should be allocated in the means for storage.
[0013] An advantage of one or more of the disclosed aspects is that the disclosed aspects allow for a reduction in latency associated with obtaining an exclusive cache coherence state for a store instruction. In some aspects, this can improve system performance and reduce wasted power associated with stalls in a computing device while waiting for a store instruction to complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A block diagram showing a computing device configured to prefetch an exclusive cache coherence state of a store instruction in accordance with certain aspects of the present disclosure.
[0015] Figure 2 A detailed block diagram showing a collection buffer and a secondary cache in accordance with certain aspects of the present disclosure.
[0016] Figure 3 A block diagram showing a method for prefetching an exclusive cache coherence state of a store instruction in accordance with certain aspects of the present disclosure.
[0017] Figure 4 A system-level diagram showing a computing device configured to prefetch an exclusive cache coherence state of a store instruction in accordance with certain aspects of the present disclosure. DETAILED DESCRIPTION
[0018] Aspects of the inventive teachings herein are disclosed in the following description and related drawings directed to particular aspects. Alternative aspects may be devised without departing from the scope of the inventive concepts herein. Additionally, well-known elements of the environment may not be described in detail or may be omitted so as not to obscure relevant details of the inventive teachings herein.
[0019] The term "exemplary" is used herein to mean "serving as an example, instance, or illustration". Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects. Similarly, the term "aspect of the invention" does not require that all aspects of the invention include the discussed feature, advantage, or mode of operation.
[0020] The terms used herein are for the purpose of describing particular aspects only and are not intended to limit aspects of the invention. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the terms "comprises", "comprising", "includes", and / or "including", when used herein, specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0021] In addition, many aspects are described in terms of sequences of actions to be performed by, for example, elements of a computing device. It will be recognized that the various actions described herein can be performed by specific circuitry (e.g., an application specific integrated circuit (ASIC)), by program instructions being executed by one or more processors, or by a combination of both. Additionally, the sequences of actions described herein can be considered to be fully embodied within any form of computer-readable storage medium having stored therein a corresponding set of computer instructions that, when executed, would cause an associated processor to perform the functionality described herein. Accordingly, the various aspects of the invention can be implemented in many different forms, all of which are considered to be within the scope of the claimed subject matter. Additionally, for each aspect described herein, any such aspect's corresponding form can be described herein as, for example, "logic configured to" perform the described action.
[0022] Figure 1 A block diagram of a computing device 100 configured to prefetch store instructions for exclusive cache coherence states in accordance with certain aspects of the present disclosure is shown. The computing device includes CPUs 110, 110a, and 110b. More particularly, CPU 110 includes a collection buffer 112. The collection buffer 112 is a logic block configured to collect outstanding memory access instructions (including store instructions) in CPU 110 and control when those store instructions are presented to the associated cache and memory system, which, in the aspect shown, can include a level 2 cache 120 coupled to a main memory 140 via a system bus 130.
[0023] CPUs 110a and 110b may also include their own collection buffers (not shown), which may be configured to operate in a manner similar to collection buffer 112. CPUs 110a and 110b may also each have an associated level 2 cache 120a and 120b, which are coupled to main memory 140 via system bus 130. In one aspect, each of level 2 caches 120a and 120b may have an architecture substantially similar to that of level 2 cache 120 and may be configured to operate in a manner similar to level 2 cache 120. In another aspect, level 2 caches 120a and 120b may have a heterogeneous architecture relative to level 2 cache 120. Those skilled in the art will recognize that many configurations of CPUs 110, 110a, and 110b and level 2 caches 120, 120a, and 120b are possible, and the teachings of the present disclosure may be applied to any system, whether heterogeneous or homogeneous, in which a subset of CPUs and processors share access to a portion of a memory hierarchy (in one aspect, main memory 140).
[0024] In operation, when a store instruction is executed in CPU 110, an entry associated with the store instruction is allocated in collection buffer 112. When the entry is allocated in collection buffer 112, information relating to the store instruction is transmitted to the associated level 2 cache 120. If the store instruction is to be allocated in level 2 cache 120 (i.e., stored in non-transitory or temporary data that will not be used by any other CPU of computing device 100), the allocation in collection buffer 112 is accompanied by a write-ahead request generated by collection buffer 112 and presented to level 2 cache 120. The purpose of the write-ahead request is to obtain an exclusive coherence state for the line associated with the store instruction for level 2 cache 120 (rather than the cache line associated with the store instruction). This has the additional advantage of allowing collection buffer 112 to know about other demands for the data contained in that cache line (as a result of keeping that cache line in an exclusive coherence state). For architectures that require writes to reach memory within a finite amount of time, this may allow computing device 100 to indefinitely postpone the write while still nominally meeting the architecture requirements.
[0025] In one aspect, this can reduce the memory or bus bandwidth associated with executing store instructions. Since the write-ahead does not fetch any data, the amount of information transferred over the memory bus or system bus is reduced compared to an operation that fetches both the exclusive cache coherence state and associated data. Additionally, since the level 2 cache 120 already has exclusive access to the cache line associated with the store instruction when the store instruction is ready to complete, other CPUs (e.g., CPU 110a or 110b) that may be waiting for the data generated by the store instruction will not be delayed by the latency involved in the level 2 cache 120, which first acquires the exclusive cache coherence state for the cache line associated with the store instruction prior to executing the store instruction and any associated memory synchronization operations (e.g., barrier instructions).
[0026] Figure 2 FIG. 200 shows a detailed block diagram of the gather buffer 112 and the level 2 cache 120 in accordance with certain aspects of the present disclosure. The gather buffer 112 can include write-ahead logic 216 and a data array 214. The data array 214 is configured to store cache lines (i.e., cache lines that include store instructions) associated with store instructions executed by the CPU 110. The data array 214 is coupled to the write-ahead logic 216, which is configured to: determine when a new cache line has been allocated in the data array 214 as a result of a store instruction executed by the CPU 110, and generate a write-ahead request 218 that is transmitted to the level 2 cache 120.
[0027] The level 2 cache 120 can include a memory array 222 that includes a plurality of individual cache lines 224a-d. Each cache line 224 can also include a cache line coherence state indicator 225 and a tag / data portion 226. The level 2 cache 120 also includes a management block 228 that is configured to service write-ahead requests and manage the cache lines 224a-d.
[0028] In operation, the gather buffer 112 may allocate row 214a in the data array in response to a store instruction executed by the CPU 110. The gather buffer 112 may also generate a write-ahead request 218 associated with the store instruction. The management block 228 receives the write-ahead request 218 from the gather buffer 112, allocates a cache line associated with the store instruction executed by the CPU 110 (e.g., cache line 224a, corresponding to row 214a in the gather buffer 112), and sends a request to the memory hierarchy (not shown) to obtain the exclusive cache coherence state of the cache line 224a. If and when the management block 228 receives the exclusive cache coherence state of the cache line 224a, the management block 228 updates the cache line coherence state indicator 225 associated with the cache line 224a to indicate the exclusive cache coherence state. However, the tag / data portion 226 of the cache line 224a is not updated in response to the write-ahead request. Since the cache line 224a of the level 2 cache 120 now has an exclusive cache coherence state, when the associated store instruction is ready to complete, the cache line 224a can be immediately updated since the relevant permission has been obtained.
[0029] Figure 3 FIG. shows a block diagram of a method 300 for prefetching an exclusive cache coherence state for a store instruction in accordance with certain aspects of the present disclosure. The method begins at block 310: determining whether a cache line associated with a store instruction should be allocated in the cache. For example, referring Figure 1 and Figure 2 , the gather buffer 112 may determine whether a store instruction executed by the CPU 110 is transient or whether the store instruction should be allocated in the level 2 cache 120.
[0030] If the store instruction should not be allocated in the cache, the method 300 ends at block 315. However, if the store instruction should be allocated in the cache, a write-ahead to the cache is performed by obtaining the exclusive cache coherence state of the cache line associated with the store instruction, and the method proceeds to block 320. For example, referring Figure 2 , the level 2 cache 120 receives the write-ahead request 218 from the gather buffer 112, allocates the cache line 224a associated with the store instruction, and obtains the exclusive cache coherence state for the cache line 224a from the memory hierarchy. In some aspects, block 320 may include setting the cache line coherence state indicator 225 associated with the cache line 224a to indicate the exclusive cache coherence state.
[0031] In one aspect, the write-ahead of the cache can be triggered by a first store instruction to a particular cache line that does not yet exist in the gather buffer. For example, for a store instruction where the associated cache line already exists in the gather buffer, the store instruction may not trigger a write-ahead because it is assumed that the gather buffer will already have triggered a write-ahead request in response to the store instruction that originally caused the cache line to be allocated in the gather buffer. Additionally, the execution of write-ahead can be enabled or disabled by software because there may be known situations or code sequences where performing a write-ahead may harm performance (e.g., situations where the interface between the CPU 110 and the level 2 cache 120 is near the capacity for write operations, especially those that write back to memory quite aggressively, in which cases the write-ahead generates additional traffic for prefetch permissions, which may overwhelm the CPU / level 2 cache interface and result in a performance loss).
[0032] In block 330, the method continues by performing a write to the cache line associated with the store instruction. For example, the gather buffer 112 provides the line 214a that has been updated with the result of the store instruction to the level 2 cache 120, and the level 2 cache updates the associated cache line 224a based on the line 214a.
[0033] Performing a write to the cache line associated with the store instruction can be triggered in various different ways, all of which are within the scope of the teachings of the present disclosure. In one aspect, a write to the cache line can be triggered in response to a snoop request for the cache line from another CPU (i.e., a request from another CPU for exclusive access to the cache line). In yet another aspect, a write to the cache line can be triggered by an architectural request or instruction, such as a barrier instruction.
[0034] Now, examples of apparatuses that can utilize aspects of the present disclosure will be discussed with respect to Figure 4 A diagram of a computing device 400 is shown that includes a structure for prefetching the exclusive cache coherence state of store instructions as described with reference to Figure 4 and Figure 1 and the computing device can operate according to the method described in Figure 2 In this aspect, the system 400 includes a processor 402 that can incorporate the CPU 110 and the gather buffer 112, the level 2 cache 120, the system bus 130, and as described with respect to Figure 3 and Figure 1 and Figure 2 The system 400 further includes a main memory 140 coupled to the processor 402 via the system bus 130. The memory 140 can also store non-transitory computer-readable instructions that, when executed by the processor 402, can perform Figure 3Method 300.
[0035] Figure 4 Optional boxes are also shown in dashed lines, such as an encoder / decoder (CODEC) 434 (e.g., an audio and / or voice CODEC) coupled to the processor 402, and the speaker 436 and the microphone 438 may be coupled to the CODEC 434; and a wireless antenna 442 coupled to the wireless controller 440, with the wireless controller coupled to the processor 402. Additionally, the system 402 also shows a display controller 426 coupled to the processor 402 and the display 428, and a wired network controller 470 coupled to the processor 402 and the network 472. In the presence of one or more of these optional boxes, in certain aspects, the processor 402, the display controller 426, the memory 432, and the wireless controller 440 may be included in a system-in-package or system-on-chip device 422.
[0036] Thus, in certain aspects, the input device 430 and the power supply 444 are coupled to the system-on-chip device 422. Additionally, in certain aspects, as Figure 4 illustrated, in the presence of one or more optional boxes, the display 428, the input device 430, the speaker 436, the microphone 438, the wireless antenna 442, and the power supply 444 are external to the system-on-chip device 422. However, each of the display 428, the input device 430, the speaker 436, the microphone 438, the wireless antenna 442, and the power supply 444 may be coupled to components of the system-on-chip device 422, such as an interface or a controller.
[0037] It should be noted that although Figure 4 a computing device is generally depicted, the processor 402 and the memory 404 may also be integrated into a mobile phone, a communication device, a computer, a server, a laptop computer, a tablet computer, a personal digital assistant, a music player, a video player, an entertainment unit, and a set-top box or other similar devices.
[0038] Those skilled in the art will appreciate that any of a variety of different technologies and techniques may be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chipsets that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0039] In addition, those skilled in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the aspects disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.
[0040] The methods, sequences, and / or algorithms described in connection with the aspects disclosed herein can be embodied directly in hardware, in software modules executed by a processor, or in combinations thereof. The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor.
[0041] While the foregoing disclosure shows illustrative aspects of the present invention, it should be noted that various changes and modifications can be made herein without departing from the scope of the present invention as defined by the appended claims. The functions, steps, and / or acts of the method claims in accordance with the aspects of the present invention described herein need not be performed in any particular order. Moreover, although the elements of the present invention may be described or claimed in the singular, the plural is also included unless expressly stated to be limited to the singular.
Claims
1. A method for prefetching an exclusive cache coherence state of a store instruction, comprising: executing a store instruction; in response to the execution of the store instruction, determining whether the store instruction should allocate a cache line associated with the store instruction in the cache; if it is determined that the cache line associated with the store instruction should be allocated in the cache, issuing a write - ahead request to the cache; and in response to the write - ahead request: allocating in the cache the cache line associated with the store instruction; and acquiring the exclusive cache coherence state of the cache line associated with the store instruction without updating the tag / data portion of the cache line.
2. The method according to claim 1, further comprising: when the store instruction is determined to be transient, determining that the cache line associated with the store instruction should not be allocated in the cache.
3. The method according to claim 1, wherein issuing the write - ahead request is selectively triggered by determining that the store instruction is a non - transient store instruction associated with a first cache line, and wherein the first cache line does not exist in the collection buffer.
4. The method according to claim 1, further comprising: selectively prohibiting the issuance of the write - ahead request by a software module.
5. The method according to claim 1, wherein acquiring the exclusive cache coherence state of the cache line associated with the store instruction comprising: retrieving from the memory hierarchy the exclusive cache coherence state for the cache line and setting an indicator associated with the cache line indicating that the cache line has an exclusive cache coherence state.
6. The method according to claim 1, further comprising: updating the tag / data portion of the cache line associated with the store instruction to complete the execution of the store instruction.
7. The method according to claim 6, wherein updating the tag / data portion of the cache line comprising: writing data associated with the store instruction from a collection buffer in a processing unit to the cache line in the cache.
8. The method according to claim 6, wherein updating the tag / data portion of the cache line is triggered in response to a snoop request for the cache line.
9. The method according to claim 6, wherein updating the tag / data portion of the cache line is triggered in response to an architecture request or an instruction.
10. The method according to claim 9, wherein the instruction is a barrier instruction.
11. An apparatus for prefetching an exclusive cache coherence state of a store instruction, comprising: a cache; and a collection buffer, coupled to the cache and configured to store a plurality of cache lines, each cache line in the plurality of cache lines being associated with a corresponding store instruction, wherein: the apparatus is configured to execute a first store instruction; the collection buffer is further configured to: In response to the execution of the first store instruction, determine whether the first store instruction should be allocated a first cache line associated with the first store instruction in the cache; and if it is determined that the first cache line associated with the first store instruction should be allocated in the cache, issue a write-ahead request to the cache; and the cache is configured to: allocate the first cache line; and acquire an exclusive cache coherence state of the first cache line associated with the first store instruction without updating the tag / data portion of the first cache line.
12. The apparatus according to claim 11, wherein the collection buffer is configured to: when the first store instruction is determined to be transient, not issue the write-ahead request.
13. The apparatus according to claim 11, wherein the collection buffer is configured to: when the first store instruction is determined to be non-transient and the first cache line associated with the first store instruction does not exist in the collection buffer, issue the write-ahead request.
14. The apparatus according to claim 11, wherein the software module is configured to selectively prohibit the collection buffer from issuing the write-ahead request.
15. The apparatus according to claim 11, further comprising: a memory hierarchy including a main memory coupled to the cache; wherein the cache further includes a plurality of cache lines, each cache line having a coherence state indicator and a tag / data portion; and wherein acquiring the exclusive cache coherence state of the first cache line associated with the first store instruction includes: retrieving the exclusive cache coherence state from the memory hierarchy and setting the coherence state indicator associated with the first cache line to indicate that the cache line has the exclusive cache coherence state, the first cache line being associated with the first store instruction.
16. The apparatus according to claim 11, wherein the collection buffer is configured to: receive data to be written as a result of the execution of the first store instruction and write the data into a cache line in the collection buffer associated with the first store instruction.
17. The apparatus according to claim 16, wherein the cache is configured to: receive data to be written as a result of the execution of the first store instruction from the collection buffer and write the data into the tag / data portion of the first cache line of the cache associated with the store instruction.
18. The apparatus according to claim 17, wherein writing the data into the first cache line of the cache is triggered in response to a snoop request for the first cache line.
19. The apparatus according to claim 17, wherein writing the data into the first cache line of the cache is triggered in response to an architecture request or instruction.
20. The apparatus according to claim 19, wherein the instruction is a barrier instruction.
21. The apparatus according to claim 11, the apparatus being integrated into a device selected from the group consisting of a mobile phone, a communication device, a computer, a server, a laptop computer, a tablet computer, a personal digital assistant, a music player, a video player, an entertainment unit, and a set-top box.
22. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to: Execute a store instruction; In response to execution of the store instruction, determine whether a cache line associated with the store instruction should be allocated in a cache; ; And If it is determined that the cache line associated with the store instruction should be allocated in the cache, issue a write-ahead request to the cache; And In response to the write-ahead request: Allocate in the cache the cache line associated with the store instruction; and Obtain an exclusive cache coherence state of the cache line associated with the store instruction without updating a tag / data portion of the cache line.
23. An apparatus for prefetching an exclusive cache coherence state of a store instruction, Comprising: Means for storing; And Means for storing memory access instructions, coupled to the means for storing, and configured to store a plurality of lines, each line being associated with a corresponding store instruction, wherein: The apparatus is configured to execute a first store instruction; The means for storing memory access instructions is further configured to: In response to execution of the first store instruction, determine whether a first line associated with the first store instruction should be allocated in the means for storing; If it is determined that the first line associated with the first store instruction should be allocated in the means for storing, issue a write-ahead request to the means for storing; and The means for storing is configured to: Allocate the first line; and Obtain an exclusive coherence state of the first line associated with the first store instruction without updating a tag / data portion of the first line.
Citation Information
Patent Citations
Method and apparatus for pipelining ordered input / output transactions in a cache coherent multi-processor system
CN1470019A
Cache migration
US20150149730A1