System architecture and techniques for fixed memory migration

Through the PMM logic's 'dual projection' operation and transactional memory copy technology, the DMA pause problem caused by the immobility of fixed memory pages is solved, improving system reliability and memory management efficiency.

CN120723151APending Publication Date: 2025-09-30INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510197294.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-30
Filing Date
2025-02-21
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

In the prior art, the immobility of fixed memory pages causes DMA operations to be suspended, leading to I/O structure backpressure, latency, and system-wide timeouts, affecting the reliability, availability, and serviceability of cloud, edge, and client systems.

Method used

A new 'dual cast' operation is generated by the PMM logic to ensure that write operations from the I/O device are observed in both the source and destination pages, and a transactional memory copy is performed to ensure data consistency to support ongoing DMA operations.

Benefits of technology

The migration of fixed memory pages is implemented, which improves system reliability, availability, and serviceability, reduces power consumption, and optimizes memory tiering and defragmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723151A_ABST
    Figure CN120723151A_ABST
Patent Text Reader

Abstract

Methods and apparatus relating to system architectures and techniques for fixed memory migration are described. In an embodiment, one or more memory devices store a source page and a destination page. Logic circuitry causes write operations from an input / output (I / O) device to be observed for both a source page and a destination page. A transactional memory copy operation is performed from a source page to a destination page. Other embodiments are also disclosed and claimed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to the field of memory and, more particularly, to system architecture and techniques for fixed memory migration. Background Art

[0002] In some implementations, memory pages used for input / output (I / O) purposes may be pinned (e.g., by system software) so that these pinned memory pages cannot be moved for a fixed duration. For example, such pinning can ensure data correctness in situations where there may be in-flight access to these memory pages. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] The detailed description is provided with reference to the accompanying drawings. In the drawings, the leftmost digit(s) of a reference number identifies the drawing in which the reference number first appears. The use of the same reference numbers in different drawings indicates similar or identical items.

[0004] Figure 1 A block diagram illustrating a system architecture according to an embodiment.

[0005] Figure 2 FIG. 1 shows a diagram of a method for Figure 1 A block diagram of some components of the PMM Agent's control interface.

[0006] Figure 3 A block diagram illustrating components for performing a dual projection operation is shown in accordance with an embodiment.

[0007] Figure 4 A flow chart illustrating a method for providing fixed storage migration according to an embodiment is shown.

[0008] Figure 5 An example computing system is illustrated.

[0009] Figure 6 A block diagram illustrates an example processor and / or System on a Chip (SoC) that may have one or more cores and an integrated memory controller.

[0010] 7(A) is a block diagram illustrating both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline, according to an example.

[0011] 7(B) is a block diagram illustrating both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to an example.

[0012] Figure 8 An example of execution unit(s) circuitry is shown. DETAILED DESCRIPTION

[0013] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the various embodiments. However, the various embodiments may be implemented without these specific details. In other instances, well-known methods, processes, components, and circuits are not described in detail to avoid obscuring the particular embodiments. Further, various aspects of the embodiments may be performed using various means such as integrated semiconductor circuits ("hardware"), computer-readable instructions organized into one or more programs ("software"), or some combination of hardware and software. For the purposes of this disclosure, references to "logic" shall mean hardware (such as, logic circuitry or more generally, circuitry or circuits), software, firmware, or some combination thereof.

[0014] As mentioned above, memory page pinning can ensure data correctness in situations where there may be ongoing accesses to pinned memory pages. In some implementations, Direct Memory Access (DMA) operations can be suspended. However, pausing DMA operations can cause I / O fabric backpressure, which can block or delay DMA requests, causing system-wide timeouts, etc.

[0015] To this end, some embodiments provide system architectures and / or techniques for fixed memory migration. In an embodiment, when there may be ongoing access (ongoing access) to one or more fixed memory pages from (one or more) I / O devices, these fixed memory pages can be migrated or moved. The inability to move fixed memory pages may affect many cloud, edge, and client systems. Therefore, some embodiments implement migration of fixed pages to allow improved operations, such as memory off-line, power consumption reduction, memory tiering, memory defragmentation, memory pooling, etc. for reliability, availability, and serviceability (RAS).

[0016] In at least one embodiment, a new system architecture implements pinned page migration in the following manner: (1) system software (e.g., an operating system (OS), a virtual machine manager (VMM), etc.) passes the addresses of the source and destination pages associated with the migration to pinned memory migration (PMM) logic (e.g., a logic circuit system); (2) the PMM logic ensures that write operations from I / O devices are observed not only in the source page / region but also in the destination page / region by generating a new "dualcast" operation to the destination page / region; and (3) the system software performs a "transactional memory copy" to ensure data consistency between the source page / region and the destination page / region to support any in-flight DMA operations. In an embodiment, the PMM logic can be provided on an integrated circuit device (such as a system on a chip ("SOC" or "SoC")), such as reference 1. Figure 1 and discussed in the following figures.

[0017] Current DMA pause techniques may involve various operations that may reduce efficiency and / or speed. For example, to implement DMA pause: (1) DMA operations on a source page may be paused; (2) the source page may then be unmapped from the IO Memory Management Unit (IOMMU) and any IO Translation Lookaside Buffers (IOTLBs) may be synchronized; (3) the source page may then be copied to a destination page; (4) the destination page may then be mapped in the (one or more) IOMMU page tables; and (5) DMA is resumed. Such DMA pause operations may be functionally fragile. For example, a DMA operation may need to be paused for a meaningful duration (e.g., including unmapping a memory page, followed by invalidation of the IOTLB, copying the memory page, and mapping the memory page). Thus, pausing a DMA operation may cause I / O fabric backpressure, which in turn blocks or delays DMA requests following the I / O fabric. As an example, an upstream completion sequenced after a suspended DMA write operation may then trigger a timeout for the Central Processing Unit (CPU) (eg, when a driver issues a Memory Mapped IO (MMIO) read operation).

[0018] Furthermore, since most I / O devices do not support page faults, they generally require system software (e.g., an operating system (OS), a virtual machine manager (VMM), etc.) to pin memory pages that are DMA targets. On an exemplary system, this can be achieved by calling an OS application programming interface (API) to pin memory pages associated with an I / O buffer before the I / O device accesses them, and by calling the OS API to unpin memory pages associated with the I / O buffer when the I / O device no longer needs to access them. On virtualized systems, because the VMM has no visibility into which pages are DMA targets versus which are not, when a virtual machine (VM) has directly assigned I / O devices (e.g., Peripheral Component Interconnect express (PCIe) physical function (PF) passthrough, PCIe single root I / O virtualization (SR-IOV) virtual function (VF) passthrough, Scalable I / O Virtualization (SIOV) device interface assignment to the VM), the VMM may end up pinning the entire memory of the VM. Therefore, pinning pages means that these pages cannot be migrated (moved) from one memory location to another, which affects many cloud / edge / client scenarios.

[0019] For example, in the following:

[0020] Memory Offline (Scenario #1): The VMM may keep a counter of the number of Correctable Machine Check Interrupts (CMCI) for a memory page and migrate the page to a different location in memory when the Correctable Machine Check Counter reaches a certain threshold. However, the VMM may not do this today for a VM with the assigned device because the VM may be using the page in question for I / O purposes (e.g., observing that the CMCI page is pinned).

[0021] Memory Offline (Scenario #2): System software can migrate pages from different memory modules to a specific memory module to enable the rest of the memory module to be lowered to a low power state (e.g., in a client / personal computer (PC) to preserve battery life) or to replace a faulty Dual Inline Memory Module (DIMM) (e.g., in a server system). System software may not do this today because the page(s) can be pinned for DMA.

[0022] Memory tiering: Consider a server system where the memory capacity consists of two memory tiers: performance (near) memory and capacity (far) memory. The capacity (far) memory could just be ComputeExpress Link (CXL) attached memory, or hardware managed two-level memory (e.g., flat two-level memory (2LM)). A VM can be provisioned with some near memory capacity and some far memory capacity. To improve performance, the VMM would use hot / cold page tracking information to promote hot (more frequently accessed) pages to near memory, or demote cold (less frequently accessed) pages to far memory. For VMs with directly assigned devices, the VMM may not implement this because the page(s) being promoted / demoted may be pinned for DMA.

[0023] Memory Defragmentation (Scenario #1): Consider a scenario where the Graphics Processing Unit (GPU) hardware only supports 64 kilobyte (KB) pages, while the CPU hardware supports 4KB, 2MB, and 1 gigabyte (GB) pages. Based on a request from the graphics subsystem, the OS would find a 2MB page to store the GPU allocation. However, without defragmenting the memory (e.g., moving the different 4KB pages to one location to open up physically contiguous 64KB / 2MB memory blocks), the OS might not find a contiguous block of 64KB / 2MB blocks. If one or more of these pages to be moved are pinned for DMA, the OS might not do this today.

[0024] Memory defragmentation (scenario #2): Consider a scenario where the OS promotes multiple 4KB pages to a single 2MB page, or multiple 2MB pages to a single 1GB page for performance reasons (e.g., in one study, performance across cloud workloads improved by 7.5% when pages were promoted). The OS may not do this today if the page(s) to be promoted are pinned for DMA (e.g., 73% of non-removable memory in these cloud workloads may be due to network I / O).

[0025] Memory Defragmentation (Scenario #3): Some Cloud Service Providers (CSPs) may use different page sizes for different VM types (e.g., 1GB pages for larger VMs, 2MB pages for smaller VMs, a mix of 1GB and 2MB pages, 4KB pages for VM containers, 2MB or 4KB pages for confidential VMs, etc.). This results in system-level memory fragmentation when VMs are started, stopped, or migrated. The VMM may be designed to defragment memory to help solve the packing problem for new VMs or to improve the performance of existing VMs. However, because the pages assigned to a VM are pinned, the VMM's ability to defragment memory may be very limited.

[0026] Memory pooling: CXL 3.0 introduces support for CXL Dynamic Memory Capacity (DCD) devices. For example, consider a scenario where a CXL DCD device is attached to multiple hosts and there is a request to detach a region of CXL DCD memory from one host and attach it to a different host. Today, system software may not allow memory to be removed from a host if pinned pages exist in the region to be removed.

[0027] To this end, at least one embodiment provides a system architecture that implements fixed page migration in the following manner: (1) system software (e.g., an operating system (OS), a virtual machine manager (VMM), etc.) passes the addresses of the source page and destination page associated with the migration to fixed memory migration (PMM) logic (e.g., a logic circuit system); (2) the PMM logic ensures that write operations from I / O devices are observed not only in the source page / region but also in the destination page / region by generating a new "dualcast" operation to the destination page / region; and (3) the system software performs "transactional memory copy" to ensure data consistency between the source page / region and the destination page / region to support any ongoing DMA operations. In an embodiment, the PMM logic can be provided on an integrated circuit device (such as a system on chip ("SOC" or "SoC")), such as reference 1. Figure 1 and discussed in the following figures.

[0028] Figure 1 A block diagram of a system architecture 100 according to an embodiment is shown. The components of the system architecture 100 may be provided on an integrated circuit device such as a SoC. As shown, (one or more) PMM agents 102-1 to 102-N (which may be implemented as logic comprising hardware) reside in the path from the I / O device to the memory 104 after translation by the IO memory management unit (IOMMU). As shown, a coherence fabric 106 couples the core, cache, and memory 104 to other components of the system (e.g., one or more host I / O controllers). In one embodiment, (one or more) PMM agents are implemented in hardware, and in other embodiments, the PMM agents are implemented via a combination of hardware and firmware and / or software. Each PMM agent may include three interfaces: (1) a control interface for configuring the PMM address range; (2) a DMA request interface for receiving DMA requests; and (3) a memory transaction interface for generating memory transactions associated with the DMA requests. As shown, the IOMMU in turn communicates with the PCIe device via one or more root ports and a root complex integrated endpoint (RCIEP).

[0029] Figure 2 FIG. 1 shows a diagram of a method for Figure 1 1. Block diagram of some components of the control interface of the PMM agent. Initially, it should be noted that some components of the system architecture 100 are omitted simply to simplify the discussion.

[0030] The system software 202 utilizes the PMM control interface 204 to program the PMM ranges 206 in the device. The PMM ranges correspond to the source memory regions and destination memory regions associated with the memory being migrated. In one embodiment, the PMM control interface 204 utilizes a set of registers (implemented via the PMM ranges 206 in the device) where the system software programs the source and destination memory ranges associated with the migration (e.g., see the discussion of register details below with reference to Tables 2 and 3): (1) a set of registers that store the address and size associated with the source region, and a "valid" bit to indicate whether the value in the register is valid (e.g., a matching address register); (2) a set of registers that store the address(es) associated with the corresponding destination region, and a "valid" bit to indicate whether the value in the register is valid (e.g., a spare address register); (3) a capability register that enumerates the capabilities of the PMM agent and the offset / location of the source / destination registers / regions (e.g., a dual projection unit capability register).

[0031] In one embodiment, after programming the PMM range(s), a drain operation is performed to drain ongoing DMA requests, thereby ensuring that the newly programmed PMM range(s) are valid. In one embodiment, the drain operation is implicitly performed by the PMM agent when the PMM range(s) are programmed. In another embodiment, the system software 202 queues an invalid wait descriptor to the IOMMU 208 to drain ongoing DMA.

[0032] In one embodiment, the memory locations for the PMM registers are reported to the system software 202 via the Advanced Control and Power Interface (ACPI) DMA Remapping (DMAR) reporting table (e.g., see Table 1 below) using a newly defined “Dual Projection Unit Report Structure”. Table 1

[0033] In one embodiment, the PMM registers can be divided into two sets: trusted registers and untrusted registers, where the trusted registers are accessible only to trusted software entities (e.g., Trusted Domain Extension (TDX) modules), while the untrusted registers are accessible to untrusted software entities (e.g., OS, VMM, etc.). In one embodiment, the trusted registers are protected from untrusted access using the SEAM SAI (Security Attributes of Initiator) policy group.

[0034] In one embodiment, the matching address registers and the backup address registers associated with a PMM range are at least 64B apart from each other so that system software 202 can efficiently program / deprogram them using streaming (direct storage) writes (e.g., memory-mapped IO (MMIO) writes using the MOVDIRI instruction). This is to avoid serialization performed by the CPU core when writes fall within a 64B cache line. In an embodiment, a PMM range is valid / active only when the "Valid" bits in both the matching address register and the backup address register are set (e.g., 0x1), and is not valid / active when at least one of the "Valid" bits is clear (e.g., 0x0). In at least one embodiment, the set of matching address registers is co-located, and the set of backup address registers is co-located to allow system software 202 to efficiently program / deprogram more than one register using larger writes (e.g., MMIO writes using MOVDIR64B).

[0035] In another embodiment, the PMM control interface is implemented as a work queue (or collection of work queues)-based interface, where system software 202 queues work descriptors to program PMM ranges associated with source and destination memory regions. In one embodiment, the work descriptors contain the source region address, destination region address, size, and other region-related attributes. In one embodiment, the work descriptors are batch descriptors or lists containing region information. In one embodiment, separate work queues are used by trusted software entities (e.g., TDX modules) and by untrusted software entities (e.g., OS, VMM, etc.). In one embodiment, the work queues are stored within the PMM agent (or otherwise implemented on the same integrated circuit device), and in another embodiment, the work queues are stored in system memory. In one embodiment, work descriptors are written to the work queues using enqueue command(s) ENQCMD / ENQMDS instructions, and in other embodiments, work descriptors are written to the work queues using the MOVDIR64B instruction. In various embodiments, new ISA instructions may be introduced for the CPU core(s) to program PMM ranges in the PMM agent.

[0036] refer to Figure 2, the PMM agent 102 also includes PMM decoder logic 210 and PMM dual-cast engine / logic 212. When a DMA request is received from the IOMMU 208 via the DMA request interface 213, the PMM decoder 210 checks whether the received DMA request falls within one of the PMM "source" ranges (206). It does this by comparing the address received in the DMA request (e.g., the Host Physical Address (HPA)) against the source memory region address and the size programmed in the PMM ranges 206 (i.e., whether the HPA falls within [source region address, source region address + region size]) to see if there is a match.

[0037] DMA (write / read / atomic operation) requests can be the result of:

[0038] The IOMMU 208 receives I / O requests (e.g., untranslated requests, Address Translation Services (ATS) translated requests, ATS translated requests, etc.) from the I / O devices 214; for example, memory read requests, memory write requests, unordered IO (UIO) read requests, UIO write requests, and / or deferred memory write (DMWr) requests.

[0039] Requests from the IOMMU 208 to read, write, or atomically update memory; for example, memory read requests for IOMMU structures (e.g., root table, context table, scalable mode root table, scalable mode context table, scalable mode Process Address Space Identifier (PASID) directory, scalable mode PASID table, interrupt remapping table, invalidation queue, page request queue, status write for invalidation wait, and / or virtual IOMMU structure), memory write requests for IOMMU structures (e.g., page request queue, status write for invalidation wait, and / or virtual IOMMU structure), memory read requests for paging structures (e.g., first stage paging structure, second stage paging structure, and / or HPT table), atomic operations on accessed / dirty bits in paging structures (“AtomicOp”), and / or atomic operations on posted interrupt descriptors (PIDs).

[0040] In one embodiment, PMM decoder 210 performs this decoding only for DMA requests that modify the source memory region (eg, DMA write requests or DMA AtomicOp requests), but skips checking for DMA requests that only read the source memory region (eg, DMA read requests).

[0041] refer to Figure 2 , the PMM dual-cast engine / logic 212 causes the generation of additional memory transactions to ensure that the pinned memory page(s) (i.e., the page(s) that may be the target of the DMA operation(s)) can be successfully migrated. Thus, the PMM dual-cast engine 212 ensures that any relevant DMA updates are observed in both the source memory region / page(s) and the destination memory region / page(s) while the migration is in progress.

[0042] More specifically, Figure 3 FIGURE 1 is a block diagram illustrating components for performing a dual projection operation according to an embodiment. Figure 3 Omitted Figure 1 and / or Figure 2 Some components are shown here just to simplify the discussion.

[0043] refer to Figures 1 to 3 To complete a dual-cast operation for a DMA write request, the PMM dual-cast engine 212 (of the PMM logic 102) performs a write to both the source and destination memory regions (i.e., performs a dual-cast). The PMM dual-cast engine 212 ensures that the write operation to the destination memory region is strongly ordered after the write operation to the source memory region. As discussed herein, "strongly ordered after the write operation" generally indicates that the data from the first write operation is guaranteed to become visible to other agents in the system no later than the data from the second subsequent write operation. The PMM dual-cast engine 212 determines a source address based on the address received on the DMA request, and the destination address is derived based on the destination region address stored in the matching PMM range 206 found based on the source address.

[0044] Consider a scenario where system software intends to migrate a page at HPA X to a new page at HPA Y, HPA X is mapped to DMA address D in the IOMMU (where D can be a Guest Physical Address (GPA), a Guest I / O Virtual Address (GIOVA), or a Guest Virtual Address (GVA) with virtualization usage or an I / O Virtual Address (IOVA) with bare-metal OS usage), and a DMA write arrives from an I / O device. The PMM agent will perform the following operations on both HPA X and HPA Y: Figure 3 In an embodiment, write operations in the last four segments (segments 00041 to 00044) may be directed to the cache.

[0045] Furthermore, upon a PMM range match found for a DMA read request, the PMM dual-cast engine 212 does not perform any dual-casting, but instead performs a read on the source memory region according to the original DMA read request. Upon a PMM range match found for a DMA atomic operation request, the PMM dual-cast engine 212 performs the atomic operation on the source memory region and then writes the resulting value to the destination memory region. Upon a PMM range match found for a DMA UIO write request, the PMM dual-cast engine 212 writes to both the source and destination memory regions and returns a UIO write completion to the I / O device when both writes are observed (e.g., reaching GO (global observability)). In one embodiment, the UIO write completion is returned to the I / O device as soon as the write to the source memory region is observed.

[0046] Additionally, upon a PMM range match found for a DMA UIO read request, the PMM dual-cast engine 212 does not perform any dual-cast operations, but instead performs a read on the source memory region according to the original DMA read request. Upon a PMM range match found for a DMWr request, the PMM dual-cast engine 212 performs a write to both the source and destination memory regions, and returns a DMWr completion to the I / O device when both writes are observed (e.g., a GO is reached). In one embodiment, the DMWr completion is returned to the I / O device as soon as the write to the source memory region is observed.

[0047] According to some embodiments, several scenarios are provided below to illustrate example operations of the PMM agent 102. In the following description, the use of ")" indicates that the referenced range includes addresses up to but not including the address immediately preceding the ")". For example, "[A, A+1000)" means that the range includes addresses from A up to but not including the address A+1000. The system software has programmed the following PMM ranges: oPMM range #0: [A, A+1000), [B, B+1000) οPMM Range #1: [X, X+1000), [Y, Y+1000) oPMM range #2: [M, M+3000), [O, O+3000) DMA write example #1 oDMA request reaches IOMMU, where address D is translated into HPA X o PMM decoder finds a match for HPA X relative to PMM range #1 o PMM dual-projection engine performs writes to both HPA X and HPA Y DMA Write Example #2 oDMA request reaches IOMMU where address E is translated to HPA M+1064 o PMM decoder finds a match for HPA M+1064 against PMM range #2 o PMM dual projection engine performs writes to both HPA M+1064 and HPA O+1064 DMA Write Example #3 oDMA request reaches IOMMU where address G is translated into HPA N o The PMM decoder could not find a match with any valid (enabled) PMM range oPMM dual projection engine performs writes only to HPA N DMA read example oDMA request reaches IOMMU, where address D is translated into HPA X oPMM decoder skips PMM range check οPMM dual projection engine performs reads only on HPA X DMA AtomicOp (atomic operation) example oDMA request reaches IOMMU, where address D is translated into HPA X o PMM decoder finds a match for HPA X relative to PMM range #1 The PMM dual-projection engine performs an AtomicOp (atomic operation) on HPA X and writes the result to HPA Y using the value obtained after the AtomicOp.

[0048] Table 2 and Table 3 below show example register definitions for the PMM control interface and dual-projection unit capabilities, respectively, according to some embodiments. Table 2 Table 3

[0049] Figure 4 A flow chart illustrating a method 400 for providing fixed memory migration according to an embodiment is shown. The method 400 is intended to migrate a memory page at HPA X to a new page at HPA Y, where HPA X is mapped to DMA address D in IOMMU 208 (e.g., D can be GPA / GIOVA / GVA for virtualization purposes or IOVA for bare-metal OS purposes). In various embodiments, the operations of method 400 can be performed by system software (or other software, e.g., software that is not part of PMM agent 102).

[0050] refer to Figures 1-4 At operation 402, the software unmaps page X for CPU accesses (except those used for transactional memory copy operations) and invalidates the CPU translation lookaside buffer (TLB), while page X continues to be mapped for I / O at DMA address D. In one embodiment, the removed mappings include Extended Page Table (EPT) mappings used to control accesses from within the VM. In one embodiment, the removed mappings include user-mode mappings used by applications.

[0051] At operation 404, software enables dual-cast operations between page X and page Y by programming the PMM range 206 in the PMM agent 102. The PMM agent copies all DMA writes to page X to page Y. In at least one embodiment, write operations to page Y (i.e., the destination page) are guaranteed to be strongly ordered after write operations to page X (i.e., the source page).

[0052] At operation 406, software causes any in-progress DMA writes that may have been queued to memory to be drained prior to the dual cast being enabled at operation 404. From this point on, all new DMA writes to X are guaranteed to be copied to Y. The drain operation may be implicitly initiated by the PMM agent when the PMM range 206 is programmed or initiated by system software that queues an IOMMU invalidate wait descriptor (e.g., inv_wait_dsc).

[0053] At operation 408, the software causes a "transactional memory copy" to be performed from page X to page Y. As discussed herein, a "transactional memory" operation generally refers to a memory operation that allows load and store operations to be performed atomically. For example, if a transactional memory copy detects a change to page X while copying from page X to page Y, it retries the copy operation. The transactional memory copy operation can be performed with any block size (e.g., 64B, ..., 512B, 1KB, etc.). In general, the larger the block size, the higher the likelihood that a conflict for a DMA write may exist during the block copy.

[0054] At operation 410, the memory page table is updated to map the IOMMU 208 (and CPU accesses) to page Y and invalidate the IOTLB. At operation 412, software disables dual-cast operations between page X and page Y, for example, by deprogramming the PMM range 206 in the PMM agent.

[0055] Furthermore, when a page is migrated by system software from source memory location X to destination memory location Y (e.g., at operation 408, there is a scenario where a DMA write may modify source page X while it is being migrated to page Y), transactional memory copying helps the system software detect dangers in data correctness and redo the copy operation.

[0056] As an example, consider the following scenario with ongoing DMA memory writes and page migration copies:

[0057] Use case #1: MemWr(X) -> DcastWr(Y) -> COPY(X,Y)

[0058] Use case #2: COPY(X,Y)->MemWr(X)->DcastWr(Y)

[0059] Use case #3: MemWr(X)->COPY(X,Y)->DcastWr(Y)

[0060] Use Case #4: MemWr(X) / DcastWr(Y) interleaved within COPY(X,Y)

[0061] Use cases 1, 2, and 3 above are safe, while use case 4 is dangerous (i.e., COPY reads the old value from X -> MemWr updates X -> PMM updates Y -> COPY overwrites Y with the old value of X). In this case, transactional replication is used to detect the danger and retry the copy operation using the updated value of X.

[0062] In various embodiments, different techniques may be used to perform the transactional memory copy. In a first embodiment, the system software utilizes the transactional memory instruction set architecture (ISA) in the CPU core to perform the transactional memory copy operation. For example, XBEGIN (X start); READ (X); WRITE (Y); XEND (X end). In a second embodiment, the system software utilizes the CPU core to perform the transactional memory copy. For example, READ (X); WRITE (Y); COMPARE (X, Y). In a third embodiment, the system software offloads the memory copy to an accelerator (such as the Data Streaming Accelerator provided by Intel Corporation of Santa Clara, California). TM (DSA TM )). For example, the system software issues a memory copy MEMCPY(X,Y) work descriptor followed by a memory compare MEMCMP(X,Y) work descriptor to the accelerator; or the system software issues a memory copy with cyclic redundancy code (CRC) MEMCPYWITHCRC(X,Y) work descriptor followed by a check CRC CHECKCRC(X) work descriptor to the accelerator.

[0063] Table 4 below illustrates an example method for performing a transactional memory copy operation according to some embodiments. It should be noted that the description of DSA in Table 4 is for illustration purposes only, and any "migration accelerator" can be utilized in Method 4. Table 4

[0064] In one embodiment, transactional replication is performed with a smaller block size to detect hazards more quickly and enable replication to be resumed. For example, the following pseudo code can be used:

[0065] Separate IO and CPU read / write permissions

[0066] In addition, many scenarios today require system software to share page tables between the CPU MMU and the IOMMU, which limits the ability to remove memory accesses to the CPU core while maintaining memory accesses to I / O devices. In some embodiments, the following extensions are introduced to the paging structure entries and IOMMU hardware to support separate I / O and CPU R / W (read / write) permissions: Define bit 62 of the second-stage paging structure entry as the I / O Write (IW) field. For example, bit 62 of the second-stage PML5 table entry, the second-stage Page Map Level 4 (PML4) table entry, the second-stage Page Directory Pointer table entry, the second-stage Page Directory entry, and the second-stage Page Table entry. Define bit 61 of the second-stage paging structure entry as the I / O Read (IR) field. For example, bit 61 of the second-stage PML5 table entry, the second-stage PMI4 table entry, the second-stage page directory pointer table entry, the second-stage page directory entry, and the second-stage page table entry. Define bit 57 in the IOMMU Extended Capabilities Register as Second-Stage I / O Read / Write Support (SSIRWS) to enumerate support for this feature oWhen 0, the IOMMU hardware does not support the I / O Read and I / O Write bits in the second stage paging entries oWhen 1, the IOMMU hardware supports the I / O Read and I / O Write bits in the second stage paging entries Define bit 7 in the IOMMU root table address register as second-stage I / O read / write enable (SSIRWE) to control the enable / disable of separate I / O read / write permissions oWhen 0, the IOMMU hardware uses bits 0 (R) and 1 (W) in the second-stage paging entry to compute the effective permissions for memory accesses. oWhen 1, the IOMMU hardware uses bits 61 (IR) and 62 (IW) in the second stage paging entry to compute the effective permissions for memory accesses. The value of this field takes effect only after the software executes the Set Root Table Pointer command.

[0067] In one embodiment, when migrating a page from X to Y, the system software programs all PMM agents on the platform. In a second embodiment, the system software will only program those PMM agents that may expect in-flight DMA for the source page being migrated. For example, consider a scenario where there are five PMM agents on the platform, but the I / O devices assigned to a VM are within the scope of only the first two PMM agents. The system software will only program the PMM scopes in the first two PMM agents.

[0068] In one embodiment, system software first uses a Move Doubleword as a DirectStore (MOVDIRI) instruction to set up the PMM range in all (or some) PMM agents, pre-performs a Store Fence (SFENCE) operation, writes an invalidation descriptor in all (or some) IOMMUs, and uses the MOVDIRI instruction to update the tail pointer associated with the invalidation queue to force draining of ongoing DMA. The drain operation then performs an SFENCE and waits for all invalidations to complete. This approach is envisioned to allow for more parallel processing with system software, so that less time is spent waiting for (one or more) MMIO write operations to complete before continuing. This may be particularly important in SoCs where there may be many PMM agents or IOMMUs.

[0069] Additionally, some embodiments may be applied to a computing system including one or more processors (eg, where the one or more processors may include one or more processor cores), such as a computing system comprising a processor core. Figure 1 The computing systems discussed in the following figures include, for example, desktop computers, workstations, computer servers, server blades, or mobile computing devices. Mobile computing devices may include smartphones, tablets, UMPCs (Ultra Mobile Personal Computers), laptops, ultrabooks, and other similar devices. TM Computing devices, wearable devices (such as smart watches, smart rings, smart bracelets or smart glasses), etc.

[0070] Example computer architecture

[0071] The following describes an example computer architecture. Other system designs and configurations known in the art for laptops, desktop computers, handheld personal computers (PCs), personal digital assistants, engineering workstations, servers, separate servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, various systems or electronic devices that can include processors and / or other execution logic as disclosed herein are generally suitable.

[0072] Figure 5 An example computing system is illustrated. Multiprocessor system 500 is an interfaced system and includes multiple processors or cores including a first processor 570 and a second processor 580 coupled via an interface 550 such as a point-to-point (PP) interconnect, structure, and / or bus. In some examples, first processor 570 and second processor 580 are isomorphic. In some examples, first processor 570 and second processor 580 are heterogeneous. Although example system 500 is shown as having two processors, the system can have three or more processors or can be a single processor system. In some examples, the computing system is a system on a chip (SoC).

[0073] Processors 570 and 580 are shown as including integrated memory controller (IMC) circuitry 572 and 582, respectively. Processor 570 also includes interface circuits 576 and 578; similarly, second processor 580 includes interface circuits 586 and 588. Processors 570, 580 can exchange information via interface 550 using interface circuits 578, 588. IMCs 572 and 582 couple processors 570, 580 to respective memories, namely, memory 532 and memory 534, which may be portions of main memory locally attached to the respective processors.

[0074] Processors 570, 580 can each use interface circuits 576, 594, 586, 598 to exchange information with a network interface (NW I / F) 590 via separate interfaces 552, 554. Network interface 590 (e.g., one or more of an interconnect, bus, and / or fabric, and in some examples, a chipset) can optionally exchange information with a coprocessor 538 via interface circuit 592. In some examples, coprocessor 538 is a special-purpose processor, such as, for example, a high-throughput processor, a network or communication processor, a compression engine, a graphics processor, a general-purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, or the like.

[0075] A shared cache (not shown) may be included in either processor 570, 580, or external to both processors but connected to these processors via an interface (such as a PP interconnect) so that if the processors are placed in a low power mode, the local cache information of either or both processors may be stored in the shared cache.

[0076] The network interface 590 can be coupled to the first interface 516 via the interface circuit 596. In some examples, the first interface 516 can be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 516 is coupled to a power control unit (PCU) 517, which can include circuitry, software, and / or firmware for performing power management operations related to the processors 570, 580, and / or the coprocessor 538. The PCU 517 provides control information to a voltage regulator (not shown) so that the voltage regulator generates an appropriate regulated voltage. The PCU 517 also provides control information to control the generated operating voltage. In various examples, the PCU 517 can include various power management logic units (circuitry) for performing hardware-based power management. Such power management may be entirely controlled by the processor (e.g., by various processor hardware, and which may be triggered by workload and / or power, thermal constraints, or other processor constraints), and / or power management may be performed in response to an external source (such as a platform or power management source or system software).

[0077] The PCU 517 is illustrated as existing as logic separate from the processor 570 and / or the processor 580. In other cases, the PCU 517 may execute on a given one or more of the cores (not shown) of the processor 570 or 580. In some cases, the PCU 517 may be implemented as a (dedicated or general-purpose) microcontroller or other control logic configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by the PCU 517 may be implemented external to the processor, such as by way of a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by the PCU 517 may be implemented within the BIOS or other system software.

[0078] Various I / O devices 514 can be coupled to the first interface 516 along with a bus bridge 518, which couples the first interface 516 to the second interface 520. In some examples, one or more additional processors 515 (such as a coprocessor, a high-throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor) are coupled to the first interface 516. In some examples, the second interface 520 can be a low pin count (LPC) interface. Various devices can be coupled to the second interface 520, including, for example, a keyboard and / or mouse 522, a communication device 527, and a storage circuit system 528. The storage circuit system 528 can be one or more non-transitory machine-readable storage media, such as a disk drive or other mass storage device, as described below, which in some examples can include instructions / code and data 530 and can implement a storage device. Further, audio I / O 524 may be coupled to second interface 520. Note that other architectures are possible besides the point-to-point architecture described above. For example, a system such as multiprocessor system 500 may implement a multi-drop interface or other such architecture rather than a point-to-point architecture.

[0079] Example core architectures, processors, and computer architectures

[0080] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) a general-purpose in-order core intended for general-purpose computing; 2) a high-performance general-purpose out-of-order core intended for general-purpose computing; 3) a specialized core intended primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) a CPU that includes one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors give rise to different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such coprocessors are sometimes referred to as dedicated logic or as dedicated cores, such as integrated graphics and / or scientific (throughput) logic); and 4) a system on a chip (SoC), which may include the described CPU (sometimes referred to as application core(s) or application processor(s), the coprocessor(s) described above, and additional functionality on the same die. An example core architecture is described next, followed by a description of an example processor and computer architecture.

[0081] Figure 6 A block diagram of an example processor and / or SoC 600 that may have one or more cores and an integrated memory controller is shown. The solid line box illustrates the processor 600 having a single core 602(A), a system agent unit circuitry 610, and a collection of one or more interface controller unit circuitry 616, while the optional addition of the dashed line box illustrates an alternative processor 600 having multiple cores 602(A)-602(N), a collection of one or more integrated memory controller unit circuitry 614 in the system agent unit circuitry 610, and dedicated logic 608, and a collection of one or more interface controller unit circuitry 616. Note that the processor 600 may be Figure 5 one of the processors 570 or 580, or the coprocessors 538 or 515.

[0082] Thus, different implementations of processor 600 may include: 1) a CPU, wherein dedicated logic 608 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and cores 602(A)-602(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of the two); 2) a coprocessor, wherein cores 602(A)-602(N) are a large number of dedicated cores intended primarily for graphics and / or scientific (throughput); and 3) a coprocessor, wherein cores 602(A)-602(N) are a large number of general-purpose in-order cores. Thus, processor 600 may be a general-purpose processor, a coprocessor, or a dedicated processor, such as, for example, a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many-integrated-core (MIC) coprocessor (including 30 or more cores), an embedded processor, or the like. The processor may be implemented on one or more chips. Processor 600 may be part of and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, complementary metal-oxide-semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0083] The memory hierarchy includes one or more levels of cache unit circuitry 604(A)-604(N) within the cores 602(A)-602(N), a set of one or more shared cache unit circuitry 606, and external memory (not shown) coupled to a set of integrated memory controller unit circuitry 614. The set of one or more shared cache unit circuitry 606 may include one or more intermediate levels of cache (such as level 2 (L2), level 3 (L3), level 4 (L4)), or other levels of cache (such as last level cache (LLC)), and / or combinations thereof. While in some examples, an interface network circuitry 612 (e.g., a ring interconnect) provides an interface to dedicated logic 608 (e.g., integrated graphics logic), the set of one or more shared cache unit circuitry 606, and the system agent unit circuitry 610, alternative examples use any number of well-known techniques for providing interfaces to such units. In some examples, coherency is maintained between the shared cache unit circuitry(s) 606 and one or more of the cores 602(A)-602(N). In some examples, the interface controller unit circuitry 616 couples the cores 602(A)-602(N) to one or more other devices 618, such as one or more I / O devices, storage, one or more communication devices (e.g., a wireless network, a wired network, etc.), and the like.

[0084] In some examples, one or more of the cores 602(A)-602(N) can be multithreaded. System agent unit circuitry 610 includes components that coordinate and operate the cores 602(A)-602(N). System agent unit circuitry 610 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be or may include the logic and components required to regulate the power state of the cores 602(A)-602(N) and / or the dedicated logic 608 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays.

[0085] The cores 602(A)-602(N) may be homogeneous in terms of instruction set architecture (ISA). Alternatively, the cores 602(A)-602(N) may be heterogeneous in terms of ISA; that is, a subset of the cores 602(A)-602(N) may be capable of executing an ISA, while other cores may only be capable of executing a subset of that ISA or another ISA.

[0086] Example Core Architecture—In-Order and Out-Of-Order Core Block Diagrams

[0087] FIG7(A) is a block diagram illustrating both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to an example. FIG7(B) is a block diagram illustrating both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to an example. The solid-line boxes in FIG7(A)-FIG7(B) illustrate the in-order pipeline and the in-order core, while the optional addition of dashed-line boxes illustrates the register renaming, out-of-order issue / execution pipeline and the core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0088] In FIG7(A), processor pipeline 700 includes a fetch stage 702, an optional length decode stage 704, a decode stage 706, an optional allocation (Alloc) stage 708, an optional rename stage 710, a dispatch (also known as dispatch or issue) stage 712, an optional register read / memory read stage 714, an execute stage 716, a writeback / memory write stage 718, an optional exception handling stage 722, and an optional commit stage 724. One or more operations may be performed in each of these processor pipeline stages. For example, during fetch stage 702, one or more instructions are fetched from an instruction memory, and during decode stage 706, the fetched one or more instructions may be decoded, an address using a forwarded register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., an immediate offset or link register (LR)) may be performed. In one example, decode stage 706 and register read / memory read stage 714 may be combined into one pipeline stage. In one example, during the execute stage 716, decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.

[0089] As an example, the example register renaming, out-of-order issue / execution architecture core of FIG. 7(B) may implement pipeline 700 as follows: 1) instruction fetch circuitry 738 executes fetch stage 702 and length decode stage 704; 2) decode circuitry 740 executes decode stage 706; 3) rename / allocator unit circuitry 752 executes allocate stage 708 and rename stage 710; 4) (one or more) scheduler circuit systems 756 perform the schedule stage 712; 5) (one or more) physical register file circuit systems 758 and memory unit circuit systems 770 perform the register read / memory read stage 714; (one or more) execution clusters 760 perform the execute stage 716; 6) memory unit circuit systems 770 and (one or more) physical register file circuit systems 758 perform the write back / memory write stage 718; 7) various circuit systems may be involved in the exception handling stage 722; and 8) retirement unit circuit system 754 and (one or more) physical register file circuit systems 758 perform the commit stage 724.

[0090] 7(B) shows a processor core 790 comprising a front-end unit circuitry 730 coupled to an execution engine unit circuitry 750, both of which are coupled to a memory unit circuitry 770. Core 790 may be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, core 790 may be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose graphics processing unit (GPGPU) core, a graphics core, or the like.

[0091] The front end unit circuitry 730 may include a branch prediction circuitry 732 coupled to an instruction cache circuitry 734, which is coupled to an instruction translation lookaside buffer (TLB) 736, which is coupled to an instruction fetch circuitry 738, which is coupled to a decode circuitry 740. In one example, the instruction cache circuitry 734 is included in the memory unit circuitry 770 rather than in the front end unit circuitry 730. The decode circuitry 740 (or decoder) may decode an instruction and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals decoded from, or otherwise reflecting, or derived from, the original instruction. The decode circuitry 740 may further include an address generation unit (AGU) circuitry (not shown). In one example, the AGU generates an LSU address using the forwarded register port and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). The decode circuitry 740 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLA), microcode read-only memories (ROM), etc. In one example, the core 790 includes a microcode ROM (not shown) or other medium (e.g., in the decode circuitry 740 or otherwise within the front-end unit circuitry 730) that stores microcode for certain macroinstructions. In one example, the decode circuitry 740 includes a micro-op or operation cache (not shown) to store / cache decoded operations, microtags, or micro-operations generated during the decode stage or other stages of the processor pipeline 700. The decode circuitry 740 may be coupled to the rename / allocator unit circuitry 752 in the execution engine circuitry 750.

[0092] Execution engine circuitry 750 includes rename / allocator unit circuitry 752, which is coupled to retirement unit circuitry 754 and a set of one or more scheduler circuitry 756. Scheduler circuitry(s) 756 represent any number of different schedulers, including reservation stations, central instruction windows, and the like. In some examples, scheduler circuitry(s) 756 may include an arithmetic logic unit (ALU) scheduler / scheduling circuitry, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuitry, an AGU queue, and the like. Scheduler circuitry(s) 756 are coupled to physical register file(s) circuitry 758. Each physical register file circuitry in the physical register file(s) circuitry 758 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, status (e.g., an instruction pointer, which is the address of the next instruction to be executed), etc. In one example, the physical register file(s) circuitry 758 includes vector register unit circuitry, write mask register unit circuitry, and scalar register unit circuitry. These register units may provide architectural vector registers, vector mask registers, general purpose registers, etc. The physical register file(s) circuitry 758 is coupled to the retirement unit circuitry 754 (also referred to as a "retire queue" or "retirement queue") to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using reorder buffer(s) (ROBs) and retirement register file(s); using future file(s), history buffer(s) and retirement register file(s); using register maps and register pools, etc.). Retirement unit circuitry 754 and physical register file(s) circuitry 758 are coupled to execution cluster(s) 760. Execution cluster(s) 760 include a set of one or more execution unit circuitry 762 and a set of one or more memory access circuitry 764. Execution unit circuitry(s) 762 may perform various arithmetic, logical, floating point, or other types of operations (e.g., shifts, additions, subtractions, multiplications) and on various data types (e.g., scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point).While some examples may include multiple execution units or execution unit circuitry dedicated to a particular function or set of functions, other examples may include only one execution unit circuitry or multiple execution units / execution unit circuitry that all perform all functions. Scheduler circuitry(s) 756, physical register file circuitry(s) 758, and execution cluster(s) 760 are shown as potentially multiple because some examples create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline each with its own scheduler circuitry, physical register file circuitry(s), and / or execution cluster—and in the case of separate memory access pipelines, certain examples are implemented where only the execution cluster of that pipeline has memory access unit circuitry(s) 764). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution, and the remaining pipelines may be in-order issue / execution.

[0093] In some examples, the execution engine unit circuit system 750 can perform load store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown), as well as address staging and writeback, data staging loads, stores, and branches.

[0094] The set of memory access circuitry 764 is coupled to memory cell circuitry 770, which includes data TLB circuitry 772, which is coupled to data cache circuitry 774, which is coupled to second level (L2) cache circuitry 776. In one example, memory access circuitry 764 may include load unit circuitry, store address unit circuitry, and store data unit circuitry, each of which is coupled to data TLB circuitry 772 in memory cell circuitry 770. Instruction cache circuitry 734 is further coupled to second level (L2) cache circuitry 776 in memory cell circuitry 770. In one example, instruction cache 734 and data cache 774 are combined into L2 cache circuitry 776, third level (L3) cache circuitry (not shown), and / or a single instruction and data cache in main memory (not shown). L2 cache circuitry 776 is coupled to one or more other levels of cache and ultimately to main memory.

[0095] Core 790 may support one or more instruction sets (e.g., the x86 instruction set architecture (optionally with some extensions that have been added with newer versions); the MIPS instruction set architecture; the ARM instruction set architecture (optionally with optional additional extensions such as NEON)) that include the instruction(s) described herein. In one example, core 790 includes logic to support packed data instruction set architecture extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.

[0096] Example Execution Unit(s) Circuitry

[0097] Figure 8 An example of execution unit circuitry (one or more) is shown, such as execution unit circuitry (one or more) 762 of FIG. 7(B) . As shown, execution unit circuitry (one or more) 762 may include one or more ALU circuits 801, optional vector / single instruction multiple data (SIMD) circuitry 803, load / store circuitry 805, branch / jump circuitry 807, and / or floating-point unit (FPU) circuitry 809. ALU circuitry 801 performs integer arithmetic and / or Boolean operations. Vector / SIMD circuitry 803 performs vector / SIMD operations on packed data (such as SIMD / vector registers). Load / store circuitry 805 executes load and store instructions to load data from memory into registers or store data from registers to memory. Load / store circuitry 805 may also generate addresses. Branch / jump circuitry 807 causes branches or jumps to memory addresses, depending on the instruction. FPU circuitry 809 performs floating-point arithmetic. The width of execution unit circuitry 762 varies depending on the example and may range, for example, from 16 bits to 1024 bits. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).

[0098] In this description, numerous specific details are set forth to provide a more thorough understanding. However, it will be apparent to one skilled in the art that the embodiments described herein can be practiced without one or more of these specific details. In other instances, well-known features are not described in order to avoid obscuring the details of the present embodiments.

[0099] The following examples relate to further embodiments. Example 1 includes an apparatus comprising: one or more memory devices for storing a source page and a destination page; and logic circuitry for causing a write operation from an input / output (I / O) device to be observed for both the source page and the destination page, wherein a transactional memory copy operation is performed from the source page to the destination page.

[0100] Example 2 includes an apparatus as in Example 1, wherein the logic circuit system is used to cause all write operations to the source page to be copied to the destination page to observe the write operation from the I / O device. Example 3 includes an apparatus as in any one of Examples 1 to 2, wherein, in order to observe the write operation from the I / O device, the logic circuit system is used to cause: the copying of all write operations to the source page to the destination page; and the draining of all ongoing write operations that have been queued for access to one or more memory devices before the copying begins. Example 4 includes an apparatus as in any one of Examples 1 to 3, wherein, in order to observe the write operation from the I / O device, the logic circuit system is used to cause an update to a fixed memory migration (PMM) range to enable the copying of all write operations to the source page to the destination page. Example 5 includes an apparatus as in any one of Examples 1 to 4, wherein the write operation includes a direct memory access (DMA) write memory operation.

[0101] Example 6 includes the apparatus of any of Examples 1 to 5, wherein the logic circuitry is configured to cause all write operations to the source page to be copied to the destination page to observe write operations from the I / O device, wherein any write operations to the destination page are sequenced after any write operations to the source page. Example 7 includes the apparatus of any of Examples 1 to 6, wherein the transactional memory copy operation is repeated in response to determining that data in the source page has changed during the transactional memory copy operation. Example 8 includes the apparatus of any of Examples 1 to 7, wherein, before causing execution of the transactional memory copy operation, the source page is unmapped for processor accesses other than those used for the transactional memory copy operation. Example 9 includes the apparatus of any of Examples 1 to 8, wherein, before causing execution of the transactional memory copy operation, the source page is unmapped for accesses by the processor other than those used for the transactional memory copy operation and the source page is invalidated in all translation lookaside buffers (TLBs) of the processor. Example 10 includes the apparatus of any of Examples 1 to 9, wherein, after execution of the transactional memory copy operation, the memory page table is updated to map future accesses by the processor to the destination page.

[0102] Example 11 includes an apparatus as in any one of Examples 1 to 10, wherein the logic circuit system is used to communicate with the memory through a host I / O controller. Example 12 includes an apparatus as in any one of Examples 1 to 11, wherein the host I / O controller is coupled to the memory via a consistency structure. Example 13 includes an apparatus as in any one of Examples 1 to 12, wherein the system on chip (SoC) includes one or more memory devices, a logic circuit system, and a processor with one or more processor cores. Example 14 includes an apparatus as in any one of Examples 1 to 13, wherein the transactional memory copy operation is migrated to an accelerator. Example 15 includes an apparatus as in any one of Examples 1 to 14, wherein the accelerator includes a data streaming accelerator (DSA).

[0103] Example 16. One or more non-transitory computer-readable media comprising one or more instructions that, when executed on a processor, configure the processor to perform one or more operations to: store a source page and a destination page; cause a write operation from an input / output (I / O) device to be observed for both the source page and the destination page at persistent memory migration (PMM) logic; and cause a transactional memory copy operation to be performed from the source page to the destination page.

[0104] Example 17 includes one or more non-transitory computer-readable media of Example 16, further comprising one or more instructions that, when executed on a processor, configure the processor to perform one or more operations to cause the PMM logic to cause all write operations to a source page to be copied to a destination page to observe write operations from an I / O device. Example 18 includes one or more non-transitory computer-readable media of any of Examples 16 to 17, further comprising one or more instructions that, when executed on a processor, configure the processor to perform one or more operations to cause the PMM logic to cause: copying all write operations to a source page to a destination page; and draining all in-progress write operations that have been queued for access to one or more memory devices before the copying begins.

[0105] Example 19 includes one or more non-transitory computer-readable media of any one of Examples 16 to 18, further comprising one or more instructions that, when executed on a processor, configure the processor to perform one or more operations to cause the PMM logic to cause an update to a PMM range to enable replication of all write operations to a source page to a destination page. Example 20 includes one or more non-transitory computer-readable media of any one of Examples 16 to 19, further comprising one or more instructions that, when executed on a processor, configure the processor to perform one or more operations to cause the PMM logic to cause replication of all write operations to a source page to a destination page to observe write operations from an I / O device, wherein any write operation to a destination page is ordered after any write operation to a source page.

[0106] Example 21 includes a method comprising: storing a source page and a destination page in one or more memory devices; causing write operations from an input / output (I / O) device to be observed for both the source page and the destination page at fixed memory migration (PMM) logic; and causing a transactional memory copy operation to be performed from the source page to the destination page. Example 22 includes the method of Example 21, further comprising causing the PMM logic to cause a copy of all write operations to the source page to the destination page to observe the write operations from the I / O device. Example 23 includes the method of any of Examples 21 to 22, further comprising causing the PMM logic to cause: a copy of all write operations to the source page to the destination page; and, before the copy begins, a drain of all in-progress write operations that have been queued for access to the one or more memory devices. Example 24 includes the method of any of Examples 21 to 23, further comprising causing the PMM logic to cause an update to a PMM range to enable the copy of all write operations to the source page to the destination page. Example 25 includes a method as in any of Examples 21 to 24, further comprising causing the PMM logic to cause a copy of all write operations to the source page to the destination page to observe the write operations from the I / O device, wherein any write operation to the destination page is sequenced after any write operation to the source page.

[0107] Example 26 includes an apparatus comprising means for performing a method as set forth in any of the preceding examples. Example 27 includes a machine-readable storage device comprising machine-readable instructions that, when executed, are used to implement a method as set forth in any of the preceding examples or to implement an apparatus as set forth in any of the preceding examples.

[0108] In each embodiment, reference Figure 1One or more operations discussed in the figures and below may be performed by one or more components (interchangeably referred to herein as "logic") discussed with reference to any of the figures.

[0109] In some embodiments, herein (e.g., with reference to Figure 1 The operations discussed in the accompanying drawings and the following figures) can be implemented as hardware (e.g., logic circuitry), software, firmware, or a combination thereof, which can be provided as a computer program product, for example, including one or more tangible (e.g., non-transitory) machine-readable or computer-readable media having instructions (or software programs) stored thereon for programming a computer to perform the processes discussed herein. The machine-readable medium may include storage devices such as those discussed with reference to the accompanying figures.

[0110] Additionally, such computer-readable media may be downloaded as a computer program product, where the program may be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a data signal provided in a carrier wave or other propagation medium, via a communication link (e.g., a bus, a modem, or a network connection).

[0111] References in this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, and / or characteristic described in connection with the embodiment may be included in at least one implementation. The appearances of the phrase "in one embodiment" in various places in this specification may or may not all refer to the same embodiment.

[0112] Furthermore, in the specification and claims, the terms "coupled" and "connected," and their derivatives, may be used. In some embodiments, "connected" may be used to indicate that two or more elements are in direct physical or electrical contact with each other. "Coupled" may mean that two or more elements are in direct physical or electrical contact. However, "coupled" may also mean that two or more elements may not be in direct contact with each other, but may still cooperate or interact with each other.

[0113] Thus, although various embodiments have been described using language specific to structural features and / or methodological acts, it will be understood that the claimed subject matter may not be limited to the specific features or acts described. Rather, the specific features and acts are disclosed as sample forms of implementing the claimed subject matter.

Claims

1. A system architecture and technology for fixed storage migration, the system comprising: one or more memory devices for storing source pages and destination pages; as well as logic circuitry for causing a write operation from an input / output (I / O) device to be observed for both the source page and the destination page, Wherein, a transactional memory copy operation is performed from the source page to the destination page.

2. The device according to claim 1, wherein The logic circuitry is to cause replication of all write operations to the source page to the destination page to observe the write operations from the I / O device.

3. The device according to claim 1, wherein To observe the write operation from the I / O device, the logic circuitry is configured to cause: Copying all write operations to the source page to the destination page; as well as Before copying begins, all ongoing write operations queued for access to the one or more memory devices are flushed.

4. The device according to claim 1, wherein To observe the write operations from the I / O device, the logic circuitry is to cause an update to a fixed memory migration PMM range to enable copying of all write operations to the source page to the destination page.

5. The device according to claim 1, wherein The write operation includes a direct memory access (DMA) write memory operation.

6. The device according to claim 1, wherein The logic circuitry is to cause replication of all write operations to the source page to the destination page for observation of the write operations from the I / O device, wherein any write operation to the destination page is sequenced after any write operation to the source page.

7. The device according to claim 1, wherein The transactional memory copy operation is repeated in response to determining that data in the source page has changed during the transactional memory copy operation.

8. The device according to claim 1, wherein Prior to causing execution of the transactional memory copy operation, the source page is unmapped for processor accesses other than those used for the transactional memory copy operation.

9. The device according to claim 1, wherein Prior to causing execution of the transactional memory copy operation, the source page is unmapped for accesses by the processor other than those used for the transactional memory copy operation, and the source page is invalidated in all translation lookaside buffers (TLBs) of the processor.

10. The device of claim 1, wherein Following execution of the transactional memory copy operation, a memory page table is updated to map future accesses by the processor to the destination page.

11. The device according to claim 1, wherein The logic circuitry is configured to communicate with the memory through a host I / O controller.

12. The device according to claim 11, wherein The host I / O controller is coupled to the memory via a coherency fabric.

13. The device of claim 1, wherein: A system on chip (SoC) includes the one or more memory devices, the logic circuitry, and a processor having one or more processor cores.

14. The device of claim 1, wherein The transactional memory copy operations are migrated to the accelerator.

15. The apparatus of claim 14, wherein: The accelerator comprises a data streaming accelerator DSA.

16. A method comprising: Store source and destination pages; causing, at fixed memory migration PMM logic, a write operation from an input / output (I / O) device to be observed for both the source page and the destination page; as well as Execution of a transactional memory copy operation is caused from the source page to the destination page.

17. The method of claim 16, further comprising the PMM logic causing replication of all write operations to the source page to the destination page to observe the write operations from the I / O device.

18. The method of claim 16, further comprising the PMM logic causing: replication of all write operations to the source page to the destination page; and Before the copying begins, all ongoing write operations queued for access to the one or more memory devices are flushed.

19. The method of claim 16, further comprising the PMM logic being operable to cause an update to a PMM extent to enable replication of all write operations to the source page to the destination page.

20. The method of claim 16, further comprising the PMM logic being operable to cause replication of all write operations to the source page to the destination page to observe the write operations from the I / O device, wherein Any write operations to the destination page are sequenced after any write operations to the source page.

21. A machine-readable medium comprising code which, when executed, causes a machine to perform the operations of any one of claims 1 to 20.

22. An apparatus comprising means for performing the operations of any one of claims 1 to 20.