Apparatus and method for remote virtual-to-physical address translation

By implementing remote TLB and PTW circuit systems at each memory endpoint, the problem of low bandwidth utilization efficiency in remote virtual to physical address translation is solved, and efficient address translation and bandwidth saturation is achieved, suitable for working flows that deal with ultrasparse graph characteristics.

CN120179581APending Publication Date: 2025-06-20INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411665485.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-11-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently optimize network and memory bandwidth utilization during remote virtual to physical address translation, especially in processing workflows that display supersparse graph characteristics.

Method used

Implement a remote translation backup buffer (TLB) and page table walk-through (PTW) circuitry at each memory endpoint to walk-through page tables when the remote TLB is missed and translate virtual addresses into physical addresses while cache translation results for efficiency.

Benefits of technology

Through this method, the translation of virtual address to physical address can be completed without returning to the publishing thread, which improves the efficiency of indirect memory operations, ensures saturation of network and memory bandwidth, and is suitable for applications with low computing strength.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179581A_ABST
    Figure CN120179581A_ABST
Patent Text Reader

Abstract

An apparatus and method for remote virtual-to-physical address translation. For example, one embodiment of a processor comprises: an interconnection network; a plurality of cores coupled to the interconnect network, a core of the plurality of cores including a core translation lookaside buffer (TLB) and a core page table walkthrough (PTW) circuit; a plurality of memory endpoint subsystems coupled to the interconnect network, each memory endpoint subsystem comprising a memory access circuit and a memory for access via the interconnect network, each memory access circuit comprising a memory endpoint TLB and a memory endpoint PTW circuit; wherein one or more of the memory endpoint subsystems are configured to perform an operation on behalf of the requester core to process an indirect memory access request, one or more of the memory endpoint subsystems are configured to read a virtual pointer address based on the indirect memory access request, translate the virtual pointer address to a physical pointer address, and transmit the physical pointer address to the requester core. The data is accessed based on the physical pointer address, and an indirect memory access response including the data is transmitted to the requester core.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This invention was made with government support under contract number W911NF-22-C-0081 awarded by the Army Research Office and IARPA. The government has certain rights in this invention. Technical Field

[0002] This invention generally relates to the field of computer processors. More specifically, this invention relates to apparatus and methods for remote virtual-to-physical address translation. Background Art

[0004] The Advanced Graphic Intelligence Logical Computing Environment (AGILE) project, funded by IARPA, aims to design systems for workflows that exhibit ultra-sparse graph characteristics. These applications can utilize specialized operations to address common behaviors in graph algorithms, saturating both network and memory bandwidth - a common performance goal for low-compute-intensity applications. Brief Description of the Drawings

[0005] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings, in which:

[0006] Figure 1 Illustrative example computer system architecture;

[0007] Figure 2 Illustrative processor including multiple cores;

[0008] FIG. 3(A) illustrates multiple stages of a processing pipeline;

[0009] FIG. 3(B) illustrates details of one embodiment of a core;

[0010] Figure 4 Illustrative execution circuitry according to one embodiment;

[0011] Figure 5 Illustrative embodiment of a register architecture;

[0012] Figure 6 Illustrative example of an instruction format;

[0013] Figure 7 Illustrative addressing techniques according to one embodiment;

[0014] Figure 8 Illustrative embodiment of an instruction prefix;

[0015] Figure 9(A)-Figure 9(D)Examples showing how to use the prefix R, X, and B fields;

[0016] Figure 10(A)-Figure 10(B) Examples showing a second instruction prefix;

[0017] Figure 11 Examples showing the payload bytes of an embodiment of an instruction prefix;

[0018] Figure 12 Examples showing instruction conversion and binary translation implementations;

[0019] Figure 13A-Figure 13B Examples showing a Transactional Integrated Global-memory system with Dynamic Routing and End-to-end flow control (TIGRE);

[0020] Figure 14 Examples showing a local atomicity buffer according to an embodiment of the present invention;

[0021] Figure 15 Examples showing an embodiment of an atomic unit (ATMU) according to an embodiment of the present invention;

[0022] Figure 16 Examples showing an architecture including an ATMU according to some embodiments;

[0023] Figure 17 Examples showing a reduced address table according to some embodiments;

[0024] Figure 18 Examples showing an embodiment with different types of request handlers;

[0025] Figure 19 Examples showing an embodiment with cache / memory accumulation;

[0026] Figure 20 Examples showing a method according to an embodiment of the present invention; and

[0027] Figure 21 Examples showing techniques for translating virtual addresses generated by an application into physical addresses and caching such translations in a TLB. DETAILED DESCRIPTION

[0028] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention described below. It will be apparent, however, to one of ordinary skill in the art that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid obscuring the fundamental principles of embodiments of the invention.

[0029] Exemplary Computer Architecture

[0030] Described below is a description of an exemplary computer architecture. Also suitable are other system designs and configurations known in the art for laptops, desktops, handheld PCs, personal digital assistants, engineering workstations, servers, network devices, network hubs, switches, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and a wide variety of other electronic devices. In general, a variety of systems or electronic devices capable of incorporating the processors and / or other execution logic disclosed herein are suitable.

[0031] Figure 1 An embodiment of an exemplary system is illustrated. The multiprocessor system 100 is a point-to-point interconnect system and includes a plurality of processors, including a first processor 170 and a second processor 180 coupled via a point-to-point interconnect 150. In some embodiments, the first processor 170 and the second processor 180 are homogeneous. In some embodiments, the first processor 170 and the second processor 180 are heterogeneous.

[0032] Processors 170 and 180 are shown as including integrated memory controller (IMC) unit circuits 172 and 182, respectively. Processor 170 also includes point-to-point (P-P) interfaces 176 and 178 as part of its interconnect controller unit; similarly, the second processor 180 includes P-P interfaces 186 and 188. Processors 170, 180 may exchange information via the P-P interfaces 178, 188 using a P-P interconnect 150. The IMCs 172 and 182 couple the processors 170, 180 to respective memories, namely memory 132 and memory 134, which may be part of the main memory locally attached to each processor.

[0033] The processors 170, 180 can each utilize point-to-point interface circuits 176, 194, 186, 198 to exchange information with the chipset 190 via respective P-P interconnections 152, 154. The chipset 190 can optionally exchange information with the coprocessor 138 via a high-performance interface 192. In some embodiments, the coprocessor 138 is a dedicated processor, such as a high-throughput MIC processor, a network or communication processor, a compression engine, a graphics processor, a GPGPU, an embedded processor, and so on.

[0034] A shared cache (not shown) can be included in either processor 170, 180, or outside both processors but connected to these processors via a P-P interconnection, such that: if a processor is placed in a low-power mode, the local cache information of either or both processors can also be stored in the shared cache.

[0035] The chipset 190 can be coupled to a first interconnection 116 via an interface 196. In some embodiments, the first interconnection 116 can be a Peripheral Component Interconnect (PCI) interconnection, or an interconnection such as a PCI Express interconnection or another I / O interconnection. In some embodiments, one of these interconnections is coupled to a power control unit (PCU) 117, and the PCU 117 can include circuitry, software, and / or firmware to perform power management operations regarding the processors 170, 180 and / or the coprocessor 138. The PCU 117 provides control information to a voltage regulator such that the voltage regulator generates an appropriate regulated voltage. The PCU 117 also provides control information to control the generated operating voltage. In various embodiments, the PCU 117 can include various power management logic units (circuits) to perform hardware-based power management. Such power management can be completely processor-controlled (e.g., controlled by various processor hardware and can be triggered by workload and / or power constraints, thermal constraints, or other processor constraints), and / or power management can be performed in response to an external source (e.g., a platform or a power management source or system software).

[0036] PCU 117 is illustrated as existing as separate logic from processor 170 and / or processor 180. In other cases, PCU 117 may execute on one or more given cores (not shown) within the core of processor 170 or 180. In some cases, PCU 117 may be implemented as a microcontroller (either dedicated or general purpose) or other control logic that is configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other embodiments, the power management operations to be performed by PCU 117 may be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other embodiments, the power management operations to be performed by PCU 117 may be implemented within the BIOS or other system software.

[0037] A variety of I / O devices 114 and an interconnect (bus) bridge 118 may be coupled to the first interconnect 116, and the interconnect (bus) bridge 118 couples the first interconnect 116 to the second interconnect 120. In some embodiments, one or more additional processors 115 are coupled to the first interconnect 116, such as a coprocessor, a high throughput MIC processor, a GPGPU, an accelerator (such as, for example, a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor. In some embodiments, the second interconnect 120 may be a low pin count (LPC) interconnect. A variety of devices may be coupled to the second interconnect 120, such devices including, for example, a keyboard and / or mouse 122, a communication device 127, and a storage unit circuit 128. The storage unit circuit 128 may be a disk drive or other mass storage device, which may include instructions / code and data 130 in some embodiments. Additionally, audio I / O 124 may be coupled to the second interconnect 120. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as multiprocessor system 100 may implement a multi-drop interconnect or other such architecture instead of a point-to-point architecture.

[0038] Exemplary Core Architectures, Processors, and Computer Architectures

[0039] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, the implementation of these cores can include: 1) general-purpose in-order cores, for general computing purposes; 2) high-performance general-purpose out-of-order cores, for general computing purposes; 3) specialized cores, mainly for graphics and / or scientific (throughput) computing purposes. The implementation of different processors can include: 1) a CPU, including one or more general-purpose in-order cores for general computing purposes and / or one or more general-purpose out-of-order cores for general computing purposes; and 2) a coprocessor, including one or more specialized cores mainly for graphics and / or scientific (throughput) purposes. These different processors result in different computer system architectures, which can include: 1) the coprocessor and the CPU on separate chips; 2) the coprocessor and the CPU on separate dies within the same package; 3) the coprocessor and the CPU on the same die (in this case, such a coprocessor is sometimes referred to as specialized logic, such as integrated graphics and / or scientific (throughput) logic, or as a specialized core); and 4) a system-on-chip, which can include the above coprocessor and additional functions on the same die as the described CPU (sometimes referred to as (one or more) application cores or (one or more) application processors). An exemplary core architecture will be described next, followed by a description of exemplary processors and computer architectures.

[0040] Figure 2 A block diagram illustrating an embodiment of an example processor 200 is shown. The processor 200 can have more than one core, can have an integrated memory controller, and can have an integrated graphics device. The processor 200 illustrated by the solid-line block diagram has a single core 202(A), a system agent 210, and a set of one or more interconnect controller unit circuits 216, while the optionally added dashed-line block diagram illustrates an alternative processor 200 as having multiple cores 202(A)-(N), a set of one or more integrated memory control unit circuits 214 in the system agent unit circuit 210, specialized logic 208, and a set of one or more interconnect controller unit circuits 216. Note that the processor 200 can be Figure 1 one of the processors 170 or 180 or the coprocessors 138 or 115.

[0041] Accordingly, different implementations of the processor 200 can include: 1) a CPU, where the dedicated logic 208 is integrated graphics and / or scientific (throughput) logic (which can include one or more cores, not shown), and the cores 202(A)-(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 202(A)-(N) are a large number of dedicated cores mainly for graphics and / or scientific (throughput) purposes; and 3) a coprocessor, where the cores 202(A)-(N) are a large number of general-purpose in-order cores. Accordingly, the processor 200 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit circuit), a high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor can be implemented on one or more chips. The processor 200 can be part of one or more substrates and / or can be implemented on one or more substrates using any of a variety of process technologies, such as BiCMOS, CMOS, or NMOS.

[0042] The memory hierarchy includes one or more levels of cache unit circuits 204(A)-(N) within the cores 202(A)-(N), a group of one or more shared cache unit circuits 206, and external memory (not shown) coupled to the group of integrated memory controller unit circuits 214. The group of one or more shared cache unit circuits 206 can include one or more intermediate-level caches, such as a level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, such as a last level cache (LLC), and / or combinations thereof. Although in some embodiments the ring-based interconnect network circuit 212 interconnects the dedicated logic 208 (e.g., integrated graphics logic), the group of shared cache unit circuits 206, and the system agent unit circuit 210, alternative embodiments use any number of well-known techniques to interconnect these units. In some embodiments, coherence is maintained between one or more of the circuits in the shared cache unit circuits 206 and the cores 202(A)-(N).

[0043] In some embodiments, one or more of cores 202(A)-(N) have multithreading capabilities. The system agent unit circuitry 210 includes those components that coordinate and operate on cores 202(A)-(N). The system agent unit circuitry 210 may include, for example, a power control unit (PCU) circuitry and / or a display unit circuitry (not shown). The PCU may be (or may include) the logic and components required to regulate the power states of cores 202(A)-(N) and / or dedicated logic 208 (e.g., integrated graphics logic). The display unit circuitry is used to drive one or more externally connected displays.

[0044] Cores 202(A)-(N) may be homogeneous or heterogeneous in terms of the architectural instruction set; that is, two or more of cores 202(A)-(N) may be capable of executing the same instruction set, while other cores may be capable of executing only a subset of that instruction set or may be capable of executing a different ISA.

[0045] Exemplary core architectures

[0046] In-order and out-of-order core block diagrams

[0047] The block diagram of FIG. 3(A) illustrates an exemplary in-order pipeline and both an exemplary register renaming, out-of-order issue / execution pipeline in accordance with an embodiment of the present invention. The block diagram of FIG. 3(B) illustrates an exemplary embodiment of an in-order architecture core and both an exemplary register renaming, out-of-order issue / execution architecture core to be included in a processor in accordance with an embodiment of the present invention. The solid boxes in FIGS. 3(A)-(B) illustrate the in-order pipeline and in-order core, while the optionally added dashed boxes illustrate the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0048] In FIG. 3(A), the processor pipeline 300 includes a fetch stage 302, an optional length decoding stage 304, a decode stage 306, an optional allocation stage 308, an optional rename stage 310, a schedule (also referred to as dispatch or issue) stage 312, an optional register read / memory read stage 314, an execution stage 316, a write-back / memory write stage 318, an optional exception handling stage 322, and an optional commit stage 324. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 302, one or more instructions are fetched from an instruction memory. During the decode stage 306, the fetched one or more instructions may be decoded, an address using a forwarding register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In one embodiment, the decode stage 306 and the register read / memory read stage 314 may be combined into one pipeline stage. In one embodiment, during the execution stage 316, the decoded instructions may be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.

[0049] As an example, an exemplary register renaming, out-of-order issue / execution core architecture may implement the pipeline 300 as follows: 1) Instruction fetch 338 performs the fetch and length decoding stages 302 and 304; 2) The decode unit circuit 340 performs the decode stage 306; 3) The rename / allocator unit circuit 352 performs the allocation stage 308 and the rename stage 310; 4) The (one or more) scheduler unit circuits 356 perform the schedule stage 312; 5) The (one or more) physical register file unit circuits 358 and the memory unit circuit 370 perform the register read / memory read stage 314; The execution cluster 360 performs the execution stage 316; 6) The memory unit circuit 370 and the (one or more) physical register file unit circuits 358 perform the write-back / memory write stage 318; 7) Various units (unit circuits) may be involved in the exception handling stage 322; and 8) The retirement unit circuit 354 and the (one or more) physical register file unit circuits 358 perform the commit stage 324.

[0050] FIG. 3(B) shows that the processor core 390 includes a front-end unit circuit 330 coupled to an execution engine unit circuit 350, and both are coupled to a memory unit circuit 370. The core 390 can be a reduced instruction set computing (RISC) core, a complex instruction set computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 390 can be a specialized core, such as a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.

[0051] The front-end unit circuit 330 may include a branch prediction unit circuit 332, which is coupled to an instruction cache unit circuit 334, which is coupled to an instruction translation lookaside buffer (TLB) 336, which is coupled to an instruction fetch unit circuit 338, which is coupled to a decoding unit circuit 340. In one embodiment, the instruction cache unit circuit 334 is included in the memory unit circuit 370 instead of the front-end unit circuit 330. The decoding unit circuit 340 (or decoder) may decode the instruction and generate one or more micro-operations, microcode entry points, micro-instructions, other instructions, or other control signals as output, which are decoded from the original instruction, or otherwise reflect the original instruction, or are derived from the original instruction. The decoding unit circuit 340 may also include an address generation unit circuit (AGU, not shown). In one embodiment, the AGU uses the forwarded register ports to generate LSU addresses and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). Various different mechanisms may be utilized to implement the decoding unit circuit 340. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc. In one embodiment, the core 390 includes a microcode ROM (not shown) or other medium that stores microcode for certain macro instructions (e.g., in the decoding unit circuit 340 or otherwise within the front-end unit circuit 330). In one embodiment, the decoding unit circuit 340 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache the decoded operations, micro-tags, or micro-operations generated during the decoding or other stages of the processor pipeline 300. The decoding unit circuit 340 may be coupled to a rename / allocator unit circuit 352 in the execution engine unit circuit 350.

[0052] The execution engine circuit 350 includes a rename / allocator unit circuit 352 that is coupled to a retirement unit circuit 354 and a set of one or more scheduler circuits 356. The scheduler circuits 356 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. In some embodiments, the (one or more) scheduler circuits 356 may include an arithmetic logic unit (ALU) scheduler / scheduling circuit, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuit, an AGU queue, and the like. The (one or more) scheduler circuits 356 are coupled to the (one or more) physical register file circuits 358. Each of the (one or more) physical register file circuits 358 represents one or more physical register files, where different physical register files among these physical register files store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., instruction pointer, i.e., the address of the next instruction to be executed), and the like. In one embodiment, the (one or more) physical register file unit circuits 358 include a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units may provide architected vector registers, vector mask registers, general-purpose registers, and the like. The (one or more) physical register unit file circuits 358 are overlapped by the retirement unit circuit 354 (also referred to as a retirement queue) to illustrate various ways that can be used to implement register renaming and out-of-order execution (e.g., using the (one or more) reorder buffer(s) (ROB) and the (one or more) retirement register files; using the (one or more) future files, the (one or more) history buffers, and the (one or more) retirement register files; using register maps and pools of registers; etc.). The retirement unit circuit 354 and the (one or more) physical register file circuits 358 are coupled to the (one or more) execution clusters 360. The (one or more) execution clusters 360 include a set of one or more execution unit circuits 362 and a set of one or more memory access circuits 364. The execution unit circuits 362 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shift, add, subtract, multiply) on various types of data (e.g., scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some embodiments may include several execution units or execution unit circuits dedicated to a particular function or set of functions, other embodiments may include only one execution unit circuit or multiple execution units / execution unit circuits that perform all functions.(One or more) scheduler circuits 356, (one or more) physical register file unit circuits 358, and (one or more) execution clusters 360 are shown as potentially multiple because some embodiments create separate pipelines for certain types of data / operations (e.g., scalar integer pipelines, scalar floating point / tight integer / tight floating point / vector integer / vector floating point pipelines, and / or memory access pipelines, each having its own scheduler circuit, (one or more) physical register file unit circuits, and / or execution cluster - and in the case of a separate memory access pipeline, in some embodiments only the execution cluster of that pipeline has (one or more) memory access unit circuits 364). It should also be understood that in cases where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution while the rest are in-order.

[0053] In some embodiments, the execution engine unit circuit 350 may perform load / store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown), as well as address phase and write-back, data phase load, store, and branch.

[0054] A set of memory access circuits 364 is coupled to a memory unit circuit 370, which includes a data TLB unit circuit 372, the data TLB circuit being coupled to a data cache circuit 374, which is coupled to a level 2 (L2) cache circuit 376. In an exemplary embodiment, the memory access unit circuit 364 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 372 in the memory unit circuit 370. An instruction cache circuit 334 is further coupled to the level 2 (L2) cache unit circuit 376 in the memory unit circuit 370. In one embodiment, the instruction cache 334 and the data cache 374 are combined into a single instruction and data cache (not shown) in the L2 cache unit circuit 376, a level 3 (L3) cache unit circuit (not shown), and / or main memory. The L2 cache unit circuit 376 is coupled to one or more other levels of cache and ultimately to main memory.

[0055] The core 390 may support one or more instruction sets (e.g., the x86 instruction set (with some extensions added with newer versions); the MIPS instruction set; the ARM instruction set (with optional additional extensions such as NEON)), which includes the (one or more) instructions described herein. In one embodiment, the core 390 includes logic to support SIMD (Single Instruction, Multiple Data) instruction set extensions (e.g., AVX1, AVX2), thus allowing operations used by many multimedia applications to be performed using SIMD data.

[0056] Exemplary Execution Unit Circuit

[0057] Figure 4 Illustrates an embodiment of the (one or more) execution unit circuits, such as the (one or more) execution unit circuits 362 of FIG. 3(B). As shown, the (one or more) execution unit circuits 362 may include one or more ALU circuits 401, vector / SIMD unit circuits 403, load / store unit circuits 405, and / or branch / jump unit circuits 407. The ALU circuit 401 performs integer arithmetic and / or Boolean operations. The vector / SIMD unit circuit 403 performs vector / SIMD operations on packed data (e.g., SIMD / vector registers). The load / store unit circuit 405 executes load and store instructions to load data from memory into registers or store data from registers to memory. The load / store unit circuit 405 may also generate addresses. The branch / jump unit circuit 407 causes a branch or jump to a certain memory address depending on the instruction. The floating-point unit (FPU) circuit 409 performs floating-point arithmetic. The width of the (one or more) execution unit circuits 362 varies depending on the embodiment and may range from 16 bits to 1024 bits. In some embodiments, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).

[0058] Exemplary Register Architecture

[0059] Figure 5 Is a block diagram of a register architecture 500 according to some embodiments. As shown, there are vector / SIMD registers 510, whose widths vary from 128 bits to 1024 bits. In some embodiments, the vector / SIMD registers 510 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some embodiments, the vector / SIMD register 510 is a 512-bit ZMM register: the lower 256 bits are used for the YMM register, and the lower 128 bits are used for the XMM register. Thus, there is register coverage. In some embodiments, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length. Scalar operations are operations performed on the lowest-order data element positions in the ZMM / YMM / XMM registers; the higher-order data element positions are either held the same as they were before the instruction or are zeroed, depending on the embodiment.

[0060] In some embodiments, the register architecture 500 includes write mask / predicate registers 515. For example, in some embodiments, there are eight write mask / predicate registers (sometimes referred to as k0 through k7), each of which may be 16 bits, 32 bits, 64 bits, or 128 bits in size. The write mask / predicate registers 515 may allow merging (e.g., allowing any set of elements in the destination to be protected from update during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask allows any set of elements in the destination to be zeroed) during the execution of any operation. In some embodiments, each data element position in a given write mask / predicate register 515 corresponds to a data element position in the destination. In other embodiments, the write mask / predicate registers 515 are scalable and consist of a set number of enable bits for a given vector element (e.g., eight enable bits for each 64-bit vector element).

[0061] The register architecture 500 includes a plurality of general-purpose registers 525. These registers may be 16 bits, 32 bits, 64 bits, etc., and are capable of being used for scalar operations. In some embodiments, these registers are named RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0062] In some embodiments, the register architecture 500 includes scalar floating-point registers 545, which are used for scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set extensions, or as MMX registers to perform operations on 64-bit packed integer data, and to hold operands for some operations performed between MMX and XMM registers.

[0063] One or more flag registers 540 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 540 may store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some embodiments, one or more flag registers 540 are referred to as program status and control registers.

[0064] The segment registers 520 contain segment points for accessing memory. In some embodiments, these registers are named CS, DS, SS, ES, FS, and GS.

[0065] The machine-specific register (MSR) 535 controls and reports processor performance. Most MSRs 535 handle system-related functions and are not accessible to applications. The machine-check register 560 consists of control, status, and error-reporting MSRs for detecting and reporting hardware errors.

[0066] One or more instruction pointer registers 530 store instruction pointer values. The control register(s) 555 (e.g., CR0-CR4) determine the operating mode of the processor (e.g., processors 170, 180, 138, 115, and / or 200) and the characteristics of the currently executing task. The debug register 550 controls and allows monitoring of the debug operations of the processor or core.

[0067] The memory management registers 565 specify the locations for data structures in protected-mode memory management. These registers can include the GDTR, IDTR, task register, and LDTR registers.

[0068] Alternative embodiments of the present invention can use wider or narrower registers. Additionally, alternative embodiments of the present invention can use more, fewer, or different register banks and registers.

[0069] Instruction set

[0070] An instruction set architecture (ISA) can include one or more instruction formats. A given instruction format can define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) on which the operation is to be performed and / or other data field(s) (e.g., mask), etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format can be defined as having different subsets of the fields of that instruction format (the included fields are generally in the same order, but at least some have different bit positions as fewer fields are included) and / or as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, using a given instruction template in that instruction format's instruction templates), and includes fields for specifying the operation and operands. For example, an exemplary ADD instruction has a specific opcode and an instruction format that includes an opcode field to specify the opcode and an operand field to select the operands (source 1 / destination and source 2); and the occurrence of this ADD instruction in the instruction stream will have specific contents in the operand fields that select the specific operands.

[0071] Exemplary Instruction Formats

[0072] Embodiments of the (one or more) instructions described herein may be implemented in different formats. Additionally, exemplary systems, architectures, and pipelines are detailed below. Embodiments of the (one or more) instructions may be executed on these systems, architectures, and pipelines, but are not limited to those detailed.

[0073] Figure 6 An embodiment of an instruction format is illustrated. As shown, an instruction may include multiple components, which include but are not limited to one or more fields for the following: one or more prefixes 601, an opcode 603, addressing information 605 (e.g., register identifiers, memory addressing information, etc.), a displacement value 607, and / or an immediate value 609. Note that some instructions utilize some or all of the fields of this format, while other instructions may use only the fields of the opcode 603. In some embodiments, the illustrated order is the order in which these fields are to be encoded, however it should be understood that in other embodiments, these fields may be encoded in a different order, combined, etc.

[0074] The (one or more) prefix fields 601 modify the instruction when used. In some embodiments, one or more prefixes are used for repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.), provide section override (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), perform a bus lock operation, and / or change operand (e.g., 0x66) and address size (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes may be considered “traditional” prefixes. Other prefixes (one or more examples of which are detailed herein) indicate and / or provide further capabilities, such as specifying particular registers, etc. These other prefixes typically follow the “traditional” prefixes.

[0075] The opcode field 603 is used to at least partially define the operation to be performed upon decoding of the instruction. In some embodiments, the length of the main opcode encoded in the opcode field 603 is 1, 2, or 3 bytes. In other embodiments, the main opcode may be of other lengths. An additional 3-bit opcode field is sometimes encoded in another field.

[0076] The addressing field 605 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 7An embodiment of the addressing field 605 is illustrated. In this illustration, an optional MOD R / M byte 702 and an optional Scale, Index, Base (SIB) byte 704 are shown. The MOD R / M byte 702 and the SIB byte 704 are used to encode up to two operands of an instruction, each operand being either a direct register or an effective memory address. Note that each of these fields is optional, i.e., not all instructions include one or more of these fields. The MOD R / M byte 702 includes a MOD field 742, a register field 744, and an R / M field 746.

[0077] The content of the MOD field 742 differentiates between memory access and non-memory access modes. In some embodiments, when the MOD field 742 has the value b11, register direct addressing mode is utilized, otherwise register indirect addressing is used.

[0078] The register field 744 can encode a destination register operand or a source register operand, or can also encode an opcode extension without being used to encode any instruction operand. The content of the register index field 744 directly specifies or specifies through address generation the location (in a register or in memory) of the source or destination operand. In some embodiments, the register field 744 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing.

[0079] The R / M field 746 can be used to encode an instruction operand that references a memory address, or can be used to encode a destination register operand or a source register operand. Note that in some embodiments, the R / M field 746 can be combined with the MOD field 742 to specify an addressing mode.

[0080] The SIB byte 704 includes a scale field 752, an index field 754, and a base field 756 for address generation. The scale field 752 indicates a scale factor. The index field 754 specifies the index register to be used. In some embodiments, the index field 754 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. The base field 756 specifies the base register to be used. In some embodiments, the base field 756 is supplemented with additional bits from a prefix (e.g., prefix 601) to allow for greater addressing. In practice, the content of the scale field 752 allows the content of the index field 754 to be scaled for memory address generation (e.g., for address generation using 2 缩放 * index + base).

[0081] Some addressing forms utilize displacement values to generate memory addresses. For example, according to 2缩放 *Index + base + displacement, index * scale + displacement, r / m + displacement, instruction pointer (RIP / EIP) + displacement, register + displacement, etc. to generate a memory address. The displacement can be a value of 1 byte, 2 bytes, 4 bytes, etc. In some embodiments, the displacement field 607 provides this value. Additionally, in some embodiments, the use of the displacement factor is encoded in the MOD field of the addressing field 605, which indicates a compressed displacement scheme for which the displacement value is calculated by multiplying disp8 with a scaling factor N determined based on the vector length, a value of b bits, and the input element size of the instruction. The displacement value is stored in the displacement field 607.

[0082] In some embodiments, the immediate field 609 specifies an immediate value for the instruction. The immediate value can be encoded as a 1 - byte value, 2 - byte value, 4 - byte value, etc.

[0083] Figure 8 An embodiment of the first prefix 601(A) is illustrated. In some embodiments, the first prefix 601(A) is an embodiment of the REX prefix. Instructions using this prefix can specify general - purpose registers, 64 - bit packed data registers (e.g., single - instruction multiple - data (SIMD) registers or vector registers), and / or control and debug registers (e.g., CR8 - CR15 and DR8 - DR15).

[0084] Instructions using the first prefix 601(A) can specify up to three registers using a 3 - bit field, depending on the format: 1) using the reg field 744 and the R / M field 746 of the MOD R / M byte 702; 2) using the MOD R / M byte 702 and the SIB byte 704, including using the reg field 744 and the base field 756 and the index field 754; or 3) using the register field of the opcode.

[0085] In the first prefix 601(A), bit positions 7:4 are set to 0100. Bit position 3 (W) can be used to determine the operand size but cannot determine the operand width alone. Thus, when W = 0, the operand size is determined by the code segment descriptor (CS.D), and when W = 1, the operand size is 64 bits.

[0086] Note that adding another bit allows addressing of 16 (2 4 ) registers, while the separate MOD R / M reg field 744 and the R / M field 746 of the MOD R / M can each address only 8 registers.

[0087] In the first prefix 601(A), bit position 2 (R) can be an extension of the reg field 744 of MOD R / M, and can be used to modify the reg field 744 of MOD R / M when this field encodes a general-purpose register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the MOD R / M byte 702 specifies other registers or defines an extended opcode, R is ignored.

[0088] Bit position 1 (X) The X bit can modify the SIB byte index field 754.

[0089] Bit position B (B) B can modify the base address in the R / M field 746 of MOD R / M or the SIB byte base address field 756; or it can modify the opcode register field for accessing a general-purpose register (e.g., general-purpose register 525).

[0090] Figures 9(A)-(D) illustrate examples of how the R, X, and B fields of the first prefix 601(A) are used. Figure 9(A) illustrates that when the SIB byte 704 is not used for memory addressing, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702. Figure 9(B) illustrates that when the SIB byte 704 is not used, R and B from the first prefix 601(A) are used to extend the reg field 744 and the R / M field 746 of the MOD R / M byte 702 (register-register addressing). Figure 9(C) illustrates that when the SIB byte 704 is used for memory addressing, R, X, and B from the first prefix 601(A) are used to extend the reg field 744, the index field 754, and the base address field 756 of the MOD R / M byte 702. Figure 9(D) illustrates that when a register is encoded in the opcode 603, B from the first prefix 601(A) is used to extend the reg field 744 of the MOD R / M byte 702.

[0091] Figures 10(A)-(B) illustrate examples of the second prefix 601(B). In some embodiments, the second prefix 601(B) is an embodiment of the VEX prefix. The second prefix 601(B) encoding allows an instruction to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 510) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 601(B) provides a syntax for three-operand (or more) operations. For example, a previous two-operand instruction performed an operation such as A = A + B, which overwrote the source operand. The use of the second prefix 601(B) enables the operands to perform non-destructive operations, such as A = B + C.

[0092] In some embodiments, the second prefix 601(B) has two forms - a two-byte form and a three-byte form. The two-byte second prefix 601(B) is mainly used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 601(B) provides a compact replacement for 3-byte opcode instructions and the first prefix 601(A).

[0093] Figure 10(A) illustrates an embodiment of the two-byte form of the second prefix 601(B). In one example, the format field 1001 (byte 0 1003) contains the value C5H. In one example, byte 1 1005 includes the value "R" in bit [7]. This value is the complement of the same value of the first prefix 601(A). Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with 2 or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form, for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.

[0094] Instructions using this prefix can use the R / M field 746 of the MOD R / M to encode instruction operands that reference memory addresses, or to encode destination register operands or source register operands.

[0095] Instructions using this prefix can use the reg field 744 of the MOD R / M to encode destination register operands or source register operands, which are treated as opcode extensions and not used to encode any instruction operands.

[0096] For instruction syntax that supports four operands, vvvv, the R / M field 746 of the MOD R / M, and the reg field 744 of the MOD R / M encode three of the four operands. Then bits [7:4] of the immediate 609 are used to encode the third source register operand.

[0097] Figure 10(B) illustrates an embodiment of the three-byte form of the second prefix 601(B). In one example, the format field 1011 (byte 0 1013) contains the value C4H. Byte 1 1015 includes "R", "X", and "B" in bits [7:5], which are the complements of these values of the first prefix 601(A). The bits [4:0] of byte 1 1015 (shown as mmmmm) include content for encoding one or more implicit leading opcode bytes as needed. For example, 00001 means a 0FH leading opcode, 00010 means a 0F38H leading opcode, 00011 means a 0F3AH leading opcode, and so on.

[0098] The use of bit [7] of byte 2 1017 is similar to that of W of the first prefix 601(A), including helping to determine the size of the operand that can be promoted. Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or a 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). The bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form and is used for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.

[0099] Instructions using this prefix can use the R / M field 746 of MOD R / M to encode the instruction operand that references a memory address, or to encode the destination register operand or the source register operand.

[0100] Instructions using this prefix can use the reg field 744 of MOD R / M to encode the destination register operand or the source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.

[0101] For the instruction syntax that supports four operands, vvvv, the R / M field 746 of MOD R / M, and the reg field 744 of MOD R / M encode three of the four operands. Then the bits [7:4] of the immediate 609 are used to encode the third source register operand.

[0102] Figure 11Illustrates an embodiment of a third prefix 601(C). In some embodiments, the first prefix 601(A) is an embodiment of an EVEX prefix. The third prefix 601(C) is a four-byte prefix.

[0103] The third prefix 601(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some embodiments, instructions that utilize a write mask / operation mask (see the discussion of registers in the previous figures, e.g., Figure 5 ) or a predicate utilize this prefix. The operation mask register allows for conditional processing or selective control. Operation mask instructions - whose source / destination operand is an operation mask register and treat the contents of the operation mask register as a single value - are encoded using the second prefix 601(B).

[0104] The third prefix 601(C) can encode functionality specific to instruction classes (e.g., packed instructions with "load + operation" semantics can support an embedded broadcast function, floating-point instructions with rounding semantics can support a static rounding function, floating-point instructions with non-rounding arithmetic semantics can support a "suppress all exceptions" function, etc.).

[0105] The first byte of the third prefix 601(C) is the format field 1111, which has a value of 62H in one example. The subsequent bytes are referred to as payload bytes 1115-1119 and together form a 24-bit value for P[23:0], providing specific capabilities in the form of one or more fields (detailed herein).

[0106] In some embodiments, P[1:0] of payload byte 1119 is the same as the two lowermost mm bits. In some embodiments, P[3:2] is reserved. Bit P[4] (R') allows access to the upper 16 vector register set when combined with P[7] and the reg field 744 of MOD R / M. When SIB type addressing is not required, P[6] can also provide access to the upper 16 vector registers. P[7:5] consists of R, X, and B, which are operand specifier modifier bits for vector registers, general registers, and memory addressing operands, and when combined with the MOD R / M register field 744 and the R / M field 746 of MOD R / M, allow access to the next set of 8 registers beyond the lower 8 registers. P[9:8] provides opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). P

[10] is a fixed value 1 in some embodiments. P[14:11], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with 2 or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, e.g., 1111b.

[0107] P

[15] is similar to the W of the first prefix 601(A) and the second prefix 611(B), and can be used as an opcode extension bit or an operand size promotion.

[0108] P[18:16] specifies the index of a register in an operation mask (write mask) register (e.g., write mask / predicate register 515). In one embodiment of the present invention, a particular value aaa = 000 has special behavior, implying that no operation mask is used for this particular instruction (which can be achieved in various ways, including using an operation mask hard-wired to all ones or hardware that bypasses the masking hardware). When combined, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base and enhanced operations); in another embodiment, the old value of each element of the destination is retained (if the corresponding mask bit has a value of 0). In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base and enhanced operations); in one embodiment, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span of the elements being modified, from the first to the last); however, the elements being modified do not have to be contiguous. Thus, the operation mask field allows for partial vector operations, including loads, stores, arithmetic, logic, etc. Although in the described embodiments of the present invention, the content of the operation mask field selects which one of several operation mask registers contains the operation mask to be used (thus the content of the operation mask field indirectly identifies the masking to be performed), alternatively or additionally, alternative embodiments allow the content of the mask write field to directly specify the masking to be performed.

[0109] P

[19] can be combined with P[14:11] to encode a second source vector register in a non-destructive source syntax that can utilize P

[19] to access the upper 16 vector registers. P

[20] encodes various functions that vary among different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P

[23] indicates support for merge-write masking (e.g., when set to 0) or support for zeroing and merge-write masking (e.g., when set to 1).

[0110] The following table details an exemplary embodiment of the encoding of registers in instructions using the third prefix 601(C).

[0111]

[0112]

[0113] Table 1: 32-register support in 64-bit mode

[0114]

[0115] Table 2: Encoding Register Designators in 32 - bit Mode

[0116]

[0117]

[0118] Table 3: Encoding of Operation Mask Register Designators

[0119] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0120] The program code can be implemented in a high - level procedural or object - oriented programming language to communicate with the processing system. If desired, the program code can also be implemented in assembly or machine language. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0121] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of these implementation approaches. Embodiments of the present invention can be implemented as a computer program or program code, executed on a programmable system including at least one processor, a storage system (including volatile and non - volatile memory and / or storage elements), at least one input device, and at least one output device.

[0122] One or more aspects of at least one embodiment can be implemented by representative instructions stored on a machine - readable medium, which represent various logics within a processor. When these instructions are read by the machine, they cause the machine to fabricate the logic for performing the techniques described herein. These representations are referred to as "IP cores" and can be stored on a tangible machine - readable medium and provided to various customers or manufacturing facilities to be loaded into the fabrication machines that actually make the logic or processor.

[0123] These machine-readable storage media may include—but are not limited to—non-transitory tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as: hard disks, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.

[0124] Accordingly, embodiments of the present invention also include non-transitory tangible machine-readable media that contain instructions or contain design data that define the structural, circuit, device, processor, and / or system features described herein, such as a Hardware Description Language (HDL). Such embodiments may also be referred to as program products.

[0125] Emulation (including binary translation, code morphing, etc.)

[0126] In some cases, an instruction converter may be used to convert instructions from a source instruction set to a target instruction set. For example, the instruction converter may translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions to one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, outside the processor, or part on the processor and part outside the processor.

[0127] Figure 12A block diagram illustrating the use of a software instruction converter to convert binary instructions in a source instruction set into binary instructions in a target instruction set according to some implementations. In the illustrated embodiment, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 12 It is shown that a program in a high-level language 1202 can be compiled using a first ISA compiler `BPL04 to generate a first ISA binary code 1206 that can be natively executed by a processor 1216 having at least one first ISA instruction set core. The processor 1216 having at least one first ISA instruction set core represents any processor that is capable of executing or otherwise processing (1) a substantial portion of the instruction set of the first ISA instruction set core or (2) executing a program in a first ISA binary code 1206 natively on a processor having at least one first ISA instruction set core. The first ISA compiler 1204 represents a compiler operable to generate first ISA binary code 1206 (e.g., object code) that can be executed on a processor 1216 having at least one first ISA instruction set core with or without additional linking processing.

[0128] Similarly, Figure 12 It is shown that a program in a high-level language 1202 can be compiled using an alternative instruction set compiler 1208 to generate an alternative instruction set binary code 1210, which can be executed natively by a processor 1214 without a first ISA core. An instruction converter 1212 is used to convert the first ISA binary code 1206 into code that can be executed natively by a processor 1214 without a first ISA instruction set core. This converted code may not be the same as the alternative instruction set binary code 1210, because an instruction converter that can do this is difficult to make; however, the converted code will implement the overall operation and be composed of instructions from the alternative instruction set. Thus, the instruction converter 1212 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have a first ISA instruction set processor or core to execute the first ISA binary code 1206 through simulation, emulation, or any other process.

[0129] Embodiments of a transactional integrated global memory system with dynamic routing and end-to-end flow control

[0130] Embodiments of the present invention include ISA and architectural support for direct memory operations for drawing data structures. Some implementations are configured outside the core cache hierarchy to provide higher efficiency through improved memory and network bandwidth utilization. The use of near-memory computing operations reduces the total latency by eliminating additional network traversals and taking the shortest total path to all physical memory locations involved in the operation.

[0131] Some implementations include a transactional integrated global memory system (TIGRE) with dynamic routing and end-to-end flow control as described herein, which is a 64-bit distributed global address space (DGAS) system solution for large-scale hybrid-mode (sparse and dense) analysis. TIGRE implements complex DMA operations that are specifically designed to handle primitives common in graph processing.

[0132] Operations on the TIGRE system can involve various subsystems, including pipeline-local DMA engines and near-memory computing at system endpoints. Additionally, an atomic locking buffer located near the memory is implemented to facilitate remote atomic locking / unlocking operations involved in DMA bit manipulation operations.

[0133] In one example, each TIGRE pipeline offloads DMA operations (e.g., exposed in the ISA) to a local memory engine (MENG), where eight TIGRE pipelines in a TIGRE pipeline are co-located with a shared cache and local SRAM scratchpads to create a TIGRE "slice". A TIGRE slice can include eight slices (e.g., 64 pipelines) and sixteen local DRAM channels. As the system scales out, multiple slices form a TIGRE "socket", and the socket count increases to scale the entire system.

[0134] Now turning to Figure 13A and Figure 13B , FIGS. 1320 of a TIGRE slice and 22 of a TIGRE slice are shown, respectively. Figure 13A and Figure 13BIllustrates the lowest level of the hierarchy of the TIGRE system. More specifically, the TIGRE slice 1320 includes a plurality of memory engines 1324 (1324a - 1324i) corresponding to a plurality of pipelines 1326 (1326a - 1326i), where each memory engine 1324 is adjacent to a pipeline among the plurality of pipelines 1326. Each TIGRE pipeline 1326 migrates DMA operations (e.g., exposed in the ISA) to the local memory engine 1324 (MENG). In the illustrated example, eight of the TIGRE pipelines 1326 in the TIGRE slice 1320 are co - located with a shared cache (not shown) and a local SRAM scratchpad 1328 to create the TIGRE slice 1320. The illustrated TIGRE tile 1322 includes eight slices 1320 - - e.g., sixty - four pipelines 1326 and sixteen local DRAM channels 1330 (1330a - 1330j). Specifically, the DMA subsystem hardware consists of units local to the pipelines 1326 and in front of all scratchpads 1328 and DRAM channel 1330 interfaces.

[0135] Atomicity units 1334 (e.g., 1334a - 1334j, not shown (e.g., ATMU)) are positioned adjacent to the scratchpad 1328 and the memory interface 1336 and handle the compute and read - lock / write - unlock functions within remote atomicity operations. Requests can be sent directly by the pipeline 1326 or by the memory engine 1324 to the ATMU 1334. The ATMU 1334 includes integer and floating - point compute units, as well as a local load - store buffer for supporting the parallel execution of instructions while also maintaining high - throughput atomic read - write requests to the DRAM channels 1330.

[0136] The memory engine 1324 (MENG) receives DMA bitmap requests from the local pipeline 1326 and initiates operations. For example, the first MENG 1324a is responsible for requesting one or more DMA bitmap manipulation operations associated with the first pipeline 1326a. Thus, with or without atomicity operations, the first MENG 1324a sends remote load - stores directly or indirectly outwards. The first MENG 1324a also tracks the remote load - stores sent and waits for all responses to return before sending the final response back to the first pipeline 1326a.

[0137] The operation engine 1332 (1332a - 1332j, not shown (e.g., OPENG (Operation engine)) is positioned adjacent to the memory interface 1336 (1336a - 1336j) and receives load / store requests from the MENG 1324. The OPENG 1332 is responsible for performing the actual memory load / store, converting the stored pointer values to physical addresses, and sending subsequent load / store or atomic requests as appropriate. Details regarding the role of the OPENG 1332 in DMA bitmap manipulation operations are provided below.

[0138] The lock buffer 1338 is positioned in front of the memory port and maintains a row lock state for memory addresses. Each lock buffer 1338 is a multi - entry buffer that allows multiple parallel lock addresses for each memory interface 1336, supports 64 - byte (byte, B) or 8B requests, handles partial row updates and write combining for partial stores, and supports "read lock" and "write unlock" requests within atomic operations ("atomicity"). The lock buffer 1338 also serves as a small cache to allow fast access to memory data for bitmap manipulation operations.

[0139] Memory System Remote Bit Map Manipulation Operation

[0140] In the memory system described herein, DMA bitmap instructions listed in Table 4 can be used to perform bitmap manipulation operations. The DMA bitmap instructions are issued from the pipeline to their corresponding local MENG 1324, which then utilizes the OPENG 1332 and the ATMU 1334 near the source and destination memory locations. In addition to direct bitmap manipulation, these instructions also implement batch bitmap manipulation (e.g., bitmap operations performed on a series of bitmaps pointed to by an initial list).

[0141]

[0142]

[0143] Table 4

[0144] Table 4 shows that the DMA operation receives a DMA_Type (DMA type) field as part of the ISA instruction. The DMA_type field contains information about the addressing mode, data type representation, and destination atomic operation (if specified). Table 5 describes the functions of the different bit fields in the DMA type modifiers.

[0145]

[0146] Table 5

[0147] Table 6 further explains the atomic operations for DMA instructions. The bit fields in the DMA_Type (DMA type) parameter accommodate operations with a relatively small number of bits and provide flexibility for future added features.

[0148]

[0149]

[0150] Table 6

[0151] Figure 14 An example of a unified memory architecture core with six pipelines 0 - 5 is illustrated (only two of the six pipelines are shown for simplicity). Some implementations may include several of these cores per die (e.g., four, eight, sixteen, etc.). The illustrated architecture scales to O(1000) nodes, where all memory endpoints are visible via a 64 - bit distributed global address space (DGAS). The storage device includes a 4MB (2 - port) scratchpad memory 1441 (e.g., static random access memory (SRAM)) within the scratchpad memory subsystem 1475, and an adjacent narrow - channel (8B) DRAM interface / memory controller 1460 within the near - memory subsystem 1470 (e.g., operable at 4GB / core), while the model - specific registers (MSR) and physical register file (PRF) can also be targeted by any pipeline. Although not shown, embodiments of the present invention can be used within various other levels of memory (such as storage - class or "far" memory (not shown)).

[0152] In some embodiments, atomic unit (ATMU) engines 1411 - 1415 located at multiple positions handle the execution of atomic operations next to the memory interface. Each ATMU engine 1411 - 1415 includes one or more integer units and one or more floating - point units, and can support parallel execution of multiple instructions to maintain high throughput.

[0153] In some implementations, the lock buffers 1431 - 1433 maintain a row lock state for memory addresses behind the scratchpad 1441 and the DRAM port 1460. The lock buffers 1431 - 1433 are multi - entry buffers that allow multiple parallel lock addresses per memory interface, handle partial row updates and write combining for partial stores, and support "read lock" and "write unlock" requests within atomicity. Requests of different sizes are supported, such as 64B or 8B requests described in detail below. The arbitration units 1490 - 1493 control access to the high - speed interconnect network 1420, which interconnects core pipelines 0 - 5 to the scratchpad memory 1441, the ATMUs 1414 - 1414, the lock buffers 1431 - 1432, and the near - memory subsystem 1470.

[0154] The on - die switch 1450 couples the near - memory subsystem 1470 to the interconnect network 1420 to provide access to the DRAM memory controller / interface 1460 and the associated lock buffer 1433 and ATMU 1415 via the arbitration unit 1494.

[0155] In some implementations, each of pipelines 0, 5 includes an engine sequence queue (ESQ) 1450A, 1450B and an atomic buffer (ABUF) 1455A, 1455B respectively. The ESQs 1450A - 1450B manage the state and order of migrated operations, including DMA, atomic, thread - context operations, and collections. In these embodiments, each of the ESQs 1450A - 1450B supports "chaining" of migrated operations to ensure sequential operations. The ESQs 1450A - 1450B provide visibility of the migrated operations to each corresponding pipeline 0, 5 by implementing wait - like instructions and / or triggering interrupts when the migrated operations are complete. For the purpose of fencing migrated operations, the ESQs 1450A - 1450B also allow their respective pipelines 0, 5 to track the state.

[0156] The atomic buffers (ABUFs) 1455A - 1455B are also pipeline - local units that track outstanding remote atomic requests sent from that pipeline. Each of the ABUF 1455A - 1455B is responsible for submitting the return value (if necessary) back to the local register file of its corresponding pipeline 0, 5, thus indicating completion to the corresponding ESQs 1450A - 1450B and reporting any exceptions that occurred during the remote atomic operation.

[0157] The remainder of this disclosure will discuss each aspect of the embodiments of the present invention, including remote atomic operations, ISA descriptions, and pipeline behavior, and then discuss the operation of the ATMUs 1413 - 1415 and the lock buffers 1431 - 1433 at each respective memory interface.

[0158] Table 4 lists a set of remote atomic instructions that can be implemented by embodiments of the present invention. These instructions can be provided as part of an ISA extension. However, note that the fundamental principles of the present invention do not require the implementation of these specific instructions.

[0159] For operations that include computation, both integer and floating-point types are supported via different instructions (i.e., xaddI and xaddF). These are combined into a single row in Table 4. All operation types have three types of instruction input parameters to support different types of addressing. All types are listed, and the differences in address generation are specified in the "Description" column.

[0160] The RVAL parameter shown in these instructions allows the programmer to specify whether a return value (e.g., for writing to the register file) is needed. If needed, the instruction will become blocking when data dependencies occur within the pipeline. If not needed, the instruction is migrated and executed in the background, becoming visible only when the software implements a fence for the atomic operation. Instructions without an RVAL (i.e., xchg) will always write the return value to the register file. The SIZE parameter specifies the target data size as 8, 16, 32, or 64 bits.

[0161]

[0162]

[0163]

[0164]

[0165] Table 7

[0166] As mentioned, each atomic buffer (ABUF) 1455A - 1455B is a pipeline local unit responsible for packing and transmitting remote atomic packets to the destination, tracking the status of remote atomic operations, and submitting the return value back to the local register file.

[0167] Figure 15Illustrated is the ATMU 1415 according to some embodiments of the present invention. Instructions (remote requests) received from the interconnect network 1420 via the ARB port 1530 are queued in the input FIFO 1505 for execution, and responses 1520 are sent via the remote RSP port 1531 (also via the arbitration unit). At the front end of the FIFO 1505, the instructions are assigned to one of "n" parallel execution threads T0, T1, … T(n-2), T(n-1), which share access to one or more pipelined ALU / FPU units 1510 and two output ports: the output local request port 1520 goes through the arbitration unit 1494 to the lock buffer 1433 for "read lock" requests, and the output port 1521 goes directly to the lock buffer 1433 for "write unlock" requests.

[0168] In operation, once an instruction packet is accepted by the ATMU 1415, the ATMU 1415 decodes the instruction and generates a "read lock" request to the local memory controller 1460 via the lock buffer 1433. The lock buffer 1433 returns the requested data, and the address remains locked in the lock buffer. The ATMU 1415 then uses its internal ALU / FPU 1510 to perform the operation specified by the instruction and submits the result of the operation back to the memory 1460 via the lock buffer 1433 using a "write unlock" request. At this point, the operation is complete, and the lock buffer 1433 unlocks the (one or more) lines. The ATMU 1415 transmits an instruction completion indication from the remote response port 1531 (which travels through the interconnect network 1420 to the ABUF 1455A) back to the source ABUF1455A. If a return value is requested, the data value is included in the response packet and can be written to the register file 1405. For example, a 64-bit value can be sent in the response packet, which the ABUF 1455A stores in the register file 1405 and notifies the ESQ 1450A of the corresponding pipeline.

[0169] The ATMU 1415 supports the highest throughput for performing remote atomic operations. Since remote atomic requests can arrive at the ATMU 1415 at a rate of one per cycle, some implementations of the ATMU 1414 support performing at least one operation per cycle.

[0170] In some embodiments, the number of execution threads (n) (T0 - T(n - 1)) is determined by the latency of a single atomic operation, making the number of execution threads present in each ATMU instance unique across each subsystem block. All other features may be consistent across the ATMUs. Due to the lower latency of SRAM access on the die, the ATMUs in front of the staging ports (such as ATMUs 1413 - 1414) may have fewer execution threads.

[0171] Apparatus and method for remote virtual to physical address translation

[0172] Some embodiments of the present invention are directed to workflows that exhibit ultra - sparse graph characteristics. These embodiments utilize specialized operations to address common behavior in graph algorithms to saturate both network and memory bandwidths - a common performance goal for low - compute - intensity applications.

[0173] The ultra - sparse graph requirements of these workflows can be addressed by implementing custom indirect load and store instructions (e.g., for facilitating A[B[i]]), and direct memory access (DMA) operations with unique graph traversal behavior and the ability to compute on near - memory targets. These DMA operations saturate network and memory bandwidths by maintaining a large number of requests to hide the latency of remote random pointer chasing behavior across the entire system. Some embodiments facilitate this through an Indirect Unit (IDU) at each memory endpoint, which can support multiple levels of indirection beyond just one level.

[0174] Figure 16The figure illustrates an example of using a single level of indirection. As illustrated, requests and responses of various components are transmitted through the interconnect network 1420. In this example, a logical processor (LP) or core 1601 generates an indirect request for data at address A (e.g., in response to one or more load instructions). The indirect request, which initially has to read a pointer to A from memory 1621A, is received and processed by the IDU 1610, which generates a read from memory 1621A that returns the pointer A to the IDU 1610. Instead of returning the pointer A to the LP / core 1601 which would then need to perform address translation and transmit the request to address A, the IDU 1610 directly transmits the pointer A to the IDU 1611 local to memory 1621B with address A. The IDU 1611 then performs a load / store operation on address A in memory 1621B, and memory 1621B transmits the data at address A (e.g., a cache line) to the IDU 1611, which returns the data to the LP / core 1601 in the form of an indirect response. Thus, the operations of reading and translating the pointer A, and reading the data at address A, are migrated from the LP / core 1601 to the IDUs 1610 - 1611.

[0175] As used herein, a logical processor (LP) includes a representation of the processor core presented to system software (e.g., an operating system). Depending on the configuration, a single LP may include a portion of a core (e.g., 1 / 2 of a hyperthreaded core), a single core, or multiple cores. The terms logical processor and core may sometimes be used interchangeably herein to indicate discrete portions of the processor's instruction processing resources. 1 / 2 cores), a single core, or multiple cores. The terms logical processor and core may sometimes be used interchangeably herein to indicate discrete portions of the processor's instruction processing resources.

[0176] As mentioned, some embodiments are integrated within the TIGRE (Transactional Integrated Global Memory System with Dynamic Routing and End - to - End Flow Control) architecture, which implements the RISC - V RV64 ISA, including the SV57 virtual addressing extension that supports a 57 - bit virtual address space and an additional level of page tables. As a virtual addressing system, this means that DMA and indirect load / store operations will require pointers read from memory to be translated to physical addresses before being sent to a second destination address. For example, in Figure 16 this case, the address A pointer is fetched and translated to a physical address in memory 1621B. This will typically require the pointer to be sent back to the issuing core 1601 for translation via the core's translation lookaside buffer (TLB). However, doing so can eliminate the efficiency gains provided by the indirect memory access operation.

[0177] To overcome this limitation and better enable the features provided by TIGRE for sparse analysis, embodiments of the present invention include a remote TLB and page table walker (PTW) circuitry (e.g., a state machine) implemented at each memory endpoint to maintain virtual-to-physical translation and walk the page table upon a remote TLB miss. The ability to perform virtual-to-physical translation efficiently anywhere in the memory system preserves the efficiency gains of the indirect memory operations of the global memory system. In these embodiments, page table information may be transmitted with each memory request such that indirect memory requests of any thread may be translated anywhere in the global memory system.

[0178] According to these embodiments, any virtual address from any issuing thread may be translated anywhere (i.e., locally to any physical memory) rather than being sent back to the issuing core for translation. While described in the context of a particular global memory system with a distributed global address space (DGAS) and a highly scalable low-diameter and high-radix network to scale to 100k sockets in a single system, the fundamental principles of the present invention may be implemented in any other architecture using virtual addressing to achieve similar capabilities and efficiency gains.

[0179] Several classes of operations may be implemented to help meet the target requirements, particularly those related to sparse memory access. Some of these operations include DMA as well as indirect load and store instructions. DMA includes functions such as gather / scatter memory to / from an address list, and indirect load and store instructions read a pointer from memory and then submit a load or store request to that pointer without returning to the issuing thread (e.g., as described with respect to Figure 16 ).

[0180] Indirect memory operations pose a challenge to virtual addressing because each indirect memory operation must translate from a virtual address to a physical address. The remote TLB and PTW mechanisms described herein address this challenge while leveraging indirect memory operations.

[0181] Some embodiments are configured to work with a stable toolchain running code written in general-purpose languages such as C++ and Python, including support for the Linux operating system. Thus, these embodiments implement RISC-V ISA, including privileged modes and 57-bit virtual addressing extensions.

[0182] As a brief overview, RISC-V page table information is stored in the Supervisor Address Translation and Protection (SATP) register, which is defined for each hardware thread in the system. Refer to Figure 17 , an embodiment of the SATP register 1700 includes a MODE (mode) field 1701 that indicates the version of virtual addressing used by the implementation (e.g., set to 0xA for SV-57 in some implementations). The address space identifier (ASID) field 1702 (16 bits in one embodiment) indicates the address space associated with a given memory access request. In some embodiments, each virtual machine or application is assigned a unique ASID. The physical page number (PPN) field 1703 (44 bits in one embodiment) stores the physical address of the root of the page table for the corresponding hardware thread.

[0183] Figure 18 The figure includes an example of an SV57 virtual address that includes five 9-bit VPN values 1800 - 1804 corresponding to the virtual address of the page, and a 12-bit page offset value 1810 corresponding to a specific byte position within the page.

[0184] In some implementations, the page table walk (PTW) translates the virtual address to a physical address using the PPN provided in the SATP register to identify the root table and potentially iterate through all five levels of the virtual address until a valid page table entry (PTE) is identified, which provides the physical address mapped to the virtual address for that ASID.

[0185] The translation lookaside buffer (TLB) caches these translations so that a full page table walk does not need to be performed every time (e.g., in response to subsequent instruction fetches or memory access requests issued by the core pipeline) the same page is accessed. Embodiments of the present invention cache address translations in each remote TLB when processing the indirect memory access requests described herein.

[0186] Figure 19FIG. illustrates circuitry / logic associated with each memory endpoint according to some embodiments. Remote translation lookaside buffer (TLB) 1920 and page table walk (PTW) circuitry 1925 are placed at each memory endpoint 1930 to facilitate the indirect memory access techniques described herein. An indirect memory access request may include a load or store operation of the form A[B[i]], which indicates that a first memory access (A[]) is required to read a pointer address from memory, and a second memory access (B[]) is required to access the data identified via the pointer.

[0187] In some embodiments, when one of these operations is initiated by the core 1901 (or logical processor (LP)) pipeline, the initial source and / or destination address is first translated from virtual to physical via the TLB 1904 within the core 1901. The core 1901 then issues a request to the interconnect network 1910, 1420, along with the ASID and PPN information provided by the SATP register of the thread. The privilege level of the hardware thread at the time of instruction execution may also be provided. This results in the drawback that more information is transmitted in each request case, specifically the virtual address (64 bits), ASID (16 bits), and PPN (44 bits), as well as the permission level of the thread (2 bits), approximately an additional 16 bytes of information. However, in some embodiments of the interconnect network 1420, 1910, no additional network microtiles are required to transfer data. In these embodiments, indirect load / store instructions generate a single request, while DMA operations may generate a certain number (N) of requests.

[0188] The above information travels through the interconnect network 1420, 1910 to the memory endpoint where the pointer resides, along with the memory operation. The corresponding indirect unit 1915 receives the request and initiates a read of the local memory 1930 to obtain the pointer. Once the indirect unit receives a response from the memory including the virtual address, it sends the request to the remote TLB 1920 for translation. If there is a TLB hit (i.e., there is an entry that maps the virtual address to a physical address), the TLB 1920 returns the corresponding physical address to the indirect unit 1915, and the indirect unit 1915 transmits the next request (which may be a load or store) to the next destination (i.e., local to this memory 1930, or local to another memory accessible via the interconnect network 1910).

[0189] Once there is a TLB miss, the TLB 1920 sends the PPN and ASID information to the page table walk circuitry 1925, which walks the page table until a valid PTE is found. The PTE provides the physical address in the local memory 1930 or in a remote memory. If no PTE is found, an exception is generated in response to the indirect request.

[0190] These embodiments provide various levels of indirect / pointer traversal. At the final level of indirection, the indirection unit 1915 will send a response to the originating core 1901.

[0191] Another advantage is that the PTW circuitry 1925 (a simple state machine in one embodiment) resides adjacent to the memory controller and the corresponding memory 1930. When the core TLB 1904 or the remote TLB 1920 has a page miss, the page information for the thread (the root of the page table provided in the PPN is the physical address) is sent to the memory endpoint 1930 that stores the page table. In this way, the iteration through the page table levels occurs locally at the memory endpoint that holds the page table. This results in a reduction in page walk time because each iteration of the page walk will occur locally at the memory endpoint rather than being sent back to the TLB (e.g., the core TLB 1910) that initiated the walk.

[0192] In some embodiments, the entries in the remote TLB 1920 will contain SV-57 page table entries, along with additional bits to indicate whether the translated physical address exists locally at this memory endpoint 1930 or remotely at another endpoint. The remote TLB 1920 and the remote PTW 1925 provide optimizations for TLB shootdown support. At the local core 1901, when virtual-to-physical translation occurs, the core TLB 1904 determines whether the target physical address of the memory request exists locally at the core, and if so, caches the translation. However, if the physical address points to the memory of a different core, the TLB 1904 marks this entry as a pointer to another TLB. Any hit to this entry in the local core 1901 will cause the request to be forwarded to the remote core, the translation and memory access will occur at the remote core, and the remote TLB will store the actual translation.

[0193] Each memory 1930 that can be accessed via the interconnect network 1910 is sometimes referred to as a "memory endpoint", and each combination of the memory endpoint 1930, the indirection unit 1915, the remote TLB 1920, and the PTW circuitry 1925 is sometimes referred to herein as a "memory endpoint subsystem". Embodiments of the present invention may include multiple such memory endpoint subsystems interconnected via the interconnect network 1910 to multiple cores and other processor / SoC components (e.g., IO interfaces). Each combination of the indirection unit 1915, the remote TLB 1920, and the corresponding PTW circuitry 1925 is sometimes referred to as a "memory access circuit" or a "local memory access circuit".

[0194] Figure 20The figure illustrates a method according to some embodiments. The method may be implemented on various architectures described herein, but is not limited to any particular processor or system architecture.

[0195] At 2001, the core executes an instruction to generate an indirect memory access request (e.g., an indirect load or store operation). At 2002, the core performs an initial address translation by performing a lookup in its local TLB and, if necessary, a page walk to determine the corresponding physical address of the pointer included in the memory access request.

[0196] Using the physical address, the core identifies the location of the memory endpoint and the corresponding IDU. For example, a table structure may be maintained to map physical address ranges to various memories / IDUs in the system. Once the IDU is identified, the core transmits the request to the IDU via an interconnect network.

[0197] At 2003, the IDU reads the virtual pointer address from its local memory using the physical address included in the request and translates the virtual pointer address to a physical pointer address using its local TLB and / or PTW circuitry. If at 2004 it is determined that the physical pointer address is associated with the memory local to the IDU, then at 2005, the IDU accesses the data from the local memory using the physical pointer address and returns a response including the data to the core.

[0198] At 2004, if the IDU determines that the physical address is associated with a different memory endpoint, then at 2006, it transmits a memory access request including the physical pointer address via the interconnect network to a second IDU corresponding to that memory endpoint. At 2007, the second IDU reads the data from its corresponding memory endpoint and returns a response with the data to the core via the interconnect network.

[0199] Embodiments for scalable TLB misses

[0200] Some embodiments of the present invention reduce the cost of TLB misses by converting TLB misses from a global operation (linear with the number of processors) to a local operation with a constant cost (independent of the number of processors).

[0201] The virtual memory system relies on a translation lookaside buffer (TLB), which is a hardware unit per logical processor (LP) that accelerates virtual-to-physical address translation by caching recent mappings. Whenever a mapping in the page table is removed or modified, it must be flushed from any TLB that may contain a copy, a process known as a "TLB miss".

[0202] TLB misses are a widely recognized scalability limiter because they effectively serialize the threads of a parallel application and their cost is high because they involve inter-processor interrupts. These costs grow proportionally with the parallelism of the application. In one extreme case, TLB misses are global operations that can span an entire parallel machine.

[0203] These embodiments of the invention reduce global operations to local operations, which greatly reduces the cost of TLB misses and maintains the cost of TLB misses rather than allowing it to grow linearly with the size of the parallel machine.

[0204] Specifically, some embodiments prevent the TLB cache from being used for virtual-to-physical mapping of memory pages associated with a memory controller (MC) that is remote from a logical processor or core (e.g., coupled to the LP / core via an interconnect network). Virtual-to-physical translations for "local" pages (e.g., using the TLB management techniques described herein) are cached and TLB accesses are performed in response to translation requests. TLB accesses for "remote" pages are sent to the logical processor / core that is directly coupled to the memory controller associated with the page. In some implementations, the remote TLB access includes a request message that includes a virtual address and a current address space identifier (ASID) (e.g., the ASID associated with the instruction that generated the request). The remote memory endpoint then uses its TLB to perform the virtual-to-physical translation and complete the request. Since each TLB is only allowed to cache virtual-to-physical translations for "local" pages, only a single TLB will be flushed for changes to any given page, resulting in a predictable and constant latency.

[0205] Accordingly, embodiments of the invention employ techniques for both translating virtual addresses generated by an application into physical addresses and caching such translations into the TLB. Figure 21 FIG. illustrates a particular implementation, including a logical processor (LP) or core 2150 that is directly coupled to local memory via a memory controller 2105 and coupled to a remote memory endpoint 2160 via an interconnect fabric 2170 (which may be the interconnect network 1910 described above).

[0206] At 2101, a load operation that indicates a virtual address and optionally a size value is generated by the LP / Core 2150 (e.g., by a load instruction being executed). At 2102, a lookup is performed on the local TLB using the virtual address. If a TLB miss is detected, then at 2103, a page miss handler is run to perform a TLB fill operation (e.g., via access to a page walk circuit system). If a hit is detected, then at 2104, a determination is made as to whether the corresponding virtual address is a local reference or a remote reference. In one embodiment, a local reference is directly translated to a physical address from the corresponding TLB entry or via a page walk of the local page table through the memory controller 2105, and the translation is then cached in the local TLB.

[0207] In these embodiments, a remote reference is immediately forwarded to the remote memory endpoint 2160 through the network fabric 2170, but instead of a physical address, the target memory of the operation is described using a combination of a virtual address and an address space identifier. For example, on an x86 architecture, this includes a pair (VA, CR3), where CR3 refers to a control register that holds the current ASID value, which is used to form the base address for the corresponding page table.

[0208] In these embodiments, the translation associated with a remote access is also cached in the local TLB as "remote". For example, the corresponding TLB entry stores an indication of the memory endpoint 2160 where the translation is stored, rather than storing the virtual-to-physical address translation. Thus, any further reference to the same page will be immediately forwarded to the remote memory endpoint 2160.

[0209] As indicated at 2115, if a load request received by the memory endpoint 2160 causes a TLB hit (e.g., the local TLB includes an entry corresponding to the VA and ASID values), then the corresponding physical address is read from the TLB and used to access the data 2120 via the memory controller 2112, and the data 2120 is returned to the LP / Core 215 via the network fabric 2170. If the load request causes a TLB miss, then at 2118, a page miss handler is initiated to perform a page walk operation, and the resulting translation is filled into the TLB at the memory endpoint 2160. The translated physical address is then used to read the data 2120 via the memory controller 2112.

[0210] The basic idea of these embodiments is to distinguish between references to local memory (i.e., page frames accommodated when referencing the integrated memory controller 2105 of the processor) and references to remote memory (i.e., page frames accommodated in the integrated memory controllers of other processors, e.g., memory controller 2112). When such untranslated references are received, any memory controller 2112 in the system performs the translation and caches it in their local TLB before responding with the data.

[0211] In this way, the underlying virtual-to-physical translation is cached only in the TLBs of the processors whose integrated memory controllers store the corresponding page frames. Thus, when such address translations need to be modified, only a small, fixed number of TLBs (the TLBs close to the memory controller storing the affected page frame) need to participate in TLB invalidation, rather than all the TLBs of all the processors that may have executed threads in that address space across the entire platform.

[0212] This transforms the cost of the TLB invalidation process from O(n) (i.e., linear with the system size) to O(1) (i.e., constant cost), thus effectively eliminating TLB invalidation as a limiting factor for scaling parallel global address space applications on such platforms, such as applications on top of threading, OpenMP, OpenCL, etc. or applications built using threading, OpenMP, OpenCL, etc. The modified address translation process performs exactly the same number of structural messages to perform any memory operation, so it will not cause congestion in the interconnect and will also not add any additional message-related latency.

[0213] In some implementations, additional page walks may be required for remote access, once at the initiator side and once at the storage side, but these will not be very frequent because their results will be cached in the corresponding TLBs and then amortized over all subsequent accesses to the same page frame. In the very common case where the TLBs at both ends have cached the mapping, the embodiments of the present invention only add the minimal cost of an additional TLB hit lookup, which is typically one of the most highly optimized paths in modern processors.

[0214] Example

[0215] The following are example implementations of different embodiments of the present invention.

[0216] Example 1. A processor includes: an interconnect network; a plurality of cores coupled to the interconnect network, wherein a requester core among the plurality of cores includes a core translation lookaside buffer (TLB) and core page table walk (PTW) circuitry; a plurality of memory endpoint subsystems coupled to the interconnect network, each memory endpoint subsystem including memory access circuitry and a memory for access via the interconnect network, each memory access circuitry including a memory endpoint TLB and a memory endpoint PTW circuitry; wherein one or more of the memory endpoint subsystems are configured to perform operations on behalf of the requester core to process an indirect memory access request, and the one or more of the memory endpoint subsystems are configured to: read a virtual pointer address based on the indirect memory access request, translate the virtual pointer address into a physical pointer address, access data based on the physical pointer address, and transmit an indirect memory access response including the data to the requester core.

[0217] Example 2. The processor of Example 1, wherein one or more of the memory endpoint subsystems include: a first memory endpoint subsystem including first memory access circuitry and a corresponding first memory, the first memory endpoint subsystem being configured to receive an indirect memory access request from the requester core via the interconnect network, and the first memory access circuitry being configured to read a virtual pointer address from the first memory at a physical address included in the indirect memory access request.

[0218] Example 3. The processor of Example 1 or 2, wherein the first memory access circuitry includes a first memory endpoint TLB and a first memory endpoint PTW circuitry, and the first memory access circuitry is configured to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuitry to translate the virtual pointer address into a physical pointer address.

[0219] Example 4. The processor of any one of Examples 1-3, wherein one or more of the memory endpoint subsystems further include: a second memory endpoint subsystem including second memory access circuitry and a corresponding second memory, wherein the first memory access circuitry is configured to determine that the physical pointer address is associated with the second memory, and the first memory access circuitry is configured to transmit a request including the physical pointer address to the second memory endpoint subsystem via the interconnect network.

[0220] Example 5. The processor of any one of Examples 1-4, wherein the second memory access circuitry is configured to read data from the second memory based on the physical pointer address, and the second memory endpoint subsystem is configured to transmit an indirect memory access response including the data to the requester core via the interconnect network.

[0221] Example 6. A processor as in any one of Examples 1-5, wherein the indirect memory access request is for transmission by a requesting core to a first memory endpoint subsystem via an interconnect network, and the indirect memory access request includes one or more first fields to indicate an address space identifier (ASID) value, a physical page number, and a permission level associated with a corresponding thread.

[0222] Example 7. A processor as in any one of Examples 1-6, wherein the request transmitted by the first memory access circuit includes one or more second fields to indicate an ASID value, a physical pointer address, and a permission level.

[0223] Example 8. A processor as in any one of Examples 1-7, wherein, to translate a virtual pointer address to a physical pointer address, the first memory access circuit is configured to perform a lookup in a first memory endpoint TLB, and in response to a TLB miss, the first memory access circuit is configured to provide the virtual pointer address and the address space identifier ASID value included in the indirect memory access request to a PTW circuitry, which is configured to perform a page walk operation using a page table stored in the first memory, the virtual pointer address, and the ASID value.

[0224] Example 9. A processor as in any one of Examples 1-8, wherein the requesting core is configured to execute one or more instructions to generate an indirect memory access request, the requesting core is configured to access at least one of a core TLB and a core PTW circuitry to translate a corresponding virtual address to a corresponding physical address, and the requesting core is configured to transmit the indirect memory access request with the corresponding physical location via the interconnect network, wherein the first memory endpoint subsystem among one or more memory endpoint subsystems is configured to read a virtual pointer address using the corresponding physical address.

[0225] Example 10. A method, comprising: executing one or more instructions on a requesting core among a plurality of cores to perform a first address translation to translate a corresponding virtual address to a corresponding physical address, the requesting core being configured to transmit an indirect memory access request to a first memory endpoint subsystem via an interconnect network; reading, by the first memory endpoint subsystem, a virtual pointer address based on the corresponding physical address; translating, at the first memory endpoint subsystem, the virtual pointer address to a physical pointer address; accessing, at a second memory endpoint subsystem, data based on the physical pointer address; and transmitting, from the second memory endpoint subsystem to the requesting core, an indirect memory access response including the data.

[0226] Example 11. The method of Example 10, wherein the first memory endpoint subsystem includes a first memory access circuit and a corresponding first memory, and the first memory access circuit is configured to read a virtual pointer address from the first memory at a physical address included in an indirect memory access request.

[0227] Example 12. The method of Example 10 or 11, wherein the first memory access circuit includes a first memory endpoint TLB and a first memory endpoint PTW circuitry, and the first memory access circuit is configured to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuitry to translate the virtual pointer address into a physical pointer address.

[0228] Example 13. The method of any one of Examples 10-12, further comprising: determining at the first memory access circuit that the physical pointer address is associated with a second memory endpoint subsystem; and transmitting the request from the first memory endpoint subsystem to the second memory endpoint subsystem, the request including the physical pointer address.

[0229] Example 14. The method of any one of Examples 10-13, wherein the second memory endpoint subsystem includes a second memory access circuit and a second memory, and the second memory access circuit is configured to read data from the second memory based on the physical pointer address.

[0230] Example 15. The method of any one of Examples 10-14, wherein the indirect memory access request is for transmission by a requesting core through an interconnect network to the first memory endpoint subsystem, and the indirect memory access request includes one or more first fields to indicate an address space identifier (ASID) value, a physical page number, and a permission level associated with a corresponding thread.

[0231] Example 16. The method of any one of Examples 10-15, wherein the request transmitted by the first memory endpoint subsystem includes one or more second fields to indicate the ASID value, the physical pointer address, and the permission level.

[0232] Example 17. The method of any one of Examples 10-16, wherein, to translate the virtual pointer address into a physical pointer address, the first memory access circuit is configured to: perform a lookup in the first memory endpoint TLB, and in response to a TLB miss, the first memory access circuit is configured to provide the virtual pointer address and an address space identifier (ASID value) included in the indirect memory access request to the PTW circuitry, and the PTW circuitry is configured to perform a page walk operation using a page table stored in the first memory, using the virtual pointer address and the ASID value.

[0233] Example 18. A machine-readable medium having program code stored thereon that, when executed by a machine, causes the machine to perform operations including: executing one or more instructions on a requesting core of a plurality of cores to perform a first address translation to translate a corresponding virtual address to a corresponding physical address, the requesting core being configured to transmit an indirect memory access request to a first memory endpoint subsystem via an interconnect network; reading, by the first memory endpoint subsystem, a virtual pointer address based on the corresponding physical address; translating, at the first memory endpoint subsystem, the virtual pointer address to a physical pointer address; accessing, at a second memory endpoint subsystem, data based on the physical pointer address; and transmitting an indirect memory access response including the data from the second memory endpoint subsystem to the requesting core.

[0234] Example 19. The machine-readable medium of Example 18, wherein the first memory endpoint subsystem includes a first memory access circuit and a corresponding first memory, the first memory access circuit being configured to read a virtual pointer address from the first memory at the physical address included in the indirect memory access request.

[0235] Example 20. The machine-readable medium of any one of Examples 18-19, wherein the first memory access circuit includes a first memory endpoint TLB and a first memory endpoint PTW circuitry, the first memory access circuit being configured to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuitry to translate the virtual pointer address to a physical pointer address.

[0236] Example 21. The machine-readable medium of any one of Examples 18-20, further comprising: determining, at the first memory access circuit, that the physical pointer address is associated with the second memory endpoint subsystem; and transmitting a request from the first memory endpoint subsystem to the second memory endpoint subsystem, the request including the physical pointer address.

[0237] Example 22. The machine-readable medium of any one of Examples 18-21, wherein the second memory endpoint subsystem includes a second memory access circuit and a second memory, the second memory access circuit being configured to read data from the second memory based on the physical pointer address.

[0238] Example 23. The machine-readable medium of any one of Examples 18-22, wherein the indirect memory access request is configured to be transmitted by the requesting core to the first memory endpoint subsystem via an interconnect network, the indirect memory access request including one or more first fields to indicate an address space identifier (ASID) value, a physical page number, and a permission level associated with a corresponding thread.

[0239] Example 24. A machine-readable medium as in any one of Examples 18 - 23, wherein the request transmitted by the first memory endpoint subsystem includes a second one or more fields to indicate an ASID value, a physical pointer address, and a permission level.

[0240] Example 25. A machine-readable medium as in Example 20, wherein, to translate a virtual pointer address into a physical pointer address, the first memory access circuit is configured to: perform a lookup in a first memory endpoint TLB, and in response to a TLB miss, the first memory access circuit is configured to provide the virtual pointer address and an address space identifier (ASID value) included in an indirect memory access request to a PTW circuitry that is configured to perform a page walk operation using a page table stored in the first memory, the virtual pointer address, and the ASID value.

[0241] Embodiments of the present invention may include the various steps described above. These steps may be embodied as machine-executable instructions that may be used to cause a general or special-purpose processor to execute these steps. Alternatively, these steps may be performed by specific hardware components that include hardwired logic for performing these steps, or by any combination of programmed computer components and custom hardware components.

[0242] As described herein, an instruction can refer to a specific configuration of hardware such as an application specific integrated circuit (ASIC) that is configured to perform certain operations or has a predetermined function or software instructions stored in a memory embodied in a non-transitory computer-readable medium. Thus, the techniques shown in the figures can be implemented using code and data stored and executed on one or more electronic devices (e.g., end stations, network elements, etc.). Such electronic devices use computer machine-readable media (internally and / or via a network with other electronic devices) to store and transmit the code and data, the computer machine-readable media such as non-transitory computer machine-readable storage media (e.g., magnetic disks; optical disks; random access memory; read-only memory; flash devices; phase change memory) and transitory computer machine-readable communication media (e.g., electrical, optical, acoustic, or other forms of propagated signals - such as, carrier waves, infrared signals, digital signals, etc.). Additionally, such electronic devices typically include a collection of one or more processors coupled to one or more other components such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., keyboards, touchscreens, and / or displays), and network connections. The coupling of the collection of processors to the other components is typically through one or more buses and bridges (also referred to as bus controllers). The storage devices and signals carrying network traffic represent one or more machine-readable storage media and machine-readable communication media, respectively. Thus, the storage devices of a given electronic device typically store code and / or data for execution on the collection of one or more processors of that electronic device. Of course, different combinations of software, firmware, and / or hardware can be used to implement one or more portions of the embodiments of the present invention. Throughout this detailed description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention can be practiced without some of these specific details. In some instances, well-known structures and functions are not described in detail so as not to obscure the subject matter of the present invention. Accordingly, the scope and spirit of the present invention should be determined according to the appended claims.

Claims

1. A processor, comprising: Interconnection network; a plurality of cores coupled to the interconnection network, wherein a requesting core among the plurality of cores comprises a core translation lookaside buffer TLB and a core page table walk PTW circuit system; a plurality of memory endpoint subsystems coupled to the interconnect network, each memory endpoint subsystem comprising memory access circuitry and memory for access via the interconnect network, each memory access circuitry comprising a memory endpoint TLB and a memory endpoint PTW circuitry; Wherein, one or more of the memory endpoint subsystems are used to perform operations on behalf of the requester core to process an indirect memory access request, and one or more of the memory endpoint subsystems are used to: reading a virtual pointer address based on the indirect memory access request, translating the virtual pointer address into a physical pointer address, and accessing data based on the physical pointer address, and An indirect memory access response including the data is transmitted to the requestor core.

2. The processor of claim 1, wherein: One or more of the memory endpoint subsystems include: A first memory endpoint subsystem, the first memory endpoint subsystem comprising a first memory access circuit and a corresponding first memory, the first memory endpoint subsystem being configured to receive the indirect memory access request from the requestor core through the interconnection network, and the first memory access circuit being configured to read the virtual pointer address from the first memory at a physical address included in the indirect memory access request.

3. The processor of claim 2, wherein: The first memory access circuit includes a first memory endpoint TLB and a first memory endpoint PTW circuit system, and the first memory access circuit is used to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuit system to translate the virtual pointer address into the physical pointer address.

4. The processor of claim 3, wherein: One or more of the memory endpoint subsystems further comprises: A second memory endpoint subsystem, the second memory endpoint subsystem includes a second memory access circuit and a corresponding second memory, wherein the first memory access circuit is used to determine that the physical pointer address is associated with the second memory, and the first memory access circuit is used to transmit a request including the physical pointer address to the second memory endpoint subsystem through the interconnection network.

5. The processor of claim 4, wherein: The second memory access circuit is used to read the data from the second memory based on the physical pointer address, and the second memory endpoint subsystem is used to transmit the indirect memory access response including the data to the requester core through the interconnection network.

6. A processor as claimed in claim 4 or 5, wherein: The indirect memory access request is used to be transmitted by the requesting core to the first memory endpoint subsystem through the interconnection network, and the indirect memory access request includes one or more first fields to indicate an address space identifier ASID value, a physical page number, and a permission level associated with a corresponding thread.

7. The processor of claim 6, wherein: The request transmitted by the first memory access circuit includes a second one or more fields to indicate the ASID value, the physical pointer address, and the permission level.

8. A processor as claimed in any one of claims 3 to 7, wherein: In order to translate the virtual pointer address into the physical pointer address, the first memory access circuit is used to perform a search in the first memory endpoint TLB, and in response to a TLB miss, the first memory access circuit is used to provide the virtual pointer address and the address space identifier ASID value included in the indirect memory access request to the PTW circuit system, and the PTW circuit system is used to use the page table stored in the first memory, the virtual pointer address and the ASID value to perform a page walk operation.

9. A processor as claimed in any one of claims 1 to 8, wherein: The requester core is used to execute one or more instructions to generate the indirect memory access request, the requester core is used to access at least one of the core TLB and the core PTW circuit system to translate the corresponding virtual address into the corresponding physical address, and the requester core is used to transmit the indirect memory access request with the corresponding physical location through the interconnection network, wherein a first memory endpoint subsystem of one or more memory endpoint subsystems is used to read the virtual pointer address using the corresponding physical address.

10. A method comprising: executing one or more instructions on a requestor core among the plurality of cores to perform a first address translation to translate a corresponding virtual address into a corresponding physical address, the requestor core being used to transmit an indirect memory access request to a first memory endpoint subsystem through an interconnect network; The first memory endpoint subsystem reads a virtual pointer address based on a corresponding physical address, translating the virtual pointer address into a physical pointer address at the first memory endpoint subsystem, accessing data at a second memory endpoint subsystem based on the physical pointer address, and An indirect memory access response including the data is transmitted from the second memory endpoint subsystem to the requestor core.

11. The method of claim 10, wherein: The first memory endpoint subsystem includes a first memory access circuit and a corresponding first memory, the first memory access circuit being configured to read the virtual pointer address from the first memory at a physical address included in the indirect memory access request.

12. The method of claim 11, wherein: The first memory access circuit includes a first memory endpoint TLB and a first memory endpoint PTW circuit system, and the first memory access circuit is used to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuit system to translate the virtual pointer address into the physical pointer address.

13. The method of claim 12, further comprising: determining, at the first memory access circuit, that the physical pointer address is associated with the second memory endpoint subsystem; as well as A request is transmitted from the first memory endpoint subsystem to the second memory endpoint subsystem, the request including the physical pointer address.

14. The method of claim 13, wherein: The second memory endpoint subsystem includes a second memory access circuit and a second memory, wherein the second memory access circuit is configured to read the data from the second memory based on the physical pointer address.

15. The method according to claim 13 or 14, wherein: The indirect memory access request is used to be transmitted by the requesting core to the first memory endpoint subsystem through the interconnection network, and the indirect memory access request includes one or more first fields to indicate an address space identifier ASID value, a physical page number, and a permission level associated with a corresponding thread.

16. The method of claim 15, wherein: The request transmitted by the first memory endpoint subsystem includes a second one or more fields to indicate the ASID value, the physical pointer address, and the permission level.

17. The method according to any one of claims 12 to 16, wherein: In order to translate the virtual pointer address into the physical pointer address, the first memory access circuit is used to: performing a lookup in the first memory endpoint TLB, and In response to a TLB miss, the first memory access circuit is configured to provide the virtual pointer address and an address space identifier ASID value included in the indirect memory access request to the PTW circuit system, The PTW circuit system is used to perform a page walk operation using the page table stored in the first memory, the virtual pointer address and the ASID value.

18. A machine-readable medium having program code stored thereon, wherein the program code, when executed by a machine, causes the machine to perform operations, the operations comprising: executing one or more instructions on a requestor core among the plurality of cores to perform a first address translation to translate a corresponding virtual address into a corresponding physical address, the requestor core being used to transmit an indirect memory access request to a first memory endpoint subsystem through an interconnect network; The first memory endpoint subsystem reads a virtual pointer address based on a corresponding physical address; translating the virtual pointer address into a physical pointer address at the first memory endpoint subsystem; accessing data based on the physical pointer address at a second memory endpoint subsystem; as well as An indirect memory access response including the data is transmitted from the second memory endpoint subsystem to the requestor core.

19. The machine-readable medium of claim 18, wherein: The first memory endpoint subsystem includes a first memory access circuit and a corresponding first memory, the first memory access circuit being configured to read the virtual pointer address from the first memory at a physical address included in the indirect memory access request.

20. The machine-readable medium of claim 19, wherein: The first memory access circuit includes a first memory endpoint TLB and a first memory endpoint PTW circuit system, and the first memory access circuit is used to access at least one of the first memory endpoint TLB and the first memory endpoint PTW circuit system to translate the virtual pointer address into the physical pointer address.