Circuits and methods for linear memory access control table switching for fine-grained partitioning

By switching the linear memory tag table in the processor, fine-grained memory management is achieved, which solves the security and efficiency problems of partition isolation in existing technologies and is suitable for hosting microservices and function-as-a-service applications.

CN120704737APending Publication Date: 2025-09-26INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510325950.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-26
Filing Date
2025-03-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing partition isolation technologies lack security, have functional limitations, are non-scalable, and have low memory efficiency in processors. In particular, they are unable to achieve fine-grained partition isolation when managing multiple sub-processes.

Method used

Memory tagging technology is adopted to achieve fine-grained partition management within the virtual address space of a single process by switching the linear memory tag table. The switching mechanism of the memory tag table and the page table is utilized to provide 16-byte granularity memory management, which is independent of the page table and TLB structure and realizes fine-grained access control to the memory.

Benefits of technology

It enables efficient management of multiple partitions within a single process, improves security and efficiency, and is suitable for hosting microservices and function-as-a-service applications, providing a more secure platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120704737A_ABST
    Figure CN120704737A_ABST
Patent Text Reader

Abstract

Circuits and methods for implementing one or more handover subprocess instructions are described. In some examples, a hardware processor (e.g., a core) includes (e.g., a coupling to) a memory management circuit to control memory access based on a memory tag stored in a memory tag data structure and a memory tag based on a pointer to the memory; decoder circuitry to decode an instruction into a decoded instruction, the instruction comprising an operand to identify a memory tag data structure for a sub-process of a plurality of memory tag data structures for a corresponding sub-process of a process and an opcode, the operand to identify a memory tag data structure for the sub-process of the process. The opcode is used for indicating the execution circuitry to switch from another memory tag data structure for another sub-process of the process to a memory tag data structure for the sub-process; and execution circuitry to execute the decoded instruction according to the opcode. The memory tag data structure may be resumed for providing access control permissions for sub-processes per memory particle.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A processor or set of processors executes instructions from an instruction set (e.g., an instruction set architecture (ISA)). An instruction set is the programming-related portion of a computer architecture and generally includes native data types, instructions, register architecture, addressing modes, memory architecture, interrupt and exception handling, and external input and output (I / O). It should be noted that the term instruction herein may refer to either macroinstructions, e.g., instructions provided to a processor for execution, or microinstructions, e.g., instructions decoded from macroinstructions by a processor's decoder. Memory tags provide a mechanism for verifying the correctness of load and / or store operations performed by processor instructions. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] Various examples according to the present disclosure will be described with reference to the accompanying drawings, in which:

[0003] Figure 1 A block diagram of a computer system according to an example of the present disclosure is illustrated, the computer system including: a system memory for storing a memory tag data structure (e.g., a memory tag table (MTT)); and a processor for executing one or more instructions to switch to a memory tag data structure for a child process from among a plurality of memory tag data structures for corresponding child processes of a process.

[0004] Figure 2 An example format of a pointer with a tag (eg, a memory corruption detection (MCD) tag) is illustrated according to examples of the present disclosure.

[0005] Figure 3 Illustrate a dual paging structure and process for changing a memory tag table mapping according to examples of the present disclosure.

[0006] Figure 4 is a flowchart illustrating the operation of a method (e.g., SWITCHSP) for switching a memory tag data structure (e.g., a memory tag table) for a process based on input of an index and a linear address (LA) according to an example of the present disclosure.

[0007] Figure 5 is a flow diagram illustrating the operation of a method (eg, SWITCHSP) for switching a memory tag data structure (eg, a memory tag table) for a process based on input of a linear range and a linear address (LA) according to examples of the present disclosure.

[0008] Figure 6 Illustrated is an example of computing hardware for processing SWITCHSP instruction(s) according to examples of the present disclosure.

[0009] Figure 7 Illustrated is an example method executed by a processor to process a SWITCHSP instruction according to examples of the present disclosure.

[0010] Figure 8 Illustrated is an example method for processing a SWITCHSP instruction using emulation or binary translation according to examples of the present disclosure.

[0011] Figure 9 Illustrated is linear address translation to 4K byte pages using 4-level paging according to some examples.

[0012] Figure 10 An example computing system is illustrated.

[0013] Figure 11 A block diagram illustrates an example processor and / or System on a Chip (SoC) that may have one or more cores and an integrated memory controller.

[0014] Figure 12 is a block diagram illustrating a computing system 1200 configured to implement one or more aspects of the examples described herein.

[0015] Figure 13A Illustrate an example of a parallel processor.

[0016] Figure 13B An example of a block diagram illustrating a partition unit is shown.

[0017] Figure 13C An example of a block diagram illustrating a processing cluster within a parallel processing unit.

[0018] Figure 13D Illustrated is an example of a graphics multiprocessor coupled with a pipeline manager of a processing cluster.

[0019] Figures 14A-14C An additional graphics multiprocessor is illustrated according to an example.

[0020] Figure 15 A parallel computing system 1500 is shown according to some examples.

[0021] Figures 16A-16B Illustrate a hybrid logical / physical view of a split parallel processor according to examples described herein.

[0022] Figure 17Ais a block diagram illustrating both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to an example.

[0023] Figure 17B is a block diagram illustrating both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to an example.

[0024] Figure 18 Illustrated are examples of execution unit circuitry, such as execution unit circuitry.

[0025] Figure 19 is a block diagram of a register architecture according to some examples.

[0026] Figure 20 An example of the instruction format is shown.

[0027] Figure 21 Figure 1 shows an example of an addressing information field.

[0028] Figure 22 An example of the first prefix is ​​shown in FIG.

[0029] Figures 23A-23D Illustrated is an example of how the R, X, and B fields of the first prefix are used.

[0030] Figures 24A-24B An example of the second prefix is ​​shown in FIG.

[0031] Figure 25 An example of the third prefix is ​​shown in FIG.

[0032] Figures 26A-26B Illustrated is thread execution logic including an array of processing elements employed in a graphics processor core according to examples described herein.

[0033] Figure 27 An additional execution unit is illustrated according to an example.

[0034] Figure 28 is a block diagram illustrating a graphics processor instruction format 2800 according to some examples.

[0035] Figure 29 is a block diagram of another example of a graphics processor.

[0036] Figure 30A is a block diagram illustrating a graphics processor command format according to some examples.

[0037] Figure 30B is a block diagram illustrating a graphics processor command queue according to an example.

[0038] Figure 31is a block diagram illustrating converting binary instructions in a source ISA to binary instructions in a target ISA using a software instruction converter according to an example.

[0039] Figure 32 is a block diagram illustrating an IP core development system 3200 that may be used to fabricate integrated circuits to perform operations, according to some examples. DETAILED DESCRIPTION

[0040] The present disclosure relates to methods, apparatus, systems, and non-transitory computer-readable storage media for switching subprocesses (SWITCHSP) (e.g., one or more SWITCHSP instructions). In some examples, memory tagging provides vulnerability detection at very fine memory granularity (e.g., 16 bytes wide) by matching the tag of a pointer with the corresponding tag of the memory particle pointed to by the pointer (e.g., a 16-byte naturally aligned sub-cache line memory region or cache line). In some examples, partitioning provides partition isolation via memory regions such as provided by segmentation, page tables, and / or virtual machine separation. Today, partitioning and memory tagging are considered different technologies and are thus implemented separately, for example, via memory paging and memory tag tables, respectively. The examples herein relate to novel methods of switching linear memory tag data structures (e.g., memory tagging tables (MTTs)) to allow the runtime to efficiently manage infinite, sparse, and / or overlapping partitions of per-object (e.g., with 16-byte aligned boundaries) granularity within the virtual address space of a single process.

[0041] As an example, a process implements a web browser for a web page that runs a first child process for a cryptographic script (and thus, is of higher importance to maintaining security) and a second child process for a secondary content script, e.g., for displaying secondary content to the user and thus, of lower importance to maintaining security. Therefore, it may be desirable to use memory tagging to provide protection for the first child process from the second child process, etc. Partitions can be created using managed runtimes, segmentation, process separation, or virtual machines. However, these approaches lack security, limit functionality, lack scalability, do not partition traditional monolithic programs, and are very inefficient for child processes (e.g., small functions used by services / microservices). Furthermore, some partitioning solutions rely on coarse-grained paging (e.g., page granularity) or memory region-based mechanisms (e.g., segmentation or hardware-assisted fault isolation (HFI)) to provide separation, which is very memory inefficient and does not align with the software-centric object granularity model.

[0042] To overcome these problems, certain examples herein relate to memory tagging techniques that utilize a child process page table switching mechanism to provide a runtime capability for efficiently creating multiple (e.g., an unlimited number) partitions to individual function and object granularities within a single sparse process address space. Examples herein allow runtime execution of a switch of the active memory tag table within a user process. This can be accomplished by switching the current linear memory tag table location (e.g., via a range register that identifies the location of the currently active table), or by switching a secondary page table from user space (e.g., via a control register (e.g., CR3)), which changes the physical memory mapping for only the memory tag table (e.g., the memory tag table includes a special tag value, e.g., indicating an inaccessible or read-only granularity). In this manner, memory is managed at a (e.g., 16-byte) memory tag granularity, while the page table(s) and translation lookaside buffer (TLB) maintain page mappings for a single process, e.g., only changing entries in the memory tag table portion of the linear address space used for the child process switch. The examples in this article provide a sub-process page table switching mechanism to enable improved partitioning, providing an optimal platform for hosting microservices and function-as-a-service (FaaS) applications, providing site isolation for browsers, and providing the most secure platform. In certain examples, by leveraging memory tagging, the runtime can manage individual objects in memory that are sparsely populated down to (e.g., 16-byte) memory tag granularity, which is not possible with other traditional partitioning techniques.

[0043] The instructions disclosed herein are improvements to the operation of a processor (e.g., of a computer) itself, as these instructions implement the above functions by electrically altering a general-purpose computer (e.g., decoder circuitry and / or execution circuitry of a general-purpose computer) by creating electrical pathways within the computer (e.g., within the decoder circuitry and / or execution circuitry of the computer). These electrical pathways create a special-purpose machine for performing specific functions.

[0044] The instructions disclosed herein are improvements to the operation of a processor (e.g., of a computer) itself. Instruction decoding circuitry (e.g., decoder circuitry 106) that does not have such instructions as part of its instruction set will not decode as discussed herein. Execution circuitry (e.g., execution circuitry 106) that does not have such instructions as part of its instruction set will not execute as discussed herein. For example, SWITCHSP instructions according to the present disclosure are improvements to the operation of a processor (e.g., of a computer) itself because they provide a sub-process page table switching mechanism, for example, to switch the active memory tag table within a user process by switching a secondary page table from user space, which changes the physical memory mapping for only the memory tag table.

[0045] Now turning to the accompanying drawings, Figure 1 A block diagram of a computer system 100 according to an example of the present disclosure is shown, the computer system 100 including: a system memory 101 (e.g., dynamic random access memory (DRAM)) for storing a memory tag data structure 110 (e.g., a memory tag table (MTT)); and a processor 102 for executing one or more instructions to switch to a memory tag data structure 110 for a child process from among a plurality of memory tag data structures for corresponding child processes of a (e.g., a single) process.

[0046] As described above, some computer systems provide partition isolation via memory regions (e.g., of system memory 101) such as provided by segmentation, page tables, and / or virtual machine separation. The following disclosure relates to "page" data structures (e.g., page tables), but it should be understood that this can be extended to other partition isolation techniques (e.g., segmentation and / or virtual machine separation). In some examples, the operating system (e.g., operating system code 103) uses address translation support called paging. In some examples, paging utilizes multiple page data structures (e.g., tables) 108. In some examples, paging utilizes these page data structures (e.g., tables) 108 to translate linear addresses (e.g., virtual addresses) used by software into corresponding physical addresses, which are used to access memory (or memory-mapped input / output (I / O) devices). In some examples, the width of the linear address is 48 bits or 57 bits. Figure 9 An example 4-level paging is depicted, e.g., a 4-level hierarchy of page data structures (e.g., tables) 108, the root of which resides at a physical address in a control register (e.g., CR3). In some examples, a process / partition ID register 122-CID (e.g., CR3) enables the processor 102 to translate a linear address into a physical address by locating the page directory and page table for the current code (e.g., task). In some examples, the high-order bits (e.g., 20 bits) of CR3 become the page directory base register (PDBR), which stores the physical address of the first page directory, and (e.g., if the PCIDE bit in CR4 is set), the low-order bits (e.g., 12 bits) are used for the process-context identifier (PCID).

[0047] Some processors support instructions (e.g., VMFUNC) that allow user-space programs to switch underlying guest page table structures from a list (e.g., an extended page table pointer (EPTP) list) pre-approved by privileged software (e.g., a virtual machine monitor (VMM) and / or an OS (e.g., a kernel). Some processors include a mechanism for performing user-space page table switches by utilizing the root page table pointed to (e.g., by CR3), while adding additional page tables to override permissions for switching child processes to the parent process. Some processors provide a secondary paging structure for kernel linear regions.

[0048] In addition to or as an alternative to these, examples herein relate to a processor that allows switching of page table entries corresponding only to the linear address region occupied by a memory tag table (MTT). In some examples, the page table structure memory mapping and page permissions remain unchanged for the process, and the child process is defined only by an additional paging structure below the memory tag table 110, e.g., only the memory tag table 100 is switched when switching partitions using the new SWITCHSP instruction(s) disclosed herein (e.g., such a switch only affects the memory tag table). In some examples, because the memory tag and tag cache are independent of the page table and TLB structures, the memory tag table 110 effectively overrides the page-granularity permissions of the TLB entries (e.g., is more restrictive than the page-granularity permissions of the TLB entries). In some examples, the processor includes (e.g., separate from any TLB) an object lookaside buffer (OLB) 111 (e.g., a tag cache) for caching memory tags for linear address objects while also providing fine-grained (e.g., 16-byte) access control to each data granule in memory (no matter how sparsely populated). Some examples herein allow a guest to switch its current tag table or override the tag table, for example, depending on which partition is running.

[0049] In some examples, the processor 102 (e.g., the memory management circuitry 110 of the processor 102) compares a memory tag for a particular portion (e.g., a cache line) of a memory (e.g., memory 101) with a corresponding tag of the memory pointed to by the pointer (e.g., a cache line), and, for example, allows memory access for a match and / or denies memory access for a mismatch.

[0050] Some examples split the tag value into a special value with no access, and / or a special tag value of read-only or read / write or execute. In some examples, these tags specify the permissions for each particle of memory, for example, so that the correct tag value can be applied when views and tables are overwritten. In some examples, for each (e.g., 16-byte) particle of memory, the (e.g., 4-bit) memory tag checked from the memory tag table allows the following to be specified: 4-bit label value license describe 0 No access Preventing access to granules 1 Read-only Allow read-only access 2 Read / Write Allow read / write access 3 implement Execution Permission 4 Pointer Label Matching relative to pointer labels 5 Pointer Label Matching relative to pointer labels 6 Pointer Label Matching relative to pointer labels … … … 15 Pointer Label Matching relative to pointer labels

[0051] Alternatively, the tag space can be split so that there is a permission bit and multiple (e.g., three) memory tag bits that match the pointer tag: 1-bit label value license describe 0 No access Preventing access to granules 1 Read / Write Allow read-only access

[0052] In some examples, increasing the tag value size allows for more permission mappings without affecting memory tags (e.g., each memory particle has an 8-bit tag). Similarly, a smaller tag size may be suitable only for access control (e.g., as shown in the table above, a 1-bit tag for access / no access, where there is only 1 tag bit for each memory particle), resulting in a smaller tag table size. Even with such limited tag size options, the underlying page table permissions can serve as a precedent, for example, allowing read access to a page only when the MTT permission allows access to the memory particle. Alternatively, examples with larger tag sizes can override the underlying page table permissions by providing more specific permissions for the memory particles.

[0053] In some examples herein, there are actually two control registers for paging (e.g., two "CR3s" in x86 parlance). In some examples, the first control register is the process / partition ID register 122-CID (e.g., CR3) which is the parent process control register (e.g., parent process CR3). In some examples, the process / partition ID register 122-CID points to a page data structure 108 containing (e.g., all) virtual-to-physical page mappings for the process. In some examples, the second control register is a child process control register, such as a memory tag data structure list (e.g., memory tag table list (MTTL)) index register 122-INDX. In some examples, the MTTL index register 122-INDX points to an element that, in turn, points to a memory tag data structure (e.g., memory tag table (MTT)) 112 that stores the virtual-to-physical page mapping for the corresponding memory tag table entry. In some examples, the child process control register (e.g., MTTL index register 122-INDX) controls only the page mapping for the linear range of the memory tag table. In some examples, the child process control register allows switching (e.g., from user space, such as via a SWITCHSP instruction) to a new memory tag table mapping, such as while retaining all other page mappings (e.g., and TLB entries) for the process from the parent process control register (e.g., root CR3).

[0054] In some examples, privileged software (e.g., an OS) populates a memory tag table list (MTTL) 112 that describes page table locations that specify alternate memory tag table mappings (e.g., for different child processes of the same process). In some examples, the physical addresses for the pages of the MTT table 112 are specified in a register accessible only to privileged software (e.g., an OS). In some examples, the register is an MTTL register 122-MTTL, which, for example, stores a value pointing to the physical address of the MTTL 112.

[0055] In some examples, a user space process (e.g., executing from user code 105) can then select between these authorized memory tag table 110 mappings by specifying the desired index in a list of memory tag data structures (e.g., a memory tag table list (MTTL)) index register 122-INDX to which the process is to switch at runtime (e.g., as specified in a SWITCHSP instruction). In some examples, one or more of the TLB entries for the child process can be tagged (e.g., with a tag other than the memory type) so that these entries can be reused when returning to the original child process and / or partition. Other examples flush all TLB entries corresponding to the linear range of the memory tag table on a child process and / or partition switch. In some examples, a unique page mapping that changes is used to switch between memory tag tables, e.g., all of which have the same overlapping contiguous linear range but have (potentially) different physical page mappings so that the memory tags can be child process / partition specific. Returning to the web browser example above, the child process control registers (e.g., MTTL index register 122-INDX) allow page table mappings (and thereby memory tags corresponding to those mappings) to be switched between the first child process and the second child process, e.g., without flushing the TLB(s) and / or without flushing the OLB 111. In some examples, tags for virtual addresses (e.g., and / or physical addresses) are stored in the OLB 111.

[0056] In some examples, memory 101 may include operating system (OS) and / or virtual machine monitor code 103, user (e.g., program) code 105, page data structure(s) 108, memory tag data structure(s) 110, memory tag information structure list 112, or any combination thereof. In some examples of computing, a virtual machine (VM) is an emulation of a computer system. In some examples, VMs are based on a specific computer architecture and provide the functionality of an underlying physical computer system. Their implementation may involve specialized hardware, firmware, software, or a combination thereof. In some examples, a virtual machine monitor (VMM) (also known as a hypervisor) is a software program that, when executed, enables the creation, management, and control of VM instances and manages the operation of a virtualized environment on a physical host machine. In some examples, the VMM is the main software behind the virtualization environment and implementation. In some examples, when installed on a host machine (e.g., a processor), the VMM facilitates the creation of VMs, e.g., each VM having a separate operating system (OS) and applications. The VMM can manage the back-end operations of these VMs by allocating the necessary computing, memory, storage, and other input / output (I / O) resources, such as, but not limited to, an input / output memory management unit (IOMMU). The VMM can provide a centralized interface for managing the complete operation, status, and availability of VMs installed on a single host machine or distributed across different interconnected hosts.

[0057] The memory 101 may be a memory separate from the core and may be a DRAM.

[0058] Coupling (eg, an input / output (I / O) fabric interface) may be included to allow communication between the accelerator core(s) 104 -A to 104 -B, memory 101 , a network interface controller, or any combination thereof.

[0059] In some examples, the hardware initialization manager (non-transitory) storage device 121 stores hardware initialization manager firmware (e.g., or software). In some examples, the hardware initialization manager (non-transitory) storage device 121 stores Basic Input / Output System (BIOS) firmware. In another example, the hardware initialization manager (non-transitory) storage device 121 stores Unified Extensible Firmware Interface (UEFI) firmware. In some examples (e.g., triggered by power-on or reboot of the processor), the computer system 100 (e.g., core 104-A) executes the hardware initialization manager firmware (e.g., or software) stored in the hardware initialization manager (non-transitory) storage device 121 to initialize the system 100 for operation (e.g., to start executing an operating system (OS)) and / or initialize and test (e.g., hardware) components of the system 100.

[0060] The depicted processor 102 includes a set of caches (a first level (L1) cache 108, a second level (L2) cache 116, and a third level (L3) cache 118) and a translation lookaside buffer (TLB) coupled to memory according to examples of the present disclosure. In some examples, the system 100 (e.g., the processor 102) includes a cache coherence circuitry 120 for maintaining cache coherence in the L1 112, L2 (e.g., MLC) 116, and / or L3 118 (e.g., a last level cache (LLC), e.g., the last cache searched before a data item is retrieved from memory 101), for example, according to a cache coherence protocol (such as, but not limited to, the MESI protocol or the MESIF protocol discussed herein). In some examples, the cache coherence circuitry 120 (or other memory circuitry) is further configured to cause TLB accesses, fills, and / or evictions. In some examples, memory management circuitry 110 (eg, including a page walker for performing page walks on misses) is used to manage memory accesses, eg, to implement the paged memory and tag memory disclosed herein.

[0061] Although Figure 1 While two cores (Core A 104-A and Core B 104-B) are depicted, a single core or more than two cores may be utilized. While multiple levels of cache are depicted, a single cache or any number of caches may be utilized. The cache(s) may be organized in any manner, for example, as a physically or logically centralized or distributed cache. Core B 104-B may include Figure 1For example, core B 104 -B may include its own registers 122 and / or OLB 111 , as an example of one or more of the components shown for core A 104 -A in FIG.

[0062] In some examples, each core (e.g., core A 104-A and core B 104-B) includes components for executing instructions. In some examples, core A 104-A includes decoder circuitry and execution circuitry 106, e.g., for decoding instructions and executing the decoded instructions, respectively. In some examples, core A 104-A includes an address generation unit (AGU) (e.g., as part of the execution circuitry), e.g., for generating virtual addresses for memory access requests (e.g., via pointers with tag 107), e.g., to allow core A 104-A to access system memory. In some examples, the AGU takes as input a data value (e.g., a register value and / or address referenced in an instruction) and outputs a (e.g., virtual) address for the data value. In some examples, the execution circuitry (e.g., an execution unit) performs arithmetic operations, such as addition, subtraction, modulo operations, or bit shifts, e.g., using its adders, multipliers, shifters, rotators, etc.

[0063] In some examples, processor 102 stores data and instructions in (e.g., system) memory 101. In some examples, access to these data and / or instructions in memory 101 is slower than the access and / or cycle time of a core access cache (e.g., a cache on processor 102).

[0064] In some examples, core A 104-A includes one or more caches (e.g., a first level (L1) cache 112, a second level (L2) cache 116, and a third level (L3) cache 118) for storing data and / or instructions (e.g., for storing information (e.g., cache lines) itself, rather than retrieving information from memory 101). In some examples, a first level instruction cache (L1I) 112-I is included to store instructions (e.g., corresponding instructions mapped to virtual addresses) and / or a first level data cache (L1D) 112-D is included to store data (e.g., corresponding data mapped to virtual addresses). In some examples, the second level (L2) cache 116 includes, for example, data and / or instructions evicted from the L1 cache(s) of core A 104-A. In some examples, the third level (L3) cache 118 includes, for example, data and / or instructions evicted from the L2 cache of core A 104-A and / or the L2 cache of core B 104-B. In some examples, if the data or instruction is not found in the cache (e.g., not a "hit"), the memory management circuitry 110 (or other memory circuitry) is used to retrieve the data or instruction from memory 101 (e.g., and subsequently store (e.g., "cache") the data or instruction in one or more levels of cache).

[0065] In some examples, cache coherence circuitry 120 is included to maintain cache coherence in the L1 112 , L2 116 , and / or L3 118 caches, eg, according to a cache coherence protocol such as, but not limited to, the MESI protocol or the MESIF protocol discussed herein.

[0066] In some examples, system 100 includes one or more corresponding translation lookaside buffers (TLBs) for cache(s), e.g., where the translation lookaside buffers (TLBs) convert virtual addresses (e.g., of system memory 101) into physical addresses. In some examples, the physical addresses are used to access the cache. In some examples, the TLBs are used to store a data structure including (e.g., most recently used) virtual-to-physical memory address translations (e.g., so that a translation does not have to be performed (e.g., from page data structure 108) for every virtual address present to obtain a physical memory address). In some examples, if the virtual address entry is not in the TLB, a processor (e.g., memory management circuitry 110) is used to perform a page walk to determine the virtual-to-physical memory address translation (e.g., and then store the translation in one or more levels of the TLB).

[0067] In some examples, a first level TLB 114 is included. In some examples, a first level (L1) instruction TLB 114-L1I is included to store virtual address to physical address translations for instructions (e.g., for data that may be stored in L1I cache 112-I). In some examples, a first level (L1) data TLB 114-L1D is included to store virtual address to physical address translations for data (e.g., for data that may be stored in L1D cache 112-D). In some examples, a second level (L2) data and instruction TLB (e.g., a shared TLB (STLB)) 114-L2 is included to store virtual address to physical address translations for data and / or instructions (e.g., for data and / or instructions that may be stored in L2 cache 116). In other examples, the TLBs and cache levels are not as Figure 1 Connected as shown, for example, L1 TLB 114 may contain a mapping for data that is not even cached at all, or is cached in L2 cache 116. Conversely, even if the data is cached, there may not be a mapping for it in the TLB.

[0068] Figure 2 Memory corruption detection (MCD) according to an example of the present disclosure is illustrated. A processing system or processor may maintain a memory tag data structure 110 (e.g., a memory tag table (MTT)) that stores an MCD value (e.g., an MCD identifier) ​​for each of a plurality of rows (e.g., rows of a predefined size (e.g., 64 bytes, although other row sizes may be utilized)) of a memory block. In one example, when a memory block is assigned to a (e.g., newly created) memory object, a unique MCD value is generated and associated with one or more rows of the block. The MCD value may be stored in one or more (e.g., metadata) table entries corresponding to the memory block assigned to the (e.g., newly created) memory object. Figure 2 In , data rows 1 and 2 are depicted as being assigned to object 1 (e.g., as data blocks), and an MCD value (shown here as "2") is associated in memory 101 (e.g., metadata 202 storage), e.g., such that each data row is associated with an entry in memory 101 (e.g., metadata 202 storage) indicating the MCD value (e.g., "2") for that block. Figure 2In FIG, data rows 3-5 are depicted as being assigned to object 2 (e.g., as a data block), and an MCD value (here shown as "7") is associated in memory 101 (e.g., metadata 202 storage), e.g., such that each data row is associated with an entry in memory 101 (e.g., metadata 202 storage) indicating the MCD value (e.g., "7") for that block. In one example, memory 101 (e.g., metadata 202 storage) has an MCD value field for each corresponding row of addressable memory 112. In some examples, metadata 202 is stored in a memory tag data structure 110 (e.g., a memory tag table (MTT)).

[0069] In some examples, the generated MCD value, or a different value that corresponds to or maps to the MCD value generated for the data block, is stored in one or more bits of a pointer, such as a pointer returned by a memory allocation routine to an application requesting the memory allocation. Figure 2 , pointer 107-1 includes an MCD value field 107A-1 having an MCD value ("2") and an address field 107B-1 having a value of a (e.g., linear) address for an object 1 memory block (e.g., the first row). Figure 2 , pointer 107-2 includes an MCD value field 107A-2 having an MCD value ("7") and an address field 107B-2 having a value of a (eg, linear) address for an object 2 memory block (eg, first row).

[0070] In some examples, in response to receiving a memory access instruction (e.g., a memory access instruction determined based on the instruction's opcode or an attempt to access memory), a processing system or processor compares an MCD value retrieved from an MCD table (e.g., for a data block to be accessed) with an MCD value from a pointer specified by the memory access instruction (e.g., extracted from a pointer specified by the memory access instruction). In one example, when the two MCD values ​​match, access to the data block is granted. In one example, when the two MCD values ​​do not match, access to the data block is denied, e.g., a page fault may be generated. In one example, the MCD table (e.g., memory 101 (e.g., metadata 202 storage)) is located in the linear address space of the memory. In one example, circuitry and / or logic for performing MCD validation checks (e.g., in a memory management unit (MMU) 106) is used to access memory, but other portions of the processor (e.g., execution units) are not used to access memory unless the MCD validation check passes (e.g., a match is true). In one example, the access request to the memory block is a load instruction. In one example, the access request to the memory block is a store instruction.

[0071] exist Figure 2 , a request to access an object 1 block in the addressable memory 204 of the memory 101 may be initiated (e.g., by a memory management unit) to read the pointer 107-1 for the MCD value ("2") in the MCD value field 107A-1 and the (e.g., linear) address in the address field 107B-1. The system (e.g., a processor) may then perform a validation check, for example, by loading the MCD value for the row or rows to be accessed in the memory 101 from the memory 101 (e.g., metadata 202 storage) and comparing it to the MCD value in the pointer 107-1 pointing to the row or rows. In some examples, if the system determines that the MCD values ​​match (e.g., both are "2" in this example), the system allows access (e.g., read and / or write) to the memory (e.g., data row 1 only or data rows 1 and 2). In some examples, if there is no match, the request is denied (e.g., the requesting instruction may be faulty). In one example, a request to access the object 1 block may include a request to access all rows in the object (data rows 1 and 2), and the system may perform a validation check on data row 1 (e.g., as discussed above) and may perform a second validation check on data row 2. For example, the system (e.g., a processor) may perform a validation check on row 2 by loading the MCD value for row 2 in memory 101 (e.g., MCD value "2") from memory 101 (e.g., metadata 202 storage) and comparing it to the MCD value in pointer 107-1. In some examples, if the system determines that the MCD values ​​match (e.g., in this example, both MCD values ​​are "2"), the system allows access (e.g., read and / or write) to the memory (e.g., data row 2).

[0072] An example "pointer with tag" format is discussed below, but it will be understood that other formats are possible. In some examples, the format of a capability includes one of the following or any combination of the following: A validity tag that tracks the validity of the capability. If, for example, it is invalid, the capability cannot be used for loading, storing, instruction fetching, or other operations. In some examples, it is still possible to extract fields (including the address of the capability) from an invalid capability. In some examples, capability-aware instructions maintain the tag (e.g., where desired) when the capability is loaded and stored, and when the capability field is accessed, manipulated, and used. A boundary field that identifies the lower and / or upper boundaries of the portion of the address space that the capability is authorized to access (e.g., load, store, instruction fetch, or other operations). An address field (e.g., a virtual address) for the address of the data (e.g., an object) protected by the capability.

[0073] In some examples, the validity tag provides integrity protection, the boundary field restricts how the value can be used (e.g., for memory access), and / or the address field is the memory address where the corresponding data (or instruction) protected by the capability is stored.

[0074] In some examples, the format of a capability includes one of the following or any combination of the following: A validity tag that tracks the validity of a capability. If, for example, it is invalid, the capability cannot be used for loading, storing, instruction fetching, or other operations. In some examples, it is still possible to extract fields (including the address of the capability) from an invalid capability. In some examples, capability-aware instructions maintain the tag (e.g., where desired) when the capability is loaded and stored, and when the capability is accessed, manipulated, and used. A boundary field that identifies the lower and / or upper boundaries of the portion of the address space (e.g., range) that the capability is authorized to access (e.g., load, store, instruction fetching, or other operations). An address field (e.g., virtual address) that is used for the address of data (e.g., object) protected by the capability. A permission field includes a value (e.g., mask) that controls how the capability can be used, for example, by restricting the loading and storing of data and / or capabilities or by preventing instruction fetching. An object type field identifying an object, for example (e.g., in a programming language (e.g., C++) that supports "struct" as a compound data type (or record) declaration that defines a list of variables physically grouped under one name in a block of memory, allowing different variables to be accessed via a single pointer or by the name of the structure declaration returning the same address), a first object type may be used for a structure of people's names, and a second object type may be used for a structure of their physical mailing addresses (e.g., as used in an employee directory). In some examples, if the object type field is not equal to a certain value (e.g., -1), then the capability (using such an object type) is "sealed" and cannot be modified or dereferenced. Sealed capabilities can be used to implement opaque pointer types, for example so that controlled non-monotonicity can be used to support fine-grained intra-address space partitioning.

[0075] In some examples, the permission field includes one or more of the following: "Load" to allow loading from memory protected by a capability; "Store" to allow storage to memory protected by a capability; "Execute" to allow execution of instructions protected by a capability; "LoadCap" to load a valid capability from memory into a register; "StoreCap" to store a valid capability from a register into memory; "Seal" to seal an unsealed capability; "Unseal" to unseal a sealed capability; "System" to access system registers and instructions; "BranchSealedPair" for use in an unsealed branch; "CompartmentID" to be used as a partition ID; "MutableLoad" to load a (e.g., capability) register with mutable permissions; and / or "User[N]" for software-defined permissions (where N is any positive integer greater than 0).

[0076] In some examples, the validity tag field provides integrity protection, the permission(s) field(s) limit the operations that can be performed on the corresponding data (or instructions) protected by the capability, the boundary field limits how the value can be used (e.g., for memory accesses), the object type field supports higher-level software encapsulation, and / or the address field is the memory address where the corresponding data (or instruction) protected by the capability is stored.

[0077] In some examples, a capability (e.g., a value) includes one or any combination of the following fields: an address value (e.g., 64 bits), a boundary (e.g., 87 bits), flags (e.g., 8 bits), an object type (e.g., 15 bits), permissions (e.g., 16 bits), tags (e.g., 1 bit), global (e.g., 1 bit), and / or executable (e.g., OS kernel) (e.g., 1 bit). In some examples, the lower 56 bits of the flags and the "capability boundary" are shared with the encoding of the "capability value."

[0078] In some examples, the format of a capability (e.g., as a pointer that has been extended with security metadata, such as bounds, permissions, and / or type information) overflows the available bits in a pointer (e.g., 64-bit) format. In some examples, to support storing capabilities in a general register file without register extension, examples herein logically group multiple registers (e.g., 4 registers for 256-bit capabilities) so that capabilities can be divided across the multiple underlying registers, e.g., so that general registers of narrower size can be utilized with capabilities of a wider format than (e.g., narrower size) pointers.

[0079] Examples that focus solely on access control usage rather than memory tagging might not require pointer tags at all, e.g., relying entirely on MTTL to determine whether memory access is allowed and with what privileges (e.g., read / write / execute, etc.).

[0080] Figure 3 1 illustrates a dual paging structure (e.g., paging data structure 108 and memory tag data structure 110) and a process for changing a memory tag table mapping according to examples of the present disclosure. In some examples, a process (e.g., as indicated by a process / partition ID register 122-CID (e.g., CR3 register)) utilizes the page data structure 108 (e.g., page table) to determine a set of linear address to physical address mappings (e.g., on a PA13 data page) (e.g., as described below with reference to FIG. Figure 9 discussed).

[0081] In some examples, a child process switch (e.g., an instruction) allows user space code (e.g., software) to select between a plurality of memory tags via the corresponding plurality of memory tag data structures 110. In some examples, a (e.g., user-level) child process of a (e.g., user-level) process is operable to switch a processor (e.g., core) to a set of one or more memory tags (e.g., different from other memory tags) of the child process by executing a child process switch (e.g., an instruction) to the corresponding (one or more) memory tag data structures 110. In some examples, for example, as described in reference Figure 4 As discussed, switching utilizes an index into the memory tag data structure 110 (e.g., from the MTT list index register 122-IDX) from a list 122 of possible memory tag tables (e.g., populated by the OS) (e.g., the list is stored (e.g., by privileged software) at a physical address indicated by the MTTL register 122-MTTL).

[0082] Thus, in some examples, for a switch, the child process stores an index indicator into the MTTL register 122-MTTL for the memory tag set to be utilized, and the index is used to select the physical address for the root of the memory tag set in the memory tag data structure 110, and then traverse the paging structure to (e.g., from Figure 3 PA13 data page 302 in determines the memory tag set used for the child process.

[0083] In other examples, multiple linear ranges of per-partition memory tag tables are utilized and switched between. For example, SWITCHSP can simply switch the linear range of the MTT to select from a set of MTTs within the linear address space. Since the MTT table offset can be fixed based on the linear address space size, an active MTT list may not be required, as the processor can implement alignment of the MTT range registers given a fixed table size (and thus SWITCHSP can specify switching to an MTT range or enumeration thereof). Here, only the OLB may need to be flushed, and the page table mappings and TL remain unaffected by the switching child process. However, in some examples (e.g., most tags will be the same between partitions), it is more efficient to perform a copy-on-write (COW) on the physical page mapping only when something changes in the different partition views (e.g., due to runtime updates to their permissions and / or tags in the memory tag table). In addition, some examples herein allow access control to specific memory locations. In some examples, some memory tag table pages have different permissions for different partitions, but some (e.g., the vast majority) of the physical pages used for the memory tag table can be shared.

[0084] In some examples, when the memory tag table is switched (e.g., from an OS-authorized option in a memory tag table list (MTTL)), the tag / object cache (e.g., OLB) may need to be flushed. Some examples herein additionally tag the tag / object cache (e.g., OLB) to allow entries for the previous partition(s) to remain in the tag / object cache and be reused when returning to the original partition. For example, the base address of the MTT can be used directly to tag the OLB entries, or it can be hashed or transformed in other ways to consume fewer bits associated with each OLB entry. If there is a possibility of conflicts between different OLB entry tags, a structure associated with the OLB can track what OLB entry tags are in use and evict previous entries that conflict with the tag value for the newly added entry.

[0085] In some examples, the TLB handles the linear range of the memory tag table differently. In some examples, under a child process / partition SWITCHSP switch, only those entries in the linear range of the memory tag table are affected, e.g., the rest of the TLB page mapping for the process (e.g., PCID) remains unchanged.

[0086] In some examples, when the root (e.g., CR3) memory tag page is not overloaded by a child page table entry, it can be unaffected in the TLB, for example, so that shared permissions across all child processes / partitions will not be affected in the TLB when switching child processes / partitions.

[0087] Some examples herein create linear memory access control tables for fine-grained partitioning. However, some examples adjust the granule size to be equal to that of a page and still provide per-partition permissions that can be switched from user space, for example, where increasing the granule size reduces the memory required for the memory tag table and where the need for a special object cache (e.g., OLB) can be eliminated by simply updating the TLB permissions on a switch child operation (e.g., a SWITCHSP instruction).

[0088] In some examples, when cross-partition edits of a label table are separate pages, there is no need to lock these edits.

[0089] In some examples, a flat memory tag data structure (e.g., a table) is maintained in linear memory for direct editing by applications in user space, for example, to allow efficient updating of permissions at a runtime executed from user space. In some examples, access control to the linear extents of the memory tag table is limited to runtime or keyed instructions only, to provide the runtime with additional privileges not available to sandboxed user space applications. Other access control alternatives allow editing access to the memory tag table only via special instructions that can be scanned in sandboxed code to ensure that sandboxed applications cannot edit the memory tag table.

[0090] Alternatively, a hierarchical tag table can be maintained in linear memory, and each non-leaf entry in the table can point to a deeper table also in linear memory. The entry can indicate (e.g., specified in a dedicated register) that the current partition ID should be incorporated into the address for the deeper table. For example, if each tag table is 4K, the 16-bit current partition ID can be inserted as bits 27:12 of the address for the deeper table.

[0091] Based on a bit in the linear range of the memory tag table, page table, or tag table, some TLB entries associated with the linear range of the memory tag table can be marked as being partition local, and these TLB entries are automatically flushed when switching to a different partition or tagged with a partition ID so that they can be preserved across partition switches. Partition IDs can be compressed (e.g., hashed) into "compressed partition IDs", and conflicts in the compressed partition ID space can be handled by evicting conflicting TLB entries and / or object lookaside buffer (OLB) entries. The partition ID can be specified in a dedicated register, a slice of the tag table pointer register, or it can be derived from the address specified in the tag table pointer register.

[0092] Figure 4 4 is a flow diagram illustrating operation 400 of a method for switching a memory tag data structure (e.g., a memory tag table) for a process based on an index and a linear address (LA) input (e.g., Switch Child Process (SWITCHSP)) according to an example of the present disclosure. Some or all of operation 400 (or other processes described herein, or variations and / or combinations thereof) are performed by one or more computer systems configured with (one or more) executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors, through hardware, or a combination thereof. The code is stored on a computer-readable storage medium, for example, in the form of a computer program including instructions that can be executed by one or more processors. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operation 400 are performed by a processor (e.g., execution circuitry) of other figures.

[0093] Operation 400 includes, at block 402, receiving a request to switch to a child process (e.g., to switch to its memory tag) (e.g., via an index value stored in the MTTL index register 122-IDX). Operation 400 further includes, at block 404, checking whether the request is valid, e.g., whether the index is a valid index, and if so, continuing to block 408; if not, continuing to block 406 to return an error (e.g., an exception or fault). Operation 400 further includes, at block 406, flushing the stored memory tags, e.g., flushing the memory tags to be changed by the switch (e.g., flushing the memory tags to be changed by the switch). Figure 1 The operation 400 further includes, at block 410, jumping to the linear address (LA) of the child process (e.g., as an entry point into the child process). The operation 400 further includes, at block 412, determining whether there is a TLB miss in the memory tag table (MTT) range (e.g., in the Figure 1If yes, then proceed to block 416, if no, proceed to block 414 to use the (e.g., CR3) TLB entry, page number, and / or permissions (e.g., the TLB entry exists (e.g., from CR3) and is valid). Operation 400 further includes, at block 416, performing a page walk (e.g., such as a page miss handler (PMH)) of the memory tag data structure(s) 110 (e.g., MTT) for the physical address (PA) referenced by the MTTL entry at the index. Figure 3 ). Operation 400 further includes, at block 418, checking whether the mapping is found, and if so, continuing to block 422; if not, continuing to block 420. Operation 400 further includes, at block 420, performing a (e.g., page miss handler (PMH)) page walk of the page data structure (e.g., table) 108 (e.g., as shown in FIG. 1 ). Figure 3 ). Operation 400 further includes, at block 422, performing a TLB fill on the mapping (e.g., and PCID) for the memory tag table page. Operation 400 further includes, at block 424, populating the object lookaside buffer (e.g., OLB 111) with the new tag table mapping, e.g., and thereby with the new memory tag(s) stored at the physical address(es) from the new tag table mapping. Operation 400 further includes, at block 426, using the new tag (e.g., permission) from the memory tag(s).

[0094] Figure 51 is a flow chart illustrating operation 500 of a method (e.g., SWITCHSP) for switching a memory tag data structure (e.g., a memory tag table) for a process based on inputs of a linear range and a linear address (LA) or an enumeration thereof, according to examples of the present disclosure. In some examples, there is a corresponding linear range for each of a plurality of tag tables. Some or all of operation 500 (or other processes described herein, or variations and / or combinations thereof) are performed by one or more computer systems configured with (one or more) executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed on one or more processors, through hardware, or a combination thereof. The code is stored, for example, in the form of a computer program including instructions executable by one or more processors on a computer-readable storage medium. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operation 500 are performed by a processor (e.g., execution circuitry) of other figures. Advantageously, selecting among MTT sets within linear memory may not require changing page table mappings or flushing the TLB.

[0095] Operation 500 includes, at block 502, receiving a request to switch to a different MTTL range for a child process, e.g., along with a linear address to an entry point in the child process. Here, the processor may assume that all MTTs are the same size and that all available MTT tables begin at a specified memory location (e.g., specified by a processor control register set by privileged software). Operation 500 further includes, at block 504, checking whether the provided linear address (LA) and MTT range are valid (e.g., according to Figure 5 If no, continue to block 506 to return an error (e.g., an exception or fault), if yes, continue to block 506. Here, SWITCHSP can specify an enumerated range (e.g., 1, 2, 3, ...) instead of a lookup table, because the size of the MTT tables are all the same and the offset of the enumerated table can be directly calculated by the processor. Operation 500 further includes: at block 506, flushing the stored memory tags, for example, flushing the memory tags to be changed by the switch (e.g., flushing Figure 1111 in the object lookaside buffer (OLB) 111). Operation 500 further includes, at block 508, using the tag table with the specified linear range. Operation 500 further includes, at block 510, jumping to the linear address (LA) of the child process (e.g., as an entry point into the child process). Operation 500 further includes, at block 512, populating the object lookaside buffer (e.g., OLB 111) with the new tag table mapping, e.g., and thereby with the new memory tag(s) stored at the physical address(es) from the new tag table mapping. Operation 500 further includes, at block 514, using the new tag (e.g., permission) from the memory tag(s).

[0096] Figure 6 Illustrated is an example of computing hardware for processing SWITCHSP instruction(s) according to examples of the present disclosure.As illustrated, storage device 603 stores SWITCHSP instruction 601 to be executed.

[0097] Instruction 601 is received by decoder circuitry 605. For example, decoder circuitry 605 receives the instruction from fetch circuitry (not shown). The instruction may be in any suitable format, such as that described below with reference to Figure 20 In some examples, the instruction includes: a field for a source operand that identifies a memory tag data structure for a child process of a (e.g., a single) process (e.g., and a linear address to an entry point in the child process); and an opcode that instructs the execution circuitry to switch from another memory tag data structure for another child process of the process to the memory tag data structure for the child process. In some examples, the source(s) are register(s), and in other examples, one or more are memory locations. In some examples, one or more of the sources may be immediate operands.

[0098] A more detailed example of at least one instruction format for the instruction is described below. Decoder circuitry 605 decodes the instruction into one or more operations. In some examples, the decoding includes generating a plurality of micro-operations to be executed by the execution circuitry. In some examples, decoder circuitry 605 also decodes the instruction prefix.

[0099] In some examples, the register renaming, register allocation, and / or scheduling circuitry 607 provides functionality for one or more of: 1) renaming logical operand values ​​to physical operand values ​​(e.g., a register alias table in some examples), 2) assigning status bits and flags to decoded instructions, and 3) scheduling decoded instructions out of an instruction pool for execution by the execution circuitry (e.g., using a reservation station in some examples).

[0100] Registers (register file) and / or memory 608 store data as operands of instructions to be operated on by execution circuitry 609. Example register types include packed data registers, general purpose registers (GPRs), and floating point registers.

[0101] The execution circuitry 609 executes the decoded instructions. The example detailed execution circuitry includes Figure 1 The execution circuit system 106 shown in FIG, and Figure 17B , etc. In some examples, execution of the decoded instruction causes the execution circuitry to perform a switch sub-process operation (eg, one or more steps) in accordance with the opcode.

[0102] In some examples, retirement / write-back circuitry 611 architecturally commits the destination register(s) to registers or memory 608 and retires the instruction.

[0103] An example of a format for a SWITCHSP instruction is OPCODE SRC0, SRC1. SRC0 (e.g., indicating a reference Figure 4 Index or reference discussed in Figure 5 2004 byte 2104 byte 1618 byte 1620 byte 1622 byte 1624 byte 1626 byte 1628 byte 1629 byte 1630 byte 1631 byte 1632 byte 1633 byte 1634 byte 1635 byte 1636 byte 1637 byte 1638 byte 1639 byte 1640 byte 1641 byte 1642 byte 1643 byte 1644 byte 1645 byte 1646 byte 1647 byte 1648 byte 1649 byte 1650 byte 1651 byte 1652 byte 1653 byte 1654 byte 1655 byte 1656 byte 1657 byte 1658 byte 1660 byte 1661 byte 1662 byte 1663 byte 1664 byte 1665 byte 1666 byte 1667 byte 1668 byte 1669 byte 1670 byte 1671 byte 1672

[0104] Figure 7 1 illustrates an example method executed by a processor to process a SWITCHSP instruction according to an example of the present disclosure. Figure 17B 、 Figure 1 、 Figure 6The processor core shown in FIG, the pipeline as described in detail below, etc. executes the method.

[0105] At 701, an instance of a single instruction is retrieved. For example, a switch subprocess (e.g., switch memory tag table) instruction is retrieved. In some examples, the instruction includes: an identifier of one or more memory tag data structures for a subprocess of a process, one of a plurality of memory tag data structures for a corresponding subprocess of the process; and an opcode for instructing the execution circuitry to switch from another memory tag data structure for another subprocess of the process to a memory tag data structure for the subprocess. In some examples, the instruction further includes a field for a write mask. In some examples, the instruction is retrieved from an instruction cache. The opcode indicates the operation to be performed.

[0106] The fetched instruction is decoded at 703. For example, the fetched SWITCHSP instruction is decoded by decoder circuitry, such as decoder circuitry 106, decoder circuitry 605, or decode circuitry 1740 described in detail herein.

[0107] At 705, when the decoded instruction is dispatched, data values ​​associated with the source operands of the decoded instruction are retrieved. For example, when one or more of the source operands are memory operands, data is retrieved from the indicated memory location.

[0108] At 707, the decoded instructions are executed by execution circuitry (hardware), such as Figure 1 The execution circuit system 106 shown in Figure 6 The execution circuit system 609 shown in Figure 17B In some examples, for a SWITCHSP instruction, execution will cause the execution circuitry to execute the instructions above (e.g., Figure 3 、 Figure 4 or Figure 5 ) described in the operation.

[0109] In some examples, at 709 , the instruction is committed or retired.

[0110] Figure 8 Illustrated is an example method for processing a SWITCHSP instruction using emulation or binary translation according to examples of the present disclosure.

[0111] For example, Figure 17B 、 Figure 1 、 Figure 6 The processor core, pipeline, and / or emulation / translation layer shown in perform aspects of the method.

[0112] At 801, an instance of a single instruction of a first instruction set architecture is retrieved. The instance of the single instruction of the first instruction set architecture includes: an identifier of a memory tag data structure(s) for a child process of a process from a plurality of memory tag data structures for corresponding child processes of the process; and an opcode for instructing the execution circuitry to switch from another memory tag data structure for another child process of the process to the memory tag data structure. In some examples, the instruction further includes a field for a write mask. In some examples, the instruction is retrieved from an instruction cache. The opcode indicates the operation to be performed.

[0113] At 802, the retrieved single instruction of the first instruction set architecture is translated into one or more instructions of the second instruction set architecture. In some examples, the translation is performed by a translation and / or emulation layer of software. In some examples, the translation is performed by a software program such as Figure 21 The instruction converter 2112 is shown performing this translation. In some examples, the translation is performed by hardware translation circuitry.

[0114] At 803, one or more translated instructions of the second instruction set architecture are decoded. For example, the translated instructions are decoded by decoder circuitry, such as decoder circuitry 605 or decode circuitry 1740 described in detail herein. In some examples, the operations of translation and decoding at 802 and 803 are combined.

[0115] At 805, data values ​​associated with source operand(s) of one or more decoded instructions of the second instruction set architecture are retrieved and the one or more instructions are scheduled. For example, when one or more of the source operands are memory operands, data is retrieved from the indicated memory location.

[0116] At 807, the decoded instruction(s) of the second instruction set architecture are executed by execution circuitry (hardware) such as Figure 1 The execution circuit system 106 shown in Figure 6 The execution circuit system 609 shown in Figure 17B 1760) to execute to perform the operation(s) indicated by the opcode of the single instruction of the first instruction set architecture. In some examples, for the SWITCHSP instruction, execution will cause the execution circuitry to perform the above (e.g., Figure 3 、 Figure 4 or Figure 5 ) described in the operation.

[0117] In some examples, at 809 , the instruction is committed or retired.

[0118] As an example of paging, Figure 9 Illustrated is the translation of linear addresses into 4K byte pages 900 using 4-level paging according to some examples. Figure 9 In this example, the offset bits indexed 11-0 of linear address 902 are memory locations within a single page, and the page address begins at (table) bits 12 and higher in linear address 902. In some examples, multiple (e.g., 4-level, 5-level, etc.) paged memories use a hierarchy of paging structures in memory (e.g., located using the contents of a control register (e.g., CR3 or MTTL index register 122-INDX) for locating the first paging structure) to translate linear addresses. In some examples, for 4-level paging, this is the PML4 table.

[0119] Some examples utilize the instruction formats described herein. Some examples are implemented in one or more computer architectures, cores, accelerators, etc. Some examples are generated or are IP cores. Some examples utilize simulation and / or translation.

[0120] At least some examples of the disclosed technology can be described in terms of the following set of examples.

[0121] In a first set of examples, a device (e.g., a processor core) includes: memory management circuitry for controlling memory access based on memory tags stored in a memory tag data structure and based on memory tags of pointers to memory; decoder circuitry for decoding an instruction into a decoded instruction, the instruction including an operand and an opcode, the operand identifying a memory tag data structure for a child process from a plurality of memory tag data structures for a corresponding child process of a process, the opcode instructing execution circuitry to switch from another memory tag data structure for another child process of the process to the memory tag data structure for the child process; and execution circuitry for executing the decoded instruction according to the opcode. In some examples, the instruction is executable in user mode. In some examples, the operand includes an index into a memory tag data structure from a plurality of memory tag data structure indexes. In some examples, the operand further includes a linear address to an entry point of the child process. In some examples, the opcode further instructs the execution circuitry to populate an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure. In some examples, the object lookaside buffer is separate from a translation lookaside buffer of the device. In some examples, the memory tag data structure includes virtual-to-physical page mappings for the stored memory tags, and the opcode is further used to instruct the execution circuitry to retain other virtual-to-physical page mappings for the process in the translation lookaside buffer.

[0122] In another set of examples, a method includes: decoding, by decoder circuitry, an instruction into a decoded instruction, the instruction including an operand and an opcode, the operand identifying a memory tag data structure for a child process from a plurality of memory tag data structures for a corresponding child process of a process, the opcode instructing execution circuitry to switch from another memory tag data structure for another child process of the process to the memory tag data structure for the child process; executing, by the execution circuitry, the decoded instruction according to the opcode; and controlling, by memory management circuitry, memory access based on memory tags stored in the memory tag data structure and based on memory tags of pointers to memory. In some examples, execution is in user mode. In some examples, the operand includes an index into a memory tag data structure from a plurality of memory tag data structure indexes. In some examples, the operand further includes a linear address to an entry point of the child process. In some examples, execution populates an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure. In some examples, the object lookaside buffer is separate from a translation lookaside buffer of a method. In some examples, the memory tag data structure includes virtual-to-physical page mappings for stored memory tags and performs other virtual-to-physical page mappings retained in a translation lookaside buffer for the process.

[0123] In another set of examples, a non-transitory machine-readable medium stores code that, when executed by a machine, causes the machine to perform a method comprising: decoding, by decoder circuitry, an instruction into a decoded instruction, the instruction comprising an operand and an opcode, the operand identifying a memory tag data structure for a child process from a plurality of memory tag data structures for a corresponding child process of a process, the opcode instructing execution circuitry to switch from another memory tag data structure for another child process of the process to the memory tag data structure for the child process; executing, by the execution circuitry, the decoded instruction according to the opcode; and controlling, by memory management circuitry, memory access based on memory tags stored in the memory tag data structure and based on memory tags of pointers to memory. In some examples, the operand comprises an index into a memory tag data structure from the plurality of memory tag data structure indexes. In some examples, the operand further comprises a linear address to an entry point of the child process. In some examples, execution populates an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure. In some examples, the object lookaside buffer is separate from a translation lookaside buffer of the method. In some examples, the memory tag data structure includes virtual-to-physical page mappings for stored memory tags and performs other virtual-to-physical page mappings retained in a translation lookaside buffer for the process.

[0124] Exemplary architectures, systems, etc. that may be used in the foregoing are detailed below. Exemplary instruction formats for the SWITCHSP instruction are detailed below. Example Architecture

[0125] The following describes an example computer architecture. Other system designs and configurations known in the art for laptops, desktop computers, handheld personal computers (PCs), personal digital assistants, engineering workstations, servers, separate servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices are also suitable. In general, various systems or electronic devices that can include processors and / or other execution logic as disclosed herein are generally suitable. Example System

[0126] Figure 10 An example computing system is illustrated. Multiprocessor system 1000 is an interfaced system and includes multiple processors or cores including a first processor 1070 and a second processor 1080 coupled via an interface 1050 such as a point-to-point (PP) interconnect, fabric, and / or bus. In some examples, first processor 1070 and second processor 1080 are isomorphic. In some examples, first processor 1070 and second processor 1080 are heterogeneous. Although example system 1000 is shown as having two processors, the system can have three or more processors or can be a single processor system. In some examples, the computing system is a system on a chip (SoC).

[0127] Processor 1070 and processor 1080 are shown as including integrated memory controller (IMC) circuitry 1072 and 1082, respectively. Processor 1070 also includes interface circuits 1076 and 1078; similarly, second processor 1080 includes interface circuits 1086 and 1088. Processors 1070, 1080 can exchange information via interface 1050 using interface circuits 1078, 1088. IMCs 1072 and 1082 couple processors 1070, 1080 to respective memories, namely, memory 1032 and memory 1034, which may be portions of main memory locally attached to the respective processors.

[0128] Processors 1070, 1080 can each exchange information with a network interface (NW I / F) 1090 via separate interfaces 1052, 1054 using interface circuits 1076, 1094, 1086, 1098. Network interface 1090 (e.g., one or more of an interconnect, bus, and / or fabric, and in some examples, a chipset) can optionally exchange information with a coprocessor 1038 via interface circuit 1092. In some examples, coprocessor 1038 is a special-purpose processor, such as, for example, a high-throughput processor, a network or communication processor, a compression engine, a graphics processor, a general-purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, or the like.

[0129] A shared cache (not shown) may be included in either processor 1070, 1080, or external to both processors but connected to these processors via an interface (such as a PP interconnect) so that if the processors are placed in a low power mode, the local cache information of either or both processors may be stored in the shared cache.

[0130] The network interface 1090 can be coupled to the first interface 1016 via the interface circuit 1096. In some examples, the first interface 1016 can be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 1016 is coupled to a power control unit (PCU) 1017, which can include circuitry, software, and / or firmware for performing power management operations related to the processors 1070, 1080, and / or the coprocessor 1038. The PCU 1017 provides control information to a voltage regulator (not shown) so that the voltage regulator generates an appropriate regulated voltage. The PCU 1017 also provides control information to control the generated operating voltage. In various examples, the PCU 1017 can include various power management logic (circuitry) for performing hardware-based power management. Such power management may be entirely controlled by the processor (e.g., by various processor hardware, and which may be triggered by workload and / or power, thermal constraints, or other processor constraints), and / or power management may be performed in response to an external source (such as a platform or power management source or system software).

[0131] The PCU 1017 is illustrated as existing as logic separate from the processor 1070 and / or the processor 1080. In other cases, the PCU 1017 may execute on a given one or more of the cores (not shown) of the processor 1070 or 1080. In some cases, the PCU 1017 may be implemented as a (dedicated or general-purpose) microcontroller or other control logic configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by the PCU 1017 may be implemented external to the processor, such as by way of a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by the PCU 1017 may be implemented within the BIOS or other system software.

[0132] Various I / O devices 1014 can be coupled to the first interface 1016 along with a bus bridge 1018, which couples the first interface 1016 to the second interface 1020. In some examples, one or more additional processors 1015 (such as a coprocessor, a high-throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor) are coupled to the first interface 1016. In some examples, the second interface 1020 can be a low pin count (LPC) interface. Various devices can be coupled to the second interface 1020, including, for example, a keyboard and / or mouse 1022, a communication device 1027, and a storage circuit system 1028. Storage circuitry 1028 may be one or more non-transitory machine-readable storage media, such as a disk drive or other mass storage device, as described below, which in some examples may include instructions / code and data 1030 and may implement storage 603. Further, audio I / O 1024 may be coupled to second interface 1020. Note that other architectures are possible besides the point-to-point architecture described above. For example, a system such as multiprocessor system 1000 may implement a multi-drop interface or other such architecture rather than a point-to-point architecture.

[0133] Exemplary Core Architectures, Processors, and Computer Architectures

[0134] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores may include: 1) a general-purpose in-order core intended for general-purpose computing; 2) a high-performance general-purpose out-of-order core intended for general-purpose computing; 3) a specialized core intended primarily for graphics and / or scientific (throughput) computing. Different processor implementations may include: 1) a CPU that includes one or more general-purpose in-order cores intended for general-purpose computing and / or one or more general-purpose out-of-order cores intended for general-purpose computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors give rise to different computer system architectures, which may include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such coprocessors are sometimes referred to as dedicated logic or as dedicated cores, such as integrated graphics and / or scientific (throughput) logic); and 4) a system on a chip (SoC), which may include the described CPU (sometimes referred to as application core(s) or application processor(s), the coprocessor(s) described above, and additional functionality on the same die. An example core architecture is described next, followed by a description of an example processor and computer architecture.

[0135] Figure 11 A block diagram of an example processor and / or SoC 1100 that may have one or more cores and an integrated memory controller is shown. The solid line box illustrates the processor 1100 having a single core 1100(A), a system agent unit circuitry 1110, and a collection of one or more interface controller unit circuitry 1116, while the optional addition of the dashed line box illustrates an alternative processor 1100 having multiple cores 1100(A)-1102(N), a collection of one or more integrated memory controller units 1114 in the system agent unit circuitry 1110, and dedicated logic 1108, and a collection of one or more interface controller unit circuitry 1116. Note that the processor 1100 may be Figure 10 processor 1070 or 1080, or one of the coprocessors 1038 or 1015.

[0136] Thus, different implementations of processor 1100 may include: 1) a CPU, wherein dedicated logic 1108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and cores 1102(A)-1102(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of the two); 2) a coprocessor, wherein cores 1102(A)-1102(N) are a large number of dedicated cores intended primarily for graphics and / or scientific (throughput); and 3) a coprocessor, wherein cores 1102(A)-1102(N) are a large number of general-purpose in-order cores. Thus, processor 1100 may be a general-purpose processor, a coprocessor, or a dedicated processor, such as, for example, a network or communications processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput many integrated core (MIC) coprocessor (including 30 or more cores), an embedded processor, etc. The processor may be implemented on one or more chips. Processor 1100 may be part of and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, complementary metal-oxide-semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0137] The memory hierarchy includes one or more levels of cache unit circuitry 1104(A)-1104(N) within cores 1102(A)-1104(N), a set of one or more shared cache unit circuitry 1106, and external memory (not shown) coupled to a set of integrated memory controller unit circuitry 1114. The set of one or more shared cache unit circuitry 1106 may include one or more intermediate levels of cache (such as level 2 (L2), level 3 (L3), level 4 (L4)), or other levels of cache (such as last level cache (LLC)), and / or combinations thereof. While in some examples, an interface network circuitry 1112 (e.g., a ring interconnect) provides an interface to dedicated logic 1108 (e.g., integrated graphics logic), the set of one or more shared cache unit circuitry 1106, and the system agent unit circuitry 1110, alternative examples use any number of well-known techniques for providing interfaces to such units. In some examples, coherency is maintained between the shared cache unit circuitry(s) 1106 and one or more of the cores 1102(A)-1102(N). In some examples, the interface controller unit circuitry 1116 couples the cores 1102(A)-1102(N) to one or more other devices 1118, such as one or more I / O devices, storage, one or more communication devices (e.g., a wireless network, a wired network, etc.), and the like. In some examples, one or more of cores 1102(A)-1102(N) can be multithreaded. System agent unit circuitry 1110 includes components that coordinate and operate cores 1102(A)-1102(N). System agent unit circuitry 1110 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may include or may include the logic and components required to regulate the power state of cores 1102(A)-1102(N) and / or dedicated logic 1108 (e.g., integrated graphics logic). Display unit circuitry is used to drive one or more externally connected displays.

[0138] The cores 1102(A)-1102(N) may be homogeneous in terms of instruction set architecture (ISA). Alternatively, the cores 1102(A)-1102(N) may be heterogeneous in terms of ISA; that is, a subset of the cores 1102(A)-1102(N) may be capable of executing an ISA, while other cores may only be capable of executing a subset of that ISA or another ISA.

[0139] Figure 12 1 is a block diagram illustrating a computing system 1200 configured to implement one or more aspects of the examples described herein. Computing system 1200 includes a processing subsystem 1201 having one or more processors 1202 and a system memory 1204 communicating via an interconnect path, which may include a memory hub 1205. Memory hub 1205 may be a separate component within a chipset component or may be integrated within one or more processors 1202. Memory hub 1205 is coupled to an I / O subsystem 1211 via a communication link 1206. I / O subsystem 1211 includes an I / O hub 1207, which may enable computing system 1200 to receive input from one or more input devices 1208. Additionally, I / O hub 1207 may enable a display controller (which may be included in one or more processors 1202) to provide output to one or more display devices 1210A. In some examples, one or more display devices 1210A coupled to I / O hub 1207 may include local, internal, or embedded display devices.

[0140] The processing subsystem 1201 includes, for example, one or more parallel processors 1212 coupled to a memory hub 1205 via a bus or other communication link 1213. The communication link 1213 can be one of any number of standards-based communication link technologies or protocols, such as, but not limited to, PCI Express, or can be a vendor-specific communication interface or communication fabric. The one or more parallel processors 1212 can form a parallel or vector processing system in a computational cluster that can include a large number of processing cores and / or processing clusters, such as an integrated multi-core (MIC) processor. For example, the one or more parallel processors 1212 form a graphics processing subsystem that can output pixels to one of one or more display devices 1210A coupled via the I / O hub 1207. The one or more parallel processors 1212 can also include a display controller and display interface (not shown) for enabling direct connection to the one or more display devices 1210B.

[0141] Within the I / O subsystem 1211, a system storage unit 1214 may be connected to the I / O hub 1207, thereby providing a storage mechanism for the computing system 1200. An I / O switch 1216 may be used to provide an interface mechanism to enable connections between the I / O hub 1207 and other components, such as a network adapter 1218 and / or a wireless network adapter 1219 that may be integrated into the platform, as well as various other devices that may be added via one or more plug-in devices 1220. The plug-in device(s) 1220 may also include, for example, one or more external graphics processor devices, graphics cards, and / or computing accelerators. The network adapter 1218 may be an Ethernet adapter or another wired network adapter. The wireless network adapter 1219 may include one or more of the following: Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more wireless radio devices.

[0142] The computing system 1200 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 1207. Figure 12 The communication paths interconnecting the various components in the system may be implemented using any suitable protocol, such as a Peripheral Component Interconnect (PCI) based protocol (e.g., PCI Express), or any other bus or point-to-point communication interface and / or protocol(s), such as NVLink high-speed interconnect, Compute Express Link, or a similar protocol. TM (ComputeExpress Link TM , CXL TM) (e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE 802.3), remote direct memory access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel Quick Path Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators The data may be copied or stored to the virtualized storage node using protocols such as 3GPP LTE (3GPP Accelerators, CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and variants thereof, or wired or wireless interconnect protocols known in the art. In some examples, data may be copied or stored to the virtualized storage node using protocols such as non-volatile memory express (NVMe) over Fabrics (NVMe-oF) or NVMe.

[0143] One or more parallel processors 1212 may include circuitry optimized for graphics and video processing (including, for example, video output circuitry) and constitute a graphics processing unit (GPU). Alternatively or additionally, as described in more detail herein, one or more parallel processors 1212 may include circuitry optimized for general-purpose processing while retaining the underlying computing architecture. Components of computing system 1200 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 1212, memory hub 1205, (one or more) processors 1202, and I / O hub 1207 may be integrated into a system-on-chip (SoC) integrated circuit. Alternatively, components of computing system 1200 may be integrated into a single package to form a system-in-package (SIP) configuration. In some examples, at least a portion of the components of computing system 1200 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system.

[0144] It will be appreciated that the computing system 1200 shown herein is illustrative and that variations and modifications are possible. The connection topology, including the number and arrangement of bridges, the number of processor(s) 1202, and the number of parallel processor(s) 1212, may be modified as desired. For example, the system memory 1204 may be connected to the processor(s) 1202 directly rather than through a bridge, while other devices communicate with the system memory 1204 via the memory hub 1205 and the processor(s) 1202. In other alternative topologies, the parallel processor(s) 1212 are connected to the I / O hub 1207 or directly to one of the processor(s) 1202 rather than to the memory hub 1205. In other examples, the I / O hub 1207 and the memory hub 1205 may be integrated into a single chip. It is also possible for two or more sets of processors 1202 to be attached via multiple sockets, which may be coupled to two or more instances of the parallel processor(s) 1212.

[0145] Some of the specific components shown herein are optional and may not be included in all implementations of the computing system 1200. For example, any number of plug-in cards or peripherals may be supported, or some components may be eliminated. Additionally, some architectures may be specific to the architecture of the computer. Figure 12 Different terminology is used for components that are similar to those illustrated in FIG. For example, in some architectures, memory hub 1205 may be referred to as a north bridge, while I / O hub 1207 may be referred to as a south bridge.

[0146] Figure 13A An example of a parallel processor 1300 is shown. The parallel processor 1300 may be a GPU, a GPGPU, or the like as described herein. The various components of the parallel processor 1300 may be implemented using one or more integrated circuit devices, such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). The illustrated parallel processor 1300 may be Figure 12 One or more of the parallel processor(s) 1212 shown in .

[0147] Parallel processor 1300 includes parallel processing unit 1302. Parallel processing unit includes an I / O unit 1304 that enables communication with other devices, including other instances of parallel processing unit 1302. I / O unit 1304 can be directly connected to other devices. For example, I / O unit 1304 connects to other devices via the use of a hub or switch interface, such as memory hub 1205. The connection between memory hub 1205 and I / O unit 1304 forms communication link 1213. Within parallel processing unit 1302, I / O unit 1304 is connected to a host interface 1306, which receives commands related to performing processing operations, and a memory crossbar switch 1316, which receives commands related to performing memory operations.

[0148] When host interface 1306 receives command buffers via I / O unit 1304, host interface 1306 can direct work operations for executing those commands to front end 1308. In some examples, front end 1308 is coupled to scheduler 1310, which is configured to dispatch commands or other work items to processing cluster array 1312. Scheduler 1310 ensures that processing cluster array 1312 is properly configured and in a valid state before tasks are dispatched to its processing clusters. Scheduler 1310 can be implemented via firmware logic executing on a microcontroller. A microcontroller-implemented scheduler 1310 can be configured to perform complex scheduling and work dispatch operations at both coarse and fine granularity, thereby enabling fast preemption and context switching of threads executing on processing cluster array 1312. Preferably, host software can validate workloads for scheduling on processing cluster array 1312 via one of a plurality of graphics processing doorbells. In other examples, polling for new workloads or interrupts may be used to identify or indicate the availability of work to be performed.The workload may then be automatically distributed across the processing cluster array 1312 by the scheduler 1310 logic within the scheduler microcontroller.

[0149] Processing cluster array 1312 may include up to "N" processing clusters (e.g., cluster 1314A, cluster 1314B, through cluster 1314N). Each cluster 1314A-1314N of processing cluster array 1312 may execute a large number of concurrent threads. Scheduler 1310 may assign work to clusters 1314A-1314N in processing cluster array 1312 using various scheduling and / or work distribution algorithms that may vary depending on the workload generated for each type of program or computation. Scheduling may be handled dynamically by scheduler 1310 or may be assisted in part by compiler logic during the compilation of program logic configured for execution by processing cluster array 1312. Optionally, different clusters 1314A-1314N of processing cluster array 1312 may be assigned to process different types of programs or to perform different types of computations.

[0150] Processing cluster array 1312 can be configured to perform various types of parallel processing operations. For example, processing cluster array 1312 is configured to perform general-purpose parallel computing operations. For example, processing cluster array 1312 can include logic for performing processing tasks including filtering of video and / or audio data, performing modeling operations including physics operations, and performing data transformations.

[0151] Processing cluster array 1312 is configured to perform parallel graphics processing operations. In such examples where parallel processor 1300 is configured to perform graphics processing operations, processing cluster array 1312 may include additional logic for supporting the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, as well as tessellation logic and other vertex processing logic. Additionally, processing cluster array 1312 may be configured to execute graphics processing-related shader programs, such as, but not limited to, vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. Parallel processing unit 1302 may transfer data from system memory via I / O unit 1304 for processing. The transferred data may be stored in on-chip memory (e.g., parallel processor memory 1322) during processing and subsequently written back to system memory.

[0152] In an example where parallel processing units 1302 are used to perform graphics processing, scheduler 1310 may be configured to divide the processing workload into tasks of approximately equal size to better distribute graphics processing operations to multiple clusters 1314A-1314N in processing cluster array 1312. In some of these examples, portions of processing cluster array 1312 may be configured to perform different types of processing. For example, a first portion may be configured to perform vertex shading and topology generation, a second portion may be configured to perform tessellation and geometry shading, and a third portion may be configured to perform pixel shading or other screen-space operations to produce a rendered image for display. Intermediate data generated by one or more of clusters 1314A-1314N may be stored in a buffer to allow the intermediate data to be transferred between clusters 1314A-1314N for further processing.

[0153] During operation, processing cluster array 1312 may receive processing tasks to be executed via scheduler 1310, which receives commands defining the processing tasks from front end 1308. For graphics processing operations, a processing task may include data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, and an index of state parameters and commands defining how the data is to be processed (e.g., what program to execute). Scheduler 1310 may be configured to retrieve an index corresponding to a task, or may receive an index from front end 1308. Front end 1308 may be configured to ensure that processing cluster array 1312 is configured in a valid state before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is initiated.

[0154] Each of the one or more instances of parallel processing unit 1302 can be coupled to parallel processor memory 1322. Parallel processor memory 1322 can be accessed via memory crossbar 1316, which can receive memory requests from processing cluster array 1312 and I / O unit 1304. Memory crossbar 1316 can access parallel processor memory 1322 via memory interface 1318. Memory interface 1318 can include a plurality of partition units (e.g., partition unit 1320A, partition unit 1320B, through partition unit 1320N), each of which can be coupled to a portion (e.g., memory cells) of parallel processor memory 1322. The number of partition units 1320A-1320N can be configured to be equal to the number of memory cells, such that the first partition unit 1320A has a corresponding first memory cell 1324A, the second partition unit 1320B has a corresponding second memory cell 1324B, and the Nth partition unit 1320N has a corresponding Nth memory cell 1324N. In other examples, the number of partition units 1320A-1320N may not be equal to the number of memory devices.

[0155] Memory units 1324A-1324N may include various types of memory devices, including dynamic random-access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. Optionally, memory units 1324A-1324N may also include 3D stacked memory, including, but not limited to, high bandwidth memory (HBM). Those skilled in the art will appreciate that the specific implementation of memory units 1324A-1324N may vary and may be selected from one of a variety of conventional designs. Render targets, such as frame buffers or texture maps, may be stored across memory units 1324A-1324N, allowing partition units 1320A-1320N to write portions of each render target in parallel to efficiently use the available bandwidth of parallel processor memory 1322. In some examples, local instances of parallel processor memory 1322 may be eliminated in favor of utilizing a unified memory design of system memory in conjunction with local cache memory.

[0156] Optionally, any of the clusters 1314A-1314N in the processing cluster array 1312 has the capability to process data to be written to any of the memory units 1324A-1324N within the parallel processor memory 1322. The memory crossbar 1316 can be configured to transmit the output of each cluster 1314A-1314N to any partition unit 1320A-1320N or to another cluster 1314A-1314N, which can perform additional processing operations on the output. Each cluster 1314A-1314N can communicate with a memory interface 1318 via the memory crossbar 1318 to read from or write to various external memory devices. In one example having a memory crossbar 1316, the memory crossbar 1316 has connections to a memory interface 1318 for communicating with the I / O unit 1304 and to a local instance of parallel processor memory 1322, thereby enabling processing units within different processing clusters 1314A-1314N to communicate with system memory or other memory that is not local to the parallel processing unit 1302. In general, the memory crossbar 1316 may be capable of separating traffic flows between the clusters 1314A-1314N and the partition units 1320A-1320N, for example, using virtual channels.

[0157] Although a single instance of parallel processing unit 1302 is illustrated within parallel processor 1300, any number of instances of parallel processing unit 1302 may be included. For example, multiple instances of parallel processing unit 1302 may be provided on a single plug-in card, or multiple plug-in cards may be interconnected. For example, parallel processor 1300 may be a plug-in device, such as Figure 12(One or more) plug-in devices 1220, which can be graphics cards (such as, discrete graphics cards including one or more GPUs, one or more memory devices, and device-to-device or network or fabric interfaces). Different instances of parallel processing units 1302 can be configured to interoperate even if different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. Optionally, some instances of parallel processing units 1302 can include floating point units with higher precision than other instances. Systems comprising one or more instances of parallel processing units 1302 or parallel processors 1300 can be implemented in various configurations and form factors, including but not limited to desktop computers, laptop computers, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems. The orchestrator can use one or more of the following to form a composite node for workload execution: separate processor resources, cache resources, memory resources, storage resources, and networking resources.

[0158] In some examples, the parallel processing unit 1302 can be partitioned into multiple instances. Those multiple instances can be configured to execute workloads associated with different clients in an isolated manner, thereby providing a predetermined quality of service to each client. For example, each cluster 1314A-1314N can be partitioned and isolated from other clusters, allowing the processing cluster array 1312 to be divided into multiple computing partitions or instances. In such a configuration, workloads executed on isolated partitions are protected from errors or errors associated with different workloads executed on different partitions. Partition units 1,320A-1,320N can be configured to enable dedicated and / or isolated paths to the memory of clusters 1,314A-1,314N associated with the corresponding computing partition. This data path isolation enables computing resources within a partition to communicate with one or more assigned memory units 1324A-1324N without being interfered with by the activities of other partitions.

[0159] Figure 13B is a block diagram of the partition unit 1320. The partition unit 1320 may be Figure 13A1320N, and a frame buffer interface 1325. As shown, the partition unit 1320 includes an L2 cache 1321, a frame buffer interface 1325, and a ROP 1326 (raster operation unit). The L2 cache 1321 is a read / write cache configured to perform load and store operations received from the memory crossbar 1316 and the ROP 1326. Read misses and urgent writeback requests are output by the L2 cache 1321 to the frame buffer interface 1325 for processing. Updates can also be sent to the frame buffer via the frame buffer interface 1325 for processing. In some examples, the frame buffer interface 1325 communicates with memory units in the parallel processor memory (such as, for example, within the parallel processor memory 1322). Figure 13A The partition unit 1320 may also additionally or alternatively interface with one of the memory units in the parallel processor memory via a memory controller (not shown).

[0160] In graphics applications, ROP 1326 is a processing unit that performs raster operations such as stenciling, z-testing, blending, and the like. ROP 1326 then outputs processed graphics data, which is stored in graphics memory. In some examples, ROP 1326 includes or is coupled to a codec (CODEC) 1327 that includes compression logic for compressing depth or color data written to memory or L2 cache 1321 and decompressing depth or color data read from memory or L2 cache 1321. The compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The type of compression performed by CODEC 1327 can vary based on the statistical characteristics of the data to be compressed. For example, in some examples, delta color compression is performed on the depth and color data on a tile-by-tile basis. In some examples, CODEC 1327 includes compression and decompression logic that can compress and decompress computational data associated with machine learning operations. CODEC 1327 can, for example, compress sparse matrix data for sparse machine learning operations. CODEC 1327 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to generate compressed and encoded sparse matrix data. The compressed and encoded sparse matrix data can be decompressed and / or decoded before being processed by the processing element, or the processing element can be configured to consume compressed, encoded, or compressed and encoded data for processing.

[0161] ROP 1326 may be included in each processing cluster (e.g., Figure 13A 1314N) rather than being included in partition unit 1320. In such an example, read and write requests for pixel data rather than pixel fragment data are routed through memory crossbar 1316. The processed graphics data may be displayed on a display device such as a Figure 12 1210B), is routed for further processing by the processor(s) 1202, or is routed for processing by the processor(s) 1202. Figure 13A One of the processing entities within the parallel processor 1300 further processes.

[0162] Figure 13C is a block diagram of a processing cluster 1314 within a parallel processing unit. For example, a processing cluster is Figure 13A 1314N。Processing cluster 1314 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a specific program executed on a specific set of input data. Optionally, single-instruction, multiple-data (SIMD) instruction issuance technology can be used to support the parallel execution of a large number of threads without providing multiple independent instruction units. Alternatively, single-instruction, multiple-thread (SIMT) technology can be used to use a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster in the processing cluster to support the parallel execution of a large number of generally synchronized threads. Unlike the SIMD execution mechanism in which all processing engines typically execute the same instruction, SIMT execution allows different threads to more easily follow divergent execution paths through a given thread program. Those skilled in the art will understand that the SIMD processing scheme represents a functional subset of the SIMT processing scheme.

[0163] The operation of the processing cluster 1314 can be controlled via a pipeline manager 1332 that distributes processing tasks to SIMT parallel processors. Figure 13A Scheduler 1310 receives instructions and manages the execution of those instructions via graphics multiprocessor 1334 and / or texture unit 1336. The illustrated graphics multiprocessor 1334 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors with different architectures may be included within processing cluster 1314. One or more instances of graphics multiprocessor 1334 may be included within processing cluster 1314. Graphics multiprocessor 1334 may process data, and data crossbar 1340 may be used to distribute the processed data to one of multiple possible destinations, including other shader units. Pipeline manager 1332 may facilitate the distribution of processed data by specifying a destination for the processed data to be distributed via data crossbar 1340.

[0164] Each graphics multiprocessor 1334 within a processing cluster 1314 can include an identical set of function execution logic (e.g., arithmetic logic units, load-store units, etc.). The function execution logic can be configured in a pipelined manner, whereby new instructions can be issued before previous instructions have completed. The function execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, Boolean operations, bit shifts, and calculations of various algebraic functions. The same functional unit hardware can be utilized to perform different operations, and any combination of functional units can be present.

[0165] Instructions transmitted to processing cluster 1314 constitute threads. A collection of threads executed across a collection of parallel processing engines is a thread group. Thread groups execute the same program on different input data. Each thread within a thread group can be assigned to a different processing engine within graphics multiprocessor 1334. A thread group can include fewer threads than the number of processing engines within graphics multiprocessor 1334. When a thread group includes fewer threads than the number of processing engines, one or more of the processing engines can be idle during cycles during which the thread group is being processed. A thread group can also include more threads than the number of processing engines within graphics multiprocessor 1334. When a thread group includes more threads than the number of processing engines within graphics multiprocessor 1334, processing can be performed in consecutive clock cycles. Optionally, multiple thread groups can be executed concurrently on graphics multiprocessor 1334.

[0166] The graphics multiprocessor 1334 may include internal cache memory to perform load and store operations. Alternatively, the graphics multiprocessor 1334 may forgo the internal cache and use cache memory within the processing cluster 1314 (e.g., the first level (L1) cache 1348). Each graphics multiprocessor 1334 also has a partition unit (e.g., Figure 13A 1,320N) are shared among all processing clusters 1314 and can be used to transfer data between threads. Graphics multiprocessor 1334 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. Any memory external to parallel processing unit 1302 can be used as global memory. In examples where processing cluster 1314 includes multiple instances of graphics multiprocessor 1334, common instructions and data can be shared, which can be stored in L1 cache 1348.

[0167] Each processing cluster 1314 may include a memory management unit (MMU) 1345 configured to map virtual addresses to physical addresses. In other examples, one or more instances of the MMU 1345 may reside in Figure 13A 1318. The MMU 1345 includes a set of page table entries (PTEs) for mapping virtual addresses to physical addresses of the slice, and optionally includes a cache line index. The MMU 1345 may include an address translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 1334 or L1 cache 1348 of the processing cluster 1314. Physical addresses are processed to distribute surface data access locality, thereby allowing efficient request interleaving between partition units. The cache line index can be used to determine whether a request for a cache line is a hit or a miss.

[0168] In graphics and compute applications, the processing clusters 1314 can be configured such that each graphics multiprocessor 1334 is coupled to a texture unit 1336 for performing texture mapping operations, such as determining texture sample locations, reading texture data, and filtering texture data. Texture data is read from an internal texture L1 cache (not shown), or in some examples, from an L1 cache within the graphics multiprocessor 1334 and retrieved from an L2 cache, local parallel processor memory, or system memory as needed. Each graphics multiprocessor 1334 outputs processed tasks to a data crossbar 1340 to provide the processed tasks to another processing cluster 1314 for further processing, or to store the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 1316. A pre-raster operations unit (preROP) 1342 is configured to receive data from the graphics multiprocessor 1334 and direct the data to ROP units that can interact with partition units (e.g., Figure 13A The preROP 1342 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0169] It will be appreciated that the core architecture described herein is illustrative and that variations and modifications are possible. Any number of processing units (e.g., graphics multiprocessor 1334, texture unit 1336, preROP 1342, etc.) may be included within processing cluster 1314. Further, while only one processing cluster 1314 is shown, a parallel processing unit as described herein may include any number of instances of processing cluster 1314. Optionally, each processing cluster 1314 may be configured to operate independently of other processing clusters 1314 using separate and distinct processing units, L1 cache, L2 cache, etc.

[0170] Figure 13D An example of a graphics multiprocessor 1334 is shown, where the graphics multiprocessor 1334 is coupled to a pipeline manager 1332 of a processing cluster 1314. The graphics multiprocessor 1334 has an execution pipeline that includes, but is not limited to, an instruction cache 1352, an instruction unit 1354, an address mapping unit 1356, a register file 1358, one or more general-purpose graphics processing unit (GPGPU) cores 1362, and one or more load / store units 1366. The GPGPU cores 1362 and the load / store units 1366 are coupled to a cache memory 1372 and a shared memory 1370 via a memory and cache interconnect 1368. The graphics multiprocessor 1334 may additionally include a tensor and / or ray tracing core 1363 that includes hardware logic for accelerating matrix and / or ray tracing operations.

[0171] The instruction cache 1352 may receive a stream of instructions to be executed from the pipeline manager 1332. Instructions are cached in the instruction cache 1352 and dispatched for execution by the instruction unit 1354. The instruction unit 1354 may dispatch instructions as thread groups (e.g., warps), where each thread in a thread group is assigned to a different execution unit within the GPGPU core 1362. Instructions may access any of the local address space, the shared address space, or the global address space by specifying an address within the unified address space. The address mapping unit 1356 may be used to translate addresses in the unified address space into different memory addresses that can be accessed by the load / store unit 1366.

[0172] The register file 1358 provides a collection of registers for the functional units of the graphics multiprocessor 1334. The register file 1358 provides temporary storage for operands of the data paths of the functional units (e.g., GPGPU core 1362, load / store unit 1366) connected to the graphics multiprocessor 1334. The register file 1358 can be divided among each of the functional units so that each functional unit is allocated a dedicated portion of the register file 1358. For example, the register file 1358 can be divided among different groups of units executed by the graphics multiprocessor 1334.

[0173] Each GPGPU core 1362 may include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing instructions for the graphics multiprocessor 1334. In some implementations, the GPGPU core 1362 may include hardware logic that may otherwise reside within the tensor core and / or ray tracing core 1363. The GPGPU cores 1362 may be architecturally similar or architecturally distinct. For example, and in some examples, a first portion of the GPGPU core 1362 includes a single-precision FPU and integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. Optionally, the FPU may implement the IEEE 754-2008 standard for floating-point arithmetic or enable variable-precision floating-point arithmetic. The graphics multiprocessor 1334 may additionally include one or more fixed-function or special-function units for performing specific functions, such as copying rectangles or pixel blending operations. One or more of the GPGPU cores may also include fixed-function or special-function logic.

[0174] GPGPU core 1362 may include SIMD logic capable of executing a single instruction on multiple sets of data. Optionally, GPGPU core 1362 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. SIMD instructions for the GPGPU core can be generated by a shader compiler at compile time, or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. Multiple threads of a program configured for a SIMT execution model can be executed via a single SIMD instruction. For example, and in some examples, eight SIMT threads can be executed in parallel via a single SIMD8 logic unit, and these eight SIMT threads perform the same or similar operations.

[0175] The memory and cache interconnect 1368 is an interconnect network that connects each of the functional units in the graphics multiprocessor 1334 to the register file 1358 and to the shared memory 1370. For example, the memory and cache interconnect 1368 is a crossbar interconnect that allows the load / store unit 1366 to perform load and store operations between the shared memory 1370 and the register file 1358. The register file 1358 can operate at the same frequency as the GPGPU core 1362, so data transfers between the GPGPU core 1362 and the register file 1358 are very low-latency. Shared memory 1370 can be used to enable communication between threads executing on the functional units of the graphics multiprocessor 1334. Cache memory 1372 can be used as a data cache, for example, to cache texture data transferred between the functional units and the texture unit 1336. Shared memory 1370 can also be used as a managed cached program. Shared memory 1370 and cache memory 1372 can be coupled to the data crossbar 1340 to enable communication with other components of the processing cluster. In addition to automatically cached data stored in cache memory 1372, threads executing on GPGPU core 1362 can also programmatically store data in shared memory.

[0176] Figures 14A-14C An additional graphics multiprocessor is illustrated according to an example. Figures 14A-14B Graphics multiprocessors 1425 and 1450 are shown in the figure. Graphics multiprocessors 1425 and 1450 are shown in the figure. Figure 13C 1425, 1450. Figure 14C A graphics processing unit (GPU) 1480 is shown that includes a dedicated set of graphics processing resources arranged into multi-core groups 1465A-1465N that correspond to graphics multiprocessors 1425, 1450. The illustrated graphics multiprocessors 1425, 1450 and multi-core groups 1465A-1465N may be streaming multiprocessors (SMs) capable of executing a large number of execution threads simultaneously.

[0177] Figure 14A The graphics multiprocessor 1425 includes relative Figure 13DThe graphics multiprocessor 1425 may include multiple additional instances of execution resource units of the graphics multiprocessor 1434. For example, the graphics multiprocessor 1425 may include multiple instances of instruction units 1432A-1432B, register files 1434A-1434B, and texture unit(s) 1444A-1444B. The graphics multiprocessor 1425 also includes multiple sets of graphics or compute execution units (e.g., GPGPU cores 1436A-1436B, tensor cores 1437A-1437B, ray tracing cores 1438A-1438B) and multiple sets of load / store units 1440A-1440B. The execution resource units have a common instruction cache 1430, texture and / or data cache memory 1442, and shared memory 1446.

[0178] The various components can communicate via an interconnect fabric 1427. Interconnect fabric 1427 may include one or more crossbar switches to facilitate communication between the various components of graphics multiprocessor 1425. Interconnect fabric 1427 may be a separate high-speed network fabric layer upon which each component of graphics multiprocessor 1425 is stacked. Components of graphics multiprocessor 1425 communicate with remote components via interconnect fabric 1427. For example, cores 1436A-1436B, 1437A-1437B, and 1438A-1438B can each communicate with shared memory 1446 via interconnect fabric 1427. Interconnect fabric 1427 may arbitrate communications within graphics multiprocessor 1425 to ensure fair bandwidth distribution between components.

[0179] Figure 14B The graphics multiprocessor 1450 includes a plurality of execution resource sets 1456A-1456D, wherein Figure 13D and Figure 14A As shown in FIG, each set of execution resources includes multiple instruction units, register files, GPGPU cores, and load-store units. Execution resources 1456A-1456D can work in conjunction with (one or more) texture units 1460A-1460D for texture operations while sharing instruction cache 1454 and shared memory 1453. For example, execution resources 1456A-1456D can share instruction cache 1454 and shared memory 1453, as well as multiple instances of texture and / or data cache memory 1458A-1458B. Each component can be connected to the CPU via a processor similar to Figure 14A The interconnect fabric 1427 communicates with the interconnect fabric 1452 .

[0180] Those skilled in the art will understand that Figure 1 、 Figures 13A-13D as well as Figures 14A-14BThe architecture described in is illustrative and not limiting with respect to the scope of the present examples. Thus, the techniques described herein may be implemented on any appropriately configured processing unit without departing from the scope of the examples described herein, including but not limited to one or more mobile application processors, one or more desktop or server central processing units (CPUs) including multi-core CPUs, one or more parallel processing units (such as Figure 13A parallel processing unit 1302), and one or more graphics processors or special processing units.

[0181] The parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU can be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink or other known, standardized, or proprietary protocols). In other examples, the GPU can be integrated on the same package or chip as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., inside the package or chip). Regardless of the manner in which the GPU is connected, the processor core can assign work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0182] Figure 14C A graphics processing unit (GPU) 1480 is shown, which includes a collection of dedicated graphics processing resources arranged into multi-core groups 1465A-1465N. Although details are provided for only a single multi-core group 1465A, it will be understood that other multi-core groups 1465B-1465N may be equipped with the same or similar collections of graphics processing resources. The details described with respect to multi-core groups 1465A-1465N may also apply to any of the graphics multiprocessors 1334, 1425, 1450 described herein.

[0183] As shown, multi-core group 1465A may include a set of graphics cores 1470, a set of tensor cores 1471, and a set of ray tracing cores 1472. Scheduler / dispatcher 1468 schedules and dispatches graphics threads for execution on the respective cores 1470, 1471, 1472. A set of register files 1469 stores operand values ​​used by cores 1470, 1471, 1472 when executing graphics threads. These register files may include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and slice registers for storing tensor / matrix values. Slice registers may be implemented as a combined set of vector registers.

[0184] One or more combined first level (L1) caches and shared memory units 1473 store graphics data locally within each multi-core group 1465A, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. One or more texture units 1474 can also be used to perform texture operations, such as texture mapping and sampling. A second level (Level 2, L2) cache 1475, shared by all multi-core groups 1465A-1465N or a subset of multi-core groups 1465A-1465N, stores graphics data and / or instructions for multiple concurrent graphics threads. As shown, the L2 cache 1475 can be shared across multiple multi-core groups 1465A-1465N. One or more memory controllers 1467 couple the GPU 1480 to memory 1466, which can be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0185] Input / output (I / O) circuitry 1463 couples GPU 1480 to one or more I / O devices 1462, such as digital signal processors (DSPs), network controllers, or user input devices. On-chip interconnects may be used to couple I / O devices 1462 to GPU 1480 and memory 1466. One or more I / O memory management units (IOMMUs) 1464 of I / O circuitry 1463 directly couple I / O devices 1462 to system memory 1466. Optionally, IOMMUs 1464 manage multiple sets of page tables used to map virtual addresses to physical addresses in system memory 1466. Consequently, I / O devices 1462, CPU(s) 1461, and GPU(s) 1480 may share the same virtual address space.

[0186] In one implementation of the IOMMU 1464, the IOMMU 1464 supports virtualization. In this case, the IOMMU 1464 can manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within system memory 1466). The base address of each of the first set of page tables and the second set of page tables can be stored in a control register and swapped out upon context switching (e.g., so that the new context is provided with access to the relevant set of page tables). Although not described in detail in the present disclosure, the first set of page tables and the second set of page tables can be used to map guest / graphics virtual addresses to guest / graphics physical addresses. Figure 14C , but each of the cores 1470, 1471, 1472 and / or multi-core groups 1465A-1465N may include a translation lookaside buffer (TLB) for caching guest virtual to guest physical translations, guest physical to host physical translations, and guest virtual to host physical translations.

[0187] (One or more) CPUs 1461, GPU 1480, and I / O devices 1462 can be integrated on a single semiconductor chip and / or chip package. The illustrated memory 1466 can be integrated on the same chip or can be coupled to a memory controller 1467 via an off-chip interface. In one implementation, memory 1466 includes GDDR6 memory that shares the same virtual address space as other physical system-level memory, but the basic principles described herein are not limited to this particular implementation.

[0188] Tensor Core 1471 may include multiple execution units specifically designed to perform matrix operations, which are fundamental computational operations for performing deep learning operations. For example, synchronized matrix multiplication operations may be used for neural network training and inference. Tensor Core 1471 may perform matrix processing using a variety of operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer words (16 bits), bytes (8 bits), and half bytes (4 bits). For example, a neural network implementation extracts features from each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.

[0189] In a deep learning implementation, parallel matrix multiplication work can be scheduled for execution on the Tensor Core 1471. Neural network training, in particular, requires a large number of matrix dot product operations. To handle the inner product formulation of an N×N×N matrix multiplication, the Tensor Core 1471 may include at least N dot product processing elements. Before the matrix multiplication begins, a complete matrix is ​​loaded into the slice register, and for each of N cycles, at least one column of the second matrix is ​​loaded. For each cycle, there are N dot products processed.

[0190] Depending on the specific implementation, matrix elements can be stored with different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for the Tensor Core 1471 to ensure that the most efficient precision is used for different workloads (e.g., such as inference workloads, which can tolerate quantization down to bytes and nibbles). Supported formats additionally include 64-bit floating point (FP64) and non-IEEE floating point formats such as the bfloat16 format (e.g., Brain floating point), a 16-bit floating point format with one sign bit, eight exponent bits, and eight significand bits (seven of which are explicitly stored). One example includes support for a reduced-precision Tensor Floating Point (TF32) mode that performs computations using the range of FP32 (8 bits) and the precision of FP16 (10 bits). Reduced-precision TF32 operations can be performed on FP32 inputs and produce FP32 outputs with higher performance relative to FP32 and increased precision relative to FP16. In some examples, one or more 8-bit floating point formats (FP8) are supported.

[0191] In some examples, the tensor core 1471 supports sparse operation modes for matrices in which the majority of values ​​are zero. The tensor core 1471 includes support for sparse input matrices encoded in sparse matrix representations (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). The tensor core 1471 also includes support for compressed sparse matrix representations where the sparse matrix representation can be further compressed. Compressed matrix data, encoded matrix data, and / or compressed and encoded matrix data and associated compression and / or encoding metadata can be read by the tensor core 1471, and non-zero values ​​can be extracted. For example, for a given input matrix A, non-zero values ​​can be loaded from a compressed and / or encoded representation of at least a portion of matrix A. Based on the position of the non-zero values ​​in matrix A (which can be determined from the index or coordinate metadata associated with the non-zero values), the corresponding values ​​in the input matrix B can be loaded. Depending on the operation to be performed (e.g., multiplication), loading values ​​from input matrix B can be bypassed if the corresponding value is a zero value. In some examples, the pairing of values ​​for certain operations (such as multiplication operations) can be pre-scanned by the scheduler logic, and only operations between non-zero inputs are scheduled. Depending on the dimensions of matrices A and B and the operation to be performed, the output matrix C can be dense or sparse. In the case where the output matrix C is sparse and depending on the configuration of the tensor core 1471, the output matrix C can be output in a compressed format, sparse coding, or compressed sparse coding.

[0192] The ray tracing core 1472 can accelerate ray tracing operations for both real-time ray tracing implementations and non-real-time ray tracing implementations. Specifically, the ray tracing core 1472 can include ray traversal / intersection circuitry for performing ray traversals and identifying intersections between rays and primitives enclosed within a bounding volume hierarchy (BVH) using a bounding volume hierarchy. The ray tracing core 1472 can also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, the ray tracing core 1472 performs traversal and intersection operations in conjunction with the image denoising techniques described herein, at least portions of which can be executed on the tensor core 1471. For example, the tensor core 1471 can implement a deep learning neural network to perform denoising on frames generated by the ray tracing core 1472. However, the CPU(s) 1461, graphics core 1470, and / or ray tracing core 1472 may also implement all or portions of the denoising and / or deep learning algorithms.

[0193] Furthermore, as described above, a distributed approach to noise reduction can be employed, wherein GPU 1480 is located in a computing device coupled to other computing devices via a network or high-speed interconnect. In this distributed approach, the interconnected computing devices can share neural network learning / training data to improve the speed at which the entire system learns to perform noise reduction for different types of image frames and / or different graphics applications.

[0194] The ray tracing cores 1472 can handle all BVH traversals and / or ray-primitive intersections, thereby freeing the graphics core 1470 from being overloaded with thousands of instructions for each ray. For example, each ray tracing core 1472 includes a first set of specialized circuits for performing bounding box tests (e.g., for traversal operations) and / or a second set of specialized circuits for performing ray-triangle intersection tests (e.g., intersecting rays that have already been traversed). Thus, for example, the multi-core group 1465A can simply start ray probing, and the ray tracing cores 1472 independently perform ray traversals and intersections and return hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. While the ray tracing cores 1472 perform traversals and intersection operations, the other cores 1470, 1471 are freed up to perform other graphics or compute work.

[0195] Optionally, each ray tracing core 1472 may include a traversal unit for performing BVH test operations and / or an intersection unit for performing ray-primitive intersection tests. The intersection unit generates a "hit," "no hit," or "multiple hits" response, which it provides to the appropriate thread. During traversal and intersection operations, execution resources of other cores (e.g., graphics core 1470 and tensor core 1471) are freed to perform other forms of graphics work.

[0196] In some examples described below, a hybrid rasterization / ray tracing approach is used in which the work is distributed between graphics core 1470 and ray tracing core 1472 .

[0197] The ray tracing core 1472 (and / or other cores 1470, 1471) may include hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, nearest hit shaders, any hit shaders, and miss shaders, which enable the assignment of a unique set of shaders and textures to each object. Another ray tracing platform that may be supported by the ray tracing core 1472, graphics core 1470, and tensor core 1471 is the Vulkan API (e.g., Vulkan version 1.1.85 and subsequent versions). However, it is noted that the basic principles described herein are not limited to any particular ray tracing ISA.

[0198] In general, each core 1472, 1471, 1470 may support a ray tracing instruction set including instructions / functions for one or more of the following: ray generation, nearest hit, any hit, ray-primitive intersection, per-primitive and hierarchy bounding box construction, misses, visits, and exceptions. More specifically, some examples include ray tracing instructions for performing one or more of the following functions: - Ray Generation - Ray generation instructions can be executed for each pixel, sample, or other user-defined work assignment. -Nearest Hit - A nearest hit instruction may be executed to locate the closest intersection of a ray with a primitive within the scene. -Any Hit - The Any Hit instruction identifies multiple intersections between rays and primitives within the scene, potentially identifying a new closest intersection point. -Intersect - The Intersect instruction performs a ray-primitive intersection test and outputs the result. - Per-primitive bounding box construction - This instruction builds a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure). -Miss - Indicates that the ray missed the scene or all geometry within the specified area of ​​the scene. -visit - Indicates the subvolumes that the ray will traverse. - Exceptions - includes various types of exception handlers (e.g., called for various error conditions).

[0199] In some examples, the ray tracing core 1472 may be adapted to accelerate general computational operations that may be accelerated using computational techniques similar to ray intersection testing. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives that perform general computational operations via the ray tracing core. Exemplary computational problems that may benefit from computational operations performed on the ray tracing core 1472 include computations involving the propagation of beams, waves, rays, or particles within a coordinate space. Interactions associated with that propagation may be computed relative to geometry or meshes within the coordinate space. For example, computations associated with the propagation of an electromagnetic signal through an environment may be accelerated using instructions or primitives that are executed via the ray tracing core. Refraction and reflection of the signal through objects in the environment may be computed as direct ray tracing simulations.

[0200] The ray tracing core 1472 can also be used to perform calculations that are not directly similar to ray tracing. For example, the ray tracing core 1472 can be used to accelerate mesh projection, mesh refinement, and volume sampling calculations. General coordinate space calculations, such as nearest neighbor calculations, can also be performed. For example, a set of points near a given point can be found by defining a bounding box around the point in coordinate space. The BVH and ray detection logic within the ray tracing core 1472 can then be used to determine the set of intersections of points within the bounding box. The intersections constitute the origin and the nearest neighbors of that origin. The calculations performed using the ray tracing core 1472 can be performed in parallel with the calculations performed on the graphics core 1472 and the tensor core 1471. The shader compiler can be configured to compile compute shaders or other general graphics processing programs into low-level primitives that can be parallelized across the graphics core 1470, the tensor core 1471, and the ray tracing core 1472.

[0201] Building increasingly larger silicon dies is challenging for a variety of reasons. As silicon dies get larger, manufacturing yields become smaller, and process technology requirements for different components can vary. On the other hand, for high-performance systems, key components should be interconnected via high-speed, high-bandwidth, low-latency interfaces. These conflicting requirements create challenges for high-performance chip development.

[0202] The examples described herein provide techniques for decomposing the architecture of a system-on-chip integrated circuit into multiple different chiplets that can be packaged onto a common substrate. In some examples, a graphics processing unit or parallel processor is composed of various silicon chiplets that are manufactured separately. A chiplet is an at least partially packaged integrated circuit that includes different logic units that can be assembled into a larger package with other chiplets. Various sets of chiplets with different IP core logic can be assembled into a single device. Additionally, the chiplets can be integrated into a base die or base chiplet using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP developed on different processes can be mixed. This avoids the complexity of converging multiple IPs onto the same process, especially for large SoCs with several flavors of IP.

[0203] Enabling the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. For customers, this means products that better suit their requirements in a cost-effective and timely manner. Additionally, separate IP can be easily modified to be independently power-gated, allowing components not in use for a given workload to be shut down, reducing overall power consumption.

[0204] Figure 15 A parallel computing system 1500 according to some examples is shown. In some examples, parallel computing system 1500 includes a parallel processor 1520, which may be a graphics processor or a computing accelerator as described herein. Parallel processor 1520 includes a global logic unit 1501, an interface 1502, a thread dispatcher 1503, a media unit 1504, a set of compute units 1505A-1505H, and a cache / memory unit 1506. In some examples, global logic unit 1501 includes global functionality for parallel processor 1520, including device configuration registers, a global scheduler, power management logic, and the like. Interface 1502 may include a front-end interface for parallel processor 1520. Thread dispatcher 1503 may receive a workload from interface 1502 and dispatch threads of the workload to compute units 1505A-1505H. If the workload includes any media operations, at least a portion of those operations may be performed by media unit 1504. The media unit may also migrate some operations to the compute units 1505A-1505H. The cache / memory unit 1506 may include cache memory (eg, L3 cache) and local memory (eg, HBM, GDDR) for the parallel processors 1520.

[0205] Figures 16A-16BIllustrate a hybrid logical / physical view of a split parallel processor according to examples described herein. Figure 16A A split parallel computing system 1600 is shown. Figure 16B A chiplet 1630 of the split parallel computing system 1600 is shown.

[0206] like Figure 16A As shown in , a split computing system 1600 may include a parallel processor 1620, wherein the components of the parallel processor SOC are distributed across multiple chiplets. Each chiplet may be a different IP core that is independently designed and configured to communicate with other chiplets via one or more common interfaces. Chiplets include, but are not limited to, a compute chiplet 1605, a media chiplet 1604, and a memory chiplet 1606. Each chiplet may be manufactured separately using a different process technology. For example, the compute chiplet 1605 may be manufactured using the smallest or most advanced process technology available at the time of manufacturing, while the memory chiplet 1606 or other chiplets (e.g., I / O, networking, etc.) may be manufactured using a larger or less advanced process technology.

[0207] The chiplets can be bonded to a base die 1610 and configured to communicate with each other and logic within the base die 1610 via an interconnect layer 1612. In some examples, the base die 1610 can include global logic 1601, which can include a scheduler 1611 and power management 1621 logic unit, an interface 1602, a dispatch unit 1603, and an interconnect fabric module 1608 coupled or integrated with one or more L3 cache banks 1609A-1609N. The interconnect fabric 1608 can be an inter-chiplet structure integrated into the base die 1610. The logic chiplets can use the fabric 1608 to relay messages between the chiplets. Additionally, the L3 cache blocks 1609A-1609N in the base die and / or the L3 cache blocks within the memory chiplet 1606 can cache data that is read from and transferred to the DRAM chiplet within the memory chiplet 1606 and transferred to the host's system memory.

[0208] In some examples, global logic 1601 is a microcontroller that can execute firmware to perform scheduler 1611 and power management 1621 functions for parallel processor 1620. The microcontroller that executes the global logic can be customized for the target use case of parallel processor 1620. Scheduler 1611 can perform global scheduling operations for parallel processor 1620. Power management 1621 functions can be used to enable or disable individual chiplets within the parallel processor when those chiplets are not in use.

[0209] The individual chiplets of the parallel processor 1620 can be designed to perform specific functions that would be integrated into a single die in existing designs. The collection of compute chiplets 1605 can include a cluster of compute units (e.g., execution units, streaming multiprocessors, etc.) that include programmable logic for executing compute or graphics shader instructions. The media chiplet 1604 can include hardware logic for accelerating media encoding and decoding operations. The memory chiplet 1606 can include volatile memory (e.g., DRAM) and one or more SRAM cache memory blocks (e.g., L3 blocks).

[0210] like Figure 16B As shown in FIG, each chiplet 1630 may include common components and specialized components. Chiplet logic 1636 within chiplet 1630 may include chiplet-specific components, such as a streaming multiprocessor, array of compute units, or execution units described herein. Chiplet logic 1636 may be coupled with an optional cache or shared local memory 1638, or a cache or shared local memory may be included within chiplet logic 1636. Chiplet 1630 may include a fabric interconnect node 1642 that receives commands via an inter-chiplet fabric. Commands and data received via fabric interconnect node 1642 may be temporarily stored in interconnect buffer 1639. Data transmitted to and received from fabric interconnect node 1642 may be stored in interconnect buffer 1640. Power control 1632 and clock control 1634 logic may also be included within the chiplet. The power control 1632 and clock control 1634 logic may receive configuration commands via the fabric and may configure dynamic voltage and frequency scaling for chiplet 1630. In some examples, each chiplet can have independent clock and power domains and can be clock-gated and power-gated independently of other chiplets.

[0211] At least some of the components within the illustrated chiplet 1630 may also be included in a Figure 16A 1610. For example, logic within a base die that communicates with the fabric may include some version of fabric interconnect nodes 1642. Base die logic that may be independently clock gated or power gated may include some version of power control 1632 and / or clock control 1634 logic.

[0212] Thus, although various examples described herein use the term SOC to describe a device or system having a processor and associated circuitry (e.g., input / output (“I / O”) circuitry, power delivery circuitry, memory circuitry, etc.) monolithically integrated into a single integrated circuit (“IC”) die or chip, the present disclosure is not limited in this respect. For example, in various examples of the present disclosure, a device or system may have one or more processors (e.g., one or more processor cores) and associated circuitry (e.g., input / output (“I / O”) circuitry, power delivery circuitry, memory circuitry, etc.) arranged in a discrete collection of discrete dies, slices, and / or chiplets (e.g., one or more discrete processor core dies arranged adjacent to one or more other dies (such as memory dies, I / O dies, etc.)). In such discrete devices and systems, the various dies, slices, and / or chiplets may be physically and / or electrically coupled together by a packaging structure including, for example, various packaging substrates, interposers, active interposers, photonic interposers, interconnect bridges, and the like. A separate collection of discrete dies, slices, and / or chiplets may also be part of a System-on-Package (“SoP”).

[0213] Example core architectures—in-order and out-of-order core block diagrams.

[0214] Figure 17A The block diagram of illustrates both an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to some examples. Figure 17B The block diagram illustrates both an example in-order architecture core and an example register renaming, out-of-order issue / execution architecture core to be included in a processor, according to an example. Figures 17A-17B The solid-line boxes in illustrate the in-order pipeline and in-order core, while the optionally added dashed-line boxes illustrate the register renaming, out-of-order issue / execution pipeline, and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0215] In FIG17(A), a processor pipeline 1700 includes a fetch stage 1702, an optional length decode stage 1704, a decode stage 1706, an optional allocate (Alloc) stage 1708, an optional rename stage 1710, a dispatch (also known as dispatch or issue) stage 1712, an optional register read / memory read stage 1714, an execute stage 1716, a writeback / memory write stage 1718, an optional exception handling stage 1722, and an optional commit stage 1724. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 1702, one or more instructions may be fetched from an instruction memory, and during the decode stage 1706, the fetched one or more instructions may be decoded, an address using a forwarding register port (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., an immediate offset or a link register (LR)) may be performed. In some examples, decode stage 1706 and register read / memory read stage 1714 can be combined into one pipeline stage. In some examples, during execute stage 1716, decoded instructions can be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface can be performed, multiplication and addition operations can be performed, arithmetic operations with branch results can be performed, and the like.

[0216] As an example, Figure 17B An example register renaming, out-of-order issue / execution architecture core may implement pipeline 1700 in the following manner: 1) instruction fetch circuitry 1738 performs fetch and length decode stages 1702 and 1704; 2) decode circuitry 1740 performs decode stage 1706; 3) rename / allocator unit circuitry 1752 performs allocate stage 1708 and rename stage 1710; 4) (one or more) scheduler circuitry 1756 performs schedule stage 1712; 5) (one or more) physical register file circuitry 1758 and memory unit circuitry 1770 performs register read / memory read stage 1714; (one or more) execution clusters 1760 perform execute stage 1716; 6) memory unit circuitry 1770 and (one or more) physical register file circuitry 1758 perform write back / memory write stage 1718; 7) various circuits may be involved in exception handling stage 1722; and 8) retirement unit circuitry 1754 and (one or more) physical register file circuitry 1758 perform commit stage 1724.

[0217] Figure 17BProcessor core 1790 is shown to include front-end unit circuitry 1730 coupled to execution engine unit circuitry 1750, and both are coupled to memory unit circuitry 1770. Core 1790 can be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, core 1790 can be a specialized core, such as a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, or the like.

[0218] Front-end unit circuitry 1730 may include branch prediction circuitry 1732 coupled to instruction cache circuitry 1734, coupled to instruction translation lookaside buffer (TLB) 1736, coupled to instruction fetch circuitry 1738, coupled to decode circuitry 1740. In some examples, instruction cache circuitry 1734 is included in memory unit circuitry 1770 rather than in front-end circuitry 1730. Decode circuitry 1740 (or decoder) may decode an instruction and generate as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals that are decoded from, or otherwise reflect, or are derived from, the original instruction. Decode circuitry 1740 may also include address generation unit (AGU) circuitry (not shown). In some examples, the AGU uses the forwarded register ports to generate LSU addresses and can further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). Decode circuitry 1740 can be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read-only memories (ROMs), and the like. In some examples, core 1790 includes a microcode ROM (not shown) or other medium that stores microcode for certain macroinstructions (e.g., in decode circuitry 1740 or otherwise within front-end circuitry 1730). In some examples, decode circuitry 1740 includes a micro-op or operation cache (not shown) to store / cache decoded operations, micro-tags, or micro-ops generated during decode or other stages of processor pipeline 1700. Decode circuitry 1740 can be coupled to rename / allocator unit circuitry 1752 in execution engine circuitry 1750.

[0219] The execution engine circuitry 1750 includes a rename / allocator unit circuitry 1752 coupled to a retirement unit circuitry 1754 and a set of one or more scheduler circuits 1756. The scheduler circuitry 1756 represents any number of different schedulers, including reservation stations, central instruction windows, and the like. In some examples, the scheduler circuitry 1756 may include an arithmetic logic unit (ALU) scheduler / scheduling circuitry, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuitry, an AGU queue, and the like. The scheduler circuitry 1756 is coupled to the physical register file(s) circuitry 1758. Each of the physical register file(s) circuitry 1758 represents one or more physical register files, different ones of which store one or more different data types, such as scalar integers, scalar floating point, packed integers, packed floating point, vector integers, vector floating point, state (e.g., an instruction pointer, i.e., the address of the next instruction to be executed), and the like. In some examples, the physical register file(s) circuitry 1758 includes vector register unit circuitry, write mask register unit circuitry, and scalar register unit circuitry. These register units can provide architectural vector registers, vector mask registers, general purpose registers, and the like. The physical register file(s) circuitry 1758 is coupled to the retirement unit circuitry 1754 (also referred to as a retirement queue) to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., utilizing reorder buffer(s) (ROBs) and retirement register file(s); utilizing future stack(s), history buffer(s), and retirement register file(s); utilizing register maps and register pools; and the like). The retirement unit circuitry 1754 and the physical register file(s) circuitry 1758 are coupled to the execution cluster(s) 1760. The execution cluster(s) 1760 includes a set of one or more execution unit circuitry 1762 and a set of one or more memory access circuitry 1764. The execution unit circuit(s) 1762 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shifts, additions, subtractions, multiplications) on various types of data (e.g., scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point). While some examples may include several execution units or execution unit circuits dedicated to a particular function or set of functions, other examples may include only one execution unit circuit or multiple execution units / execution unit circuits that all perform all functions.Scheduler circuit(s) 1756, physical register file(s) circuit(s) 1758, and execution cluster(s) 1760 are shown as potentially multiple because some examples create separate pipelines for certain types of data / operations (e.g., a scalar integer pipeline, a scalar floating point / packed integer / packed floating point / vector integer / vector floating point pipeline, and / or a memory access pipeline, each with its own scheduler circuit(s), physical register file(s) circuit(s), and / or execution cluster(s)—and in the case of a separate memory access pipeline, in some examples only the execution cluster of that pipeline has memory access unit(s) circuit(s) 1764). It should also be understood that where separate pipelines are used, one or more of these pipelines may be out-of-order issue / execution, while the rest are in-order.

[0220] In some examples, the execution engine unit circuit 1750 can perform load store unit (LSU) address / data pipelining to an advanced microcontroller bus (AMB) interface (not shown), as well as address stage and write back, data stage loads, stores, and branches.

[0221] A set of memory access circuits 1764 is coupled to memory cell circuitry 1770, which includes a data TLB circuit 1772, which is coupled to a data cache circuit 1774, which is coupled to a level 2 (L2) cache circuit 1776. In some examples, memory access circuitry 1764 may include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 1772 in memory cell circuitry 1770. Instruction cache circuitry 1734 is further coupled to level 2 (L2) cache circuitry 1776 in memory cell circuitry 1770. In some examples, instruction cache 1734 and data cache 1774 are combined into an L2 cache circuit 1776, a level 3 (L3) cache circuit (not shown), and / or a single instruction and data cache in main memory (not shown). L2 cache circuitry 1776 is coupled to one or more other levels of cache and ultimately to main memory.

[0222] Core 1790 can support one or more instruction sets (e.g., the x86 instruction set architecture (optionally with some extensions that have been added with newer versions); the MIPS instruction set architecture; the ARM instruction set architecture (optionally with optional additional extensions, such as NEON)) that include the instruction(s) described herein. In some examples, core 1790 includes logic to support packed data instruction set architecture extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data. Example execution unit circuit(s).

[0223] Figure 18 An example of an execution unit circuit (s) is shown, such as Figure 17B Execution unit(s) circuitry 1762. As shown, execution unit(s) circuitry 1762 may include one or more ALU circuits 1801, optional vector / single instruction multiple data (SIMD) circuitry 1803, load / store circuitry 1805, branch / jump circuitry 1807, and / or floating-point unit (FPU) circuitry 1809. ALU circuitry 1801 performs integer arithmetic and / or Boolean operations. Vector / SIMD circuitry 1803 performs vector / SIMD operations on packed data (e.g., SIMD / vector registers). Load / store circuitry 1805 executes load and store instructions to load data from memory into registers, or store data from registers into memory. Load / store circuitry 1805 may also generate addresses. Branch / jump circuitry 1807 causes a branch or jump to a memory address, depending on the instruction. FPU circuitry 1809 performs floating-point arithmetic. The width of execution unit circuit(s) 1762 varies depending on the example and can range, for example, from 16 bits to 1024 bits. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).

[0224] Example register architecture.

[0225] Figure 191 is a block diagram of a register architecture 1900 according to some examples. As shown, register architecture 1900 includes vector / SIMD registers 1910, whose widths vary from 128 bits to 1024 bits. In some examples, vector / SIMD registers 1910 are physically 512 bits, and depending on the mapping, only some of the low-order bits are used. For example, in some examples, vector / SIMD registers 1910 are 512-bit ZMM registers: the low-order 256 bits are used for YMM registers, and the low-order 128 bits are used for XMM registers. Thus, there is a register overlap. In some examples, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length. Scalar operations are performed on the lowest-order data element position in a ZMM / YMM / XMM register; higher-order data element positions are either maintained as they were before the instruction or reset to zero, depending on the example.

[0226] In some examples, the register architecture 1900 includes write mask / predicate registers 1915. For example, in some examples, there are 8 write mask / predicate registers (sometimes referred to as k0 to k7), each of which is 16 bits, 32 bits, 64 bits, or 128 bits in size. The write mask / predicate registers 1915 can allow merging (e.g., allowing any set of elements in the destination to be protected from updates during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask allows any set of elements in the destination to be zeroed during the execution of any operation). In some examples, each data element position in a given write mask / predicate register 1915 corresponds to a data element position in the destination. In other examples, the write mask / predicate registers 1915 are scalable and consist of a set number of enable bits for a given vector element (e.g., 8 enable bits for each 64-bit vector element).

[0227] The register architecture 1900 includes a plurality of general purpose registers 1925. These registers can be 16 bits, 32 bits, 64 bits, etc., and can be used for scalar operations. In some examples, these registers are referred to by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 to R15.

[0228] In some examples, register architecture 1900 includes a scalar floating-point (FP) register file 1945, which is used to perform scalar floating-point operations on 32 / 64 / 80-bit floating-point data using the x87 instruction set architecture extension, or as an MMX register to perform operations on 64-bit packed integer data, and to store operands for some operations performed between MMX and XMM registers.

[0229] One or more flag registers 1940 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 1940 can store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some examples, one or more flag registers 1940 are referred to as program status and control registers.

[0230] Segment registers 1920 contain segment points for accessing memory. In some examples, these registers are referred to by the names CS, DS, SS, ES, FS, and GS.

[0231] Model-specific registers or machine-specific registers (MSRs) 1935 control and report processor performance. Most MSRs 1935 handle system-related functions and are not accessible to applications. For example, MSRs may provide control of one or more of the following: performance monitoring counters, debug extensions, memory type range registers, thermal and power management, instruction-specific support, and / or processor feature / mode support. Machine check registers 1960 consist of control, status, and error reporting MSRs that are used to detect and report hardware errors. (One or more) control registers 1955 (e.g., CR0-CR4) determine the operating mode of the processor (e.g., processors 1070, 1080, 1038, 1015, and / or 1100) and the characteristics of the currently executing task. In some examples, MSRs 1935 are a subset of control registers 1955.

[0232] One or more instruction pointer registers 1930 store instruction pointer values. Debug registers 1950 control and enable monitoring of debug operations of the processor or core.

[0233] Memory (mem) management registers 1965 specify the location of data structures used in protected mode memory management. These registers may include the global descriptor table register (GDTR), the interrupt descriptor table register (IDTR), the task register, and the local descriptor table register (LDTR).

[0234] Alternative examples may use wider or narrower registers. In addition, alternative examples may use more, fewer, or different register files and registers. Register architecture 1900 may be used, for example, in register file / memory 608 or in physical register file circuit 1758.

[0235] Instruction set architecture.

[0236] An instruction set architecture (ISA) may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, bit positions) to specify the operation to be performed (e.g., opcode) and the operand(s) on which the operation is to be performed and / or other data field(s) (e.g., mask), etc. Some instruction formats are further decomposed through the definition of instruction templates (or subformats). For example, instruction templates for a given instruction format may be defined to have different subsets of the fields of the instruction format (the included fields are generally in the same order, but at least some have different bit positions because fewer fields are included) and / or to have given fields interpreted differently. Thus, each instruction of an ISA is expressed using a given instruction format (and, if defined, expressed using a given instruction template from the instruction format) and includes fields for specifying the operation and the operand. For example, an example ADD instruction has a specific opcode and instruction format, the instruction format including an opcode field to specify the opcode and an operand field to select the operand (source 1 / destination and source 2); and the appearance of this ADD instruction in an instruction stream will have specific content in the operand field that selects the specific operand. In addition, although the following description is made in the context of the x86 ISA, it is within the knowledge of those skilled in the art to apply the teachings of this disclosure to other ISAs.

[0237] Example instruction format.

[0238] The examples of the instruction(s) described herein may be implemented in different formats. In addition, example systems, architectures, and pipelines are described in detail below. The examples of the instruction(s) may be executed on these systems, architectures, and pipelines, but are not limited to those described in detail.

[0239] Figure 20 An example of an instruction format is illustrated. As shown, an instruction may include multiple components, including, but not limited to, one or more fields for: one or more prefixes 2001, an opcode 2003, addressing information 2005 (e.g., register identifiers, memory addressing information, etc.), a displacement value 2007, and / or an immediate value 2009. Note that some instructions utilize some or all of the fields of the format, while other instructions may only use the opcode 2003 field. In some examples, the order illustrated is the order in which these fields are to be encoded, however, it should be understood that in other examples, these fields may be encoded in another order, combined, etc.

[0240] (One or more) prefix fields 2001 are modified when being used to instruction.In some examples, one or more prefixes are used to repeat string instruction (for example, 0xF0, 0xF2, 0xF3, etc.), provide section override (sectionoverride) (for example, 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), perform bus lock operation and / or change operation object (for example, 0x66) and address size (for example, 0x67).Some instruction requires mandatory prefix (for example, 0x66, 0xF2, 0xF3, etc.).Some in these prefixes can be considered to be " tradition " prefix.Other prefixes (one or more examples of which are described in detail herein) indicate and / or provide further ability, for example, specify specific register etc..These other prefixes usually follow after the " tradition " prefix.

[0241] The opcode field 2003 is used to at least partially define the operation to be performed when the instruction is decoded. In some examples, the length of the primary opcode encoded in the opcode field 2003 is one, two, or three bytes. In other examples, the primary opcode can be of other lengths. The additional 3-bit opcode field is sometimes encoded in another field.

[0242] The addressing information field 2005 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 21An example of an addressing information field 2005 is illustrated. In this illustration, an optional MOD R / M byte 2102 and an optional Scale, Index, Base (SIB) byte 2104 are shown. The MOD R / M byte 2102 and SIB byte 2104 are used to encode up to two operands of the instruction, each operand being a direct register or an effective memory address. Note that both of these fields are optional, i.e., not all instructions include one or more of these fields. The MOD R / M byte 2102 includes a MOD field 2142, a register (reg) field 2144, and an R / M field 2146.

[0243] The content of the MOD field 2142 distinguishes between memory access and non-memory access modes. In some examples, when the MOD field 2142 has a binary value of 11 (11b), register direct addressing mode is utilized, otherwise register indirect addressing mode is used.

[0244] Register field 2144 may encode a destination register operand or a source register operand, or may encode an opcode extension and not be used to encode any instruction operand. The contents of register field 2144 specify the location (in a register or in memory) of the source or destination operand directly or through address generation. In some examples, register field 2144 is supplemented with additional bits from a prefix (e.g., prefix 2001) to allow for greater addressing.

[0245] The R / M field 2146 may be used to encode an instruction operand that references a memory address, or may be used to encode a destination register operand or a source register operand. Note that in some examples, the R / M field 2146 may be combined with the MOD field 2142 to specify an addressing mode.

[0246] The SIB byte 2104 includes a scale field 2152, an index field 2154, and a base address field 2156 for address generation. The scale field 2152 indicates the scale factor. The index field 2154 specifies the index register to be used. In some examples, the index field 2154 is supplemented with extra bits from the prefix (e.g., prefix 2001) to allow for larger addressing. The base address field 2156 specifies the base address register to be used. In some examples, the base address field 2156 is supplemented with extra bits from the prefix (e.g., prefix 2001) to allow for larger addressing. In practice, the contents of the scale field 2152 allow the contents of the index field 2154 to be scaled for memory address generation (e.g., for memory addresses using 2 缩放 *Address generation of index+base address).

[0247] Some addressing forms use displacement values ​​to generate memory addresses. For example, you can use 2 缩放 Memory addresses are generated by using the addressing information field 2005 as a memory address. The memory address is ...

[0248] In some examples, the immediate value field 2009 specifies an immediate value for the instruction. The immediate value can be encoded as a 1-byte value, a 2-byte value, a 4-byte value, etc.

[0249] Figure 22 An example of a first prefix 2001(A) is shown. In some examples, the first prefix 2001(A) is an example of a REX prefix. Instructions using this prefix can specify general registers, 64-bit packed data registers (e.g., single instruction multiple data (SIMD) registers or vector registers), and / or control registers and debug registers (e.g., CR8-CR15 and DR8-DR15).

[0250] Instructions using the first prefix 2001(A) can specify up to three registers using a 3-bit field, depending on the format: 1) using the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102; 2) using the MOD R / M byte 2102 with the SIB byte 2104, including using the reg field 2144 as well as the base field 2156 and the index field 2154; or 3) using the register field of the opcode.

[0251] In the first prefix 2001(A), bit positions 7:4 of the payload byte are set to 0100. Bit position 3(W) can be used to determine the operand size, but cannot determine the operand width alone. Thus, when W=0, the operand size is determined by the code segment descriptor (CS.D), and when W=1, the operand size is 64 bits.

[0252] Note that adding another bit allows for 16(2 4 ) registers, while the separate MOD R / M reg field 2144 and the R / M field 2146 of MOD R / M can only address 8 registers each.

[0253] In the first prefix 2001(A), bit position 2 (R) may be an extension of the MOD R / M's reg field 2144 and may be used to modify the MOD R / M's reg field 2144 when that field encodes a general register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the MOD R / M byte 2102 specifies another register or defines an extended opcode, R is ignored.

[0254] Bit position 1 (X) may modify the SIB Byte Index field 2154 .

[0255] Bit position 0 (B) may modify the base address in the R / M field 2146 or SIB byte base address field 2156 of MOD R / M; or it may modify the opcode register field used to access a general register (eg, general register 1925).

[0256] Figures 23A-23D An example of how the R, X, and B fields of the first prefix 2001(A) are used is illustrated. Figure 23A It is illustrated that when the SIB byte 2104 is not used for memory addressing, the R and B from the first prefix 2001 (A) are used to extend the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102. Figure 23B It is illustrated that when the SIB byte 2104 is not used, the R and B from the first prefix 2001 (A) are used to extend the reg field 2144 and the R / M field 2146 of the MOD R / M byte 2102 (register-register addressing). Figure 23C It illustrates that when the SIB byte 2104 is used for memory addressing, the R, X, and B from the first prefix 2001 (A) are used to extend the reg field 2144 as well as the index field 2154 and the base field 2156 of the MOD R / M byte 2102. Figure 23D It is illustrated that when a register is encoded in the opcode 2003 , the B from the first prefix 2001 (A) is used to extend the reg field 2144 of the MODR / M byte 2102 .

[0257] Figures 24A-24BAn example of a second prefix 2001(B) is illustrated. In some examples, the second prefix 2001(B) is an example of a VEX prefix. The second prefix 2001(B) encoding allows instructions to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 1910) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 2001(B) provides a syntax for three operands (or more). For example, a previous two-operand instruction performs an operation such as A=A+B, which overwrites the source operand. The use of the second prefix 2001(B) enables operands to perform non-destructive operations, such as A=B+C.

[0258] In some examples, the second prefix 2001(B) has two forms: a two-byte form and a three-byte form. The two-byte second prefix 2001(B) is primarily used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 2001(B) provides a compact alternative to 3-byte opcode instructions and the first prefix 2001(A).

[0259] Figure 24A An example of a two-byte form of the second prefix 2001 (B) is illustrated. In some examples, the format field 2401 (byte 0 2403) contains the value C5H. In some examples, byte 1 2405 includes an "R" value in bit [7]. This value is the complement of the "R" value of the first prefix 2001 (A). Bit [2] is used to specify the length (L) of the vector (wherein a value of 0 is a scalar or 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used to: 1) encode the first source register operand, which is specified in inverted (one's complement) form and is valid for instructions with 2 or more source operands; 2) encode the destination register operand, which is specified in one's complement form for certain vector shifts; or 3) not encode any operand, in which case the field is reserved and should contain some value, such as 1111b.

[0260] Instructions using this prefix may use the R / M field 2146 of MOD R / M to encode an instruction operand that references a memory address, or to encode a destination register operand or a source register operand.

[0261] Instructions using this prefix may use the reg field 2144 of MOD R / M to encode a destination register operand or a source register operand, or be treated as an opcode extension and not used to encode any instruction operand.

[0262] For instruction syntax supporting four operands, vvvv, the R / M field of MOD R / M 2146, and the reg field of MOD R / M 2144 encode three of the four operands. Bits [7:4] of the immediate value field 2009 are then used to encode the third source register operand.

[0263] Figure 24B An example of a three-byte form of the second prefix 2001(B) is shown. In some examples, the format field 2411 (byte 0 2413) contains the value C4H. Byte 1 2415 includes "R," "X," and "B" in bits [7:5], which are the complements of these values ​​of the first prefix 2001(A). Bits [4:0] (shown as mmmmm) of byte 1 2415 include content that encodes one or more implicit leading opcode bytes as needed. For example, 00001 means a leading opcode of 0FH, 00010 means a leading opcode of 0F38H, 00011 means a leading opcode of 0F3AH, and so on.

[0264] Bit [7] of byte 2 2417 is used similarly to W of the first prefix 2001 (A), including helping determine the size of operands that can be promoted. Bit [2] is used to specify the length (L) of the vector (where a value of 0 is a scalar or 128-bit vector, and a value of 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used to: 1) encode the first source register operand, which is specified in inverted (one's complement) form, valid for instructions with 2 or more source operands; 2) encode the destination register operand, which is specified in one's complement form, for certain vector shifts; or 3) not encode any operand, in which case the field is reserved and should contain a value such as 1111b.

[0265] Instructions using this prefix may use the R / M field 2146 of MOD R / M to encode an instruction operand that references a memory address, or to encode a destination register operand or a source register operand.

[0266] Instructions using this prefix may use the reg field 2144 of MOD R / M to encode a destination register operand or a source register operand, or be treated as an opcode extension and not used to encode any instruction operand.

[0267] For instruction syntax supporting four operands, vvvv, the R / M field of MOD R / M 2146, and the reg field of MOD R / M 2144 encode three of the four operands. Bits [7:4] of the immediate value field 2009 are then used to encode the third source register operand.

[0268] Figure 25 An example of a third prefix 2001(C) is shown. In some examples, the third prefix 2001(C) is an example of an EVEX prefix. The third prefix 2001(C) is a four-byte prefix.

[0269] The third prefix 2001(C) can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some examples, write masks / operation masks are used (see the discussion of registers in the previous figures, e.g., Figure 19 ) or predicate instructions utilize this prefix. The operation mask register allows conditional processing or selection control. Operation mask instructions, whose source / destination operands are operation mask registers and treat the contents of the operation mask register as a single value, are encoded using the second prefix 2001 (B).

[0270] The third prefix 2001(C) may encode features specific to the instruction class (e.g., packed instructions with "load+operation" semantics may support embedded broadcast functionality, floating-point instructions with rounding semantics may support static rounding functionality, floating-point instructions with non-rounding arithmetic semantics may support "suppress all exceptions" functionality, etc.).

[0271] The first byte of the third prefix 2001(C) is the format field 2511, which in some examples has a value of 62 H. The subsequent bytes are referred to as payload bytes 2515-2519 and together form the 24-bit value of P[23:0], providing specific capabilities in the form of one or more fields (described in detail herein).

[0272] In some examples, P[1:0] of payload byte 2519 is identical to the lower two mm bits. In some examples, P[3:2] are reserved. Bit P[4] (R') allows access to the upper 16 vector registers when combined with P[7] and the reg field 2144 of MOD R / M. P[6] can also provide access to the upper 16 vector registers when SIB-type addressing is not required. P[7:5] consists of R, X, and B, which are operand specifier modifier bits for vector registers, general registers, and memory addressing, and when combined with the MOD R / M register field 2144 and the R / M field 2146 of MOD R / M, allow access to the next set of 8 registers beyond the lower 8 registers. P[9:8] provides opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). P

[10] is fixed to 1 in some examples. P[14:11], shown as vvvv, can be used to: 1) encode the first source register operand, which is specified in inverted (one's complement) form, valid for instructions with 2 or more source operands; 2) encode the destination register operand, which is specified in one's complement form, used for certain vector shifts; or 3) not encode any operand, the field is reserved and should contain some value, such as 1111b.

[0273] P

[15] is similar to W of the first prefix 2001(A) and the second prefix 2011(B), and can be used as an opcode extension bit or an operand size promotion.

[0274] P[18:16] specifies the index of a register in the operation mask (write mask) register (e.g., write mask / predicate register 1915). In some examples, the special value aaa=000 has special behavior, implying that no operation mask is used for this particular instruction (this can be achieved in a variety of ways, including using an operation mask hardwired to all ones or hardware that bypasses the masking hardware). When merging, the vector mask allows any set of elements in the destination to be protected from updates during the execution of any operation (specified by the basic operation and the enhanced operation); in other examples, the old value of each element of the destination is retained (if the corresponding mask bit has a value of 0). In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the basic operation and the enhanced operation); in some examples, elements of the destination are set to 0 if the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the length of the vector of operations being performed (i.e., the span of elements modified, from first to last); however, the modified elements do not have to be contiguous. Thus, the operation mask field enables a subset of vector operations, including loads, stores, arithmetic, logical, etc. Although in the described examples, the contents of the operation mask field select which of several operation mask registers contains the operation mask to be used (so that the contents of the operation mask field indirectly identify the masking to be performed), alternative examples alternatively or additionally allow the contents of the mask write field to directly specify the masking to be performed.

[0275] P

[19] can be combined with P[14:11] to encode a second source vector register in a non-destructive source syntax that can utilize P

[19] to access the upper 16 vector registers. P

[20] encodes various features that vary across different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P

[23] indicates support for merge-write masking (e.g., when set to 0) or support for zeroing and merge-write masking (e.g., when set to 1).

[0276] The following table details an example of the encoding of registers in instructions using the third prefix 2001 (C). Table 1: 32-register support in 64-bit mode Table 2: Encoding register designators in 32-bit mode Table 3: Operation Mask Register Specifier Encoding

[0277] Graphics Execution Unit

[0278] Figures 26A-26B Thread execution logic 2600 is illustrated including an array of processing elements employed in a graphics processor core according to examples described herein. Figures 26A-26B Elements having the same reference numbers (or names) as elements of any other figures herein may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto. Figure 26A represents the execution unit within a general-purpose graphics processor, and Figure 26B Represents an execution unit that can be used within a computing accelerator.

[0279] As in Figure 26A , in some examples, thread execution logic 2600 includes a shader processor 2602, a thread dispatcher 2604, an instruction cache 2606, a scalable execution unit array including a plurality of execution units 2608A-2608N, a sampler 2610, a shared local memory 2611, a data cache 2612, and a data port 2614. In some examples, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., execution units 2608A, 2608B, 2608C, 2608D through any of 2608N-1 and 2608N) based on the computational requirements of the workload. In some examples, the included components are interconnected via an interconnect fabric that links to each of the components. In some examples, thread execution logic 2600 includes one or more connections to memory (such as system memory or cache memory) through instruction cache 2606, data port 2614, sampler 2610, and one or more of execution units 2608A-2608N. In some examples, each execution unit (e.g., 2608A) is a self-contained programmable general-purpose computing unit capable of executing multiple simultaneous hardware threads while processing multiple data elements in parallel for each thread. In various examples, the array of execution units 2608A-2608N is scalable to include any number of separate execution units.

[0280] In some examples, execution units 2608A-2608N are primarily used to execute shader programs. Shader processor 2602 can process various shader programs and can dispatch execution threads associated with the shader programs via thread dispatcher 2604. In some examples, the thread dispatcher includes logic for arbitrating thread initiation requests from the graphics pipeline and the media pipeline and instantiating the requested threads on one or more execution units in execution units 2608A-2608N. For example, the geometry pipeline can dispatch vertex shaders, tessellation shaders, or geometry shaders to thread execution logic for processing. In some examples, thread dispatcher 2604 can also handle runtime thread generation requests from executed shader programs.

[0281] In some examples, execution units 2608A-2608N support an instruction set that includes native support for many standard 3D graphics shader instructions, allowing shader programs from graphics libraries (e.g., Direct 3D and OpenGL) to be executed with minimal translation. These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general processing (e.g., compute and media shaders). Each execution unit in execution units 2608A-2608N is capable of multi-issue single instruction multiple data (SIMD) execution, and multi-threaded operation enables an efficient execution environment in the face of high-latency memory accesses. Each hardware thread within each execution unit has a dedicated high-bandwidth register file and associated independent thread state. For pipelines capable of integer operations, single-precision floating-point operations, double-precision floating-point operations, SIMD branching capabilities, logical operations, transcendental operations, and other miscellaneous operations, execution is multi-issued per clock. While waiting for data from memory or one of the shared functions, dependency logic within execution units 2608A-2608N puts the waiting thread to sleep until the requested data has been returned. While the waiting thread is sleeping, hardware resources can be dedicated to processing other threads. For example, during a delay associated with a vertex shader operation, an execution unit can perform operations for a pixel shader, a fragment shader, or another type of shader program including a different vertex shader. Each example can be applied to use execution utilizing single instruction multiple threads (SIMT) as an alternative to the use of SIMD, or in addition to the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to a combination of SIMD and SIMT.

[0282] Each execution unit in execution units 2608A-2608N operates on an array of data elements. The number of data elements is the "execution size," or the number of lanes used for the instruction. An execution lane is a logical unit used to perform data element access, masking, and flow control within an instruction. The number of lanes may be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) used for a particular graphics processor. In some examples, execution units 2608A-2608N support integer and floating point data types.

[0283] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers by compact data types, and the execution unit will process various elements based on the data size of element. For example, when a 256-bit wide vector is operated, the 256 bits of vector are stored in registers, and the execution unit operates the vector into four independent 64-bit compact data elements (quad-word (Quad-Word, QW) size data elements), eight independent 32-bit compact data elements (double-word (Double Word, DW) size data elements), sixteen independent 16-bit compact data elements (word (Word, W) size data elements) or thirty-two independent 8-bit data elements (byte (byte, B) size data elements). However, different vector widths and register sizes are possible.

[0284] In some examples, one or more execution units can be combined into a fused execution unit 2609A-2609N, which has a thread control logic (2607A-2607N) common to the fused EUs. Multiple EUs can be fused into an EU group. Each EU in the fused EU group can be configured to execute a separate SIMD hardware thread. The number of EUs in the fused EU group can vary depending on the example. Additionally, various SIMD widths can be executed EU by EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 2609A-2609N includes at least two execution units. For example, the fused execution unit 2609A includes a first EU 2608A, a second EU 2608B, and a thread control logic 2608A common to the first EU 2607A and the second EU 2608B. Thread control logic 2607A controls threads executing on fused graphics execution unit 2609A, allowing each EU within fused execution units 2609A-2609N to execute using a common instruction pointer register.

[0285] One or more internal instruction caches (e.g., 2606) are included in thread execution logic 2600 to cache thread instructions for the execution unit. In some examples, one or more data caches (e.g., 2612) are included to cache thread data during thread execution. Threads executing on execution logic 2600 may also store explicitly managed data in shared local memory 2611. In some examples, a sampler 2610 is included to provide texture sampling for 3D operations and media sampling for media operations. In some examples, sampler 2610 includes specialized texture or media sampling functionality to process texture data or media data during the sampling process before providing the sampled data to the execution unit.

[0286] During execution, the graphics pipeline and media pipeline send thread initiation requests to the thread execution logic 2600 via thread generation and dispatch logic. Once the group of geometric objects has been processed and rasterized into pixel data, pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 2602 is invoked to further calculate output information and cause the results to be written to output surfaces (e.g., color buffer, depth buffer, stencil buffer, etc.). In some examples, the pixel shader or fragment shader calculates values ​​for each vertex attribute, which are interpolated across the rasterized objects. In some examples, the pixel processor logic within the shader processor 2602 then executes the pixel shader program or fragment shader program provided by the application programming interface (API). To execute the shader program, the shader processor 2602 dispatches the thread to an execution unit (e.g., 2608A) via the thread dispatcher 2604. In some examples, the shader processor 2602 uses texture sampling logic in the sampler 2610 to access texture data stored in a texture map in memory. Arithmetic operations on texture data and input geometry data compute pixel color data for each geometry fragment, or discard one or more pixels without further processing.

[0287] In some examples, data port 2614 provides a memory access mechanism for thread execution logic 2600 to output processed data to memory for further processing on the graphics processor output pipeline. In some examples, data port 2614 includes or is coupled to one or more cache memories (e.g., data cache 2612) to cache data for memory access via the data port.

[0288] In some examples, the execution logic 2600 may also include a ray tracer 2605 that may provide ray tracing acceleration functionality. The ray tracer 2605 may support a ray tracing instruction set that includes instructions / functions for ray generation.

[0289] Figure 26B 26. The following illustrates exemplary internal details of an execution unit 2608 according to an example. Graphics execution unit 2608 may include an instruction fetch unit 2637, a general register file (GRF) array 2624, an architectural register file (ARF) array 2626, a thread arbiter 2622, a dispatch unit 2630, a branch unit 2632, a set of SIMD floating point units (FPUs) 2634, and in some examples, a set of dedicated integer SIMD ALUs 2635. GRF 2624 and ARF 2626 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in graphics execution unit 2608. In some examples, per-thread architectural state is maintained in ARF 2626, while data used during thread execution is stored in GRF 2624. The execution state of each thread, including an instruction pointer for each thread, may be stored in thread-specific registers in ARF 2626.

[0290] In some examples, graphics execution unit 2608 has an architecture that is a combination of simultaneous multi-threading (SMT) and fine-grained interleaved multi-threading (IMT). The architecture has a modular configuration that can be fine-tuned at design time based on the target number of simultaneous threads and the number of registers per execution unit, where execution unit resources are divided across the logic used to execute multiple simultaneous threads. The number of logical threads that can be executed by graphics execution unit 2608 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.

[0291] In some examples, graphics execution unit 2608 may issue multiple instructions in concert, each of which may be a different instruction. A thread arbiter 2622 of a graphics execution unit thread 2608 may dispatch the instruction to one of the following for execution: a dispatch unit 2630, a branch unit 2632, or a SIMD FPU(s) 2634. Each execution thread may access 128 general purpose registers within GRF 2624, each of which may store 32 bytes accessible as a SIMD 8-element vector with 32-bit data elements. In some examples, each execution unit thread may have access to 4 kilobytes within GRF 2624, but examples are not limited thereto, and more or fewer register resources may be provided in other examples. In some examples, graphics execution unit 2608 is partitioned into seven hardware threads that may independently execute computational operations, but the number of threads per execution unit may vary depending on the example. For example, in some examples, a maximum of 16 hardware threads may be supported. In an example where seven threads can access 4 kilobytes, GRF 2624 can store a total of 28 kilobytes. Where 16 threads can access 4 kilobytes, GRF 2624 can store a total of 64 kilobytes. Flexible addressing modes allow registers to be addressed together, effectively creating wider registers or representing strided rectangular block data structures.

[0292] In some examples, memory operations, sampler operations, and other longer latency system communications are dispatched via "send" instructions executed through message passing send unit 2630. In some examples, branch instructions are dispatched to dedicated branch unit 2632 to facilitate SIMD scatter and eventual convergence.

[0293] In some examples, graphics execution unit 2608 includes one or more SIMD floating point units (FPUs) 2634 for performing floating point operations. In some examples, FPU(s) 2634 also support integer computations. In some examples, FPU(s) 2634 can SIMD perform up to M 32-bit floating point (or integer) operations, or SIMD perform up to 2M 16-bit integer or 16-bit floating point operations. In some examples, at least one of the FPU(s) provides extended math capabilities that support high-throughput transcendental math functions and double-precision 64-bit floating point. In some examples, a set of 8-bit integer SIMD ALUs 2635 is also present and can be specifically optimized to perform operations associated with machine learning computations.

[0294] In some examples, an array of multiple instances of graphics execution unit 2608 can be instantiated in graphics sub-core groupings (e.g., sub-slices). For scalability, product architects can choose the exact number of execution units per sub-core grouping. In some examples, execution unit 2608 can execute instructions across multiple execution lanes. In further examples, each thread executed on graphics execution unit 2608 is executed on a different lane.

[0295] Figure 27 FIG2 illustrates an additional execution unit 2700 according to an example. In some examples, the execution unit 2700 includes a thread control unit 2701, a thread state unit 2702, an instruction fetch / prefetch unit 2703, and an instruction decode unit 2704. The execution unit 2700 additionally includes a register file 2706 that stores registers that can be assigned to hardware threads within the execution unit. The execution unit 2700 additionally includes a dispatch unit 2707 and a branch unit 2708. In some examples, the dispatch unit 2707 and the branch unit 2708 can be configured to execute commands in the same manner as the thread control unit 2701. Figure 26B The sending unit 2630 and the branching unit 2632 of the graphics execution unit 2608 operate in a similar manner.

[0296] Execution unit 2700 also includes a compute unit 2710, which includes multiple different types of functional units. In some examples, compute unit 2710 includes an ALU unit 2711, which includes an array of arithmetic logic units. ALU unit 2711 can be configured to perform 64-bit, 32-bit, and 16-bit integer and floating-point operations. Integer and floating-point operations can be performed simultaneously. Compute unit 2710 also includes a systolic array 2712 and a math unit 2713. Systolic array 2712 includes a wide (W) and deep (D) network of data processing units that can be used to perform vector or other data-parallel operations in a systolic manner. In some examples, systolic array 2712 can be configured to perform matrix operations, such as matrix dot product operations. In some examples, systolic array 2712 supports 16-bit floating-point operations as well as 8-bit and 4-bit integer operations. In some examples, systolic array 2712 can be configured to accelerate machine learning operations. In such examples, systolic array 2712 may be configured with support for the bfloat 16-bit floating point format. In some examples, math unit 2713 may be included to perform a specific subset of math operations in an efficient and lower power manner than ALU unit 2711. Math unit 2713 may include a variation of math logic found in the shared function logic of the graphics processing engine provided by other examples. In some examples, math unit 2713 may be configured to perform 32-bit and 64-bit floating point operations.

[0297] The thread control unit 2701 includes logic for controlling the execution of threads within the execution units. The thread control unit 2701 may include thread arbitration logic for starting, stopping, and preempting the execution of threads within the execution units 2700. The thread state unit 2702 may be used to store thread states for threads assigned to execute on the execution units 2700. Storing thread states within the execution units 2700 enables quick preemption of threads when they become blocked or idle. The instruction fetch / prefetch unit 2703 may fetch instructions from the instruction cache of higher level execution logic (e.g., Figure 26A The instruction fetch / prefetch unit 2703 may also issue prefetch requests for instructions to be loaded into the instruction cache based on analysis of the currently executing thread. The instruction decode unit 2704 may be used to decode instructions to be executed by the compute unit. In some examples, the instruction decode unit 2704 may be used as a secondary decoder to decode complex instructions into their constituent micro-operations.

[0298] Execution unit 2700 additionally includes a register file 2706 that can be used by hardware threads executing on execution unit 2700. Registers in register file 2706 can be divided across logic used to execute multiple simultaneous threads within compute unit 2710 of execution unit 2700. The number of logical threads that can be executed by graphics execution unit 2700 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread. The size of register file 2706 can vary across examples based on the number of hardware threads supported. In some examples, register renaming can be used to dynamically assign registers to hardware threads.

[0299] Figure 28 2 is a block diagram illustrating a graphics processor instruction format 2800 according to some examples. In one or more examples, a graphics processor execution unit supports an instruction set having instructions in multiple formats. Solid-line boxes illustrate components that are typically included in execution unit instructions, while dashed lines include components that are optional or included only in a subset of instructions. In some examples, the described and illustrated instruction format 2800 are macroinstructions because they are instructions supplied to the execution unit, as opposed to micro-operations that result from instruction decoding once the instruction is processed.

[0300] In some examples, the graphics processor execution unit natively supports instructions in the 128-bit instruction format 2810. Based on the selected instruction, instruction options, and number of operands, a 64-bit compact instruction format 2830 may be used for some instructions. The native 128-bit instruction format 2810 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 2830. The native instructions available in the 64-bit format 2830 vary depending on the example. In some examples, instructions are partially compressed using a set of index values ​​in the index field 2813. The execution unit hardware references a set of compression tables based on the index values ​​and uses the compression table output to reconstruct the native instruction in the 128-bit instruction format 2810. Instructions of other sizes and formats may be used.

[0301] For each format, the instruction opcode 2812 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an addition instruction, the execution unit performs a synchronous addition operation across each color channel representing a texture element or picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some examples, the instruction control field 2814 enables control of certain execution options, such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 2810, the execution size field 2816 limits the number of data channels to be executed in parallel. In some examples, the execution size field 2816 is not available for the 64-bit compact instruction format 2830.

[0302] Some execution unit instructions have up to three operands, including two source operands src0 2820 and src1 2822, and one destination 2818. In some examples, the execution unit supports dual-destination instructions, where one of the two destinations is implicit. Data manipulation instructions may have a third source operand (e.g., src2 2824 (source 2 2824)), where the instruction opcode 2812 determines the number of source operands. The last source operand of an instruction may be an immediate (e.g., hard-coded) value passed with the instruction.

[0303] In some examples, the 128-bit instruction format 2810 includes an access / addressing mode field 2826 that specifies, for example, whether direct register addressing mode or indirect register addressing mode is used. When direct register addressing mode is used, the register addresses of one or more operands are provided directly by bits in the instruction.

[0304] In some examples, the 128-bit instruction format 2810 includes an access / addressing mode field 2826 that specifies the addressing mode and / or access mode of the instruction. In some examples, the access mode is used to define the data access alignment for the instruction. Some examples support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, wherein the byte alignment of the access mode determines the access alignment of the instruction operands. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands.

[0305] In some examples, the addressing mode portion of the access / addressing mode field 2826 determines whether the instruction uses direct or indirect addressing. When direct register addressing mode is used, bits in the instruction directly provide the register address of one or more operands. When indirect register addressing mode is used, the register address of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.

[0306] In some examples, instructions are grouped based on the opcode 2812 bit field to simplify opcode decoding 2840. For 8-bit opcodes, bits 4, 5, and 6 allow the execution unit to determine the type of opcode. The exact opcode grouping shown is only an example. In some examples, the move and logic opcode group 2842 includes data movement and logic instructions (e.g., move (mov), compare (cmp)). In some examples, the move and logic group 2842 shares five most significant bits (MSB), wherein the move (mov) instruction adopts the form of 0000xxxxb, and the logic instruction adopts the form of 0001xxxxb. The flow control instruction group 2844 (e.g., call (call), jump (jmp)) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 2846 includes a mixture of instructions, including synchronization instructions (e.g., wait (wait), send (send)) in the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 2848 includes component-wise arithmetic instructions (e.g., add, multiply (mul)) of the form 0100xxxxb (e.g., 0x40). The parallel math group 2848 performs arithmetic operations in parallel across the data lanes. The vector math group 2850 includes arithmetic instructions (e.g., dp4) of the form 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In some examples, the illustrated opcode decode 2840 can be used to determine which portion of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by a systolic array. Other instructions, such as ray tracing instructions (not shown), can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic.

[0307] Graphics pipeline

[0308] Figure 29 is a block diagram of another example of a graphics processor 2900 . Figure 29 Elements having the same reference numbers (or names) as elements of any other figures herein may operate or function in any manner similar to that described elsewhere herein, but are not limited thereto.

[0309] In some examples, graphics processor 2900 includes a geometry pipeline 2920, a media pipeline 2930, a display engine 2940, thread execution logic 2950, ​​and a render output pipeline 2970. In some examples, graphics processor 2900 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued to the graphics processor 2900 via a ring interconnect 2902. In some examples, ring interconnect 2902 couples graphics processor 2900 to other processing components, such as other graphics processors or general-purpose processors. Commands from ring interconnect 2902 are interpreted by command stream converter 2903, which supplies instructions to various components of geometry pipeline 2920 or media pipeline 2930.

[0310] In some examples, command stream converter 2903 directs the operation of vertex fetcher 2905, which reads vertex data from memory and executes vertex processing commands provided by command stream converter 2903. In some examples, vertex fetcher 2905 provides vertex data to vertex shader 2907, which performs coordinate space transformation and lighting operations on each vertex. In some examples, vertex fetcher 2905 and vertex shader 2907 execute vertex processing instructions by dispatching execution threads to execution units 2952A-2952B via thread dispatcher 2931.

[0311] In some examples, execution units 2952A-2952B are arrays of vector processors with instruction sets for performing graphics operations and media operations. In some examples, execution units 2952A-2952B have an attached L1 cache 2951 that is dedicated to each array or shared between arrays. The cache can be configured as a data cache, an instruction cache, or a single cache partitioned into different partitions containing data and instructions.

[0312] In some examples, the geometry pipeline 2920 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some examples, the programmable hull shader 2911 configures the tessellation operations. The programmable domain shader 2917 provides back-end evaluation of the tessellation output. The tessellation controller 2913 operates under the direction of the hull shader 2911 and contains dedicated logic for generating a detailed set of geometric objects based on a coarse geometric model that is provided as input to the geometry pipeline 2920. In some examples, if tessellation is not used, the tessellation components (e.g., the hull shader 2911, the tessellation controller 2913, and the domain shader 2917) can be bypassed.

[0313] In some examples, the complete geometric object may be processed by the geometry shader 2919 via one or more threads dispatched to the execution units 2952A-2952B, or may proceed directly to the clipper 2929. In some examples, the geometry shader operates on the entire geometric object, rather than on vertices or patches of vertices as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 2919 receives input from the vertex shader 2907. In some examples, the geometry shader 2919 is programmable by the geometry shader program to perform geometry tessellation when the tessellation unit is disabled.

[0314] Before rasterization, the clipper 2929 processes the vertex data. The clipper 2929 can be a fixed-function clipper or a programmable clipper with clipping and geometry shader functionality. In some examples, the rasterizer and depth test component 2973 in the render output pipeline 2970 dispatches a pixel shader to convert the geometric objects into a pixel-by-pixel representation. In some examples, the pixel shader logic is included in the thread execution logic 2950. In some examples, the application can bypass the rasterizer and depth test component 2973 and access the unrasterized vertex data via the outflow unit 2923.

[0315] The graphics processor 2900 has an interconnect bus, interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the main components of the processor. In some examples, execution units 2952A-2952B and associated logic units (e.g., L1 cache 2951, sampler 2954, texture cache 2958, etc.) are interconnected via data ports 2956 to perform memory accesses and communicate with the processor's rendering output pipeline components. In some examples, sampler 2954, caches 2951, 2958, and execution units 2952A-2952B each have a separate memory access path. In some examples, texture cache 2958 can also be configured as a sampler cache.

[0316] In some examples, the render output pipeline 2970 includes a rasterizer and depth test component 2973 that converts vertex-based objects into associated pixel-based representations. In some examples, the rasterizer logic includes a windower / masker unit for performing fixed-function triangle and line rasterization. In some examples, an associated render buffer 2978 and depth buffer 2979 are also available. A pixel operation component 2977 performs pixel-based operations on data, but in some instances, pixel operations associated with 2D operations (e.g., using mixed bit block image transfers) are performed by the 2D engine 2941 or, when displayed, by the display controller 2943 using an overlay display plane instead. In some examples, a shared L3 cache 2975 is available to all graphics components, allowing data to be shared without using main system memory.

[0317] In some examples, the graphics processor media pipeline 2930 includes a media engine 2937 and a video front end 2934. In some examples, the video front end 2934 receives pipeline commands from the command streamer 2903. In some examples, the media pipeline 2930 includes a separate command streamer. In some examples, the video front end 2934 processes the media commands before sending them to the media engine 2937. In some examples, the media engine 2937 includes thread generation functionality for generating threads for dispatching to the thread execution logic 2950 via the thread dispatcher 2931.

[0318] In some examples, graphics processor 2900 includes a display engine 2940. In some examples, display engine 2940 is external to processor 2900 and is coupled to the graphics processor via ring interconnect 2902, or some other interconnect bus or structure. In some examples, display engine 2940 includes a 2D engine 2941 and a display controller 2943. In some examples, display engine 2940 contains dedicated logic capable of operating independently of the 3D pipeline. In some examples, display controller 2943 is coupled to a display device (not shown), which can be a system-integrated display device such as in a laptop or an external display device attached via a display device connector.

[0319] In some examples, the geometry pipeline 2920 and the media pipeline 2930 can be configured to perform operations based on multiple graphics and media programming interfaces and are not specific to any one application programming interface (API). In some examples, driver software for the graphics processor converts API calls specific to a particular graphics or media library into commands that can be processed by the graphics processor. In some examples, support is provided for the Open Graphics Library (OpenGL), the Open Computing Language (OpenCL), and / or the Vulkan graphics and compute APIs, all from the Khronos Group. In some examples, support may also be provided for the Direct3D library from Microsoft. In some examples, a combination of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). If a mapping can be performed from the pipeline of a future API to the pipeline of the graphics processor, future APIs with compatible 3D pipelines will also be supported.

[0320] Graphics pipeline programming

[0321] Figure 30A is a block diagram illustrating a graphics processor command format 3000 according to some examples. Figure 30B is a block diagram illustrating a graphics processor command sequence 3010 according to an example. Figure 30A Solid-line boxes in illustrate components that are typically included in the graphics commands, while dashed lines include components that are optional or included only in a subset of the graphics commands. Figure 30A An exemplary graphics processor command format 3000 includes a data field for identifying the client 3002, a command operation code (opcode) 3004, and data for the command 3006. A sub-opcode 3005 and a command size 3008 are also included in some commands.

[0322] In some examples, client 3002 specifies a client unit of a graphics device that processes command data. In some examples, a graphics processor command parser examines the client field of each command to adjust further processing of the command and route the command data to the appropriate client unit. In some examples, a graphics processor client unit includes a memory interface unit, a rendering unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 3004 and sub-opcode 3005 (if present) to determine the operation to be performed. The client unit uses the information in the data field 3006 to execute the command. For some commands, an explicit command size 3008 is expected to specify the size of the command. In some examples, the command parser automatically determines the size of at least some of the commands based on the command opcode. In some examples, commands are aligned via multiples of double words. Other command formats may be used.

[0323] Figure 30B The flowchart in FIG. 30 illustrates an exemplary graphics processor command sequence 3010. In some examples, software or firmware of a data processing system featuring an example graphics processor uses a version of the illustrated command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for exemplary purposes only, as the examples are not limited to these specific commands or command sequences. Furthermore, commands can be issued as batches in the command sequence so that the graphics processor processes the command sequence at least partially concurrently.

[0324] In some examples, graphics processor command sequence 3010 may begin with a pipeline flush command 3012 to cause any active graphics pipelines to complete their currently pending commands. In some examples, 3D pipeline 3022 and media pipeline 3024 do not operate concurrently. A pipeline flush is performed to cause active graphics pipelines to complete any pending commands. In response to the pipeline flush, the command parser for the graphics processor will suspend command processing until the active paint engines complete pending operations and the associated read caches are invalidated. Optionally, any data marked as "dirty" in the render cache may be flushed to memory. In some examples, pipeline flush command 3012 may be used for pipeline synchronization or before placing the graphics processor in a low-power state.

[0325] In some examples, when a command sequence requires the graphics processor to explicitly switch between pipelines, a pipeline select command 3013 is used. In some examples, a pipeline select command 3013 is required only once in an execution context before issuing a pipeline command, unless the context is issuing commands for both pipelines. In some examples, a pipeline flush command 3012 is required immediately before a pipeline switch via a pipeline select command 3013.

[0326] In some examples, pipeline control commands 3014 configure the graphics pipeline for operation and are used to program 3D pipeline 3022 and media pipeline 3024. In some examples, pipeline control commands 3014 configure the pipeline state of the active pipeline. In some examples, pipeline control commands 3014 are used for pipeline synchronization and are used to flush data from one or more cache memories within the active pipeline before processing a batch of commands.

[0327] In some examples, the return buffer state command 3016 is used to configure a set of return buffers for the corresponding pipeline to write data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which the operation writes intermediate data during processing. In some examples, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some examples, the return buffer state command 3016 includes selecting the size and number of return buffers to be used for the set of pipeline operations.

[0328] The remaining commands in the command sequence differ based on the active pipeline for operation.Based on pipeline decision 3020 , the command sequence is tailored for the 3D pipeline 3022 starting at 3D pipeline state 3030 or the media pipeline 3024 starting at media pipeline state 3040 .

[0329] The commands used to configure the 3D pipeline state 3030 include 3D state setup commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that are to be configured before processing 3D primitive commands. The values ​​of these commands are determined at least in part based on the specific 3D API in use. In some examples, the 3D pipeline state 3030 commands can also selectively disable or bypass certain pipeline elements if those elements are not to be used.

[0330] In some examples, a 3D primitive 3032 command is used to submit a 3D primitive to be processed by the 3D pipeline. The command and associated parameters passed to the graphics processor via the 3D primitive 3032 command are forwarded to a vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 3032 command data to generate a vertex data structure. The vertex data structure is stored in one or more return buffers. In some examples, a 3D primitive 3032 command is used to perform vertex operations on the 3D primitive via a vertex shader. To process the vertex shader, the 3D pipeline 3022 dispatches a shader execution thread to a graphics processor execution unit.

[0331] In some examples, the 3D pipeline 3022 is triggered via an execute 3034 command or event. In some examples, a register write triggers the command execution. In some examples, the execution is triggered via a "go" or "kick" command in the command sequence. In some examples, the command execution is triggered using a pipeline synchronization command to flush the command sequence through the graphics pipeline. The 3D pipeline will perform geometry processing for the 3D primitives. Once the operation is completed, the resulting geometry object is rasterized and the pixel engine shades the resulting pixels. For those operations, additional commands for controlling pixel shading and pixel backend operations may also be included.

[0332] In some examples, when performing media operations, the graphics processor command sequence 3010 follows the media pipeline 3024 path. Generally speaking, the specific purpose and manner of programming the media pipeline 3024 depends on the media or compute operation to be performed. During media decoding, certain media decoding operations can be migrated to the media pipeline. In some examples, the media pipeline can also be bypassed and the media decoding can be performed in whole or in part using resources provided by one or more general-purpose processing cores. In some examples, the media pipeline also includes elements for general-purpose graphics processor unit (GPGPU) operations, wherein the graphics processor is configured to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0333] In some examples, the media pipeline 3024 is configured in a similar manner to the 3D pipeline 3022. A set of commands for configuring the media pipeline state 3040 is dispatched or placed into a command queue before the media object commands 3042. In some examples, the commands for the media pipeline state 3040 include data for configuring the media pipeline elements that will be used to process the media objects. This includes data for configuring the video decoding and video encoding logic within the media pipeline, such as encoding or decoding formats. In some examples, the commands for the media pipeline state 3040 also support the use of one or more pointers to "indirect" state elements that contain batches of state settings.

[0334] In some examples, media object commands 3042 supply a pointer to a media object for processing by the media pipeline. The media object includes a memory buffer that contains the video data to be processed. In some examples, all media pipeline states must be valid before issuing media object commands 3042. Once the pipeline state is configured and media object commands 3042 are queued, media pipeline 3024 is triggered via an execute command 3044 or an equivalent execute event (e.g., a register write). The output from media pipeline 3024 can then be post-processed by operations provided by 3D pipeline 3022 or media pipeline 3024. In some examples, GPGPU operations are configured and executed in a manner similar to media operations.

[0335] Program code can be applied to input information to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microprocessor, or any combination thereof.

[0336] Program code can be implemented in a process-oriented or object-oriented high-level programming language to communicate with the processing system. If desired, program code can also be implemented in assembly or machine language. In fact, the mechanism described herein is not limited to any specific programming language in scope. In any case, the language can be a compiled language or an interpreted language.

[0337] Examples of the mechanisms disclosed herein may be implemented in hardware, software, firmware, or a combination of these implementation approaches. Examples may be implemented as computer programs or program codes that execute on a programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0338] These machine-readable storage media may include, but are not limited to, non-transitory tangible arrangements of articles manufactured or formed by a machine or apparatus, including storage media such as a hard disk, any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), compact disk rewritable (CD-RW), and magneto-optical disks), semiconductor devices (e.g., read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM), phase change memory (PCM)), magnetic or optical cards, or any other type of medium suitable for storing electronic instructions.

[0339] Therefore, examples also include non-transitory tangible machine-readable media containing instructions or design data, such as hardware description languages ​​(HDL), that define the features of the structures, circuits, devices, processors, and / or systems described herein. Such examples may also be referred to as program products.

[0340] Emulation (including binary translation, code deformation, etc.).

[0341] In some cases, an instruction converter can be used to convert instructions from a source instruction set architecture to a target instruction set architecture. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), deform, simulate, or otherwise convert instructions to one or more other instructions to be processed by the core. The instruction converter can be implemented in software, hardware, firmware, or a combination thereof. The instruction converter can be on the processor, off the processor, or partly on the processor and partly off the processor.

[0342] Figure 31 The block diagram illustrates the use of a software instruction converter according to an example, which is used to convert binary instructions in a source ISA into binary instructions in a target ISA. In the illustrated example, the instruction converter is a software instruction converter, but alternatively, the instruction converter can be implemented in software, firmware, hardware, or various combinations thereof. Figure 31 It is shown that a program in a high-level language 3102 can be compiled using a first ISA compiler 3104 to generate first ISA binary code 3106 that can be natively executed by a processor having at least one first ISA core 3116. The processor having at least one first ISA core 3116 represents any processor that is capable of either (1) compatibly executing or otherwise processing a substantial portion of the first ISA or (2) executing a program in a first ISA on a processor having at least one first ISA core. The target code version of an application or other software running on a processor is used to perform substantially the same functions as an Intel processor having at least one first ISA core, so as to achieve substantially the same results as a processor having at least one first ISA core. The first ISA compiler 3104 represents a compiler that is operable to generate first ISA binary code 3106 (e.g., target code) that can be executed on a processor having at least one first ISA core 3116 with or without additional linking processing. Similarly, Figure 31It is shown that a program in a high-level language 3102 can be compiled using an alternative ISA compiler 3108 to generate an alternative ISA binary code 3110 that can be executed natively by a processor 3114 without a first ISA core. An instruction converter 3112 is used to convert the first ISA binary code 3106 into code that can be executed natively by a processor 3114 without a first ISA core. This converted code will not necessarily be identical to the alternative ISA binary code 3110; however, the converted code will implement the overall operation and be composed of instructions from the alternative ISA. Thus, the instruction converter 3112 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device without a first ISA processor or core to execute the first ISA binary code 3106 through simulation, emulation, or any other process.

[0343] IP core implementation

[0344] One or more aspects of at least some examples may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit (such as a processor). For example, a machine-readable medium may include instructions representing various logic within a processor. When read by a machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as "IP cores") are reusable units of logic for an integrated circuit that can be stored on a tangible, machine-readable medium as a hardware model that describes the organization of the integrated circuit. The hardware model can be supplied to each customer or manufacturing facility that loads the hardware model on a manufacturing machine that manufactures the integrated circuit. The integrated circuit can be manufactured so that the circuit performs the operations described in association with any of the examples described herein.

[0345] Figure 323 is a block diagram illustrating an IP core development system 3200 that can be used to manufacture integrated circuits to perform operations, according to some examples. IP core development system 3200 can be used to generate modular, reusable designs that can be incorporated into larger designs or used to build entire integrated circuits (e.g., SoC integrated circuits). Design facility 3230 can generate a software simulation 3210 of the IP core design in a high-level programming language (e.g., C / C++). Software simulation 3210 can be used to design, test, and verify the behavior of the IP core using simulation model 3212. Simulation model 3212 can include functional simulation, behavioral simulation, and / or timing simulation. A register transfer level (RTL) design 3215 can then be created or synthesized from simulation model 3212. RTL design 3215 is an abstraction of the behavior of the integrated circuit (including associated logic executed using the modeled digital signals) that models the flow of digital signals between hardware registers. In addition to RTL design 3215, lower-level designs at the logic or transistor level can also be created, designed, or synthesized. Thus, specific details of initial design and simulation may vary.

[0346] The RTL design 3215 or an equivalent solution can be further synthesized by the design facility into a hardware model 3220, which can be in hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. Non-volatile memory 3240 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 3265. Alternatively, the IP core design can be transmitted via a wired connection 3250 or a wireless connection 3260 (e.g., via the Internet). The manufacturing facility 3265 can then manufacture an integrated circuit based at least in part on the IP core design. The manufactured integrated circuit can be configured to perform operations according to at least some of the examples described herein.

[0347] References to "some examples," "examples," etc., indicate that the examples being described may include a particular feature, structure, or characteristic, but not every example may necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same example. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in conjunction with other examples, whether or not explicitly described.

[0348] Furthermore, in the examples described above, unless specifically stated otherwise, separating language such as the phrases "at least one of A, B, or C" or "A, B and / or C" are intended to be understood to mean A, B, or C, or any combination thereof (i.e., A and B, A and C, B and C, and A, B, and C).

[0349] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the disclosure as set forth in the claims.

Claims

1. A device comprising: memory management circuitry for controlling memory access based on memory tags stored in a memory tag data structure and memory tags based on pointers to the memory; a decoder circuitry configured to decode an instruction into a decoded instruction, the instruction comprising an operand identifying a memory tag data structure for a sub-process of a process among a plurality of memory tag data structures for corresponding sub-processes of the process, and an opcode instructing execution circuitry to switch from another memory tag data structure for another sub-process of the process to the memory tag data structure for the sub-process; as well as Execution circuitry is provided for executing the decoded instruction according to the opcode.

2. The device according to claim 1, wherein The instructions are executable in user mode.

3. The device according to claim 1, wherein The operation object includes a memory tag data structure index among a plurality of memory tag data structure indexes.

4. The device according to claim 1, wherein The operation object further includes a linear address to an entry point of the child process.

5. The device according to any one of claims 1 to 4, wherein: The opcode is further for instructing the execution circuitry to populate an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure.

6. The device according to claim 5, wherein The object lookaside buffer is separate from a translation lookaside buffer of the device.

7. The device according to claim 6, wherein The memory tag data structure includes a virtual-to-physical page mapping for the stored memory tags, and the opcode is further for instructing the execution circuitry to retain other virtual-to-physical page mappings for the process in the translation lookaside buffer.

8. A method comprising: decoding, by the decoder circuitry, the instruction into a decoded instruction, the instruction comprising an operand and an opcode, the operand identifying a memory tag data structure for a sub-process of a process among a plurality of memory tag data structures for corresponding sub-processes of the process, the opcode instructing the execution circuitry to switch from another memory tag data structure for another sub-process of the process to the memory tag data structure for the sub-process; executing, by the execution circuitry, the decoded instruction according to the opcode; as well as Memory access is controlled by memory management circuitry based on memory tags stored in the memory tag data structure and based on memory tags of pointers to memory.

9. The method of claim 8, wherein: The execution is in user mode.

10. The method of claim 8, wherein: The operation object includes a memory tag data structure index among a plurality of memory tag data structure indexes.

11. The method of claim 8, wherein: The operation object further includes a linear address to an entry point of the child process.

12. The method according to any one of claims 8 to 11, wherein: The executing populates an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure.

13. The method of claim 12, wherein: The object lookaside buffer is separate from the translation lookaside buffer.

14. The method of claim 13, wherein: The memory tag data structure includes a virtual-to-physical page mapping for the stored memory tags, and the execution retains other virtual-to-physical page mappings for the process in the translation lookaside buffer.

15. A non-transitory machine-readable medium storing code that, when executed by a machine, causes the machine to perform a method comprising: decoding, by the decoder circuitry, the instruction into a decoded instruction, the instruction comprising an operand and an opcode, the operand identifying a memory tag data structure for a sub-process of a process among a plurality of memory tag data structures for corresponding sub-processes of the process, the opcode instructing the execution circuitry to switch from another memory tag data structure for another sub-process of the process to the memory tag data structure for the sub-process; executing, by the execution circuitry, the decoded instruction according to the opcode; as well as Memory access is controlled by memory management circuitry based on memory tags stored in the memory tag data structure and based on memory tags of pointers to memory.

16. The non-transitory machine-readable medium of claim 15, wherein: The execution is in user mode.

17. The non-transitory machine-readable medium of claim 15, wherein: The operation object includes a memory tag data structure index among a plurality of memory tag data structure indexes.

18. The non-transitory machine-readable medium of claim 15, wherein: The operation object further includes a linear address to an entry point of the child process.

19. The non-transitory machine-readable medium of any one of claims 15-18, wherein: The executing populates an object lookaside buffer with the stored memory tags for the child process from the memory tag data structure.

20. The non-transitory machine-readable medium of claim 19, wherein: The object lookaside buffer is separate from the machine's translation lookaside buffer.

21. The non-transitory machine-readable medium of claim 20, wherein: The memory tag data structure includes a virtual-to-physical page mapping for the stored memory tags, and the execution retains other virtual-to-physical page mappings for the process in the translation lookaside buffer.