Apparatus and method for attack-resistant encryption and decryption

By randomly selecting matching value pairs in memory rows, symmetry is broken, and the problems of reducing key search space and large overhead of anti-attack measures in the face of error injection attacks are solved, achieving more secure and efficient encryption performance.

CN120197233APending Publication Date: 2025-06-24INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411906106.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-23
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the face of error injection attacks, existing AES hardware accelerators lead to a reduced key search space, and anti-attack measures cause significant silicon and power overhead, limiting their application.

Method used

By randomly selecting match value pairs in memory rows, symmetry is broken, even when using symmetric passwords. This method uses high probability to find matching value pairs, change the resulting ciphertext, and ensure that different ciphertexts are generated when the same plaintext is encrypted.

Benefits of technology

Effectively resist error injection attacks, reduce key search space, reduce silicon and power overhead for attack-resistant measures, while providing safer encryption performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197233A_ABST
    Figure CN120197233A_ABST
Patent Text Reader

Abstract

An apparatus and method for attack resistant encryption and decryption. For example, one embodiment of an apparatus includes execution circuitry to execute instructions and generate a memory access request, the memory access request including a load request to read data from a memory and a store request to store data to the memory; and cryptographic circuitry for performing a plurality of rounds of encryption or decryption to encrypt or decrypt data, respectively, the cryptographic circuitry for performing one or more redundant rounds for a corresponding one or more rounds of the plurality of rounds, the one or more redundant rounds include spatial or temporal differences relative to the corresponding one or more rounds; the cryptographic circuitry is used to generate an error upon detection of a mismatch between the output of the redundant round and the output of the corresponding round.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technical Field

[0001] The present invention generally relates to the field of computer processors. More specifically, the present invention relates to apparatuses and methods for attack-resistant encryption and decryption. Background Art

[0002] Fault-injection attack (FIA) as a practical physical attack on the keys used by AES hardware accelerators is becoming increasingly common. FIA reduces the key search space by using differential fault analysis (DFA) on ciphertexts corrupted by external stimuli. FIA on a conventional AES-256 engine typically targets the last three rounds of iterations to trigger bit flips on intermediate circuit nodes using external stimuli. Due to the avalanche effect of AES encryption, errors injected during the upstream AES rounds (0-10) cascade rapidly, making first-order DFA impractical for realistic attacks. However, errors injected during the computations of rounds 13, 12, and 11 propagate down the AES logic cone, probabilistically corrupting 1, 4, and 16 ciphertext bytes respectively. Current AES attack countermeasures for error injection incur significant silicon and power overheads, limiting the adoption of such countermeasures. Brief Description of the Drawings

[0003] A better understanding of the present invention can be obtained from the following detailed description in conjunction with the following drawings, in which:

[0004] Figure 1 A block diagram of a computer system including a processor, a memory controller circuit, and a cryptographic circuit system according to an example of the present disclosure.

[0005] Figure 2 A format of a data row including locator values for encoding repeated values according to an example of the present disclosure.

[0006] Figure 3 A format of a data row from Figure 2 including additional locator bits for encoding of repeated values according to an example of the present disclosure.

[0007] Figure 4 An example of operations of a method for performing a read from a memory with repeated value encoding according to an example of the present disclosure.

[0008] Figure 5 An example of operations of a method for performing a write to a memory with repeated value encoding according to an example of the present disclosure.

[0009] Figure 6Illustrates another example of operations of a method for performing a read from a memory encoded with duplicate values according to an example of the present disclosure.

[0010] Figure 7 Illustrates an example computing system.

[0011] Figure 8 Illustrates a block diagram of an example processor and / or system on a chip (SoC) that may have one or more cores and an integrated memory controller.

[0012] Figure 9 Is a block diagram of a computing system 900 configured to implement one or more aspects of the examples described herein.

[0013] Figure 10A Illustrates an example of a parallel processor.

[0014] Figure 10B Illustrates an example of a block diagram of a partitioning unit.

[0015] Figure 10C Illustrates an example of a block diagram of a processing cluster within a parallel processing unit.

[0016] Figure 10D Illustrates an example of a graphics multiprocessor, where the graphics multiprocessor is coupled to a pipeline manager of a processing cluster.

[0017] Figures 11A - 11C Illustrates an additional graphics multiprocessor according to an example.

[0018] Figure 12 Shows a parallel computing system 1200 according to some examples.

[0019] Figures 13A - 13B Illustrates a hybrid logical / physical view of a discrete parallel processor according to examples described herein.

[0020] Figure 14A Is a block diagram illustrating an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline according to an example.

[0021] Figure 14B Is a block diagram illustrating an example in-order architecture core to be included in a processor and an example register renaming, out-of-order issue / execution architecture core according to an example.

[0022] Figure 15 Illustrates an example of (one or more) execution unit circuitry, such as (one or more) execution unit circuitry.

[0023] Figure 16 Is a block diagram of a register architecture according to some examples.

[0024] Figure 17 Examples of the illustrated instruction format.

[0025] Figure 18 Examples of the illustrated addressing information fields.

[0026] Figure 19 Examples of the illustrated first prefix.

[0027] Figures 20A - 20D Examples of how to use the R, X, and B fields of the first prefix.

[0028] Figures 21A - 21B Examples of the illustrated second prefix.

[0029] Figure 22 Examples of the illustrated third prefix.

[0030] Figures 23A - 23B Illustrates the thread execution logic according to the examples described herein, the thread execution logic including an array of processing elements employed in a graphics processing unit core.

[0031] Figure 24 Illustrates additional execution units according to the examples.

[0032] Figure 25 Is a block diagram illustrating a graphics processing unit instruction format according to some examples.

[0033] Figure 26 Is a block diagram of another example of a graphics processing unit.

[0034] Figure 27A Is a block diagram illustrating a graphics processing unit command format according to some examples.

[0035] Figure 27B Is a block diagram illustrating a graphics processing unit command sequence according to the examples.

[0036] Figure 28 Is a block diagram illustrating the conversion of binary instructions in a source ISA to binary instructions in a target ISA using a software instruction converter according to the examples.

[0037] Figure 29 Is a block diagram illustrating an IP core development system that can be used to fabricate an integrated circuit to perform operations according to some examples.

[0038] Figure 30 Illustrates an architecture including a cryptographic circuit system according to an embodiment of the present invention.

[0039] Figure 31 Illustrates a cryptographic circuit system according to some embodiments of the present invention.

[0040] Figure 32 An example is illustrated in which cipher rounds and redundant cipher rounds are interleaved.

[0041] Figure 33 An example of a cipher circuit system based on composite field representation is illustrated.

[0042] Figure 34 An example of a circuit system affected by an injected error is illustrated.

[0043] Figure 35 Results according to some embodiments of the present invention are illustrated.

[0044] Figure 36 A method according to some embodiments of the present invention is illustrated. Detailed Description

[0045] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of embodiments of the invention described below. However, it will be apparent to one of ordinary skill in the art that embodiments of the invention may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form to avoid obscuring the fundamental principles of embodiments of the invention.

[0046] The present disclosure relates to methods, apparatuses, systems, and non-transitory computer-readable storage media for matching pairs of asymmetric encryption in a computing system with a higher probability of finding matching pairs based on the Birthday Problem. Examples herein relate to memory controller circuitry and methods for data encryption that select between sets of matching tuples (e.g., pairs or triples of bytes with the same value) to be encoded to create different ciphertexts across the encryption of the same input plaintext. This confounds an adversary who expects to see the same ciphertext when using symmetric cryptography encryption given the same plaintext. Additionally, examples herein detect ciphertext corruption and prevent software replay attacks. Examples herein relate to a novel data encryption mode that selects from a set of matching pairs (e.g., within the birthday bound) that are encoded to create different ciphertexts across the encryption of the same input plaintext, e.g., where such novel data encryption mode is an addition or alternative to the XOR (exclusive OR)-encrypt-XOR (XEX) tweakable block ciphertext stealing (XTS) mode, Electronic Code Book (ECB) mode, Cipher Block Chaining (CBC) mode, etc. of a computer system (e.g., memory controller circuitry). Memory operations of a processor on an external system memory can be protected via encryption and integrity, e.g., where integrity uses additional metadata for storing integrity tags.

[0047] In some examples, the processor includes an AES-XTS mode for memory encryption (e.g., an XEX-based tweakable ciphertext mode with ciphertext stealing), e.g., including Intel Total Memory Encryption (TME), Intel Software Guard Extensions, and Intel Trust Domain Extensions (TDX), and other storage device encryption solutions. In some examples, encryption in the XTS mode uses the memory address as a tweak to create different ciphertexts for different memory locations, even though the input data is the same, while some examples in the ECB mode produce the same ciphertext for the same plaintext. However, the technical problem is that the same plaintext at the same address encrypted with the same key will still produce the same ciphertext, e.g., allowing an adversary (e.g., an attacker) to use this symmetry to attempt to circumvent the encryption.

[0048] In some examples, a memory controller circuitry (e.g., a memory encryption engine) uses a version tree in the memory to give a unique counter value for each data encryption when storing data into the memory. This makes each write have a unique ciphertext, but the version tree takes up about 25% of the memory for metadata and reduces the performance to about one-third. The high memory overhead and performance impact prevent the real-world use of such examples.

[0049] To overcome the above technical problems, examples herein utilize the seemingly paradoxical high probability of finding pairs of matching values (e.g., bytes) in a random set of matching values in cache lines and / or memory lines to break symmetry, e.g., even when using symmetric cryptography. In some examples, when multiple pairs occur, (e.g., only) one pair is selected to be encoded within a memory line or a portion thereof. In some examples, the pair to be encoded is randomly selected to vary the resulting ciphertext, e.g., even when the input plaintext is the same. While many pairs are typical in computer data that does not occur randomly (e.g., unencrypted and / or uncompressed data, such as but not limited to code, pictures, text files, memory initialized to zero, etc.), which creates a large number of choices for the encoded (e.g., byte) pairs, pairs also occur in random data (e.g., data that has been encrypted and / or compressed). In some examples, these choices allow different encodings, such that repeated encryption of the same plaintext results in different ciphertexts. In some examples, (one or more) rules are applied to detect when an incorrect pair is encoded, e.g., thus providing integrity or authenticity. In some examples, the matching pair pattern of data encryption (e.g., the "birthday pair" pattern) provides integrity, replay prevention, and / or ciphertext differentiation, e.g., as opposed to patterns (e.g., XTS) that do not provide integrity, replay prevention, and / or ciphertext differentiation and thus allow adversarial ciphertexts to be generated or ciphertext corruption to go undetected. In some examples, a processor (e.g., a memory controller circuit) having the matching pair pattern of data encryption disclosed herein (e.g., the "birthday pair" pattern) provides more secure encryption that is even resistant to hardware adversaries with optimal performance and at the lowest cost. In some examples, the physical circuit (e.g., a memory controller circuit) for memory encryption (e.g., and decryption) includes the matching pair pattern of data encryption disclosed herein (e.g., the "birthday pair" pattern). In some examples, a memory controller circuit in the matching pair pattern of data encryption (e.g., the "birthday pair" pattern) will (e.g., with high probability) result in different ciphertexts when writing the same data to memory. In some examples, a memory controller circuit in the matching pair pattern of data encryption (e.g., the "birthday pair" pattern) is capable of encrypting (e.g., encoding) more than about 98% of the memory lines, e.g., and even about 70% of the random data can be encoded, e.g., effectively eliminating isolated storage devices and access thereto, providing lower (e.g., XTS-like) overhead and better performance of security attributes by constantly varying the ciphertext. A memory controller circuit (e.g., operating according to the matching pair pattern disclosed herein) cannot be practically performed by speculation (or by pen and paper).

[0050] Turning now to the drawings, Figure 1Block diagram of a computer system 100 according to an example of the present disclosure, the computer system 100 including a processor 101, a memory controller circuit 116, and cryptographic circuitry (e.g., cryptographic circuitry 114, cryptographic circuitry 116B, and / or cryptographic circuitry 134).

[0051] The core can be any hardware processor core, e.g., an instance of core 1490 as in Figure 14B Although multiple cores are shown, the processor 101 can have a single or any number of cores (e.g., where N is any positive integer greater than 1).

[0052] The computer system 100 includes registers 110. In some examples, the registers 110 (e.g., for a particular core) include one or any combination of the following: one or more control / capability registers 110A, a shadow stack pointer register 110B, an instruction pointer (IP) register 110C, and / or a key identifier (key ID) register 110D.

[0053] In some examples, each control / capability register 110A in the one or more control / capability registers 110A of core 102 includes the same data as the corresponding one or more control / capability registers of other cores (e.g., core_N). In some examples, the control / capability registers store control values and / or capability indication values for cryptographic circuitry (e.g., encryption circuitry and / or decryption circuitry) or one or more other components. For example, the one or more capability registers store one or more values (e.g., provided by the execution of the hardware initialization manager storage 138) indicating the functions that the corresponding cryptographic circuitry (e.g., cryptographic circuitry 114, cryptographic circuitry 116B, and / or cryptographic circuitry 134) can have, and / or the one or more control registers store values that control the corresponding cryptographic circuitry (e.g., cryptographic circuitry 114, cryptographic circuitry 116B, and / or cryptographic circuitry 134).

[0054] In some examples, the memory 120 is used to store a (e.g., data) stack 122 and / or a shadow stack 124. In some examples, the shadow stack 124 stores the context of a thread, e.g., the context includes, for example, a shadow stack pointer for the context. The shadow stack pointer can be an address, e.g., a linear address or other value indicating the value of the stack pointer. In some examples, each corresponding linear address specifies a different byte in the memory (e.g., in the stack). In some examples, the current shadow stack pointer is stored in the shadow stack pointer register 110B.

[0055] In some examples, a request to switch context (e.g., push and / or pop a stack pointer), e.g., from a thread operating at user-level privilege, can be received. In some examples, the request to switch context includes pushing one or more additional data items onto stack 122 or popping one or more additional data items from stack 122 in addition to the stack pointer. In some examples, program code (e.g., software) executing at the user level can request a push or pop of stack 122 (e.g., a non-shadow stack). In some examples, the request is an instruction issued to the processor for decoding and / or execution. For example, a request to pop the stack pointer from stack 122 can include executing a restore stack pointer instruction. For example, a request to push the stack pointer onto stack 122 can include executing a save stack pointer instruction. In some examples, shadow stack 124 is a second separate stack that "shadows" (e.g., for a program call) stack 122.

[0056] In some examples, a function loads return addresses, e.g., from both call stack 122 and shadow stack 124, and the processor 101 compares them, and if the two records of the return address are different, an attack is detected (e.g., and an exception is reported to the OS), and if they match, access (e.g., push or pop) is allowed to continue.

[0057] In some examples, the instruction pointer (IP) register 110C is used to store the (e.g., current) IP value, e.g., the RIP value for 64-bit addressing mode or the EIP value for 32-bit addressing mode.

[0058] In some examples, a memory access (e.g., store or load) request to memory 120 is generated by the processor 101 (e.g., a core), e.g., a memory access request generated by the execution circuitry 106 of core 102 (e.g., caused by the execution of an instruction decoded by decoder circuitry 104), and / or the memory access request can be generated by the execution circuitry of another core_N. In some examples, the memory address for the memory access is generated by the address generation unit (AGU) 108 of the execution circuitry 106.

[0059] In some examples, the memory access request is serviced by a cache (e.g., a cache within a core and / or cache 112 shared by multiple cores). Additionally or alternatively (e.g., for a cache miss), the memory access request can be serviced by memory 120 separate from the cache. In some examples, the memory access request is to load data from memory 120 into the cache of the processor (e.g., cache 112). In some examples, the memory access request is to store data from the processor (e.g., the cache of the processor) (e.g., cache 112) to memory 120.

[0060] In some examples, computer system 100 includes cryptographic circuitry (e.g., which uses encryption to store encrypted information and decryption to decrypt the stored encrypted information). In some examples, the cryptographic circuitry is included within processor 101. In some examples, cryptographic circuitry 116B is included within memory controller circuitry 116. In some examples, the encryption circuitry is included between levels of the cache hierarchy. In some examples, cryptographic circuitry 134 is included within network interface controller (NIC) circuitry 132 (e.g., NIC circuitry 132 for controlling the transmission and / or reception of data over a network). In some examples, a single cryptographic circuitry is used for two (e.g., all) cores of computer system 100. In some examples, the cryptographic circuitry includes controls to set it to a particular mode (e.g., mode 114A) to set cryptographic circuitry 114 to a particular mode (e.g., a matching pair mode such as data encryption and / or decryption discussed herein (e.g., “birthday pair” mode)) or similarly for other cryptographic circuitry.

[0061] Some systems (e.g., processors) use encryption and decryption of data to provide security. In some examples, the cryptographic circuitry is separate from the processor core, e.g., as migratory circuitry controlled by commands sent from the processor core, e.g., cryptographic circuitry 114 is separate from any core. Cryptographic circuitry 114 may receive memory access (e.g., store) requests from one or more of its cores (e.g., from address generation unit 108 of execution circuitry 106). In some examples, the cryptographic circuitry is used to perform encryption on an input such as a destination address and text (e.g., plaintext) (e.g., and a key) to be encrypted to generate ciphertext (e.g., encrypted data). Then, the ciphertext may be stored in a storage device, e.g., stored in memory 120. In some examples, the cryptographic circuitry performs a decryption operation on a memory load request, for example. The cryptographic circuitry may include a fine-tuned operation mode (such as AES-XTS) that uses the memory address as a fine-tuning for the cryptographic operation, e.g., to ensure that different ciphertexts are obtained even when encrypting the same data for different addresses. Other modes (such as AES-CBC) may be used to extend across an entire memory row larger than a single data block, e.g., to allow for the distribution of the ciphertext of an initial locator value for encoding across the entire memory row.

[0062] In some examples, a processor (e.g., as an instruction set architecture (ISA) extension) supports total memory encryption (TME) (e.g., memory encryption using a single short-term key) and / or multi-key TME (TME-MK or MKTME) (e.g., memory encryption that supports using multiple keys for page-granularity memory encryption, e.g., with additional support for keys supplied to software).

[0063] In some examples, TME provides the ability to encrypt an entire physical memory of a system. For example, with minor changes to the hardware initialization manager code (e.g., Basic Input / Output System (BIOS) firmware, such as stored in storage device 138), this ability is enabled at a very early stage of the boot process. In some examples, once TME is configured and locked, TME will encrypt all data on the external memory bus of computer system 100 using an encryption standard / algorithm (e.g., Advanced Encryption Standard (AES), such as but not limited to AES using a 128-bit key). In some examples, the encryption key for TME uses a hardware random number generator implemented in the computer system (e.g., the processor), and the key(s) (e.g., to be stored in data structure 126) are not accessible by software or by using an external interface to the computer system (e.g., a system-on-a-chip (SoC)). In some examples, the TME capability provides encryption protection to the external memory bus and / or memory.

[0064] In some examples, a multi-key TME (TME-MK) adds support for multiple encryption keys. In some examples, a computer system implementation supports a fixed number of encryption keys, and software can configure the computer system to use a subset of the available keys. In some examples, software manages the use of keys, and each of the available keys can be used for MK to allow for page-granularity encryption of memory, where the physical address specifies a key ID (KeyID). In some examples (e.g., by default), unless software explicitly specifies otherwise, a cryptographic circuit system (e.g., TME-MK) uses an encryption key (e.g., TME). In addition to supporting short-term keys generated by a processor (e.g., a central processing unit (CPU)) that are not accessible by software or through an external interface to the computer system, examples of TME-MK also support keys provided by software. In some examples, keys provided by software are used with non-volatile memory or when combined with a proof mechanism and / or with a key provisioning service. In some examples, fine-tuning keys for TME-MK are supplied by software. Some examples herein (e.g., platforms) use TME and / or TME-MK to prevent an attacker with physical access to the machine from reading the memory (e.g., and stealing any confidential information therein). In one example, the AES-XTS standard is used as the encryption algorithm to provide the desired security.

[0065] In some examples, each page of a memory page 128 includes a key for encrypting information, for example, and can thus be used to decrypt the encrypted information. In some examples, a key ID register is used with a page table (e.g., an extended and / or non-extended page table). In some examples, for instance, in the case where a cryptographic engine (e.g., a cryptographic circuit system) is part of a processor pipeline, the key ID register specifies the key itself. In some examples, the key ID register provides a key ID, for example, where a page table entry does not provide a key ID.

[0066] In some examples, the TME-MK cryptographic (e.g., encryption) circuitry maintains an internal key table that is not accessible by software and stores information (e.g., keys and encryption modes) associated with each key ID (e.g., the corresponding key ID for a corresponding encrypted memory block / page), where the key ID is incorporated into the physical address (e.g., incorporated in the page table) and also incorporated in each other storage location such as caches and TLBs. In one example, each key ID is associated with one of three encryption modes: (i) encrypted using a specified key, (ii) not encrypted at all (e.g., the memory will be in plaintext), or (iii) encrypted using the TME key. In some examples, unless otherwise specified by software, TME (e.g., TME-MK) defaults to using a hardware-generated short-term key that is not accessible by, for example, software or an external interface, and TME-MK also supports keys provided by software.

[0067] In some examples, PCONFIG is used to program the key ID attributes for TME-MK.

[0068] Table 1 below indicates an example TME-MK key table: Key ID Key Encryption mode (Item 1) (Item 1) (Item 1) (Item 2) (Item 2) (Item 2)

[0069] Table 2 below indicates an example PCONFIG leaf encoding:

[0070] Table 3 below indicates an example PCONFIG target (e.g., TME-MK encryption circuitry):

[0071] In a virtualization scenario, some examples herein allow a virtual machine monitor (VMM) or hypervisor to manage the use of keys to transparently support (e.g., traditional) operating systems without any changes (e.g., such that TME-MK can also be considered TME virtualization in such deployment scenarios). In some examples, the operating system (OS) is enabled to additionally utilize the TME-MK capabilities in both native and virtualized environments. In some examples, TME-MK is available to each guest OS in a virtualized environment, and the guest OS can utilize TME-MK in the same way as the native OS.

[0072] In some examples, computer system 100 includes memory controller circuit 116. In one example, a single memory controller circuit is used for multiple cores of computer system 100. Memory controller circuit 116 of processor 101 may receive an address for a memory access request (e.g., and for a store request, also receive payload data (e.g., ciphertext) to be stored at that address), and then perform a corresponding access into memory 120, e.g., via one or more memory buses 118. Each memory controller (MC) may have an identification value, e.g., "MC ID". The memory and / or (one or more) memory buses (e.g., its memory channels) may have an identification value, e.g., "channel ID". Each memory device (e.g., non-volatile memory 120 device) may have its own channel ID. Each processor (e.g., socket) (e.g., of a single SoC) may have an identification value, e.g., "socket ID". In some examples, memory controller circuit 116 includes a direct memory access engine 116A, e.g., for performing memory accesses into memory 120. The memory may be volatile memory (e.g., DRAM), non-volatile memory (e.g., non-volatile DIMM or non-volatile DRAM), and / or auxiliary (e.g., external) memory (e.g., not directly accessible by the processor), e.g., disk and / or solid state drive (e.g., Figure 7 the memory cells 728 therein). In some examples, memory controller circuit 116 is used to perform compression and / or decompression of data, e.g., where multiple bits (e.g., one or more bytes) of data that are repeated in a data row are removed to allow for compression based on such repetition (e.g., repetition-based compression / decompression).

[0073] In some examples, computer system 100 includes NIC circuit 132, e.g., for transmitting data over a network. In some examples, NIC circuit 132 includes cryptographic circuitry 134 (e.g., encryption and / or decryption circuitry) for encrypting (and / or decrypting) data, e.g., but without a core of a processor (e.g., processor die) and / or encryption (or decryption) circuitry that performs the encryption (or decryption). In cases where the NIC circuit is provided by a different vendor (e.g., manufacturer) than the socket (e.g., processor), in some examples, the NIC circuit is considered a security risk to the vendor (e.g., manufacturer) of the socket. In some examples, the encryption (and decryption) performed by NIC circuit 132 (e.g., via a request sent by the socket) is enabled or disabled. In some examples, NIC circuit 132 includes a remote DMA engine 136 for sending data over the network.

[0074] In one example, the hardware initialization manager (non-transitory) storage device 138 stores hardware initialization manager firmware (e.g., or software). In one example, the hardware initialization manager (non-transitory) storage device 138 stores Basic Input / Output System (BIOS) firmware. In another example, the hardware initialization manager (non-transitory) storage device 138 stores Unified Extensible Firmware Interface (UEFI) firmware. In certain examples (e.g., triggered by power-on or reboot of the processor), the computer system 100 (e.g., core 102) executes the hardware initialization manager firmware (e.g., or software) stored in the hardware initialization manager (non-transitory) storage device 138 to initialize the system 100 for operation (e.g., to start executing an operating system (OS)), and / or to initialize and test the (e.g., hardware) components of the system 100.

[0075] In certain examples, data is stored in the memory 120 as a single unit, e.g., a first data segment 130-1 stored on a first memory page and (e.g., at least in part) a second data segment 130-N stored on a second memory page (e.g., where N is any integer greater than 1).

[0076] In certain examples, the computer system (e.g., its memory controller circuit) implements a matching pair pattern (e.g., a “birthday pair” pattern) of data encryption and / or decryption. The following examples (e.g., patterns or sub-patterns) refer to the cache line width of data memory rows that can be utilized. In some examples, the cache line or memory row can be larger or smaller. Certain examples herein modify the input plaintext according to one or more of the examples (e.g., patterns or sub-patterns) herein to generate a modified plaintext. Certain examples herein use circuitry in the birthday pattern to modify the same input plaintext in different ways (e.g., when the same plaintext is to be encoded) to generate different outputs (e.g., ciphertexts) in multiple encryptions of the same input plaintext. In certain examples, a locator value (e.g., 8 bits / 1 byte wide) is used within a data row (e.g., a cache line), e.g., instead of within separate metadata or additional memory. In certain examples, the locator value (e.g., 8 bits / 1 byte wide) is used for: (i) identifying the positions of duplicate values within the modified plaintext (e.g., the modified plaintext including the locator value); and (ii) identifying the positions of duplicate values not within the modified plaintext to make room for the locator value.

[0077] Figure 2The figure illustrates the format of a data row 200 including a locator value 202 for encoding duplicate values according to an example of the present disclosure. Although the locator value is shown at the start at the first (e.g., leftmost) end of the data row 200, it should be understood that other positions may be used, e.g., where the pattern indicates to the memory controller circuit where the locator value is to be located in the modified data row (e.g., modified cache line). In some examples, the data row 200 (e.g., 512 bits) includes a first half 200A (e.g., upper 256 bits) and a second half 200B (e.g., lower 256 bits). In some examples, the data row 200 (e.g., the first half 200A) includes a first quadrant 200-1 (e.g., upper 128 bits in the first half 200A), a second quadrant 200-2 (e.g., lower 128 bits in the first half 200A), a third quadrant 200-3 (e.g., upper 128 bits in the second half 200B), and a fourth quadrant 200-4 (e.g., lower 128 bits in the second half 200B). Note that the quadrants are shown spaced apart to illustrate their boundaries, but it should be understood that all four quadrants are concatenated together within the data row 200.

[0078] In some examples, the data row 200 includes a plurality of elements (e.g., the 512-bit data row 200 includes 64 elements, where each element is 8 bits / 1 byte wide).

[0079] In some examples, the memory controller circuit (e.g., Figure 1 the memory controller circuit 116 in ) is configured to: receive a data row 200 (e.g., a single cache line) for writing to a memory (e.g., memory 120); search for duplicate values in the data row; determine that the duplicate values in the data row are identifiable using a locator value 202 for the duplicate values in the data row; in response to the determination, generate a locator value for the duplicate values in the data row; remove a second instance of the duplicate values from the data row and insert the locator value into the data row to generate a modified data row (e.g., modified plaintext); encrypt the modified data row (e.g., modified plaintext) into an encrypted data row; and cause the encrypted data row to be written to the memory.

[0080] Pattern

[0081] In some examples, one value of a duplicate value pair in the data row 200 is removed to make room (e.g., space) for the locator value in the modified data row. In some examples, the format of the locator value is based on one or more examples in the examples herein (e.g., a pattern or sub-pattern) to generate a modified data row (e.g., modified plaintext).

[0082] In some examples, in a first mode (e.g., a first sub - mode) (e.g., a first algorithm), a certain number (e.g., 16) of bits of a data row (e.g., plaintext) are encoded based on two sets of repeating values (e.g., bytes) (e.g., any "birthday pair"). In some examples, those numbers of bits (e.g., 16b / 2B) that are recovered (e.g., removed) are used for a locator value. In some examples, the locator value includes two bits to indicate a first block position and a second block position (e.g., 16B), e.g., the first or second block (e.g., quarter) indicated by the first bit of the locator value (e.g., respectively) set to 0 or 1, and the third or fourth block (e.g., quarter) indicated by the second bit of the locator value (e.g., respectively) set to 0 or 1.

[0083] In some examples, the locator value includes the four - bit position of the first byte in a block and the three - bit offset position of the second byte in the same block (e.g., can extend to wrap - around or adjacent blocks, extending to adjacent blocks can give more options as these are random bytes).

[0084] In some examples, the locator value includes four bits and three bits for identifying a byte in a second identified block (e.g., the last block can wrap around to the first block).

[0085] In some examples, if there are not two sets of valid repeating values (e.g., pairs encodable according to the format of the locator value), the first value of the locator value is set to indicate that no encoding is done using an invalid (e.g., 16b) locator value (e.g., xFFFF). In some examples, the memory controller circuit uses an error correction code (ECC) to correct this replaced (e.g., 16b) value as if it were corrupted data. In some examples, the original (e.g., byte) value that was replaced can alternatively be stored in an isolated memory (e.g., Figure 1 in the data structure 126 used for conflict resolution in it) such that the memory row can be restored to its original value.

[0086] In some examples, it is assumed that across all four blocks (e.g., quarters), there are typically more than 2 pairs of duplicate values, e.g., even for random data (e.g., where byte values are duplicated within a quarter approximately 40% of the time). In some examples, the memory controller circuit utilizes multiple sets of duplicate values for asymmetric encryption, because upon writing, the memory controller circuit (e.g., randomly) selects a first set of matching values (e.g., a first pair) for encoding and leaves the second (or third, fourth, etc.) set of matching values for the next selection, e.g., where the selection results in a ciphertext that is different / asymmetric compared to the ciphertext read.

[0087] In some examples, for instance, the modified data (e.g., modified plaintext) is encoded based on a domain key to prevent an adversary from performing a controlled replay across domains.

[0088] In some examples, in a second mode (e.g., second sub - mode) (e.g., second algorithm), a data row (e.g., plaintext) (e.g., 64 bytes) is divided into four equal - sized quarters (e.g., each having 128b / 16B), and the memory controller circuit (e.g., its encoding algorithm) searches for conflicts of values in the first quarter with those in the second quarter (e.g., at a single - byte granularity), and a value (e.g., one byte) in the third quarter with a value (e.g., one byte) in the fourth quarter. In some examples, the memory controller circuit then compresses the data row by two bytes (16 bits) and thus uses four - by - four bits to locate matching bytes in the quarters. In some examples, the memory controller circuit will on average find one matching pair for every two bytes of encoding (e.g., 16 * 16 / 2 = 128). In some examples, a location value (e.g., 0xFFFF) is obtained (e.g., reserved) to indicate that no matching value was found and no encoding will occur. In some examples, a location value is recycled as a locator position.

[0089] In some examples, in a third mode (e.g., third sub - mode) (e.g., third algorithm), the memory controller circuit is only used to encode one pair (e.g., one byte) in a half of a data row (e.g., plaintext) (e.g., 64 bytes), e.g., in one mode, the two duplicate values of a single pair need to be in the same half of the data row (e.g., and the locator value is included in that half of the data row).

[0090] In some examples, in a fourth mode (e.g., a fourth sub-mode) (e.g., a fourth algorithm), the memory controller circuit is used to expand such a pair of encodings of a third mode across data lines (e.g., 64B cache lines). In some examples, for single-byte encodings (e.g., single-byte locator values), having multiple pairs gives a choice of which pair to encode. When multiple alternative pairs are available (e.g., encoding one pair but not knowing which half it is in, there are two possible positions), this choice can also carry information. For example, in some examples, the highest byte value is always chosen for the encoded pair. This means that, on a read, when the memory controller circuit determines that there are multiple (e.g., unencoded) pairs, there are two alternative positions for the encoded byte, e.g., assuming the correct alternative encoded position is the one with the larger byte value among the two possible positions. Some examples herein choose to encode pairs with this property on a write. When there are multiple pairs, the examples herein further increase the efficiency of single-pair encoding to cover an entire data line (e.g., cache line) (e.g., covering all four quarters). In some examples, the remaining unencoded pairs indicate to the memory controller circuit which encoded position (e.g., half) is the correct position.

[0091] In some examples, for data that does not follow a uniform distribution, a probability distribution is assumed that results in a minimum number of collisions. In some examples, if the data follows any other probability distribution, e.g., similar to the characters in English text, many conflicts can be expected to occur (e.g., the space character is frequently repeated). This means that for non-random data, the proportion of cache lines that can be encoded rises, e.g., according to the examples herein, 98% of the lines are encodable, thus minimizing the need to access the conflict table and avoiding any associated performance impact.

[0092] In some examples, the memory controller circuit determines that data row 200 includes only one set of matching values (e.g., one pair), and all of these matching values are in the first half 200A of the data row. In some examples, such encoding is implemented using an eight-bit locator, e.g., such that the first five bits of the locator indicate which one of 32 different bytes within 256 bytes (e.g., 8 bits per slot × 32 slots = 256 bits) includes the first instance of the matching value that is still within the modified plaintext, and the other three bits of the locator indicate the offset (e.g., three-bit offset) within that half (e.g., within that quarter) of the second instance of the removed matching value (e.g., using a shift as discussed herein, using the removed space to store the locator value 202). In some examples, such decoding is implemented by the memory controller circuit because it does not detect other pairs, and it uses, for example, only the eight-bit encoding of one pair in the first half and uses the locator value to recreate the single pair in the first half. In some examples, the locator value is selected to indicate any bit split for absolute or relative indexing, e.g., an eight-bit locator value for cumulatively identifying two different byte positions, e.g., (i) using five bits to identify the first byte and using three bits to identify the second byte (e.g., the offset from the second byte), or (ii) using six bits to identify one byte out of 64 different bytes and using two bits to identify the second byte (e.g., the offset from the second byte) (e.g., a two-byte relative offset to that byte).

[0093] In some examples, the memory controller circuit determines that data row 200 does not include matching values, or includes only one set of matching values (e.g., one pair) and all of these matching values are in the second half 200B of the data row. In some examples, for example, such encoding is not implemented using an eight-bit locator format, and the locator value field 202 indicates that the memory row address is to be used as an index into Figure 1 the data structure 126 for conflict resolution in to determine the data value of the original plaintext that was removed (e.g., in the same bit position) (e.g., overwritten) by the locator value field 202 (e.g., the index stored in this field 202 into the data structure 126). In some examples, for example, such decoding is not implemented using an eight-bit locator format, and the locator value field 202 is instead used to store a value indicating that no compression of the plaintext was performed, e.g., indicating that the memory row address is into Figure 1The value of the index in the data structure 126 for conflict resolution, where the data structure 126 stores the data value that is removed (e.g., in the same bit position) (e.g., overwritten) by the locator / conflict value field 202 (e.g., the index into the data structure 126 stored in the field 202). In some examples, the locator conflict value is followed by an index into the data structure 126 for conflict resolution, e.g., where the conflict value and the index replace the original data currently stored in the data structure 126, thereby allowing the restoration of the full memory row while optimizing the memory usage for the data structure 126.

[0094] In certain examples, e.g., according to a pattern, the format for encoding the encoding (e.g., and the locator value) is the same as the format for decoding.

[0095] (One or more) additional locator bits

[0096] In certain examples, it is desirable to use additional locator bits (e.g., the 9th bit), however, the removal of a single value in a pair of repeated values (e.g., an octet / byte) only creates that amount (e.g., an octet) of space in the modified data row (e.g., the modified plaintext). In certain examples, the memory controller circuit includes a pattern that utilizes the additional locator bits.

[0097] In certain examples, when there are two or more pairs during a write, the additional locator bit (e.g., the 9th bit) is used to deterministically locate the encoded pair by identifying which half the encoded pair is in. In certain examples, the additional locator bit overlaps with more data, so the memory controller circuit will reconstruct the original data according to a rule, e.g., where the rule is: if the original data bit is 1, the largest or highest pair (e.g., the larger value in the two pairs of repeated values or the pair located farthest / highest from the start of the data row) is encoded; otherwise, the smallest or lowest pair (e.g., the smallest byte value or the pair closest to the start of the data row). In certain examples, if there are more than two pairs, for a 1 in this bit position (e.g., the 9th bit) in the original data (e.g., the unmodified data), the encoded pair is in the upper half, and for a 0 in this bit position (e.g., the 9th bit) in the original data (e.g., the unmodified data), the encoded pair is in the lower half.

[0098] Figure 3 Illustrating the format of a data row 200 including an additional locator bit 302 that is conditionally used for encoding repeated values according to an example of the present disclosure from Figure 2 In certain examples, the additional locator bit 302 (e.g., the 9th bit) is adjacent to the locator value 202 (e.g., bits 1 - 8) of the modified data row (e.g., the modified plaintext).

[0099] In some examples, if there are multiple duplicate value pairs (e.g., a first pair with a duplicate byte value of 6, and a second pair with a duplicate byte value of 0 in the pair), the additional locator bit 302 (e.g., the ninth bit) determines which half of the data row (e.g., cache line) the encoded pair is in (e.g., otherwise assuming a single pair is in the first half encoded with only the locator value 202 for locating the duplicate byte value and the three-bit offset value 306 of the locator value 202, e.g., with wraparound within that quarter to locate the byte replaced by the locator). This allows any one of the pairs within any quarter to be encoded.

[0100] In some examples, the memory controller circuit (e.g., in the "additional locator bit" mode) determines that there are two matching value pairs (e.g., a first pair with a first matching value and a second pair with a second matching value), and thus, due to the presence of two or more pairs, the memory controller circuit will use the additional locator bit 302 (e.g., the 9th bit). In some examples, the reason two or more pairs are needed to use the locator (e.g., 9th) bit is that the choice of which pair to encode is used to recover the data bit replaced by the ninth locator bit.

[0101] There are two (or more) pairs in the first half (no pairs in the second half)

[0102] In some examples, during a memory write, if the original (e.g., ninth) data bit (e.g., at position 302 in the data row) is zero, the memory controller circuit will encode the lowest pair (e.g., the pair at a lower relative position compared to another pair), and if the pair is in the first half of cache line 200A, the ninth bit 302 is set to zero, otherwise the ninth bit 302 is set to 1, indicating that the encoded pair is in the second half of cache line 200B, and if the original data bit is 1, the memory controller circuit will encode the highest pair (e.g., the pair at a higher relative position), and if the encoded pair is in the first half of cache line 200A, the ninth bit 302 is set to 0, otherwise if the highest pair is in the second half of cache line 200B, the ninth bit 302 is set to 1. In this way, the ninth bit locates which half the encoded pair is in, and the original data bit replaced by the ninth bit is determined by which pair (e.g., higher position or lower position) is encoded.

[0103] In some examples, decoding of a modified data row (e.g., modified plaintext) using an additional locator value 302 includes the memory controller circuit determining whether the modified data row includes an uncoded value pair (e.g., a matching byte value pair visible within the same quarter), and thus the memory controller circuit will assume that if the modified data row includes an uncoded value pair, the additional locator value 302 is used, thereby allowing coded pairs to be positioned across two halves of a cache line.

[0104] In some examples, upon decoding, an additional locator value 302 set to zero indicates to the memory controller circuit that a matching (e.g., repeated) value pair encoded by a locator value 202 is within the first half of the data row. In some examples, the memory controller circuit determines that the pair encoded by the locator value 202 and the ninth bit 302 is in a lower position relative to a second (e.g., uncoded) pair, and sets the bit position that previously stored the additional locator value 302 to zero, or else sets the bit to 1, e.g., to generate the original plaintext (e.g., where the memory controller circuit will further recover the data (e.g., bytes) encoded by the locator value 202).

[0105] Two (or more) pairs in the second half (and no pairs in the first half)

[0106] In some examples, if the original (e.g., ninth) data bit (in, e.g., a data row) is 0, the memory controller circuit will encode the minimum pair (e.g., the pair at the lower relative position), and if the encoded pair is in the second half, overwrite the data bit with 1, and if the original data bit is 1, the memory controller circuit will encode the highest pair (e.g., the pair at the higher relative position), and if the encoded pair is in the second half, overwrite the data bit with 1.

[0107] In some examples, decoding of a modified data row (e.g., modified plaintext) using an additional locator value 302 includes the memory controller circuit determining whether the modified data row includes an uncoded value pair (e.g., a matching byte value pair visible within the same quarter), and thus the memory controller circuit will assume that if the modified data row includes an uncoded value pair, the additional locator value 302 is used, thereby allowing coded pairs to be positioned across two halves of a cache line.

[0108] In some examples, an additional locator value 302 set to 1 indicates to the memory controller circuit that a matching (e.g., repeated) value pair encoded by the locator value 202 is within the second half of the data row. In some examples, the memory controller circuit determines that the pair encoded by the locator value 202 is in a lower position relative to a second (e.g., unencoded) pair and sets the bit 302 that previously stored the additional locator value to zero, or otherwise sets the bit to 1, e.g., to generate the original plaintext (e.g., where the memory controller circuit will further recover the data (e.g., bytes) encoded by the locator value 202).

[0109] Two pairs, one pair in each half

[0110] In some examples, if the original data bit 302 (e.g., in the data row) is 0, the memory controller circuit encodes the pair in the lower half (e.g., lower position), sets the ninth bit to 0, and if the original data bit is 1, the memory controller circuit encodes the pair in the upper half (e.g., higher position), sets the ninth bit to 1.

[0111] In some examples, decoding of the modified data row (e.g., modified plaintext) using the additional locator value 302 includes the memory controller circuit determining whether the modified data row includes, e.g., an unencoded value pair, and thus the memory controller circuit will assume that if the modified data row includes an unencoded value pair, the additional locator value 302 is used.

[0112] In some examples, an additional locator value 302 set to 0 indicates to the memory controller circuit that a matching (e.g., repeated) value pair encoded by the locator value 202 is within the first half of the data row, and an additional locator value 302 set to 1 indicates to the memory controller circuit that a matching (e.g., repeated) value pair encoded by the locator value 202 is within the second half of the data row.

[0113] In some examples, the memory controller circuit determines that the pair encoded by the locator value 202 and the additional locator (e.g., ninth) bit 302 is in a lower position (e.g., lower half) and sets the bit that previously stored the additional locator value 302 to zero, or otherwise sets the bit to 1, e.g., to generate the original plaintext (e.g., where the memory controller circuit will further recover the data (e.g., bytes) encoded by the locator value 202).

[0114] In some examples, this form can be extended to data rows (e.g., cache lines) having three or more pairs of data. For example, in cases with more pairs, more choices can be made. For instance, if the encoded pair is in the upper half set of all pair positions, an additional (e.g., 9th) data bit is restored to 1, and if the encoded pair is in the lower half set of all pair positions, the additional (e.g., 9th) data bit is restored to 0. In some examples, the memory controller circuit (e.g., during the creation of the modified data row) can select any pair from the upper or lower set of pair positions. For example, all pairs can be in the same quarter, yet half will be in the upper set of pair positions and half in the lower set of pair positions.

[0115] In some examples, the solution to the off-by-one problem relies on maintaining the quarters. For example, moving the quarter having the compressed / encoded pairs to the front (e.g., adjacent to the locator as the Figure 2 and Figure 3 first byte in

[0116] In some examples, since the position of the encoded pair is deterministic due to the utilization of the ninth bit, which quarter should be moved to the start of the row is also deterministic. For example, if quarter 1 is the position where the encoded pair is located, no quarter needs to be moved. For example, if quarter 2 is the position where the encoded pair is located, the positions of quarter 2 and quarter 1 are swapped. For example, if quarter 3 is the position where the encoded pair is located, the positions of quarter 3 and quarter 1 are swapped. For example, if quarter 4 is the position where the encoded pair is located, the positions of quarter 4 and quarter 1 are swapped, thus ensuring that the locator value is always located at the start (or in the same position) of the memory row together with the affected quarter.

[0117] In some examples, assuming all data appears random, if a conflict indicator in a stream of matching pairs patterns (e.g., "birthday pair" pattern) is set outside of data encryption, this will improve access control (e.g., detecting memory accesses using the wrong key). In some examples, the memory lookup step may also store the correct key ID, key hash, or integrity value used to initially encrypt the stored cache line, which matches the key currently used to access the memory line. In some examples, if a row cannot be encoded, the data is "stamped" with a conflict indicator value. In some examples, the indicator overwrites the data, so the conflict table is used to store the original data (e.g., and in some examples, this results in a performance impact because the memory is now accessed twice: once for the data row and once to obtain the original data from the conflict table). Some examples in this document use the (e.g., physical) address of the data row as an index into this conflict table (e.g., as an indexed array) to find the correct entry. In addition to storing the data overwritten by the conflict indicator in the conflict table, some examples also store the key ID used to encrypt the data (or store the key hash or integrity hash). In some examples that perform these two memory operations, the value (e.g., key ID, key hash, or integrity value) can also be used to check the access control for the data row (e.g., to check if the stored key ID in the conflict table for the address of the data row matches the key ID used to access the data row).

[0118] Some examples in this document (e.g., of a memory controller circuit) detect access control violations by the following operations: decoding the data row and observing whether the encoding rules are not followed (e.g., this is based on which pair is selected for encoding, e.g., if there are three pairs when using the ninth bit algorithm during encoding, the pair at the highest or lowest position should have been encoded, but if a pair in the middle position is found to be encoded during decoding, an access control violation or ciphertext corruption can be detected), or noticing that the row could not originally be encoded (e.g., so a memory lookup is performed anyway, which can also perform an access control check).

[0119] In some examples, the data row is all zeros, and the memory controller circuit matches the encoding of the all-zero row because each byte can be paired since they are all the same value (0), and the encoding rate is 100% (and there are always pairs to encode). Randomly picking the byte pairs to be encoded results in different ciphertexts for the same plaintext (all zeros). In some examples, using the ninth-bit algorithm, one can pick from the halves of the pair positions corresponding to the encoding of the ninth bit, again allowing 255 possible encodings for the all-zero row, resulting in 255 different possible ciphertexts. Also note that if an encrypted zero row is corrupted or read with the wrong key, it will decrypt to a random situation where access control checks can be applied in some examples. In some examples, there is a threshold on the number of pairs (e.g., three pairs) to determine when to use access control (e.g., where it only applies to random or corrupted data, e.g., because decrypted data revealing many matching pairs is unlikely to be corrupted).

[0120] In some examples, a modified data row (e.g., a modified cache row) including a locator value is then encrypted, for example, according to a key as discussed herein, and then the encrypted version of the modified data row is stored. In some examples, the encrypted version of the modified data row is decrypted, for example, according to the examples discussed herein (e.g., sub-patterns), and then the modified data row is returned to the original data row (e.g., plaintext) by the memory controller circuit in a matching pair pattern (e.g., "birthday pair" pattern).

[0121] In some examples, the modified data row (e.g., the entire data row) is encrypted by a block cipher (e.g., a symmetric-key tweakable block cipher, e.g., the Threefish cipher). In some examples, the memory address can also be used as a tweak. In some examples, the block cipher diffuses the changes caused by the alternating pair encoding across the entire memory row, and for any change in the encoded pairs, the result is a completely different ciphertext. In some examples, the CBC mode diffuses the encoding completely across the entire memory row. The CBC mode can also include the memory row address to further locate the ciphertext. In some examples, additional bits beyond the ninth bit can be encoded similarly, e.g., when the cache row is 128 bytes long, when 4 or more pairs are available for recombining the original ninth and tenth data bit values, the tenth bit can be used to determine on which side of the row the encoded pairs are located, and so on.

[0122] Rules for data integrity:

[0123] In the case of multiple pairs, there can be rules that also serve as access control and / or integrity without any additional encoding. For example, if the rule is that the highest value pair is the encoded value pair, then on an invalid read (e.g., using the wrong key or reading a corrupted written line from memory), if the encoded byte value is lower than another encodable pair, this is a violation of the rule and is detected as a violation of access control and / or integrity. In some examples of the nine-bit algorithm, if there are three pairs at the time of encoding, the rule is to encode the highest or lowest pair position. This means that at the time of decoding, if the middle pair position is found to be encoded, an access control violation or data corruption is detected. In some examples, when many pairs are detected at the time of decoding, the data is assumed to be legitimate because an incorrectly decrypted ciphertext should result in randomly decrypted data with a minimum number of matching pairs.

[0124] To cover the entire 64-byte cache line, embodiments can also use a nine-bit locator that permutes nine bits of the duplicate data. Byte alignment can still be maintained, where the first six bits of the nine-bit locator position the byte-aligned duplicate nine-bit value within the 64-byte cache line, and the remaining three bits of the nine-bit locator identify the byte-aligned position within the same quarter (with wraparound) of the duplicate nine-bit value to be replaced by the nine-bit locator. The locator can then be positioned at the start (or in an embodiment, the end) of the cache line, concatenating (shifting) all the remaining bits together to fill the gap left by the duplicate nine-bit value removed to make room for the nine-bit locator. For the special case of adjacency, where the last bit of the first byte-aligned nine-bit value overlaps the first bit of the duplicate byte-aligned nine-bit value, it is assumed that the second nine-bit value is not byte-aligned but shifted by one bit so as not to overlap the last bit of the first duplicate nine-bit value. Similarly, if the six bits of the locator identify the last byte position within a quarter as the position of the first duplicate nine-bit value, it can be assumed that the last bit of the duplicate nine-bit value wraps around to the start of the quarter in which it is located. In this way, for all four quarters, based on the birthday bound probability (about 20%) of nine-bit value collisions within a quarter, an encoding rate of about 60% can be achieved for even random data while maintaining the byte alignment typical of computer data. Similar embodiments exist for 10-bit locators, 11-bit locators, etc., allowing encoding to cover larger-sized cache lines.

[0125] Key Refresh:

[0126] In some examples, a matching pair pattern (e.g., a "birthday pair" pattern) is used with key refreshing. For example, in the matching pair pattern, the memory encryption key is changed periodically. In some examples, because the matching pair pattern (e.g., the "birthday pair" pattern) can generate many (e.g., 100) alternative ciphertexts for the same plaintext, it fills the gap between periodic key refreshes. In some examples, when the encryption key is changed, an entirely new ciphertext is generated even for the exact same plaintext.

[0127] Figure 4 FIG. illustrates an example of operation 400 of a method for performing a read from a memory with duplicate value encoding in accordance with an example of the present disclosure. Some or all of operation 400 (or other processes described herein, or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. The code is stored, for example, in the form of a computer program including instructions executable by one or more processors on a computer-readable storage medium. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operation 400 are performed by a component (e.g., memory controller circuit 116) in other figures.

[0128] Operation 400 includes: at block 402, retrieving a data row from the memory given a particular (e.g., physical) address. Operation 400 further includes: at block 404, decrypting the data row (e.g., using a specified key ID, identified key, and / or tweak). Operation 400 further includes: at block 406, checking whether a portion of the data row is a conflict indicator value (e.g., a lookup indicator (IL)), and if so, proceeding to block 408, and if not, proceeding to block 412. Some examples place the conflict indicator test 406 before the decryption 404 of the row. Operation 400 further includes: at block 408, reading a conflict resolution data structure, for example, by using the address of the data row as an index into an array structure, to determine the corresponding (e.g., original) value and replacing the conflict indicator value with the correct value to reproduce the original data row. Operation 400 further includes: at block 410, forwarding the data to a cache (e.g., Figure 1 cache 112 in or other figures (e.g., Figure 8one or more caches in)). Operation 400 further includes: at block 412, searching the decrypted data row for encodable pairs (e.g., duplicate values (e.g., duplicate byte values), e.g., duplicates within a single quarter). Operation 400 further includes: at block 414, checking whether there are one or more encodable pairs (e.g., one or more sets of duplicate values within a quarter), and if so, proceeding to block 420, and if not (no encodable pairs), proceeding to block 416, e.g., remembering which half of the data row the encoded pair is in. Operation 400 further includes: at block 416, using the first part of the locator value within the data row (e.g., Figure 3 the 5-bit locator value 304 of the locator value 202 in data row 200 in) to identify the encoded (e.g., byte) value position (e.g., the position of the duplicate value in the data row, the value of which will be copied / inserted to refill the deleted instance of that value). Operation 400 further includes: at block 418, using the second part of the locator value within the data row (e.g., Figure 3 the 3-bit locator value 306 of the locator value 202 in data row 200 in) to identify the position of the missing (e.g., byte) value (e.g., the position in the data row where the duplicate value is to be copied / inserted to refill the deleted instance of that value within the quarter, to restore the original data row to be forwarded to the cache at 410). Operation 400 further includes: at block 420, checking additional locator values (e.g., Figure 3whether the additional locator value 302 (e.g., the 9th bit) in [it] is zero, and if so, proceed to block 422, and if not, proceed to block 424. Operation 400 further includes: at block 422, if the additional locator value is zero, determining that the encoded pair is in the first half of data row 200A. Operation 400 further includes: at block 424, if the additional locator value is not zero, determining that the encoded pair is in the second half of data row 200B. Operation 400 further includes: at block 426, checking whether the encoded pair (e.g., the encoded position determined by the locator value and the 9th bit) is in the upper half (e.g., the highest position) of all pair positions in the data row, and if not (compared to all other encodable pairs found at step 412, the encoded pair is not the highest or not in the upper half of the pair position), proceed to block 428, and if so, proceed to block 430. An example of implementing access control may further check whether there are two encodable pairs and the position of the third encodable pair is in an intermediate position, neither the highest nor the lowest position, thereby triggering an access control violation error by, for example, contaminating a cache line. Operation 400 further includes: at block 428, if the check at 426 is no, setting the additional locator bit position 302 to 0 (e.g., restoring the original data value to 0), then, remembering which half of the data row contains the encoded pair (422 or 424), and sending the modified data row to block 416. Operation 400 further includes: at block 430, if the check at 426 is yes, setting the additional locator bit position 302 to 1 (e.g., restoring the original data value to 1), then, remembering which half of the data row contains the encoded pair (422 or 424), and sending the modified data row to block 416.

[0129] Figure 5 FIG. illustrates an example of operation 500 of a method for performing a write to a memory with duplicate value encoding according to an example of the present disclosure. Some or all of operation 500 (or other processes described herein, or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. The code is stored, for example, in the form of a computer program including instructions executable by one or more processors, on a computer-readable storage medium. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operation 500 are performed by (one or more) components in other figures (e.g., memory controller circuit 116).

[0130] Operation 500 includes: at block 502, receiving a data row for writing to memory (e.g., from a processor or processor cache). Operation 500 further includes: at block 504, searching the decrypted data row for encodable pairs (e.g., repeating values (e.g., repeating byte values), e.g., repeating within a single quarterword (for all quarterwords)). Operation 500 further includes: at block 506, checking if there is at least one encodable pair (e.g., a set of repeating values within a quarterword), and if so, proceeding to block 514, and if not, proceeding to block 508. Operation 500 further includes: at block 508, storing the original value of the data row (e.g., a value in the same position and having the same width as the locator value) into a data structure (e.g., a conflict table indexed by the memory row address) (e.g., Figure 1 data structure 126 in Figure 3 ), and replacing the original value in the data row with a locator value that identifies a conflict with the encoding scheme and that the data cannot be encoded due to the absence of a matching pair. Operation 500 further includes: at block 510, encrypting the modified data row (e.g., using a specified key ID, identified key, and / or tweak (e.g., the physical address of the data row as a tweak)). Some examples alternatively set a conflict indicator after encrypting the data row and store the encrypted portion corresponding to the locator value's position and size in the conflict table after 510, such that the conflict table does not need to be additionally encrypted and access control is improved. Operation 500 further includes: at block 514, checking if there are multiple encodable pairs (e.g., multiple corresponding sets of repeating values within one or more quarterwords), and if so, proceeding to block 524, and if not, proceeding to block 516 if there is only one encodable pair. Operation 500 further includes: at block 516, checking if there is a pair in the first half of the data row, and if so, proceeding to block 518, and if not, proceeding to block 508 since there is no encoding for only one pair in the second half of the data row. Operation 500 further includes: at block 518, knowing in which half the encodable pair is, locating the first instance of the repeating value of the encodable pair within that half (e.g., of the repeating value), and generating a first part of the locator value (e.g., Figure 3The locator value 306 of the 3-bit locator value 202 in the data row 200 in (e.g., bytes) to identify the position of the value to be deleted (and in some examples, swap the quarter with the encoded pair with the first quarter 200-1). Operation 500 further includes: at block 520, shift (e.g., shift right) the leftmost bit to remove the second instance of the duplicate value of the encodable pair, e.g., thus deleting the second instance of the duplicate value of the encodable pair, thereby making room for the locator value in the data row (or the quarter with the encoded pair therein). Operation 500 further includes: at block 522, insert (e.g., concatenate) the locator value (e.g., 8-bit wide) into the space created by the shift at block 520, e.g., to generate a modified data row including the locator value (e.g., where the modified data row has the same width as the data row retrieved at block 502 (e.g., 512 bits)). The example may further swap the quarter with the encoded pair with the first quarter to only shift the bytes of the affected quarter. Operation 500 further includes: at block 524, read the data row for additional locator bits (e.g., Figure 3The data bit value (e.g., the 9th bit) of the additional locator value 302 (e.g., the 9th bit) in [ ]. Operation 500 further includes: at block 526, checking whether the data bit is zero, and if so, proceeding to block 528, and if not, proceeding to block 530. Operation 500 further includes: at block 528, if the check at 526 is yes, selecting the pair to be encoded from the lower set (e.g., the lowest set) of encodable pair positions. Operation 500 further includes: at block 530, if the check at 526 is no, selecting the pair to be encoded (e.g., randomly) from the higher set (e.g., the highest set) of encodable pair positions. In some examples, a list of encodable pairs is generated at block 504 (e.g., in the list, each pair is within the same quarter), and then the list is sorted based on the positions of the pairs. In some examples, the sorted list is split in half. In some examples, (e.g., 512b) the data row has 64 elements with indices from 1 to 64, and the list generated at block 504 indicates a first matching value pair [3,6] at indices 3 and 6, and a second matching value pair [4,11] at indices 4 and 11, and thus the first indices from these two pairs are used to sort the list {3,4}. In some examples, the pairs selected from the set of encodable pairs for encoding are changed for the same plaintext, e.g., to generate different ciphertexts for multiple encodings of the same plaintext. As another example, for the "9th bit" algorithm, the memory controller circuit will encode which original data bit the 9th bit is replacing, so if the data bit is 0, it encodes the lowest set (e.g., the pairs starting from index 3 in the above example), and if the data bit is 1, it encodes the highest set (e.g., the pairs starting from index 4 in the above example). Examples that wish to perform access control may additionally check whether there are three encodable pairs, and as a rule, encode only the highest or lowest pair positions based on the value of the 9th data bit. Operation 500 further includes: at block 532, checking whether the encoded pair is in the first half (e.g., 200A) or the second half (e.g., 200B) of the data row, and if so (it is in the first half), proceeding to block 534, and if not in the first half, proceeding to block 536. Operation 500 further includes: at block 534, if the check at 532 is yes, setting the additional locator bit position (e.g., the 9th bit) to 0, e.g., indicating that the encoded pair is in the first half 200A, and then sending the modified data row to block 518, remembering which half it is in. Operation 500 further includes: at block 536, if the check at 532 is no, setting the additional locator bit (e.g., the 9th bit) to 1, e.g., indicating that the encoded pair is in the second half 200B, and then sending the modified data row to block 518.

[0131] In some examples, there are triples with the same value (e.g., 3 elements (e.g., bytes) with the same value within a quarter section), and these triples are encoded as multiple pairs. For example, the first value and the middle value produce a locator value, and the middle value and the last value produce different locator values that become alternating pairs. In some examples, the first value and the last value can be the third pair.

[0132] Figure 6 FIG. illustrates another example of operation 600 of a method for performing a read from a memory with repeated value encoding according to an example of the present disclosure. Some or all of operation 600 (or other processes described herein, or variations and / or combinations thereof) are performed under the control of one or more computer systems configured with executable instructions and implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or a combination thereof. The code is stored on a computer-readable memory executable by one or more processors. The computer-readable storage medium is non-transitory. In some examples, one or more (or all) of operation 600 are performed by one or more components (e.g., memory controller circuit 116) in other figures.

[0133] Operation 600 includes: at block 602, executing instructions by execution circuitry to generate a memory request to read a data row from the memory. Operation 600 further includes: at block 604, decrypting the data row to a decrypted data row by the memory controller circuit. Operation 600 further includes: at block 606, determining by the memory controller circuit that a field of the decrypted data row is set to a locator value for a repeated value. Operation 600 further includes: at block 608, identifying by the memory controller circuit a first position of a first instance of the repeated value in the decrypted data row based on the locator value. Operation 600 further includes: at block 610, reading the repeated value from the first position in the decrypted data row by the memory controller circuit. Operation 600 further includes: at block 612, identifying by the memory controller circuit a second position of a second instance of the repeated value in the decrypted data row based on the locator value. Operation 600 further includes: at block 614, shifting the decrypted data row by the memory controller circuit to remove the locator value from the decrypted data row and create space for the repeated value to be inserted into the second position. Operation 600 further includes: at block 616, inserting the repeated value into the space within the decrypted data row by the memory controller circuit to generate a result data row.

[0134] Some examples utilize the instruction formats described herein. Some examples are implemented in one or more computer architectures, cores, accelerators, etc. Some examples are generated or are IP cores. Some examples utilize simulation and / or translation.

[0135] At least some examples of the disclosed technology may be described in accordance with the following examples.

[0136] In a set of examples, a device (e.g., a hardware processor) includes: execution circuitry to execute instructions to generate a memory request to read a data row from a memory; and a memory controller circuit to: decrypt the data row into a decrypted data row, determine that a field of the decrypted data row is set to a locator value for a duplicate value, identify a first location of a first instance of the duplicate value in the decrypted data row based on the locator value, read the duplicate value from the first location in the decrypted data row, identify a second location of a second instance of the duplicate value in the decrypted data row based on the locator value, shift the decrypted data row to remove the locator value from the decrypted data row and create space for the duplicate value to be inserted into the second location, and insert the duplicate value into the space within the decrypted data row to generate a result data row. In some examples, the memory controller circuit is to shift bits in the decrypted data row to the left of the second location by a width of the duplicate value to remove the locator value and create space for the duplicate value to be inserted into the second location, and not shift bits in the decrypted data row to the right of the second location. In some examples, the memory controller circuit is to determine that a field of the decrypted data row is not set to a conflict indicator value, and in response to determining that the field of the decrypted data row is not set to a conflict indicator value, perform the following operations: identify the first location, read, identify the second location, shift, and insert. In some examples, the locator value includes a first value and a second value, the first value for indicating a first location of a first instance of the duplicate value within a first appropriate subset of the decrypted data row, and the second value for indicating an offset within a second appropriate subset of the decrypted data row. In some examples, the memory controller circuit is further to check another locator bit of the decrypted data row, wherein a bit set to the first value indicates to the memory controller circuit that the first location and the second location of the duplicate value are in a first half of the decrypted data row, and a bit set to the second value indicates to the memory controller circuit that the first location and the second location of the duplicate value are in a second half of the decrypted data row. In some examples, the memory controller circuit is further to: receive a second data row for writing to the memory; search for duplicate values in the second data row; use a second locator value for the duplicate values in the second data row to determine that the duplicate values in the second data row are identifiable; in response to the determination, generate a second locator value for the duplicate values in the second data row, remove a second instance of the duplicate value from the second data row, and insert the second locator value into the second data row; encrypt the second data row including the second locator value into an encrypted data row; and cause the encrypted data row to be written to the memory.In some examples, the memory controller circuit is further configured to: before encryption, in response to a first instance and a second instance of a repeated value in a second data row being located in a first half of the second data row, set another locator bit of the second data row to a first value, and in response to the first instance and the second instance of the repeated value in the second data row being located in a second half of the second data row, set the another locator bit of the second data row to a second value.

[0137] In another set of examples, a method includes: executing instructions by an execution circuitry to generate a memory request to read a data row from a memory; decrypting the data row to a decrypted data row by a memory controller circuitry; determining by the memory controller circuitry that a field of the decrypted data row is set to a locator value for a duplicate value; identifying by the memory controller circuitry a first position of a first instance of the duplicate value in the decrypted data row based on the locator value; reading by the memory controller circuitry the duplicate value from the first position in the decrypted data row; identifying by the memory controller circuitry a second position of a second instance of the duplicate value in the decrypted data row based on the locator value; shifting by the memory controller circuitry the decrypted data row to remove the locator value from the decrypted data row and to create space for the duplicate value to be inserted into the second position; and inserting by the memory controller circuitry the duplicate value into the space within the decrypted data row to generate a result data row. In some examples, shifting includes: shifting bits in the decrypted data row by a width of the duplicate value to the left of the second position to remove the locator value and to create space for the duplicate value to be inserted into the second position, and not shifting bits in the decrypted data row to the right of the second position. In some examples, the method includes: determining by the memory controller circuitry that a field of the decrypted data row is not set to a conflict indicator value, and in response to determining that the field of the decrypted data row is not set to the conflict indicator value, performing: identifying the first position, reading, identifying the second position, shifting, and inserting. In some examples, the locator value includes a first value and a second value, the first value for indicating a first position of a first instance of the duplicate value within a first appropriate subset of the decrypted data row, and the second value for indicating an offset within a second appropriate subset of the decrypted data row. In some examples, the method includes: checking by the memory controller circuitry another locator bit of the decrypted data row, wherein a bit set to the first value indicates to the memory controller circuitry that the first position and the second position of the duplicate value are in a first half of the decrypted data row, and a bit set to the second value indicates to the memory controller circuitry that the first position and the second position of the duplicate value are in a second half of the decrypted data row.In some examples, the method includes: receiving, by a memory controller circuit, a second data row for writing to a memory; searching the second data row for duplicate values; using, by the memory controller circuit, a second locator value for the duplicate values in the second data row to determine that the duplicate values in the second data row are identifiable; in response to the determination, generating, by the memory controller circuit, a second locator value for the duplicate values in the second data row, removing a second instance of the duplicate values from the second data row, and inserting the second locator value into the second data row; encrypting, by the memory controller circuit, the second data row including the second locator value into an encrypted data row; and causing, by the memory controller circuit, the encrypted data row to be written to the memory. In some examples, the method includes: before encryption, setting, by the memory controller circuit, another locator bit of the second data row to a first value in response to a first instance and a second instance of the duplicate values in the second data row being located in a first half of the second data row, and setting, by the memory controller circuit, another locator bit of the second data row to a second value in response to the first instance and the second instance of the duplicate values in the second data row being located in a second half of the second data row.

[0138] In yet another set of examples, a system includes: a memory; execution circuitry for executing instructions to generate a memory request to read a data row from the memory; and a memory controller circuit for: decrypting the data row into a decrypted data row, determining that a field of the decrypted data row is set to a locator value for a duplicate value, identifying a first location of a first instance of the duplicate value in the decrypted data row based on the locator value, reading the duplicate value from the first location in the decrypted data row, identifying a second location of a second instance of the duplicate value in the decrypted data row based on the locator value, shifting the decrypted data row to remove the locator value from the decrypted data row and create space for the duplicate value to be inserted into the second location, and inserting the duplicate value into the space within the decrypted data row to generate a resultant data row. In some examples, the memory controller circuit is used to shift bits in the decrypted data row to the left of the second location by the width of the duplicate value to remove the locator value and create space for the duplicate value to be inserted into the second location, and not shift bits in the decrypted data row to the right of the second location. In some examples, the memory controller circuit is used to determine that a field of the decrypted data row is not set to a conflict indicator value, and in response to determining that the field of the decrypted data row is not set to a conflict indicator value, perform the following operations: identify the first location, read, identify the second location, shift, and insert. In some examples, the locator value includes a first value and a second value, the first value for indicating a first location of a first instance of the duplicate value within a first appropriate subset of the decrypted data row, and the second value for indicating an offset within a second appropriate subset of the decrypted data row. In some examples, the memory controller circuit is further used to check another locator bit of the decrypted data row, wherein a bit set to the first value indicates to the memory controller circuit that the first location and the second location of the duplicate value are in the first half of the decrypted data row, and a bit set to the second value indicates to the memory controller circuit that the first location and the second location of the duplicate value are in the second half of the decrypted data row. In some examples, the memory controller circuit is further used to: receive a second data row for writing to the memory; search for duplicate values in the second data row; use a second locator value for the duplicate values in the second data row to determine that the duplicate values in the second data row are identifiable; in response to the determination, generate a second locator value for the duplicate values in the second data row, remove a second instance of the duplicate value from the second data row, and insert the second locator value into the second data row; encrypt the second data row including the second locator value into an encrypted data row; and cause the encrypted data row to be written to the memory.In some examples, the memory controller circuit is further configured to: before encryption, in response to a first instance and a second instance of a repeated value in a second data row being located in a first half of the second data row, set another locator bit of the second data row to a first value, and in response to the first instance and the second instance of the repeated value in the second data row being located in a second half of the second data row, set the another locator bit of the second data row to a second value.

[0139] Exemplary architectures, systems, etc. that can be used above are described in detail below.

[0140] Exemplary architecture

[0141] A description of an exemplary computer architecture is detailed below. Other system designs and configurations of laptops, desktops, handheld personal computers (PCs), personal digital assistants, engineering workstations, servers, blade servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices known in the art are also suitable. In general, various systems or electronic devices capable of incorporating a processor and / or other execution logic as disclosed herein are generally suitable.

[0142] Exemplary system

[0143] Figure 7 An illustrated exemplary computing system. The multi-processor system 700 is a system provided with interfaces and includes a plurality of processors or cores, the plurality of processors or cores including a first processor 770 and a second processor 780 coupled via an interface 750 such as a point-to-point (P-P) interconnect, a fabric, and / or a bus. In some examples, the first processor 770 and the second processor 780 are homogeneous. In some examples, the first processor 770 and the second processor 780 are heterogeneous. Although the exemplary system 700 is shown as having two processors, the system can have three or more processors, or can be a single-processor system. In some examples, the computing system is a system-on-chip (SoC).

[0144] Processors 770 and 780 are shown as including integrated memory controller (IMC) circuitry 772 and 782, respectively. Processor 770 also includes interface circuits 776 and 778; similarly, second processor 780 includes interface circuits 786 and 788. Processors 770, 780 may exchange information via interface circuits 778, 788, through interface 750. IMCs 772 and 782 couple processors 770, 780 to respective memories, namely memory 732 and memory 734, which may be portions of main memories locally attached to the respective processors.

[0145] Processors 770, 780 may each exchange information with network interface (NW I / F) 790 via interface circuits 776, 794, 786, 798 through respective interfaces 752, 754. Network interface 790 (e.g., one or more of an interconnect, bus, and / or fabric, and in some examples, a chipset) may optionally exchange information with coprocessor 738 via interface circuit 792. In some examples, coprocessor 738 is a specialized processor, such as, for example, a high throughput processor, a network or communication processor, a compression engine, a graphics processor, a general purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, and so on.

[0146] A shared cache (not shown) may be included in either processor 770, 780, or external to both processors but connected to these processors via an interface (such as, a P-P interconnect), such that if a processor is placed in a low power mode, local cache information of either or both processors may be stored in the shared cache.

[0147] The network interface 790 can be coupled to the first interface 716 via the interface circuit 796. In some examples, the first interface 716 can be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 716 is coupled to a power control unit (PCU) 717, which can include circuitry, software, and / or firmware for performing power management operations related to the processors 770, 780, and / or the coprocessor 738. The PCU 717 provides control information to a voltage regulator (not shown) to cause the voltage regulator to generate an appropriate regulated voltage. The PCU 717 also provides control information to control the generated operating voltage. In various examples, the PCU 717 can include various power management logic units (circuitry) for performing hardware-based power management. Such power management can be fully controlled by the processor (e.g., controlled by various processor hardware, and it can be triggered by workload and / or power, thermal constraints, or other processor constraints), and / or the power management can be performed in response to an external source (such as a platform or power management source or system software).

[0148] The PCU 717 is illustrated as existing as a logic separate from the processor 770 and / or the processor 780. In other cases, the PCU 717 can execute on a given one or more cores in the core (not shown) of the processor 770 or 780. In some cases, the PCU 717 can be implemented as a (dedicated or general-purpose) microcontroller or other control logic configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by the PCU 717 can be implemented external to the processor, such as by a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by the PCU 717 can be implemented within the BIOS or other system software.

[0149] A variety of I / O devices 714 can be coupled to a first interface 716 along with a bus bridge 718, and the bus bridge 718 couples the first interface 716 to a second interface 720. In some examples, one or more additional processors 715 (such as, a coprocessor, a high-throughput many integrated core (MIC) processor, a GPGPU, an accelerator (such as a graphics accelerator or a digital signal processing (DSP) unit), a field programmable gate array (FPGA), or any other processor) are coupled to the first interface 716. In some examples, the second interface 720 can be a low pin count (LPC) interface. A variety of devices can be coupled to the second interface 720, including, for example, a keyboard and / or a mouse 722, a communication device 727, and a storage circuitry 728. The storage circuitry 728 can be one or more non-transitory machine-readable storage media as described below, such as a disk drive or other mass storage device, which can include instructions / code and data 730 in some examples and can implement instruction storage. Further, an audio I / O 724 can be coupled to the second interface 720. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as the multi-processor system 700 can implement a multi-drop interface or other such architectures instead of the point-to-point architecture.

[0150] Example core architectures, processors, and computer architectures.

[0151] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores can include: 1) general-purpose in-order cores intended for general computing; 2) high-performance general-purpose out-of-order cores intended for general computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors can include: 1) a CPU that includes one or more general-purpose in-order cores intended for general computing and / or one or more general-purpose out-of-order cores intended for general computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors give rise to different computer system architectures, which can include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such a coprocessor is sometimes referred to as specialized logic or as a specialized core, such specialized logic such as, integrated graphics and / or scientific (throughput) logic); and 4) a system-on-chip (SoC) that can include the described CPU (sometimes referred to as the (one or more) application core or (one or more) application processor), the above-described coprocessor, and additional functionality on the same die.

[0152] An example core architecture is then described, followed by example processors and computer architectures.

[0153] Figure 8 A block diagram of an example processor and / or SoC 800 that can have one or more cores, and an integrated memory controller. The solid box diagram illustrates a processor 800 having a set with a single core 802A, a system agent unit circuitry 810, and a set of one or more interface controller unit circuitries 816, while the optional addition of the dashed box illustrates an alternative processor 800 having a set with multiple cores 802A - 802N, a set of one or more integrated memory controller units 814 in the system agent unit circuitry 810, and specialized logic 808 and a set of one or more interface controller unit circuitries 816. Note that the processor 800 can be Figure 7 one of the processors 770 or 780, or the coprocessors 738 or 715.

[0154] Accordingly, different implementations of the processor 800 may include: 1) a CPU, where the dedicated logic 808 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores, not shown), and the cores 802A - 802N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, a combination of both); 2) a coprocessor, where the cores 802A - 802N are a large number of dedicated cores designed primarily for graphics and / or scientific (throughput); and 3) a coprocessor, where the cores 802A - 802N are a large number of general-purpose in-order cores. Accordingly, the processor 800 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor can be implemented on one or more chips. The processor 800 can be part of one or more substrates and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, complementary metal-oxide-semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0155] The memory hierarchy includes one or more levels of (one or more) cache unit circuitry 804A - 804N within the cores 802A - 802N, a collection of one or more shared cache unit circuitry 806, and an external memory (not shown) coupled to a collection of (one or more) integrated memory controller unit circuitry 814. The collection of one or more shared cache unit circuitry 806 may include one or more intermediate levels of cache (such as level 2 (L2), level 3 (L3), level 4 (L4)) or other levels of cache (such as last level cache (LLC)) and / or combinations thereof. Although in some examples, the interface network circuitry 812 (e.g., ring interconnect) provides an interface to the dedicated logic 808 (e.g., integrated graphics logic), the collection of (one or more) shared cache unit circuitry 806, and the system agent unit circuitry 810, alternative examples use any number of well - known techniques for providing an interface to such units. In some examples, coherence is maintained between the collection of (one or more) shared cache unit circuitry 806 and one or more of the cores 802A - 802N. In some examples, the interface controller unit circuitry 816 couples the cores 802A - 802N to one or more other devices 818, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless network, wired network, etc.).

[0156] In some examples, one or more of the cores 802A - 802N are capable of implementing multithreading. The system agent unit circuitry 810 includes those components that coordinate and operate the cores 802A - 802N. The system agent unit circuitry 810 may include, for example, power control unit (PCU) circuitry and / or display unit circuitry (not shown). The PCU may be the logic and components required to regulate the power states of the cores 802A - 802N and / or the dedicated logic 808 (e.g., integrated graphics logic), or may include such logic and components. The display unit circuitry is used to drive one or more externally connected displays.

[0157] The cores 802A - 802N may be homogeneous in terms of instruction set architecture (ISA). Alternatively, the cores 802A - 802N may be heterogeneous in terms of ISA; that is, a subset of the cores 802A - 802N may be capable of executing the ISA, while other cores may be capable of executing only a subset of that ISA or another ISA.

[0158] Figure 9FIG. is a block diagram of a computing system 900 configured to implement one or more aspects of the examples described herein. The computing system 900 includes a processing subsystem 901 having one or more processors 902 and a system memory 904 communicating via an interconnect path that may include a memory hub 905. The memory hub 905 may be a separate component within a chipset component or may be integrated within one or more of the processors 902. The memory hub 905 is coupled to an I / O subsystem 911 via a communication link 906. The I / O subsystem 911 includes an I / O hub 907 that enables the computing system 900 to receive input from one or more input devices 908. Additionally, the I / O hub 907 enables a display controller (which may be included within one or more of the processors 902) to provide output to one or more display devices 190A. In some examples, one or more of the display devices 190A coupled to the I / O hub 907 may include a local, internal, or embedded display device.

[0159] The processing subsystem 901 includes, for example, one or more parallel processors 912 coupled to the memory hub 905 via a bus or other communication link 913. The communication link 913 may be one of any number of standard-based communication link technologies or protocols, such as, but not limited to, PCI Express (PCIe), or may be a vendor-specific communication interface or fabric. The one or more parallel processors 912 may form a parallel or vector processing system within a computing concentration that may include a large number of processing cores and / or processing clusters, such as, for example, a many integrated core (MIC) processor. For example, the one or more parallel processors 912 form a graphics processing subsystem that may output pixels to a display device of one of the one or more display devices 910A coupled via the I / O hub 907. The one or more parallel processors 912 may also include a display controller and a display interface (not shown) for enabling a direct connection to one or more display devices 910B.

[0160] Within the I / O subsystem 911, the system storage unit 914 may be connected to the I / O hub 907 to provide a storage mechanism for the computing system 900. The I / O switch 916 may be used to provide an interface mechanism to enable connections between the I / O hub 907 and other components, such as network adapter 918 and / or wireless network adapter 919 that may be integrated into the platform, and various other devices that may be added via one or more plug-in devices 120. The (one or more) plug-in devices 920 may also include, for example, one or more external graphics processing units, graphics cards, and / or computing accelerators. The network adapter 918 may be an Ethernet adapter or another wired network adapter. The wireless network adapter 919 may include one or more of the following: Wi-Fi, Bluetooth, near field communication (NFC), or other network devices including one or more wireless radio devices.

[0161] The computing system 900 may include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 907. The communication paths interconnecting the Figure 9 various components may be implemented using any suitable protocol, such as a PCI (Peripheral Component Interconnect)-based protocol (e.g., PCI Express), or any other bus or point-to-point communication interface and / or protocol, such as NVLink high-speed interconnect, Compute Express Link TM (Compute Express Link TM , CXL TM)(e.g., CXL.mem), Infinity Fabric (IF), Ethernet (IEEE802.3), remote direct memory access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Transmission Control Protocol (TCP), User Datagram Protocol (UDP), quick UDP Internet Connections (QUIC), RDMA over Converged Ethernet (RoCE), Intel QuickPath Interconnect (QPI), Intel Ultra Path Interconnect (UPI), Intel On-Chip System Fabric (IOSF), Omnipath, HyperTransport, Advanced Microcontroller Bus Architecture (AMBA) interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G and its variants, or a wired or wireless interconnect protocol known in the art. In some examples, a protocol such as non-volatile memory express over Fabrics (NVMe-oF) or NVMe may be used to copy or store data to a virtualized storage node.

[0162] One or more parallel processors 912 may include circuitry optimized for graphics and video processing (including, for example, video output circuitry), and form a graphics processing unit (GPU). Alternatively or additionally, as described in more detail herein, one or more parallel processors 912 may include circuitry optimized for general purpose processing while retaining the underlying computational architecture. The components of the computing system 900 may be integrated with one or more other system elements on a single integrated circuit. For example, one or more parallel processors 912, the memory hub 905, the processor(s) 902, and the I / O hub 907 may be integrated into a system-on-a-chip (SoC) integrated circuit. Alternatively, the components of the computing system 900 may be integrated into a single package to form a system-in-package (SIP) configuration. In some examples, at least portions of the components of the computing system 900 may be integrated into a multi-chip module (MCM), which may be interconnected with other multi-chip modules into a modular computing system.

[0163] It will be appreciated that the computing system 900 shown herein is illustrative, and variations and modifications are possible. The connection topology may be modified as needed, including the number and arrangement of bridges, the number of processor(s) 902, and the number of parallel processor(s) 912. For example, system memory 904 may be connected directly to the processor(s) 902 rather than through a bridge, and other devices communicate with system memory 904 via the memory hub 905 and the processor(s) 902. In other alternative topologies, the parallel processor(s) are connected to the I / O hub 907 or directly to one of the processor(s) 902 rather than to the memory hub 905. In other examples, the I / O hub 907 and the memory hub 905 may be integrated into a single chip. It is also possible for two or more sets of processors 902 to be attached via multiple sockets, which may be coupled to two or more instances of the parallel processor(s) 912.

[0164] Some of the specific components shown herein are optional and may not be included in all implementations of the computing system 900. For example, any number of plug-in cards or peripherals may be supported, or some components may be eliminated. Additionally, some architectures may use different terms for components similar to those illustrated Figure 9 herein. For example, the memory hub 905 may be referred to as a north bridge in some architectures, while the I / O hub 907 may be referred to as a south bridge.

[0165] Figure 10AAn example of a parallel processor 1000 is illustrated. The parallel processor 1000 can be a GPU, GPGPU, etc. as described herein. The various components of the parallel processor 1000 can be implemented using one or more integrated circuit devices, such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). The illustrated parallel processor 1000 can be one or more of the parallel processors 912 shown in Figure 9 one or more of (one or more) of those shown in .

[0166] The parallel processor 1000 includes a parallel processing unit 1002. The parallel processing unit includes an I / O unit 1004 that enables communication with other devices, including other instances of the parallel processing unit 1002. The I / O unit 1004 can be directly connected to other devices. For example, the I / O unit 1004 is connected to other devices via a hub or switch interface, such as a memory hub 905. The connection between the memory hub 905 and the I / O unit 1004 forms a communication link 913. Within the parallel processing unit 1002, the I / O unit 1004 is connected to a host interface 1006 and a memory crossbar 1016, where the host interface 1006 receives commands related to performing processing operations, and the memory crossbar 1016 receives commands related to performing memory operations.

[0167] When the host interface 1006 receives a command buffer via the I / O unit 1004, the host interface 1006 can direct the work operations for executing those commands to a front end 1008. In some examples, the front end 1008 is coupled to a scheduler 1010 that is configured to distribute commands or other work items to an array of processing clusters 1012. The scheduler 1010 ensures that the array of processing clusters 1012 is properly configured and in an active state before tasks are distributed to the processing clusters within the array of processing clusters 1012. The scheduler 1010 can be implemented via firmware logic executed on a microcontroller. The microcontroller-implemented scheduler 1010 can be configured to perform complex scheduling and work distribution operations at both a coarse-grained and fine-grained level, enabling fast preemption and context switching of threads executing on the array of processing clusters 1012. Preferably, the host software can authenticate the workload scheduled on the array of processing clusters 1012 via one of a plurality of graphics processing doorbells. In other examples, polling for new workloads or interrupts can be used to identify or indicate the availability of work to be performed. The workload can then be automatically distributed across the array of processing clusters 1012 by the scheduler 1010 logic within the scheduler microcontroller.

[0168] The processing cluster array 1012 may include up to "N" processing clusters (e.g., cluster 1014A, cluster 1014B to cluster 1014N). Each of the clusters 1014A - 1014N in the processing cluster array 1012 can execute a large number of concurrent threads. The scheduler 1010 can use various scheduling and / or work distribution algorithms to allocate work to the clusters 1014A - 1014N in the processing cluster array 1012, and these scheduling and / or work distribution algorithms can vary depending on the workload generated for each type of program or computation. Scheduling can be handled dynamically by the scheduler 1010, or can be assisted in part by compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 1012. Optionally, different clusters 1014A - 1014N in the processing cluster array 1012 can be assigned to process different types of programs or to perform different types of computations.

[0169] The processing cluster array 1012 can be configured to perform various types of parallel processing operations. For example, the processing cluster array 1012 is configured to perform general - purpose parallel computing operations. For example, the processing cluster array 1012 may include logic for performing processing tasks including filtering of video and / or audio data, performing modeling operations including physical operations, and performing data transformations.

[0170] The processing cluster array 1012 is configured to perform parallel graphics processing operations. In such examples where the parallel processor 1000 is configured to perform graphics processing operations, the processing cluster array 1012 may include additional logic to support the execution of such graphics processing operations, including but not limited to texture sampling logic for performing texture operations, and tessellation logic and other vertex processing logic. Additionally, the processing cluster array 1012 can be configured to execute graphics - processing - related shader programs, such as but not limited to vertex shaders, tessellation shaders, geometry shaders, and pixel shaders. The parallel processing unit 1002 can transfer data from the system memory via the I / O unit 1004 for processing. The transferred data can be stored in on - chip memory (e.g., parallel processor memory 1022) during processing and then written back to the system memory.

[0171] In an example where a parallel processing unit 1002 is used to perform graphics processing, the scheduler 1010 can be configured to divide the processing workload into tasks of approximately equal size to better enable the distribution of graphics processing operations to multiple clusters 1014A - 1014N in the processing cluster array 1012. In some of these examples, portions of the processing cluster array 1012 can be configured to perform different types of processing. For example, a first portion can be configured to perform vertex shading and topology generation, a second portion can be configured to perform tessellation and geometry shading, and a third portion can be configured to perform pixel shading or other screen space operations to produce a rendered image for display. Intermediate data generated by one or more of the clusters 1014A - 1014N can be stored in a buffer to allow the intermediate data to be transferred between the clusters 1014A - 1014N for further processing.

[0172] During operation, the processing cluster array 1012 can receive processing tasks to be executed via the scheduler 1010, which receives commands defining the processing tasks from the front end 1008. For graphics processing operations, the processing tasks can include data to be processed and indices of status parameters and commands defining how the data will be processed (e.g., what program will be executed), such as surface (patch) data, primitive data, vertex data, and / or pixel data. The scheduler 1010 can be configured to fetch the index corresponding to the task or can receive the index from the front end 1008. The front end 1008 can be configured to ensure that the processing cluster array 1012 is in an effective state before the workload specified by the incoming command buffer (e.g., batch buffer, push buffer, etc.) is initiated.

[0173] Each instance of one or more instances of the parallel processing unit 1002 can be coupled to the parallel processor memory 1022. The parallel processor memory 1022 can be accessed via a memory crossbar 1016, which can receive memory requests from the processing cluster array 1012 as well as the I / O unit 1004. The memory crossbar 1016 can access the parallel processor memory 1022 via a memory interface 1018. The memory interface 1018 can include a plurality of partition units (e.g., partition unit 1020A, partition unit 1020B, up to partition unit 1020N), each of which can be coupled to a portion (e.g., a memory cell) of the parallel processor memory 1022. The number of partition units 1020A - 1020N can be configured to be equal to the number of memory cells, such that the first partition unit 1020A has a corresponding first memory cell 1024A, the second partition unit 1020B has a corresponding second memory cell 1024B, and the Nth partition unit 1020N has a corresponding Nth memory cell 1024N. In other examples, the number of partition units 1020A - 1020N may not be equal to the number of memory devices.

[0174] The memory cells 1024A - 1024N can include various types of memory devices, including dynamic random-access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM), including graphics double data rate (GDDR) memory. Optionally, the memory cells 1024A - 1024N can also include 3D stacked memory, including but not limited to high bandwidth memory (HBM). Those skilled in the art will appreciate that the specific implementation of the memory cells 1024A - 1024N can vary and can be selected from one of various conventional designs. Rendering targets such as frame buffers or texture maps can be stored across the memory cells 1024A - 1024N, allowing the partition units 1020A - 1020N to write portions of each rendering target in parallel to efficiently utilize the available bandwidth of the parallel processor memory 1022. In some examples, local instances of the parallel processor memory 1022 can be excluded to favor a unified memory design that utilizes system memory in conjunction with local cache memory.

[0175] Optionally, any one of the clusters 1014A - 1014N in the processing cluster array 1012 has the ability to process data to be written to any one of the memory cells 1024A - 1024N within the parallel processor memory 1022. The memory cross - switch 1016 can be configured to transfer the output of each cluster 1014A - 1014N to any of the partition units 1020A - 1020N or to another cluster 1014A - 1014N, which can perform additional processing operations on the output. Each cluster 1014A - 1014N can communicate with the memory interface 1018 through the memory cross - switch 1016 to read from or write to various external memory devices. In some examples of the example with the memory cross - switch 1016, the memory cross - switch 1016 has a connection to the memory interface 1018 to communicate with the I / O unit 1004 and has a connection to a local instance of the parallel processor memory 1022, enabling processing units within different processing clusters 1014A - 1014N to communicate with system memory or other memory not local to the parallel processing unit 1002. Generally, the memory cross - switch 1016 can, for example, be able to use virtual channels to separate the traffic flow between the clusters 1014A - 1014N and the partition units 1020A - 1020N.

[0176] Although a single instance of the parallel processing unit 1002 is illustrated within the parallel processor 1000, any number of instances of the parallel processing unit 1002 can be included. For example, multiple instances of the parallel processing unit 1002 can be provided on a single plug - in card, or multiple plug - in cards can be interconnected. For example, the parallel processor 1000 can be a plug - in device, such as, Figure 9 the plug - in device 920, which can be a graphics card (such as a discrete graphics card including one or more GPUs, one or more memory devices, and device - to - device or network or fabric interfaces). Different instances of the parallel processing unit 1002 can be configured to interoperate even if the different instances have different numbers of processing cores, different amounts of local parallel processor memory, and / or other configuration differences. Optionally, some instances of the parallel processing unit 1002 can include higher - precision floating - point units relative to other instances. Systems containing one or more instances of the parallel processing unit 1002 or the parallel processor 1000 can be implemented in a variety of configurations and form factors, including but not limited to, desktop computers, laptop computers, or handheld personal computers, servers, workstations, game consoles, and / or embedded systems. The orchestrator can use one or more of the following to form composite nodes for workload execution: decomposed processor resources, cache resources, memory resources, storage resources, and networking resources.

[0177] In some examples, the parallel processing unit 1002 can be partitioned into multiple instances. Those multiple instances can be configured to execute workloads associated with different clients in an isolated manner, such that a predetermined quality of service is provided for each client. For example, each cluster 1014A - 1014N can be partitioned and isolated from other clusters, allowing the cluster array 1012 of processors to be divided into multiple computing partitions or instances. In such a configuration, workloads executed on the isolated partitions are protected from errors or inaccuracies associated with different workloads executed on different partitions. The partitioning units 1020A - 1020N can be configured to enable dedicated and / or isolated paths to the memories of the clusters 1014A - 1014N associated with the respective computing partitions. This data path isolation enables the computing resources within a partition to communicate with one or more assigned memory units 1024A - 1024N without being disturbed by the activities of other partitions.

[0178] Figure 10B is a block diagram of the partitioning unit 1020. The partitioning unit 1020 can be Figure 10A an instance of one of the partitioning units 1020A - 1020N. As illustrated, the partitioning unit 1020 includes an L2 cache 1021, a frame buffer interface 1025, and a ROP 1026 (raster operation unit). The L2 cache 1021 is a read / write cache configured to perform load and store operations received from the memory crossbar 1016 and the ROP 1026. Read misses and urgent write-back requests are output by the L2 cache 1021 to the frame buffer interface 1025 for processing. Updates can also be sent via the frame buffer interface 1025 to the frame buffer for processing. In some examples, the frame buffer interface 1025 interfaces with a memory unit in the parallel processor memory, such as, for example, one of the memory units 1024A - 1024N Figure 10A within the parallel processor memory 1022). The partitioning unit 1020 can also additionally or alternatively interface with a memory unit in the parallel processor memory via a memory controller (not shown).

[0179] In a graphics application, the ROP 1026 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. The ROP 1026 then outputs the processed graphics data, which is stored in the graphics memory. In some examples, the ROP 1026 includes or is coupled to a codec (CODEC) 1027, which includes compression logic for compressing depth or color data written to the memory or L2 cache 1021 and decompressing depth or color data read from the memory or L2 cache 1021. The compression logic can be lossless compression logic that utilizes one or more of a variety of compression algorithms. The type of compression performed by the CODEC 1027 can vary based on the statistical characteristics of the data to be compressed. For example, in some examples, delta color compression is performed on depth and color data on a per-tile basis. In some examples, the CODEC 1027 includes compression and decompression logic that can compress and decompress computational data associated with machine learning operations. The CODEC 1027 can, for example, compress sparse matrix data for sparse machine learning operations. The CODEC 1027 can also compress sparse matrix data encoded in a sparse matrix format (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.) to generate compressed and encoded sparse matrix data. The compressed and encoded sparse matrix data can be decompressed and / or decoded before being processed by a processing element, or the processing element can be configured to consume the compressed, encoded, or compressed and encoded data for processing.

[0180] The ROP 1026 can be included within each processing cluster (e.g., Figure 10A clusters 1014A - 1014N) rather than being included within the partitioning unit 1020. In such examples, read and write requests for pixel data rather than pixel fragment data are conveyed via the memory crossbar 1016. The processed graphics data can be displayed on a display device (such as, Figure 9 one of the one or more display devices 910A - 910B), routed for further processing by one or more processors 902, or routed for further processing by Figure 10A one of the processing entities within the parallel processor 1000.

[0181] Figure 10Cis a block diagram of processing cluster 1014 within a parallel processing unit. For example, the processing cluster is Figure 10A an instance of one of processing clusters 1014A - 1014N. Processing cluster 1014 can be configured to execute many threads in parallel, where the term "thread" refers to an instance of a particular program executed on a particular set of input data. Optionally, single-instruction, multiple-data (SIMD) instruction issue techniques can be used to support parallel execution of a large number of threads without providing multiple independent instruction units. Alternatively, single-instruction, multiple-thread (SIMT) techniques can be used to support parallel execution of a large number of generally synchronized threads using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster in the processing cluster. Different from SIMD execution mechanisms where all processing engines typically execute the same instruction, SIMT execution allows different threads to more easily follow divergent execution paths through a given thread program. Those skilled in the art will understand that SIMD processing mechanisms represent a functional subset of SIMT processing mechanisms.

[0182] The operation of processing cluster 1014 can be controlled via pipeline manager 1032 that distributes processing tasks to the SIMT parallel processors. Pipeline manager 1032 receives instructions from Figure 10A scheduler 1010 and manages the execution of those instructions via graphics multiprocessor 1034 and / or texture unit 1036. The illustrated graphics multiprocessor 1034 is an exemplary instance of a SIMT parallel processor. However, various types of SIMT parallel processors with different architectures can be included within processing cluster 1014. One or more instances of graphics multiprocessor 1034 can be included within processing cluster 1014. Graphics multiprocessor 1034 can process data, and data crossbar 1040 can be used to distribute the processed data to one of a number of possible destinations, including other shader units. Pipeline manager 1032 can facilitate the distribution of the processed data by specifying the destination for the processed data to be distributed via data crossbar 1040.

[0183] Each graphics multiprocessor 1034 within processing cluster 1014 may include the same set of functional execution logic (e.g., arithmetic logic units, load-store units, etc.). The functional execution logic can be configured in a pipelined manner, in which new instructions can be issued before previous instructions are completed. The functional execution logic supports a variety of operations, including integer and floating-point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. Different operations can be performed using the same functional unit hardware, and any combination of functional units can exist.

[0184] Instructions transmitted to processing cluster 1014 constitute a thread. A set of threads executed across a collection of parallel processing engines is a thread group. The thread group executes the same program on different input data. Each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 1034. The thread group may include fewer threads than the number of processing engines within graphics multiprocessor 1034. When the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during the cycles in which the thread group is being processed. The thread group may also include more threads than the number of processing engines within graphics multiprocessor 1034. When the thread group includes more threads than the number of processing engines within graphics multiprocessor 1034, processing can be performed in consecutive clock cycles. Optionally, multiple thread groups can be executed concurrently on graphics multiprocessor 1034.

[0185] Graphics multiprocessor 1034 may include an internal cache memory to perform load and store operations. Optionally, graphics multiprocessor 1034 may forego the internal cache and use the cache memory within processing cluster 1014 (e.g., level 1 (L1) cache 1048). Each graphics multiprocessor 1034 also has access to a second-level (L2) cache within a partition unit (e.g., Figure 10A partition units 1020A - 1020N), which are shared among all processing clusters 1014 and can be used to transfer data between threads. Graphics multiprocessor 1034 can also access off-chip global memory, which may include one or more of local parallel processor memory and / or system memory. Any memory external to parallel processing unit 1002 can be used as global memory. In an example where processing cluster 1014 includes multiple instances of graphics multiprocessor 1034, common instructions and data can be shared, and the common instructions and data can be stored in L1 cache 1048.

[0186] Each processing cluster 1014 may include an MMU 1045 (memory management unit) configured to map virtual addresses to physical addresses. In other examples, one or more instances of the MMU 1045 may reside within Figure 10A the memory interface 1018. The MMU 1045 includes a set of page table entries (PTEs) for mapping virtual addresses to the physical addresses of the die, and optionally includes cache line indices. The MMU 1045 may include a translation lookaside buffer (TLB) or cache that may reside within the graphics multiprocessor 1034 or L1 cache 1048 of the processing cluster 1014. The physical addresses are processed to distribute surface data locality, allowing for efficient request interleaving between partition units. The cache line index may be used to determine whether a request to a cache line is a hit or a miss.

[0187] In graphics and compute applications, the processing cluster 1014 may be configured such that each graphics multiprocessor 1034 is coupled to a texture unit 1036 for performing texture mapping operations, e.g., determining texture sample locations, reading texture data, and filtering texture data. The texture data is read from an internal texture L1 cache (not shown), or in some examples, from the L1 cache within the graphics multiprocessor 1034, and is fetched from the L2 cache, local parallel processor memory, or system memory as needed. Each graphics multiprocessor 1034 outputs the processed tasks to the data crossbar 1040 to provide the processed tasks to another processing cluster 1014 for further processing, or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 1016. The preROP 1042 (pre-raster operations unit) is configured to receive data from the graphics multiprocessor 1034 and direct the data to ROP units that may be located with partition units (e.g., Figure 10A partition units 1020A - 1020N) as described herein. The preROP 1042 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0188] It will be appreciated that the core architecture described herein is illustrative and that variations and modifications are possible. Any number of processing units (e.g., graphics multiprocessor 1034, texture unit 1036, preROP 1042, etc.) may be included within processing cluster 1014. Further, although only one processing cluster 1014 is shown, the parallel processing unit as described herein may include any number of instances of processing cluster 1014. Optionally, each processing cluster 1014 may be configured to operate independently of other processing clusters 1014 using separate and distinct processing units, L1 caches, L2 caches, etc.

[0189] Figure 10D An example of a graphics multiprocessor 1034 is shown, where the graphics multiprocessor 1034 is coupled to the pipeline manager 1032 of processing cluster 1014. The graphics multiprocessor 1034 has an execution pipeline that includes, but is not limited to, an instruction cache 1052, an instruction unit 1054, an address mapping unit 1056, a register file 1058, one or more general-purpose graphics processing unit (GPGPU) cores 1062, and one or more load / store units 1066. The GPGPU cores 1062 and the load / store units 1066 are coupled to cache memory 1072 and shared memory 1070 via a memory and cache interconnect 1068. The graphics multiprocessor 1034 may additionally include tensor / or ray tracing cores 1063, which include hardware logic for accelerating matrix and / or ray tracing operations.

[0190] The instruction cache 1052 may receive a stream of instructions to be executed from the pipeline manager 1032. The instructions are cached in the instruction cache 1052 and dispatched for execution by the instruction unit 1054. The instruction unit 1054 may dispatch the instructions as a thread group (e.g., a warp), where each thread in the thread group is assigned to a different execution unit within the GPGPU core 1062. Instructions may access any one of a local address space, a shared address space, or a global address space by specifying an address within the unified address space. The address mapping unit 1056 may be used to translate an address in the unified address space into a different memory address that can be accessed by the load / store unit 1066.

[0191] The register file 1058 provides a collection of registers for the functional units of the graphics multiprocessor 1034. The register file 1058 provides temporary storage for the operands of the data paths connected to the functional units (e.g., GPGPU cores 1062, load / store units 1066) of the graphics multiprocessor 1034. The register file 1058 can be partitioned among each of the functional units such that each functional unit is allocated a dedicated portion of the register file 1058. For example, the register file 1058 can be partitioned among different groups of units executed by the graphics multiprocessor 1034.

[0192] The GPGPU cores 1062 can each include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) for executing the instructions of the graphics multiprocessor 1034. In some implementations, the GPGPU cores 1062 can include hardware logic that would otherwise reside within the tensor and / or ray tracing cores 1063. The GPGPU cores 1062 can be architecturally similar or architecturally different. For example and in some examples, a first portion of the GPGPU core 1062 includes a single-precision FPU and an integer ALU, while a second portion of the GPGPU core includes a double-precision FPU. Optionally, the FPU can implement the IEEE 754-2008 standard for floating point arithmetic or enable variable precision floating point arithmetic. The graphics multiprocessor 1034 can additionally include one or more fixed-function or special-function units for performing specific functions such as copy rectangle or pixel blend operations. One or more of the GPGPU cores can also include fixed-function or special-function logic.

[0193] The GPGPU cores 1062 can include SIMD logic capable of executing a single instruction on multiple sets of data. Optionally, the GPGPU cores 1062 can physically execute SIMD4, SIMD8, and SIMD16 instructions and logically execute SIMD1, SIMD2, and SIMD32 instructions. The SIMD instructions for the GPGPU cores can be generated at compile time by a shader compiler or automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. Multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example and in some examples, eight SIMT threads that perform the same or similar operations can be executed in parallel via a single SIMD8 logic unit.

[0194] The memory and cache interconnect 1068 is an interconnect network that connects each functional unit in the functional units of the graphics multiprocessor 1034 to the register file 1058 and to the shared memory 1070. For example, the memory and cache interconnect 1068 is a crossbar interconnect that allows the load / store unit 1066 to implement load and store operations between the shared memory 1070 and the register file 1058. The register file 1058 can operate at the same frequency as the GPGPU core 1062, so data transfer between the GPGPU core 1062 and the register file 1058 is of very low latency. The shared memory 1070 can be used to enable communication between threads executing on the functional units within the graphics multiprocessor 1034. The cache memory 1072 can be used as a data cache, for example, to cache texture data passed between the functional units and the texture unit 1036. The shared memory 1070 can also be used as a managed cached program. The shared memory 1070 and the cache memory 1072 can be coupled to the data crossbar 1040 to enable communication with other components of the processing cluster. Threads executing on the GPGPU core 1062 can also programmatically store data in the shared memory in addition to the automatically cached data stored in the cache memory 1072.

[0195] Figures 11A - 11C Illustrates an additional graphics multiprocessor according to an example. Figures 11A - 11B Illustrates graphics multiprocessors 1125, 1150, which are related to the graphics multiprocessor 1034 of Figure 10C and can be used in place of one of those graphics multiprocessors. Accordingly, any disclosure of a feature in connection with the graphics multiprocessor 1034 herein also discloses the corresponding combination with the graphics multiprocessors 1125, 1150, but is not limited thereto. Figure 11C Illustrates a graphics processing unit (GPU) 1180 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 1165A - 1165N, which correspond to the graphics multiprocessors 1125, 1150. The illustrated graphics multiprocessors 1125, 1150 and the multi-core groups 1165A - 1165N can be streaming multiprocessors (SMs) capable of executing a large number of execution threads simultaneously.

[0196] Figure 11A The graphics multiprocessor 1125 of Figure 10DMultiple additional instances of execution resource units of the graphics multiprocessor 1034. For example, the graphics multiprocessor 1125 may include multiple instances of instruction units 1132A - 1132B, register files 1134A - 1134B, and (one or more) texture units 1144A - 1144B. The graphics multiprocessor 1125 also includes multiple sets of graphics or compute execution units (e.g., GPGPU cores 1136A - 1136B, tensor cores 1137A - 1137B, ray tracing cores 1138A - 1138B) and multiple sets of load / store units 1140A - 1140B. The execution resource units have a common instruction cache 1130, texture and / or data cache memory 1142, and shared memory 1146.

[0197] Each component can communicate via the interconnect structure 1127. The interconnect structure 1127 may include one or more crossbars to enable communication between the components of the graphics multiprocessor 1125. The interconnect structure 1127 is a separate, high-speed network structure layer on which each component of the graphics multiprocessor 1125 is stacked. The components of the graphics multiprocessor 1125 communicate with remote components via the interconnect structure 1127. For example, cores 1136A - 1136B, 1137A - 1137B, and 1138A - 1138B can each communicate with the shared memory 1146 via the interconnect structure 1127. The interconnect structure 1127 can arbitrate communications within the graphics multiprocessor 1125 to ensure fair bandwidth allocation between components.

[0198] Figure 11B The graphics multiprocessor 1150 includes multiple sets of execution resources 1156A - 1156D, where, as Figure 10D and Figure 11A illustrated, each set of execution resources includes multiple instruction units, register files, GPGPU cores, and load / store units. The execution resources 1156A - 1156D can work in cooperation with (one or more) texture units 1160A - 1160D for texture operations while sharing the instruction cache 1154 and the shared memory 1153. For example, the execution resources 1156A - 1156D can share the instruction cache 1154 and the shared memory 1153 as well as multiple instances of texture and / or data cache memories 1158A - 1158B. Each component can communicate via an interconnect structure 1152 similar to the interconnect structure 1127 of Figure 11A

[0199] Those skilled in the art will understand that Figure 9 、 Figures 10A - 10D and Figures 11A - 11BThe architecture described herein is descriptive and not restrictive in terms of the scope of the current example. Thus, the techniques described herein may be implemented on any suitably configured processing unit without departing from the scope of the examples described herein, including but not limited to: one or more mobile application processors; one or more desktop or server central processing units (CPUs), including multi-core CPUs; one or more parallel processor units such as, Figure 10A the parallel processing unit 1002 and one or more graphics processors or dedicated processing units.

[0200] The parallel processors or GPGPUs described herein may be communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. The GPU may be communicatively coupled to the host processor / core via a bus or another interconnect (e.g., a high-speed interconnect such as PCIe, NVLink, or other known protocols, standardized protocols, or proprietary protocols). In other examples, the GPU may be integrated on the same package or die as the core and communicatively coupled to the core via an internal processor bus / interconnect (i.e., within the package or die). Regardless of the manner in which the GPU is connected, the processor core may allocate work to the GPU in the form of a sequence of commands / instructions contained in a work descriptor. The GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0201] Figure 11C Illustrated is a graphics processing unit (GPU) 1180 that includes a collection of dedicated graphics processing resources arranged as multi-core groups 1165A - 1165N. While details are provided for only a single multi-core group 1165A, it will be appreciated that the other multi-core groups 1165B - 1165N may be equipped with the same or similar collections of graphics processing resources. The details described with respect to multi-core groups 1165A - 1165 also apply to any graphics multiprocessor 1034, 1125, 1150 described herein.

[0202] As shown, the multi-core group 1165A can include a set 1170 of graphics cores, a set 1171 of tensor cores, and a set 1172 of ray tracing cores. The scheduler / dispatcher 1168 schedules and dispatches graphics threads for execution on the respective cores 1170, 1171, 1172. The set of register files 1169 stores operand values used by the cores 1170, 1171, 1172 when executing graphics threads. These register files can include, for example, integer registers for storing integer values, floating-point registers for storing floating-point values, vector registers for storing packed data elements (integer and / or floating-point data elements), and tile registers for storing tensor / matrix values. The tile registers can be implemented as a combined set of vector registers.

[0203] One or more combined level-1 (L1) cache and shared memory units 11711 store graphics data locally within each multi-core group 1165A, such as texture data, vertex data, pixel data, ray data, bounding volume data, etc. One or more texture units 1174 can also be used to perform texture operations, such as texture mapping and sampling. The level-2 (L2) cache 1175 shared by all multi-core groups 1165A - 1165N or a subset of multi-core groups 1165A - 1165N stores graphics data and / or instructions for multiple concurrent graphics threads. As shown, the L2 cache 1175 can be shared across multiple multi-core groups 1165A - 1165N. One or more memory controllers 1167 couple the GPU 1180 to the memory 1166, which can be system memory (e.g., DRAM) and / or dedicated graphics memory (e.g., GDDR6 memory).

[0204] The input / output (I / O) circuitry 1163 couples the GPU 1180 to one or more I / O devices 1162, such as a digital signal processor (DSP), a network controller, or a user input device. The on-chip interconnect can be used to couple the I / O devices 1162 to the GPU 1180 and the memory 1166. One or more I / O memory management units (IOMMUs) 1164 of the I / O circuitry 1163 directly couple the I / O devices 1162 to the system memory 1166. Optionally, the IOMMU 1164 manages a set of multiple page tables for mapping virtual addresses to physical addresses in the system memory 1166. The I / O devices 1162, the (one or more) CPUs 1161, and the (one or more) GPUs 1180 can then share the same virtual address space.

[0205] In one implementation of the IOMMU 1164, the IOMMU 1164 supports virtualization. In this case, the IOMMU 1164 can manage a first set of page tables for mapping guest / graphics virtual addresses to guest / graphics physical addresses and a second set of page tables for mapping guest / graphics physical addresses to system / host physical addresses (e.g., within the system memory 1166). The base address of each of the first set of page tables and the second set of page tables can be stored in a control register and swapped out during a context switch (e.g., such that a new context is provided access to the relevant set of page tables). Although not illustrated in Figure 11C , each of the cores 1170, 1171, 1172, and / or multi-core groups 1165A - 1165N may include translation lookaside buffers (TLBs) for caching guest virtual-to-guest physical translations, guest physical-to-host physical translations, and guest virtual-to-host physical translations.

[0206] (One or more) CPUs 1161, GPUs 1180, and I / O devices 1162 may be integrated on a single semiconductor chip and / or chip package. The illustrated memory 1166 may be integrated on the same chip or may be coupled to the memory controller 1167 via an off-chip interface. In one implementation, the memory 1166 includes GDDR6 memory that shares the same virtual address space as other physical system-level memories, but the basic principles described herein are not limited to this particular implementation.

[0207] The tensor core 1171 may include a plurality of execution units specifically designed to perform matrix operations, which are fundamental computational operations for performing deep learning operations. For example, synchronous matrix multiplication operations can be used for neural network training and inference. The tensor core 1171 can perform matrix processing using various operand precisions, including single-precision floating point (e.g., 32 bits), half-precision floating point (e.g., 16 bits), integer (16 bits), byte (8 bits), and nibble (4 bits). For example, a neural network implementation extracts features of each rendered scene, potentially combining details from multiple frames to construct a high-quality final image.

[0208] In a deep learning implementation, schedulable parallel matrix multiplication operations can be used for execution on tensor core 1171. Training of neural networks in particular requires a large number of matrix dot product operations. To handle the inner product formulation of an N x N x N matrix multiplication, tensor core 1171 can include at least N dot product processing elements. Before matrix multiplication begins, a complete matrix is loaded into the scratch register, and for each of the N loops, at least one column of the second matrix is loaded. Depending on the specific implementation, matrix elements can be stored in different precisions, including 16-bit words, 8-bit bytes (e.g., INT8), and 4-bit nibbles (e.g., INT4). Different precision modes can be specified for tensor core 1171 to ensure that the most efficient precision is used for different workloads (e.g., inference workloads, which can tolerate quantization down to bytes and nibbles). The supported formats additionally include 64-bit floating point (FP64) and non-IEEE floating point formats, such as the bfloat16 format (e.g., Brain floating point), a 16-bit floating point format with one sign bit, eight exponent bits, and eight significand bits (seven of which are explicitly stored). Some examples include support for a reduced precision tensor float (TF112) mode, which performs calculations using the range of FP32 (8 bits) and the precision of FP16 (10 bits). Reduced precision TF32 operations can be performed on FP32 inputs with higher performance relative to FP32 and increased precision relative to FP16 and produce FP32 outputs. In some examples, one or more 8-bit floating point formats (FP32) are supported.

[0209] In some examples, tensor core 1171 supports a sparse operation mode for matrices in which the vast majority of values are zero. Tensor core 1171 includes support for sparse input matrices encoded in a sparse matrix representation (e.g., coordinate list encoding (COO), compressed sparse row (CSR), compressed sparse column (CSC), etc.). Tensor core 1171 also includes support for a compressed sparse matrix representation in cases where the sparse matrix representation can be further compressed. Compressed matrix data, encoded matrix data, and / or compressed and encoded matrix data, as well as associated compression and / or encoding metadata, can be read by tensor core 1171, and non-zero values can be extracted. For example, for a given input matrix A, non-zero values can be loaded from at least a portion of the compressed and / or encoded representation of matrix A. Based on the positions of the non-zero values in matrix A (which can be determined from the indices or coordinate metadata associated with the non-zero values), the corresponding values in input matrix B can be loaded. Depending on the operation to be performed (e.g., multiplication), if the corresponding value is a zero value, loading the value from input matrix B can be bypassed. In some examples, the pairing of values for certain operations (such as multiplication operations) can be pre-scanned by the scheduler logic, and only operations between non-zero inputs are scheduled. Depending on the dimensions of matrix A and matrix B and the operation to be performed, output matrix C can be dense or sparse. In the case where output matrix C is sparse and depending on the configuration of tensor core 1171, output matrix C can be output in a compressed format, sparse encoding, or compressed sparse encoding.

[0210] Ray tracing core 1172 can accelerate ray tracing operations for both real-time ray tracing implementations and non-real-time ray tracing implementations. Specifically, ray tracing core 1172 can include a ray traversal / intersection circuitry that uses a bounding volume hierarchy (BVH) to perform ray traversal and identify intersections between rays enclosed within the BVH volume and primitives. Ray tracing core 1172 can also include circuitry for performing depth testing and culling (e.g., using a Z-buffer or similar arrangement). In one implementation, ray tracing core 1172 performs traversal and intersection operations in cooperation with the image denoising techniques described herein, at least a portion of which can be performed on tensor core 1171. For example, tensor core 1171 can implement a deep learning neural network to perform denoising of frames generated by ray tracing core 1172. However, (one or more) CPUs 1161, graphics core 1170, and / or ray tracing core 1172 can also implement all or portions of the denoising and / or deep learning algorithms.

[0211] In addition, as described above, a distributed approach to noise reduction can be employed, where the GPU 1180 is in a computing device coupled to other computing devices via a network or high-speed interconnect. According to this distributed approach, the interconnected computing devices can share neural network learning / training data to improve the speed at which the overall system learns to perform noise reduction for different types of image frames and / or different graphics applications.

[0212] The ray tracing core 1172 can handle all BVH traversals and / or ray-primitive intersections, freeing the graphics core 1170 from being overloaded with thousands of instructions per ray. For example, each ray tracing core 1172 includes a first set of specialized circuitry for performing bounding box tests (e.g., for traversal operations) and / or a second set of specialized circuitry for performing ray-triangle intersection tests (e.g., intersecting the traversed rays). Thus, for example, the multi-core group 1165A can simply initiate a ray probe, and the ray tracing core 1172 independently performs ray traversal and intersection and returns hit data (e.g., hit, no hit, multiple hits, etc.) to the thread context. While the ray tracing core 1172 performs traversal and intersection operations, the other cores 1170, 1171 are freed up to perform other graphics or computational work.

[0213] Optionally, each ray tracing core 1172 can include a traversal unit for performing BVH test operations and / or an intersection unit for performing ray-primitive intersection tests. The intersection unit generates "hit", "no hit", or "multiple hits" responses, which the intersection unit provides to the appropriate thread. During traversal and intersection operations, the execution resources of other cores (e.g., the graphics core 1170 and the tensor core 1171) are freed up to perform other forms of graphics work.

[0214] In some examples described below, a hybrid rasterization / ray tracing method is used in which work is distributed between the graphics core 1170 and the ray tracing core 1172.

[0215] The ray tracing core 1172 (and / or other cores 1170, 1171) may include hardware support for a ray tracing instruction set, such as Microsoft's DirectX Ray Tracing (DXR), which includes the DispatchRays command; and ray generation shaders, closest hit shaders, any hit shaders, and miss shaders, which enable the assignment of a unique set of shaders and textures to each object. Another ray tracing platform that may be supported by the ray tracing core 1172, the graphics core 1170, and the tensor core 1171 is the Vulkan API (e.g., Vulkan version 1.1.85, or later versions). However, note that the basic principles described herein are not limited to any particular ray tracing ISA.

[0216] Generally, the respective cores 1172, 1171, 1170 may support a ray tracing instruction set that includes instructions / functions for one or more of the following: ray generation, closest hit, any hit, ray-primitive intersection, per-primitive and hierarchical bounding box construction, miss, traverse, and exception. More specifically, some examples include ray tracing instructions for performing one or more of the following functions:

[0217] Ray generation - Ray generation instructions may be executed for each pixel, sample, or other user-defined work assignment.

[0218] Closest hit - Closest hit instructions may be executed to locate the closest intersection of a ray with a primitive in the scene.

[0219] Any hit - Any hit instructions identify multiple intersections between a ray and primitives in the scene, potentially identifying a new closest intersection.

[0220] Intersection - Intersection instructions perform a ray-primitive intersection test and output the result.

[0221] Per-primitive bounding box construction - This instruction builds a bounding box around a given primitive or group of primitives (e.g., when building a new BVH or other acceleration data structure).

[0222] Miss - Indicates that a ray has missed all geometries in the scene or a specified region of the scene.

[0223] Traverse - Indicates the child volumes that a ray will traverse.

[0224] Exception - Includes various types of exception handlers (e.g., called for various error conditions).

[0225] In some examples, the ray tracing core 1172 may be adapted to accelerate general computing operations that may use computational techniques similar to ray intersection tests. A computational framework may be provided that enables shader programs to be compiled into low-level instructions and / or primitives for performing general computing operations via the ray tracing core. Exemplary computational problems that may benefit from computational operations performed on the ray tracing core 1172 include computations involving the propagation of light beams, waves, rays, or particles within a coordinate space. Interactions associated with that propagation may be computed relative to geometries or meshes within the coordinate space. For example, computations associated with the propagation of electromagnetic signals through an environment may be accelerated via the use of instructions or primitives executed via the ray tracing core. Refraction and reflection of signals off objects in the environment may be computed as a direct ray tracing simulation.

[0226] The ray tracing core 1172 may also be used to perform computations that are not directly similar to ray tracing. For example, the ray tracing core 1172 may be used to accelerate mesh projection, mesh refinement, and volume sampling computations. General coordinate space computations may also be performed, such as nearest neighbor computations. For example, a set of points near a given point may be found by defining a bounding box around the point in the coordinate space. The BVH and ray tracing logic within the ray tracing core 1172 may then be used to determine the set of point intersections within the bounding box. The intersections constitute the origin and the nearest neighbors of that origin. Computations performed using the ray tracing core 1172 may be executed in parallel with computations performed on the graphics core 1172 and the tensor core 1171. The shader compiler may be configured to compile compute shaders or other general graphics processing programs into low-level primitives that can be parallelized across the graphics core 1170, the tensor core 1171, and the ray tracing core 1172.

[0227] Building increasingly large silicon die is challenging for various reasons. As the silicon die gets larger, the manufacturing yield gets smaller, and the process technology requirements for different components may vary. On the other hand, for a high-performance system, key components should be interconnected via high-speed, high-bandwidth, low-latency interfaces. These conflicting requirements pose challenges to high-performance chip development.

[0228] The examples described herein provide techniques for decomposing the architecture of a system-on-chip integrated circuit into multiple different die that can be packaged onto a common substrate. In some examples, a graphics processing unit or parallel processor is composed of packaged integrated circuits that include different logic units capable of being assembled with other die into a larger package. Various collections of die with different IP core logic can be assembled into a single device. Additionally, die can be integrated into a base die or base die using active interposer technology. The concepts described herein enable interconnection and communication between different forms of IP within a GPU. IP development on different processes can be mixed. This avoids the complexity of bringing multiple IPs to the same process, especially for large SoCs with several styles of IP.

[0229] Allowing the use of multiple process technologies improves time to market and provides a cost-effective way to create multiple product SKUs. For customers, this means getting products that are better suited to their requirements in a cost-effective and timely manner. Additionally, discrete IP is easier to modify to be independently power gated, and components that are not in use for a given workload can be turned off, thus reducing overall power consumption.

[0230] Figure 12 FIG. 1200 shows a parallel computing system according to some examples. In some examples, the parallel computing system 1200 includes a parallel processor 1220, which can be a graphics processor or a computing accelerator as described herein. The parallel processor 1220 includes a global logic unit 1201, an interface 1202, a thread dispatcher 1203, a media unit 1204, a set of computing units 1205A - 1205H, and a cache / memory unit 1206. In some examples, the global logic unit 1201 includes global functions for the parallel processor 1220, including device configuration registers, a global scheduler, power management logic, and the like. The interface 1202 can include a front-end interface for the parallel processor 1220. The thread dispatcher 1203 can receive a workload from the interface 1202 and dispatch threads of the workload to the computing units 1205A - 1205H. If the workload includes any media operations, at least some of those operations can be performed by the media unit 1204. The media unit can also migrate some operations to the computing units 1205A - 1205H. The cache / memory unit 1206 can include cache memory (e.g., L3 cache) and local memory (e.g., HBM, GDDR) for the parallel processor 1220.

[0231] Figures 13A - 13B FIG. shows a hybrid logic / physical view of a discrete parallel processor according to examples described herein. Figure 13AIllustrate the discrete parallel computing system 1300. Figure 13B Illustrate the die 1330 of the discrete parallel computing system 1300.

[0232] As Figure 13A As shown, the discrete computing system 1300 may include a parallel processor 1320, wherein the components of the parallel processor SOC are distributed across multiple dies. Each die may be a different IP core that is independently designed and configured to communicate with other dies via one or more common interfaces. Dies include, but are not limited to, a compute die 1305, a media die 1304, and a memory die 1306. Each die may be manufactured separately using different process technologies. For example, the compute die 1305 may be manufactured using the smallest or most advanced process technology available at the time of manufacture, while the memory die 1306 or other dies (e.g., I / O, networking, etc.) may be manufactured using a larger or less advanced process technology.

[0233] Each die may be bonded to a base die 1310 and configured to communicate with each other and with the logic within the base die 1310 via an interconnect layer 1312. In some examples, the base die 1310 may include global logic 1301, which may include a scheduler 1311 and a power management 1321 logic unit, an interface 1302, a dispatcher unit 1303, and an interconnect structure module 1308 coupled or integrated with one or more L3 cache blocks 1309A - 1309N. The interconnect structure 1308 may be an inter-die structure integrated into the base die 1310. Logic dies may use the structure 1308 to relay messages between the respective dies. Additionally, the L3 cache blocks 1309A - 1309N in the base die and / or the L3 cache blocks within the memory die 1306 may cache data read from the DRAM dies within the memory die 1306 and data transmitted to the DRAM dies within the memory die 1306 and to the system memory of the host.

[0234] In some examples, the global logic 1301 is a microcontroller that may execute firmware to perform the scheduler 1311 and power management 1321 functions for the parallel processor 1320. The microcontroller that executes the global logic may be customized for the target use case of the parallel processor 1320. The scheduler 1311 may perform global scheduling operations for the parallel processor 1320. The power management 1321 function may be used to enable or disable the respective dies within the parallel processor when they are not in use.

[0235] The individual dies of the parallel processor 1320 can be designed to perform specific functions that would be integrated into a single die in existing designs. The set of compute dies 1305 can include clusters of compute units (e.g., execution units, streaming multiprocessors, etc.), where the clusters of compute units include programmable logic for performing compute or graphics shader instructions. The media die 1304 can include hardware logic for accelerating media encoding and decoding operations. The memory die 1306 can include volatile memory (e.g., DRAM) and one or more SRAM cache memory blocks (e.g., L3 blocks).

[0236] As Figure 13B shown, each die 1330 can include common components and specialized components. The die logic 1330 within die 1336 can include the specific components of the die, such as an array of streaming multiprocessors, compute units, or execution units described herein. The die logic 1336 can be coupled to an optional cache or shared local memory 1338, or can include a cache or shared local memory within the die logic 1336. The die 1330 can include a fabric interconnect node 1342 for receiving commands via the inter-die fabric. Commands and data received via the fabric interconnect node 1342 can be temporarily stored in the interconnect buffer 1339. Data transmitted to and received from the fabric interconnect node 1342 can be stored in the interconnect cache 1340. Power control 1332 and clock control 1334 logic can also be included within the die. The power control 1332 and clock control 1334 logic can receive configuration commands via the fabric and can configure dynamic voltage and frequency scaling for the die 1330. In some examples, each die can have an independent clock domain and power domain and can be clock-gated and power-gated independently of other dies.

[0237] At least portions of the components within the illustrated die 1330 can also be included within the logic embedded within Figure 13A the base die 1310. For example, the logic within the base die that communicates with the fabric can include a version of the fabric interconnect node 1342. The base die logic that can be independently clock-gated or power-gated can include a version of the power control 1332 and / or clock control 1334 logic.

[0238] Thus, while the various examples described herein use the term SOC to describe devices and systems having a processor and associated circuitry (e.g., input / output (“I / O”) circuitry, power delivery circuitry, memory circuitry, etc.) monolithically integrated into a single integrated circuit (“IC”) die or chip, the present disclosure is not limited to this aspect. For example, in various examples of the present disclosure, a device or system may have one or more processors (e.g., one or more processor cores) and associated circuitry (e.g., input / output (“I / O”) circuitry, power delivery circuitry, etc.) disposed in a separate collection of discrete dies, chips, and / or dielets (e.g., one or more discrete processor core dies arranged adjacent to one or more other dies such as memory dies, I / O dies, etc.). In such discrete devices and systems, the respective dies, chips, and / or dielets may be physically and / or electrically coupled together via a packaging structure that includes, for example, respective package substrates, interposers, active interposers, photonic interposers, interconnect bridges, and the like. The separate collection of discrete dies, chips, and / or dielets may also be part of a system-on-package (“SoP”).

[0239] Example core architectures—In-order and out-of-order core block diagrams.

[0240] Figure 14A is a block diagram illustrating an example in-order pipeline and an example register renaming, out-of-order issue / execution pipeline, both according to an example. Figure 14B is a block diagram illustrating an example in-order architecture core to be included in a processor and an example register renaming, out-of-order issue / execution architecture core, both according to an example. Figures 14A - 14B The solid boxes in illustrate the in-order pipeline and in-order core, while the optionally added dashed boxes illustrate the register renaming, out-of-order issue / execution pipeline and core. Given that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0241] In Figure 14AIn [the figure], the processor pipeline 1400 includes a fetch stage 1402, an optional length decoding stage 1404, a decoding stage 1406, an optional allocation (Alloc) stage 1408, an optional renaming stage 1410, a scheduling (also referred to as dispatch or issue) stage 1412, an optional register read / memory read stage 1414, an execution stage 1416, a write-back / memory write stage 1418, an optional exception handling stage 1422, and an optional commit stage 1424. One or more operations may be performed in each of these processor pipeline stages. For example, during the fetch stage 1402, one or more instructions are fetched from the instruction memory, and during the decoding stage 1406, the one or more fetched instructions may be decoded, an address using the forwarded register ports (e.g., a load store unit (LSU) address) may be generated, and branch forwarding (e.g., immediate offset or link register (LR)) may be performed. In some examples, the decoding stage 1406 and the register read / memory read stage 1414 may be combined into one pipeline stage. In some examples, during the execution stage 1416, the decoded instructions may be executed, LSU address / data pipelining to the Advanced Microcontroller Bus (AMB) interface may be performed, multiplication and addition operations may be performed, arithmetic operations with branch results may be performed, and so on.

[0242] As an example, Figure 14B An example register renaming, out-of-order issue / execution architecture core may implement the pipeline 1400 as follows: 1) The instruction fetch circuitry 1438 performs the fetch stage 1402 and the length decoding stage 1404; 2) The decoding circuitry 1440 performs the decoding stage 1406; 3) The rename / allocator unit circuitry 1452 performs the allocation stage 1408 and the renaming stage 1410; 4) The scheduler circuitry(ies) 1456 performs the scheduling stage 1412; 5) The physical register file circuitry(ies) 1458 and the memory unit circuitry 1470 perform the register read / memory read stage 1414; The execution cluster(s) 1460 performs the execution stage 1416; 6) The memory unit circuitry 1470 and the physical register file circuitry(ies) 1458 perform the write-back / memory write stage 1418; 7) Various circuitries may be involved in the exception handling stage 1422; and 8) The retirement unit circuitry 1454 and the physical register file circuitry(ies) 1458 perform the commit stage 1424.

[0243] Figure 14BThe processor core 1490 is shown, which includes a front-end unit circuitry 1430 coupled to an execution engine unit circuitry 1450, and both are coupled to a memory unit circuitry 1470. The core 1490 can be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As yet another option, the core 1490 can be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.

[0244] The front-end unit circuit system 1430 may include a branch prediction circuit system 1432 coupled to an instruction cache circuit system 1434, the instruction cache circuit system 1434 being coupled to a translation lookaside buffer (TLB) 1436, the translation lookaside buffer 1436 being coupled to an instruction fetch circuit system 1438, and the instruction fetch circuit system 1438 being coupled to a decoding circuit system 1440. In some examples, the instruction cache circuit system 1434 is included in a memory unit circuit system 1470 rather than in the front-end circuit system 1430. The decoding circuit system 1440 (or decoder) may decode the instructions and generate, as output, one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals decoded from, otherwise reflecting, or derived from the original instructions. The decoding circuit system 1440 may further include address generation unit (AGU, not shown) circuitry. In some examples, the AGU generates LSU addresses using the forwarded register ports and may further perform branch forwarding (e.g., immediate offset branch forwarding, LR register branch forwarding, etc.). The decoding circuit system 1440 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, look-up tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc. In some examples, the core 1490 includes a microcode ROM (not shown) or other medium (e.g., in the decoding circuit system 1440 or otherwise within the front-end circuit system 1430) storing microcode for certain macroinstructions. In some examples, the decoding circuit system 1440 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache decoded operations, microtags, or micro-operations generated during the decoding or other stages of the processor pipeline 1400. The decoding circuit system 1440 may be coupled to a rename / allocator unit circuit system 1452 in an execution engine circuit system 1450.

[0245] The execution engine circuitry 1450 includes a rename / allocator unit circuitry 1452 that is coupled to a retirement unit circuitry 1454 and a collection of one or more scheduler circuitries 1456. The one or more scheduler circuitries 1456 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. In some examples, the one or more scheduler circuitries 1456 may include an arithmetic logic unit (ALU) scheduler / scheduling circuitry, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuitry, an AGU queue, and so on. The one or more scheduler circuitries 1456 are coupled to the one or more physical register file circuitries 1458. Each of the one or more physical register file circuitries 1458 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integer, scalar floating point, packed integer, packed floating point, vector integer, vector floating point, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and the like.

[0246] In some examples, the (one or more) physical register file circuitry 1458 includes vector register unit circuitry, write mask register unit circuitry, and scalar register unit circuitry. These register units may provide architectural vector registers, vector mask registers, general-purpose registers, and the like. The (one or more) physical register file circuitry 1458 is coupled to the retirement unit circuitry 1454 (also referred to as a retirement queue (“retire queue” or “retirement queue”)) to illustrate various ways in which register renaming and out-of-order execution may be implemented (e.g., using the (one or more) reorder buffers (ROB) and the (one or more) retirement register files; using the (one or more) future files, the (one or more) history buffers, and the (one or more) retirement register files; using register mapping and register pooling, etc.). The retirement unit circuitry 1454 and the (one or more) physical register file circuitry 1458 are coupled to the (one or more) execution clusters 1460. The (one or more) execution clusters 1460 include a set of one or more execution unit circuitry 1462 and a set of one or more memory access circuitry 1464. The (one or more) execution unit circuitry 1462 may perform various arithmetic, logical, floating-point, or other types of operations (e.g., shift, add, subtract, multiply) and may operate on various data types (e.g., scalar integer, scalar floating-point, packed integer, packed floating-point, vector integer, vector floating-point). Although some examples may include multiple execution units or execution unit circuitry dedicated to a particular function or set of functions, other examples may include only one execution unit circuitry or multiple execution units / execution unit circuitry that all perform all functions. The (one or more) scheduler circuitry 1456, the (one or more) physical register file circuitry 1458, and the (one or more) execution clusters 1460 are shown as potentially multiple because certain examples create separate pipelines for certain types of data / operations (e.g., scalar integer pipeline, scalar floating-point / packed integer / packed floating-point / vector integer / vector floating-point pipeline, and / or a memory access pipeline each having its own scheduler circuitry, (one or more) physical register file circuitry, and / or execution cluster—and in the case of a separate memory access pipeline, some examples in which only the execution cluster of that pipeline has the (one or more) memory access circuitry 1464). It should also be understood that in the case of using separate pipelines, one or more of these pipelines may be out-of-order issue / execution, and the remaining pipelines may be in-order issue / execution.

[0247] In some examples, the execution engine unit circuitry 1450 may perform load store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown), as well as address phase and write-back, data phase load, store, and branch.

[0248] A set of memory access circuitry 1464 is coupled to memory unit circuitry 1470, which includes data TLB circuitry 1472, which is coupled to data cache circuitry 1474, which is coupled to a second-level (L2) cache circuitry 1476. In some examples, the memory access circuitry 1464 may include load unit circuitry, store address unit circuitry, and store data unit circuitry, each of which is coupled to the data TLB circuitry 1472 in the memory unit circuitry 1470. The instruction cache circuitry 1434 is further coupled to the second-level (L2) cache circuitry 1476 in the memory unit circuitry 1470.

[0249] In some examples, the instruction cache 1434 and the data cache 1474 are combined into a single instruction and data cache (not shown) in the L2 cache circuitry 1476, a third-level (L3) cache circuitry (not shown), and / or main memory. The L2 cache circuitry 1476 is coupled to one or more other levels of cache and ultimately to main memory.

[0250] The core 1490 may support one or more instruction sets (e.g., x86 instruction set architecture (optionally with some extensions added with more recent versions); MIPS instruction set architecture; ARM instruction set architecture (optionally with optional additional extensions such as NEON)), including the (one or more) instructions described herein. In some examples, the core 1490 includes logic for supporting SIMD instruction set architecture extensions (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using SIMD data.

[0251] Example(s) of execution unit circuitry.

[0252] Figure 15 Illustrates examples of (one or more) execution unit circuitry, such as Figure 14BThe execution unit circuitry 1462 (one or more). As shown, the execution unit circuitry 1462 (one or more) may include one or more ALU circuits 1501, optional vector / single instruction multiple data (SIMD) circuits 1503, load / store circuits 1505, branch / jump circuits 1507, and / or floating-point unit (FPU) circuits 1509. The ALU circuits 1501 perform integer arithmetic and / or boolean operations. The vector / SIMD circuits 1503 perform vector / SIMD operations on packed data (such as SIMD / vector registers). The load / store circuits 1505 execute load and store instructions to load data from memory into registers or store data from registers into memory. The load / store circuits 1505 may also generate addresses. The branch / jump circuits 1507 cause a branch or jump to a memory address depending on the instruction. The FPU circuits 1509 perform floating-point arithmetic. The width of the execution unit circuitry 1462 (one or more) varies depending on the example and may be in the range of, for example, from 16 bits to 1024 bits. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).

[0253] Example register architecture.

[0254] Figure 16 is a block diagram of a register architecture 1600 according to some examples. As shown, the register architecture 1600 includes vector / SIMD registers 1610, the width of which varies from 128 bits to 1024 bits. In some examples, the vector / SIMD registers 1610 are physically 512 bits, and depending on the mapping, only some of the lower bits are used. For example, in some examples, the vector / SIMD registers 1610 are 512-bit ZMM registers: the lower 256 bits are used for YMM registers, and the lower 128 bits are used for XMM registers. Thus, there is register coverage. In some examples, the vector length field selects between a maximum length and one or more other shorter lengths, where each such shorter length is half the length of the previous length. Scalar operations are operations performed on the lowest-order data element positions in the ZMM / YMM / XMM registers; depending on the example, the higher-order data element positions remain the same as before the instruction or are zeroed.

[0255] In some examples, the register architecture 1600 includes write mask / predicate registers 1615. For example, in some examples, there are 8 write mask / predicate registers (sometimes referred to as k0 through k7), each sized at 16 bits, 32 bits, 64 bits, or 128 bits. The write mask / predicate registers 1615 can allow for merging (e.g., allowing any set of elements in the destination to be exempt from update during the execution of any operation) and / or zeroing (e.g., a zeroing vector mask allows any set of elements in the destination to be zeroed during the execution of any operation). In some examples, each data element position in a given write mask / predicate register 1615 corresponds to a data element position in the destination. In other examples, the write mask / predicate registers 1615 are scalable and consist of a set number of enable bits for a given vector element (e.g., 8 enable bits for each 64-bit vector element).

[0256] The register architecture 1600 includes a plurality of general-purpose registers 1625. These registers can be 16 bits, 32 bits, 64 bits, etc., and are available for scalar operations. In some examples, these registers are referenced by the names RAX, RBX, RCX, RDX, RBP, RSI, RDI, RSP, and R8 through R15.

[0257] In some examples, the register architecture 1600 includes a scalar floating point (FP) register file 1645 for scalar floating point operations on 32 / 64 / 80-bit floating point data using the x87 instruction set architecture extensions, or for performing operations on 64-bit packed integer data as MMX registers, and for saving operands for some operations performed between MMX and XMM registers.

[0258] One or more flag registers 1640 (e.g., EFLAGS, RFLAGS, etc.) store status and control information for arithmetic, comparison, and system operations. For example, one or more flag registers 1640 can store condition code information such as carry, parity, auxiliary carry, zero, sign, and overflow. In some examples, one or more flag registers 1640 are referred to as program status and control registers.

[0259] The segment registers 1620 contain segment points for use when accessing memory. In some examples, these registers are referenced by the names CS, DS, SS, ES, FS, and GS.

[0260] Model-specific registers or machine specific registers (MSRs) 1635 control and report processor performance. Most MSRs 1635 handle system-related functions and are not accessible by applications. For example, an MSR can provide control over one or more of the following: performance monitoring counters, debug extensions, memory type range registers, thermal and power management, instruction-specific support, and / or processor feature / mode support. Machine check registers 1660 consist of control, status, and error-reporting MSRs that are used to detect and report hardware errors. One or more control registers 1655 (e.g., CR0-CR4) determine the operating mode of the processor (e.g., processors 770, 780, 738, 715, and / or 800) and the characteristics of the currently executing task. In some examples, MSR 1635 is a subset of control register 1655.

[0261] One or more instruction pointer registers 1630 store instruction pointer values. Debug registers 1650 control and allow monitoring of debug operations of the processor or core.

[0262] Memory (mem) management registers 1665 specify the locations of data structures used for protected-mode memory management. These registers can include global descriptor table register (GDTR), interrupt descriptor table register (IDTR), task register, and local descriptor table register (LDTR) registers.

[0263] Alternative examples can use wider or narrower registers. Additionally, alternative examples can use more, fewer, or different register banks and registers. The register architecture 1600 can be, for example, in register bank 110 or one or more physical register bank circuitry 1458.

[0264] Instruction set architecture.

[0265] An instruction set architecture (ISA) may include one or more instruction formats. A given instruction format may define various fields (e.g., number of bits, position of bits) to specify the operation to be performed (e.g., opcode) and the operand(s) on which the operation is to be performed and / or other data field(s) (e.g., mask), etc. Some instruction formats are further decomposed by the definition of instruction templates (or sub-formats). For example, an instruction template of a given instruction format may be defined as having different subsets of the fields of that instruction format (the included fields are typically in the same order, but at least some fields have different bit positions since fewer fields are included), and / or as having a given field interpreted in a different way. Thus, each instruction of the ISA is expressed using a given instruction format (and if defined, using a given instruction template within that instruction format's instruction templates), and includes fields for specifying the operation and operands. For example, an example ADD (addition) instruction has a specific opcode and instruction format, and the specific instruction format includes an opcode field for specifying the opcode and an operand field for selecting the operands (source1 / destination and source2); and the appearance of the ADD instruction in the instruction stream will result in specific contents in the operand field for selecting specific operands. Additionally, although the following description is in the context of the x86 instruction set architecture, applying the teachings of the present disclosure to another ISA is within the knowledge of those skilled in the art.

[0266] Example instruction format.

[0267] Examples of the (one or more) instructions described herein may be embodied in different formats. Additionally, example systems, architectures, and pipelines are detailed below. Examples of the (one or more) instructions may be executed on such systems, architectures, and pipelines, but are not limited to those detailed.

[0268] Figure 17 Example of an illustrated instruction format. As shown, an instruction may include multiple components, including but not limited to one or more fields for: one or more prefixes 1701, an opcode 1703, addressing information 1705 (e.g., register identifier, memory addressing information, etc.), a displacement value 1707, and / or an immediate value 1709. Note that some instructions utilize some or all of the fields in the format, while other instructions may use only the field for the opcode 1703. In some examples, the illustrated order is the order in which these fields are to be encoded, however, it should be understood that in other examples, these fields may be encoded, combined, etc. in a different order.

[0269] (One or more) prefix fields 1701 modify the instruction when in use. In some examples, one or more prefixes are used for repeat string instructions (e.g., 0xF0, 0xF2, 0xF3, etc.), provide section override (e.g., 0x2E, 0x36, 0x3E, 0x26, 0x64, 0x65, 0x2E, 0x3E, etc.), perform a bus lock operation, and / or change the operand (e.g., 0x66) and address size (e.g., 0x67). Certain instructions require mandatory prefixes (e.g., 0x66, 0xF2, 0xF3, etc.). Some of these prefixes can be considered "traditional" prefixes. Other prefixes (one or more examples of which are detailed herein) indicate and / or provide further capabilities, such as specifying a particular register, etc. These other prefixes typically follow the "traditional" prefixes.

[0270] The opcode field 1703 is used to at least partially define the operation to be performed when decoding the instruction. In some examples, the length of the primary opcode encoded in the opcode field 1703 is 1, 2, or 3 bytes. In other examples, the primary opcode can be of a different length. An additional 3-bit opcode field is sometimes encoded in another field.

[0271] The addressing information field 1705 is used to address one or more operands of the instruction, such as a location in memory or one or more registers. Figure 18 An example of the addressing information field 1705 is illustrated. In this illustration, an optional MOD R / M byte 1802 and an optional Scale, Index, Base (SIB) byte 1804 are shown. The MOD R / M byte 1802 and the SIB byte 1804 are used to encode up to two operands of the instruction, each of which is either a direct register or a valid memory address. Note that both of these fields are optional, i.e., not all instructions include one or more of these fields. The MOD R / M byte 1802 includes a MOD field 1842, a register (reg) field 1844, and an R / M field 1846.

[0272] The content of the MOD field 1842 distinguishes between memory access modes and non-memory access modes. In some examples, when the MOD field 1842 has a binary value of 11 (11b), register direct addressing mode is used, otherwise register indirect addressing mode is used.

[0273] Register field 1844 can encode a destination register operand or a source register operand, or can encode an opcode extension without encoding any instruction operand. The content of register field 1844 directly specifies or specifies through address generation the location of the source or destination operand (in a register or in memory). In some examples, register field 1844 is supplemented with additional bits from a prefix (e.g., prefix 1701) to allow for greater addressing.

[0274] The R / M field 1846 can be used to encode an instruction operand that references a memory address, or can be used to encode a destination register operand or a source register operand. Note that in some examples, the R / M field 1846 can be combined with the MOD field 1842 to specify an addressing mode.

[0275] The SIB byte 1804 includes a scale field 1852, an index field 1854, and a base field 1856 for address generation. The scale field 1852 indicates a scale factor. The index field 1854 specifies the index register to be used. In some examples, the index field 1854 is supplemented with additional bits from a prefix (e.g., prefix 1701) to allow for greater addressing. The base field 1856 specifies the base register to be used. In some examples, the base field 1856 is supplemented with additional bits from a prefix (e.g., prefix 1701) to allow for greater addressing. In practice, the content of the scale field 1852 allows the content of the index field 1854 to be scaled for memory address generation (e.g., for address generation using (2 缩放 *index + base).

[0276] Some addressing forms utilize displacement values to generate a memory address. For example, a memory address can be generated according to (2 缩放 *index + base + displacement), (index * scale + displacement), (r / m + displacement), (instruction pointer (RIP / EIP) + displacement), (register + displacement), etc. The displacement can be a value of 1 byte, 2 bytes, 4 bytes, etc. In some examples, the displacement field 1707 provides this value. Additionally, in some examples, the displacement factor is encoded in the MOD field of the addressing information field 1705, which indicates a compressed displacement scheme for which the displacement value is calculated and stored in the displacement field 1707.

[0277] In some examples, the immediate value field 1709 specifies an immediate value for an instruction. The immediate value can be encoded as a 1-byte value, a 2-byte value, a 4-byte value, etc.

[0278] Figure 19Illustrate an example of the first prefix 1701A. In some examples, the first prefix 1701A is an example of a REX prefix. Instructions using this prefix can specify general registers, 64-bit packed data registers (e.g., single instruction multiple data (SIMD) registers or vector registers), and / or control and debug registers (e.g., CR8-CR15 and DR8-DR15).

[0279] Depending on the format, instructions using the first prefix 1701A can use a 3-bit field to specify up to three registers: 1) using the reg field 1844 and the R / M field 1846 of the MOD R / M byte 1802; 2) using the MOD R / M byte 1802 and the SIB byte 1804, including using the reg field 1844, the base field 1856, and the index field 1854; or 3) using the register field of the opcode.

[0280] In the first prefix 1701A, bit positions 7:4 of the payload byte are set to 0100. Bit position 3 (W) can be used to determine the operand size, but may not be able to determine the operand width alone. Thus, when W = 0, the operand size is determined by the code segment descriptor (CS.D), and when W = 1, the operand size is 64 bits.

[0281] Note that adding another bit allows addressing 16 (2 4 ) registers, while the individual MOD R / M reg field 1844 and MOD R / M R / M field 1846 can each address only 8 registers.

[0282] In the first prefix 1701A, bit position 2 (R) can be an extension of the MOD R / M reg field 1844 and can be used to modify the MOD R / M reg field 1844 when this field encodes a general register, a 64-bit packed data register (e.g., an SSE register), or a control or debug register. When the MOD R / M byte 1802 specifies other registers or defines an extended opcode, R is ignored.

[0283] Bit position 1 (X) can modify the SIB byte index field 1854.

[0284] Bit position O (B) can modify the base address in the MOD R / M R / M field 1846 or the SIB byte base field 1856; or it can modify the opcode register field used to access a general register (e.g., general register 1625).

[0285] Figures 20A - 20DAn example showing how to use the R, X, and B fields of the first prefix 1701A. Figure 20A An illustration showing that when the SIB byte 1804 is not used for memory addressing, the R and B from the first prefix 1701A are used to extend the reg field 1844 and the R / M field 1846 of the MOD R / M byte 1802. Figure 20B An illustration showing that when the SIB byte 1804 is not used, the R and B from the first prefix 1701A are used to extend the reg field 1844 and the R / M field 1846 of the MOD R / M byte 1802 (for register-register addressing). Figure 20C An illustration showing that when the SIB byte 1804 is used for memory addressing, the R, X, and B from the first prefix 1701A are used to extend the reg field 1844, the index field 1854, and the base field 1856 of the MOD R / M byte 1802. Figure 20D An illustration showing that when a register is encoded in the opcode 1703, the B from the first prefix 1701A is used to extend the reg field 1844 of the MOD R / M byte 1802.

[0286] Figures 21A - 21B An example showing the second prefix 1701B. In some examples, the second prefix 1701B is an example of a VEX prefix. The encoding of the second prefix 1701B allows an instruction to have more than two operands and allows SIMD vector registers (e.g., vector / SIMD register 1610) to be longer than 64 bits (e.g., 128 bits and 256 bits). The use of the second prefix 1701B provides a syntax for three-operand (or more operand) instructions. For example, a previous two-operand instruction performed an operation such as A = A + B, which overwrote the source operand. The use of the second prefix 1701B enables operands to perform non-destructive operations such as A = B + C.

[0287] In some embodiments, the second prefix 1701B has two forms - a two-byte form and a three-byte form. The two-byte second prefix 1701B is mainly used for 128-bit, scalar, and some 256-bit instructions; while the three-byte second prefix 1701B provides a compact replacement for the first prefix 1701A and 3-byte opcode instructions.

[0288] Figure 21AIllustrates an example of the two - byte form of the second prefix 1701B. In some examples, the format field 2101 (byte 0 2103) contains the value C5H. In some examples, byte 1 2105 includes an "R" value in bit [7]. This value is the complement of the "R" value of the first prefix 1701A. Bit [2] is used to specify the length (L) of the vector (where the value 0 is a scalar or a 128 - bit vector, and the value 1 is a 256 - bit vector). Bits [1:0] provide opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, specified in inverted (ones - complement) form, and is valid for instructions with two or more source operands; 2) encoding the destination register operand, specified in ones - complement form, for certain vector shifts; or 3) not encoding any operand, the field is reserved and should contain a value such as 1111b.

[0289] Instructions using this prefix can use the R / M field 1846 of MOD R / M to encode instruction operands that reference a memory address, or to encode the destination register operand or the source register operand SIB.

[0290] Instructions using this prefix can use the reg field 1844 of MOD R / M to encode the destination register operand or the source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.

[0291] For instruction syntax that supports four operands, vvvv, the R / M field 1846 of MOD R / M, and the reg field 1844 of MOD R / M encode three of the four operands. Then, bits [7:4] of the immediate value field 1709 are used to encode the third source register operand.

[0292] Figure 21B Illustrates an example of the three - byte form of the second prefix 1701B. In some examples, the format field 2111 (byte 0 2113) contains the value C4H. Byte 1 2115 includes "R", "X", and "B" in bits [7:5], which are the complements of the same values of the first prefix 1701A detailed later. Bits [4:0] of byte 1 2115 (shown as mmmmm) include content for encoding one or more implicit leading opcode bytes as needed. For example, 00001 implies a 0FH leading opcode, 00010 implies a 0F38H leading opcode, 00011 implies a 0F3AH leading opcode, and so on.

[0293] Bit [7] of byte 22117 is used similarly to W of the first prefix 1701A, including helping to determine the size of the operand that can be promoted. Bit [2] is used to specify the length (L) of the vector (where the value 0 is a scalar or a 128-bit vector, and the value 1 is a 256-bit vector). Bits [1:0] provide opcode extensibility, equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). Bits [6:3], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with two or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form and is used for certain vector shifts; or 3) not encoding any operand, and this field is reserved and should contain a certain value, such as 1111b.

[0294] Instructions using this prefix can use the R / M field 1846 of MOD R / M to encode the instruction operand that references a memory address or to encode the destination register operand or the source register operand.

[0295] Instructions using this prefix can use the reg field 1844 of MOD R / M to encode the destination register operand or the source register operand, or be treated as an opcode extension and not be used to encode any instruction operand.

[0296] For the instruction syntax that supports four operands vvvv, the R / M field 1846 of MOD R / M and the reg field 1844 of MOD R / M encode three of the four operands. Then, bits [7:4] of the immediate value field 1709 are used to encode the third source register operand.

[0297] Figure 22 An example of the third prefix 1701C is illustrated. In some examples, the third prefix 1701C is an example of an EVEX prefix. The third prefix 1701C is a four-byte prefix.

[0298] The third prefix 1701C can encode 32 vector registers (e.g., 128-bit, 256-bit, and 512-bit registers) in 64-bit mode. In some examples, instructions that utilize a write mask / operation mask (see the discussion of the registers in the previous figures (such as Figure 16 )) or predicates utilize this prefix. The operation mask register allows conditional processing or selection control. Operation mask instructions - whose source / destination operand is an operation mask register and treats the content of the operation mask register as a single value - are encoded using the second prefix 1701B.

[0299] The third prefix 1701C can encode functions specific to instruction classes (e.g., packed instructions with “load + operation” semantics can support an embedded broadcast function; floating-point instructions with rounding semantics can support a static rounding function; floating-point instructions with non-rounding arithmetic semantics can support an “all exceptions suppressed” function; and so on).

[0300] The first byte of the third prefix 1701C is the format field 2211, which has a value of 62H in some examples. The subsequent bytes are referred to as payload bytes 2215 - 2219 and together form a 24-bit value of P[23:0], providing specific capabilities in the form of one or more fields (detailed herein).

[0301] In some examples, P[1:0] of the payload byte 2219 is the same as the two low-order mm bits. In some examples, P[3:2] is reserved. Bit P[4] (R’) allows access to the high 16 vector register set when combined with P[7] and the reg field 1844 of MOD R / M. When SIB type addressing is not required, P[6] can also provide access to the high 16 vector registers. P[7:5] consists of R, X, and B, which are operand specifier modifiers for vector registers, general-purpose registers, and memory addressing, and allow access to the next set of 8 registers beyond the low 8 registers when combined with the MOD R / M register field 1844 and the R / M field 1846 of MOD R / M. P[9:8] provides opcode extensibility equivalent to some traditional prefixes (e.g., 00 = no prefix, 01 = 66H, 10 = F3H, and 11 = F2H). In some examples, P

[10] is a fixed value of 1. P[14:11], shown as vvvv, can be used for: 1) encoding the first source register operand, which is specified in inverted (ones' complement) form and is valid for instructions with 2 or more source operands; 2) encoding the destination register operand, which is specified in ones' complement form for certain vector shifts; or 3) not encoding any operand, in which case this field is reserved and should contain a value, e.g., 1111b.

[0302] P

[15] is similar to the W of the first prefix 1701A and the second prefix 1701B and can be used as an opcode extension bit or an operand size promotion.

[0303] P[18:16] specifies the index of a register in an operation mask (write mask) register (e.g., write mask / predicate register 1615). In some examples, a particular value aaa = 000 has a special behavior that implies no operation mask is used for a particular instruction (which can be implemented in various ways, including using an operation mask hardwired to all instructions or hardware that bypasses the mask hardware). When merged, the vector mask allows any set of elements in the destination to be protected from update during the execution of any operation (specified by the base operation and the augmentation operation); in some other examples, the old value of each element of the destination where the corresponding mask bit has 0 is maintained. In contrast, when zeroing, the vector mask allows any set of elements in the destination to be zeroed during the execution of any operation (specified by the base operation and the augmentation operation); in some examples, the elements of the destination are set to 0 when the corresponding mask bit has a value of 0. A subset of this functionality is the ability to control the vector length of the operation being performed (i.e., the span from the first to the last element to be modified), however, the elements being modified do not have to be contiguous. Thus, the operation mask field allows partial vector operations, including loads, stores, arithmetic, logic, etc. Although examples have been described in which the content of the operation mask field selects one of the operation mask registers that contains the operation mask to be used among multiple operation mask registers (and thus the content of the operation mask field indirectly identifies the masking to be performed), alternative examples instead or additionally allow the content of the mask write field to directly specify the masking to be performed.

[0304] P

[19] can be combined with P[14:11] to encode a second source vector register in non-destructive source syntax, which can access the upper 16 vector registers using P

[19] . P

[20] encodes multiple functions, which vary across different classes of instructions and can affect the meaning of the vector length / rounding control specifier field (P[22:21]). P

[23] indicates support for merge-write masks (e.g., when set to 0) or support for zeroing and merge-write masks (e.g., when set to 1).

[0305] Examples of encoding registers in an instruction using the third prefix 1701C are detailed in the table below. Table 4: 32-register support in 64-bit mode Table 5: Encoded register specifiers in 32-bit mode Table 6: Operation mask register specifier encoding

[0306] Graphics Execution Unit

[0307] Figures 23A - 23B FIG. illustrates thread execution logic 2300 according to an example described herein, the thread execution logic 2300 including an array of processing elements employed in a graphics processor core. Figures 23A - 23B Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to the manner described elsewhere herein, but are not limited thereto. Figure 23A represents an execution unit within a general-purpose graphics processor, while Figure 23B represents an execution unit that can be used within a computing accelerator.

[0308] As illustrated in Figure 23A , in some examples, thread execution logic 2300 includes a shader processor 2302, a thread dispatcher 2304, an instruction cache 2306, a scalable execution unit array including a plurality of execution units 2308A - 2308N, a sampler 2310, a shared local memory 2311, a data cache 2312, and a data port 2314. In some examples, the scalable execution unit array can be dynamically scaled by enabling or disabling one or more execution units (e.g., any of execution units 2308A, 2308B, 2308C, 2308D through 2308N - 1 and 2308N) based on the computational requirements of the workload. In some examples, the components included are interconnected via an interconnect structure that links to each of the components. In some examples, thread execution logic 2300 includes one or more connections to memory (such as system memory or cache memory) via the instruction cache 2306, the data port 2314, the sampler 2310, and one or more of the execution units 2308A - 2308N. In some examples, each execution unit (e.g., 2308A) is an independent programmable general-purpose computing unit capable of executing multiple synchronous hardware threads and processing multiple data elements in parallel for each thread. In examples, the array of execution units 2308A - 2308N is scalable to include any number of individual execution units.

[0309] In some examples, execution units 2308A - 2308N are mainly used to execute shader programs. The shader processor 2302 can process various shader programs and can dispatch execution threads associated with the shader programs via the thread dispatcher 2304. In some examples, the thread dispatcher includes logic for arbitrating requests for threads from the graphics pipeline and the media pipeline and instantiating the requested threads on one or more of the execution units 2308A - 2308N. For example, the geometry pipeline can dispatch a vertex shader, a tessellation shader, or a geometry shader to the thread execution logic for processing. In some examples, the thread dispatcher 2304 can also process runtime thread generation requests from the executed shader programs.

[0310] In some examples, the execution units 2308A - 2308N support an instruction set that includes native support for many standard 3D graphics shader instructions, enabling shader programs from graphics libraries (e.g., Direct3D and OpenGL) to be executed with minimal translation. These execution units support vertex and geometry processing (e.g., vertex programs, geometry programs, vertex shaders), pixel processing (e.g., pixel shaders, fragment shaders), and general - purpose processing (e.g., compute and media shaders). Each of the execution units 2308A - 2308N is capable of multi - issue single instruction multiple data (SIMD) execution, and multi - threading operations enable an efficient execution environment in the face of higher - latency memory accesses. Each hardware thread within each execution unit has a dedicated high - bandwidth register file and an associated independent thread state. Execution is multi - issued per clock for pipelines that can perform integer operations, single - precision floating - point operations, and double - precision floating - point operations, can have SIMD branch capabilities, can perform logical operations, can perform transcendental operations, and can perform other miscellaneous operations. When waiting for data from one of the shared functions in memory or a shared function, the dependency logic within the execution units 2308A - 2308N puts the waiting threads to sleep until the requested data has been returned. While the waiting threads are sleeping, the hardware resources can be dedicated to processing other threads. For example, during the latency associated with vertex shader operations, the execution units can perform operations for pixel shaders, fragment shaders, or another type of shader program that includes a different vertex shader. Examples can be applied to use execution that utilizes single instruction multiple thread (SIMT) as an alternative to, or in addition to, the use of SIMD. References to SIMD cores or operations can also apply to SIMT, or to a combination of SIMD and SIMT.

[0311] Each execution unit in execution units 2308A - 2308N operates on an array of data elements. The number of data elements is the "execution size", or the number of lanes for the instruction. An execution lane is a logical unit for data element access, masking, and flow control execution within an instruction. The number of lanes can be independent of the number of physical arithmetic logic units (ALUs) or floating point units (FPUs) for a particular graphics processor. In some examples, execution units 2308A - 2308N support integer and floating point data types.

[0312] The execution unit instruction set includes SIMD instructions. Various data elements can be stored in registers as packed data types, and the execution units will process the various elements based on the data size of the elements. For example, when operating on a 256 - bit wide vector, the 256 bits of the vector are stored in a register, and the execution unit operates on the vector as four separate 64 - bit packed data elements (Quad - Word (QW) size data elements), eight separate 32 - bit packed data elements (Double Word (DW) size data elements), sixteen separate 16 - bit packed data elements (Word (W) size data elements), or thirty - two separate 8 - bit data elements (byte (B) size data elements). However, different vector widths and register sizes are possible.

[0313] In some examples, one or more execution units can be combined into fused execution units 2309A - 2309N, which have thread control logic (2307A - 2307N) common to the fused EUs to execute separate SIMD hardware threads. The number of EUs in a fused EU group can vary according to the example. Additionally, various SIMD widths can be executed per EU, including but not limited to SIMD8, SIMD16, and SIMD32. Each fused graphics execution unit 2309A - 2309N includes at least two execution units. For example, fused execution unit 2309A includes a first EU 2308A, a second EU 2308B, and thread control logic 2307A common to the first EU 2308A and the second EU 2308B. Thread control logic 2307A controls the threads executed on fused graphics execution unit 2309A, allowing each EU within fused execution units 2309A - 2309N to execute using a common instruction pointer register.

[0314] One or more internal instruction caches (e.g., 2306) are included in the thread execution logic 2300 to cache thread instructions for the execution units. In some examples, one or more data caches (e.g., 2312) are included to cache thread data during thread execution. Threads executing on the execution logic 2300 may also store data that is explicitly managed in the shared local memory 2311. In some examples, a sampler 2310 is included to provide texture sampling for 3D operations and media sampling for media operations. In some examples, the sampler 2310 includes specialized texture or media sampling functionality to process texture data or media data during the sampling process before providing the sampled data to the execution units.

[0315] During execution, the graphics pipeline and the media pipeline send thread initiation requests to the thread execution logic 2300 via the thread generation and dispatch logic. Once a group of geometric objects has been processed and rasterized into pixel data, the pixel processor logic (e.g., pixel shader logic, fragment shader logic, etc.) within the shader processor 2302 is called to further compute output information and cause the results to be written to an output surface (e.g., color buffer, depth buffer, stencil buffer, etc.). In some examples, the pixel shader or fragment shader computes the values of vertex attributes, and the values of the vertex attributes are interpolated across the rasterized objects. In some examples, the pixel processor logic within the shader processor 2302 then executes a pixel shader program or a fragment shader program supplied by an application programming interface (API). To execute the shader program, the shader processor 2302 dispatches a thread to an execution unit (e.g., 2308A) via the thread dispatcher 2304. In some examples, the shader processor 2302 uses the texture sampling logic in the sampler 2310 to access texture data in a texture map stored in memory. Arithmetic operations on the texture data and the input geometric data compute the pixel color data for each geometric fragment or discard one or more pixels without further processing.

[0316] In some examples, the data port 2314 provides a memory access mechanism for the thread execution logic 2300 to output processed data to memory for further processing on the graphics processor output pipeline. In some examples, the data port 2314 includes or is coupled to one or more cache memories (e.g., the data cache 2312) to cache data for memory access via the data port.

[0317] In some examples, the execution logic 2300 may further include a ray tracer 2305 that can provide ray tracing acceleration capabilities. The ray tracer 2305 may support a ray tracing instruction set that includes instructions / functions for ray generation.

[0318] Figure 23B FIG. illustrates exemplary internal details of an execution unit 2308 according to an example. The graphics execution unit 2308 may include an instruction fetch unit 2337, a general register file (GRF) array 2324, an architectural register file (ARF) array 2326, a thread arbiter 2322, a dispatch unit 2330, a branch unit 2332, a set of SIMD floating-point units (FPUs) 2334, and in some examples, a set of integer SIMD ALUs 2335. The GRF 2324 and the ARF 2326 include a set of general register files and architectural register files associated with each simultaneous hardware thread that can be active in the graphics execution unit 2308. In some examples, the per-thread architectural state is maintained in the ARF 2326, while the data used during thread execution is stored in the GRF 2324. The execution state of each thread, including the instruction pointer for each thread, may be saved in thread-specific registers in the ARF 2326.

[0319] In some examples, the graphics execution unit 2308 has an architecture that is a combination of Simultaneous Multi-Threading (SMT) and Fine-Grained Interleaved Multi-Threading (IMT). This architecture has a modular configuration that can be fine-tuned at design time based on the target number of synchronized threads and the number of registers per execution unit, where the execution unit resources are divided across the logic for executing multiple synchronized threads. The number of logical threads that can be executed by the graphics execution unit 2308 is not limited to the number of hardware threads, and multiple logical threads can be assigned to each hardware thread.

[0320] In some examples, the graphics execution unit 2308 may issue multiple instructions in concert, and these instructions may each be different instructions. The thread arbiter 2322 of the graphics execution unit 2308 may dispatch the instructions to one of the send unit 2330, the branch unit 2332, or the (one or more) SIMD FPUs 2334 for execution. Each execution thread may access 128 general-purpose registers within the GRF 2324, where each register may store 32 bytes that can be accessed as a SIMD 8-element vector with 32-byte data elements. In some examples, each execution unit thread has access to 4 kilobytes within the GRF 2324, but the examples are not limited thereto, and more or fewer register resources may be provided in other examples. In some examples, the graphics execution unit 2308 is partitioned into seven hardware threads that can independently perform computational operations, but the number of threads per execution unit may also vary according to the example. For example, in some examples, up to 16 hardware threads are supported. In an example where seven threads can access 4 kilobytes, the GRF 2324 may store a total of 28 kilobytes. In the case where 16 threads can access 4 kilobytes, the GRF 2324 may store a total of 64 kilobytes. Flexible addressing modes may permit addressing of registers together, thereby effectively creating wider registers or representing strided rectangular block data structures.

[0321] In some examples, memory operations, sampler operations, and other longer-latency system communications are dispatched via "send" instructions executed by the messaging send unit 2330. In some examples, branch instructions are dispatched to a dedicated branch unit 2332 to facilitate SIMD scatter and eventual gather.

[0322] In some examples, the graphics execution unit 2308 includes one or more SIMD floating-point units ((one or more) FPUs) 2334 for performing floating-point operations. In some examples, the (one or more) FPUs 2334 also support integer computations. In some examples, the (one or more) FPUs 2334 may perform up to M 32-bit floating-point (or integer) operations in SIMD, or perform up to 2M 16-bit integer or 16-bit floating-point operations in SIMD. In some examples, at least one of the (one or more) FPUs provides extended mathematical capabilities that support high-throughput transcendental math functions and double-precision 64-bit floating point. In some examples, a set of 8-bit integer SIMD ALUs 2335 also exists and may be specifically optimized to perform operations associated with machine learning computations.

[0323] In some examples, an array of multiple instances of the graphics execution unit 2308 may be instantiated in graphics sub-core groupings (e.g., sub-slices). For scalability, a product architect may choose the exact number of execution units per sub-core grouping. In some examples, the execution unit 2308 may execute instructions across multiple execution lanes. In further examples, each thread executing on the graphics execution unit 2308 executes on a different lane.

[0324] Figure 24 FIG. illustrates an additional execution unit 2400 according to an example. In some examples, the execution unit 2400 includes a thread control unit 2401, a thread state unit 2402, an instruction fetch / prefetch unit 2403, and an instruction decode unit 2404. The execution unit 2400 additionally includes a register file 2406 that stores registers that may be assigned to hardware threads within the execution unit. The execution unit 2400 additionally includes a dispatch unit 2407 and a branch unit 2408. In some examples, the dispatch unit 2407 and the branch unit 2408 can operate in a manner similar to the dispatch unit 2330 and the branch unit 2332 of the Figure 23B graphics execution unit 2308.

[0325] The execution unit 2400 further includes a compute unit 2410 that includes multiple different types of functional units. In some examples, the compute unit 2410 includes an ALU unit 2411 that includes an array of arithmetic logic units. The ALU unit 2411 may be configured to perform 64-bit, 32-bit, and 16-bit integer and floating-point operations. Integer and floating-point operations may be performed simultaneously. The compute unit 2410 may further include a systolic array 2412 and a math unit 2413. The systolic array 2412 includes a wide W and deep D network of data processing units that may be used to perform vector or other data parallel operations in a systolic manner. In some examples, the systolic array 2412 may be configured to perform matrix operations such as matrix dot product operations. In some examples, the systolic array 2412 supports 16-bit floating-point operations as well as 8-bit and 4-bit integer operations. In some examples, the systolic array 2412 may be configured to accelerate machine learning operations. In such examples, the systolic array 2412 may be configured with support for the bfloat 16-bit floating-point format. In some examples, the math unit 2413 may be included to perform a specific subset of math operations in an efficient and lower power manner than the ALU unit 2411. The math unit 2413 may include a variant of the math logic (e.g., the math logic of the shared functional logic) that may be found in the shared functional logic provided by other examples. In some examples, the math unit 2413 may be configured to perform 32-bit and 64-bit floating-point operations.

[0326] The thread control unit 2401 includes logic for controlling the execution of threads within the execution unit. The thread control unit 2401 may include thread arbitration logic for starting, stopping, and pre-empting the execution of threads within the execution unit 2400. The thread state unit 2402 may be used to store thread states for threads assigned to execute on the execution unit 2400. Storing the thread states within the execution unit 2400 enables rapid pre-emption of those threads when they become blocked or idle. The instruction fetch / prefetch unit 2403 may fetch instructions from an instruction cache of a higher-level execution logic (e.g., the instruction cache 2306 as in Figure 23A ). The instruction fetch / prefetch unit 2403 may also issue prefetch requests for instructions to be loaded into the instruction cache based on an analysis of the currently executing thread. The instruction decoding unit 2404 may be used to decode instructions to be executed by the compute units. In some examples, the instruction decoding unit 2404 may be used as a secondary decoder to decode complex instructions into constituent micro-operations.

[0327] The execution unit 2400 additionally includes a register file 2406 that may be used by hardware threads executing on the execution unit 2400. The registers in the register file 2406 may be partitioned across the logic for multiple concurrent threads within the compute units 2410 that execute the execution unit 2400. The number of logical threads that may be executed by the graphics execution unit 2400 is not limited to the number of hardware threads, and multiple logical threads may be assigned to each hardware thread. Based on the number of supported hardware threads, the size of the register file 2406 may vary across examples. In some examples, register renaming may be used to dynamically allocate registers to hardware threads.

[0328] Figure 25 is a block diagram illustrating a graphics processor instruction format 2500 according to some examples. In one or more examples, the graphics processor execution units support an instruction set having instructions in multiple formats. The solid boxes illustrate components typically included in the execution unit instructions, while the dashed boxes include optional or components included only in a subset of the instructions. In some examples, the described and illustrated instruction format 2500 is a macro-instruction in that they are instructions supplied to the execution unit, as opposed to micro-operations resulting from instruction decoding once the instruction is processed.

[0329] In some examples, a graphics processing unit execution unit may natively support instructions in a 128-bit instruction format 2510. Based on the selected instructions, instruction options, and number of operands, a 64-bit compact instruction format 2530 may be used for some instructions. The native 128-bit instruction format 2510 provides access to all instruction options, while some options and operations are restricted in the 64-bit format 2530. The native instructions available in the 64-bit format 2530 vary by example. In some examples, a set of index values in an index field 2513 is used to partially compress the instruction. The execution unit hardware references a set of compression tables based on the index values and uses the compression table output to reconstruct the native instruction in the 128-bit instruction format 2510. Instructions of other sizes and formats may be used.

[0330] For each format, an instruction opcode 2512 defines the operation to be performed by the execution unit. The execution unit executes each instruction in parallel across multiple data elements of each operand. For example, in response to an add instruction, the execution unit performs a synchronous add operation across each color channel representing a texture element or picture element. By default, the execution unit executes each instruction across all data channels of the operand. In some examples, an instruction control field 2514 enables control of certain execution options such as channel selection (e.g., predication) and data channel order (e.g., swizzle). For instructions in the 128-bit instruction format 2510, an execution size field 2516 limits the number of data channels that will be executed in parallel. In some examples, the execution size field 2516 is not available for the 64-bit compact instruction format 2530.

[0331] Some execution unit instructions have a maximum of three operands, including two source operands src0 2520, src1 2522, and one destination 2518. In some examples, the execution unit supports dual-destination instructions, where one of the destinations is implicit. Data manipulation determines the number of source operands. The last source operand of an instruction may be an immediate (e.g., hard-coded) value passed with the instruction.

[0332] In some examples, the 128-bit instruction format 2510 includes an access / addressing mode field 2526 that, for example, specifies whether to use direct register addressing mode or indirect register addressing mode. When using direct register addressing mode, the register addresses of one or more operands are directly provided by bits in the instruction.

[0333] In some examples, the 128-bit instruction format 2510 includes an access / addressing mode field 2526 that specifies the addressing mode and / or access mode of the instruction. In some examples, the access mode is used to define the data access alignment of the instruction. Some examples support access modes including a 16-byte aligned access mode and a 1-byte aligned access mode, where the byte alignment of the access mode determines the access alignment of the instruction operand. For example, when in a first mode, the instruction may use byte-aligned addressing for source and destination operands, and when in a second mode, the instruction may use 16-byte aligned addressing for all source and destination operands.

[0334] In some examples, the addressing mode portion of the access / addressing mode field 2526 determines whether the instruction is to use direct addressing or indirect addressing. When using the direct register addressing mode, the bits in the instruction directly provide the register addresses of one or more operands. When using the indirect register addressing mode, the register addresses of one or more operands can be calculated based on the address register value and the address immediate field in the instruction.

[0335] In some examples, instructions are grouped based on the opcode 2512-bit field to simplify opcode decoding 2540. For an 8-bit opcode, bits 4, 5, and 6 allow the execution unit to determine the type of the opcode. The exact opcode grouping shown is merely an example. In some examples, the move and logic opcode group 2542 includes data move and logic instructions (e.g., move (mov), compare (cmp)). In some examples, the move and logic group 2542 shares the five most significant bits (MSB), where the move (mov) instruction takes the form of 0000xxxxb and the logic instruction takes the form of 0001xxxxb. The flow control instruction group 2544 (e.g., call, jump) includes instructions in the form of 0010xxxxb (e.g., 0x20). The miscellaneous instruction group 2546 includes a mix of instructions, including synchronization instructions (e.g., wait, send) in the form of 0011xxxxb (e.g., 0x30). The parallel math instruction group 2548 includes per-component arithmetic instructions (e.g., add, multiply (mul)) in the form of 0100xxxxb (e.g., 0x40). The parallel instruction group 2548 performs arithmetic operations in parallel across data channels. The vector math group 2550 includes arithmetic instructions (e.g., dp4) in the form of 0101xxxxb (e.g., 0x50). The vector math group performs arithmetic on vector operands, such as dot product calculations. In some examples, the illustrated opcode decoding 2540 can be used to determine which part of the execution unit will be used to execute the decoded instruction. For example, some instructions can be designated as systolic instructions to be executed by a systolic array. Other instructions (such as ray tracing instructions (not shown)) can be routed to a ray tracing core or ray tracing logic within a slice or partition of the execution logic.

[0336] Graphics pipeline

[0337] Figure 26 is a block diagram of another example of the graphics processor 2600. Figure 26 Elements having the same reference numerals (or names) as elements in any other figure herein can operate or function in any manner similar to those described elsewhere herein, but are not limited thereto.

[0338] In some examples, the graphics processor 2600 includes a geometry pipeline 2620, a media pipeline 2630, a display engine 2640, thread execution logic 2650, and a render output pipeline 2670. In some examples, the graphics processor 2600 is a graphics processor within a multi-core processing system that includes one or more general-purpose processing cores. The graphics processor is controlled by register writes to one or more control registers (not shown) or via commands issued through the ring interconnect 2602 to the graphics processor 2600. In some examples, the ring interconnect 2602 couples the graphics processor 2600 to other processing components, such as other graphics processors or general-purpose processors. Commands from the ring interconnect 2602 are interpreted by a command stream converter 2603, which supplies instructions to the various components of the geometry pipeline 2620 or the media pipeline 2630.

[0339] In some examples, the command stream converter 2603 directs the operation of a vertex fetcher 2605, which reads vertex data from memory and executes vertex processing commands provided by the command stream converter 2603. In some examples, the vertex fetcher 2605 provides vertex data to a vertex shader 2607, which performs coordinate space transformation and lighting operations on each vertex. In some examples, the vertex fetcher 2605 and the vertex shader 2607 execute vertex processing instructions by dispatching execution threads to execution units 2652A - 2652B via a thread dispatcher 2631.

[0340] In some examples, the execution units 2652A - 2652B are an array of vector processors having instruction sets for performing graphics operations and media operations. In some examples, the execution units 2652A - 2652B have attached L1 caches 2651 dedicated to each array or shared between the arrays. The caches can be configured as data caches, instruction caches, or a single cache partitioned to contain data and instructions in different partitions.

[0341] In some examples, the geometry pipeline 2620 includes a tessellation component for performing hardware-accelerated tessellation of 3D objects. In some examples, a programmable hull shader 2611 configures the tessellation operation. A programmable domain shader 2617 provides backend evaluation of the tessellation output. The tessellator 2613 operates under the direction of the hull shader 2611 and includes specialized logic for generating a detailed set of geometric objects based on a coarse geometric model that is provided as input to the geometry pipeline 2620. In some examples, if tessellation is not used, the tessellation components (e.g., the hull shader 2611, the tessellator 2613, and the domain shader 2617) can be bypassed.

[0342] In some examples, the complete geometric object can be processed by the geometry shader 2619 via one or more threads dispatched to execution units 2652A - 2652B, or can proceed directly to the clipper 2629. In some examples, the geometry shader operates on the entire geometric object, rather than on vertices or patches of vertices as in previous stages of the graphics pipeline. If tessellation is disabled, the geometry shader 2619 receives input from the vertex shader 2607. In some examples, the geometry shader 2619 can be programmed by a geometry shader program to perform geometric tessellation in the case where the tessellation unit is disabled.

[0343] Before rasterization, the clipper 2629 processes vertex data. The clipper 2629 is a fixed - function clipper or a programmable clipper with clipping and geometry shader functionality. In some examples, the rasterizer and depth test component 2673 in the render output pipeline 2670 dispatches pixel shaders to convert the geometric object into a per - pixel representation. In some examples, the pixel shader logic is included in the thread execution logic 2650. In some examples, an application can bypass the rasterizer and depth test component 2673 and access the un - rasterized vertex data via the outflow unit 2623.

[0344] The graphics processor 2600 has an interconnect bus, an interconnect fabric, or some other interconnect mechanism that allows data and messages to be passed between the major components of the processor. In some examples, the execution units 2652A - 2652B and associated logic units (e.g., L1 cache 2651, sampler 2654, texture cache 2658, etc.) are interconnected via data ports 2656 to perform memory accesses and communicate with the render output pipeline components of the processor. In some examples, the sampler 2654, caches 2651, 2658, and execution units 2652A - 2652B each have separate memory access paths. In some examples, the texture cache 2658 can also be configured as a sampler cache.

[0345] In some examples, the rendering output pipeline 2670 includes a rasterizer and depth test component 2673 that converts vertex-based objects to their associated pixel-based representations. In some examples, the rasterizer logic includes a windower / masker unit for performing fixed functions available in some examples. The pixel operations component 2677 performs pixel-based operations on the data, but in some instances, pixel operations associated with 2D operations (e.g., bit blit image transfer with blending) are performed by the 2D engine 2641 or, when displaying, by the display controller 2643 using an overlay display plane instead. In some examples, a shared L3 cache 2675 is available to all graphics components, allowing data to be shared without using the main system memory.

[0346] In some examples, the graphics processor media pipeline 2630 includes a media engine 2637 and a video front end 2634. In some examples, the video front end 2634 receives pipeline commands from the command stream converter 2603. In some examples, the media pipeline 2630 includes a separate command stream converter. In some examples, the video front end 2634 processes the media commands before sending them to the media engine 2637. In some examples, the media engine 2637 includes a thread generation function for generating threads to be dispatched to the thread execution logic 2650 via the thread dispatcher 2631.

[0347] In some examples, the graphics processor 2600 includes a display engine 2640. In some examples, the display engine 2640 is external to the processor 2600 and is coupled to the graphics processor via a ring interconnect 2602, or some other interconnect bus or fabric. In some examples, the display engine 2640 includes a 2D engine 2641 and a display controller 2643. In some examples, the display engine 2640 contains dedicated logic capable of operating independently of the 3D pipeline. In some examples, the display controller 2643 is coupled to a display device (not shown), which may be a system-integrated display device such as in a laptop computer or an external display device attached via a display device connector.

[0348] In some examples, the geometry pipeline 2620 and the media pipeline 2630 may be configured to perform operations based on multiple graphics and media programming interfaces and are not dedicated to any one application programming interface (API). In some examples, driver software for the graphics processor translates API calls that are specific to a particular graphics or media library into commands that are processed by the graphics processor. In some examples, support is provided for all of the Open Graphics Library (OpenGL), Open Computing Language (OpenCL), and / or Vulkan graphics and compute APIs from the Khronos Group. In some examples, support may also be provided for the Direct3D library from Microsoft Corporation. In some examples, combinations of these libraries may be supported. Support may also be provided for the Open Source Computer Vision Library (OpenCV). Future APIs with compatible 3D pipelines will also be supported if a mapping can be made from the pipelines of the future API to the pipelines of the graphics processor.

[0349] Graphics Pipeline Programming

[0350] Figure 27A is a block diagram illustrating a graphics processor command format 2700 according to some examples. Figure 27B is a block diagram illustrating a graphics processor command sequence 2710 according to an example. Figure 27A The solid box diagrams in generally represent components that are typically included in a graphics command, while the dashed lines include optional components or components that are only included in a subset of graphics commands. Figure 27A The exemplary graphics processor command format 2700 of includes data fields for a client 2702 that identifies the command, a command operation code (opcode) 2704, and data 2706. A sub-opcode 2705 and a command size 2708 are also included in some commands.

[0351] In some examples, client 2702 designates a client unit of a graphics device that processes command data. In some examples, a graphics processor command parser examines the client field of each command to condition further processing of the command and routes the command data to an appropriate client unit. In some examples, the graphics processor client units include a memory interface unit, a render unit, a 2D unit, a 3D unit, and a media unit. Each client unit has a corresponding processing pipeline for processing commands. Once a command is received by a client unit, the client unit reads the opcode 2704 and the sub-opcode 2705 (if present) to determine the operation to be performed. The client unit uses the information in the data field 2706 to execute the command. For some commands, an explicit command size 2708 is expected to specify the size of the command. In some examples, the command parser automatically determines the size of at least some of the commands in the command based on the command opcode. In some examples, commands are aligned via multiples of a doubleword. Other command formats may be used.

[0352] Figure 27B The flow diagram in FIG. 2710 illustrates an exemplary graphics processor command sequence. In some examples, software or firmware of a data processing system characterized by an example graphics processor uses some version of the illustrated command sequence to establish, execute, and terminate a set of graphics operations. The sample command sequence is shown and described for illustrative purposes only and the examples are not limited to these particular commands or this command sequence. Additionally, commands may be issued in the command sequence as a batch of commands such that the graphics processor will process the command sequence in at least a partially concurrent manner.

[0353] In some examples, the graphics processor command sequence 2710 can begin with a pipeline flush clear command 2712 to cause any active graphics pipelines to complete the current outstanding commands for the pipeline. In some examples, the 3D pipeline 2722 and the media pipeline 2724 may not operate concurrently. Performing the pipeline flush clear causes the active graphics pipelines to complete any outstanding commands. In response to the pipeline flush clear, the command parser for the graphics processor will pause command processing until the active drawing engines have completed the outstanding operations and the associated read caches have been invalidated. Optionally, any data marked "dirty" in the render cache may be flushed to memory. In some examples, the pipeline flush clear command 2712 can be used for pipeline synchronization or can be used before putting the graphics processor into a low power state.

[0354] In some examples, the pipeline select command 2713 is used when a command sequence requires the graphics processor to explicitly switch between pipelines. In some examples, the pipeline select command 2713 is only required once in the execution context before issuing pipeline commands, unless the context is issuing commands for both pipelines. In some examples, the pipeline dump clear command 2712 is required immediately before a pipeline switch via the pipeline select command 2713.

[0355] In some examples, the pipeline control command 2714 configures the graphics pipeline for operation and is used to program the 3D pipeline 2722 and the media pipeline 2724. In some examples, the pipeline control command 2714 configures the pipeline state for the active pipeline. In some examples, the pipeline control command 2714 is used for pipeline synchronization and for clearing data from one or more cache memories within the active pipeline before processing a batch of commands.

[0356] In some examples, the return buffer status command 2716 is used to configure the set of return buffers for the corresponding pipeline for writing data. Some pipeline operations require the allocation, selection, or configuration of one or more return buffers into which intermediate data is written during processing. In some examples, the graphics processor also uses one or more return buffers to store output data and perform cross-thread communication. In some examples, the return buffer status 2716 includes selecting the size and number of return buffers to be used for a set of pipeline operations.

[0357] The remaining commands in the command sequence differ based on the active pipeline for the operation. Based on the pipeline determination 2720, the command sequence is customized for the 3D pipeline 2722 starting with the 3D pipeline state 2730, or for the media pipeline 2724 starting at the media pipeline state 2740.

[0358] Commands for configuring the 3D pipeline state 2730 include 3D state setting commands for vertex buffer state, vertex element state, constant color state, depth buffer state, and other state variables that will be configured before processing 3D primitive commands. The values of these commands are determined at least in part based on the specific 3D API in use. In some examples, the 3D pipeline state 2730 commands can also selectively disable or bypass certain pipeline elements if they will not be used.

[0359] In some examples, the 3D primitive 2732 commands are used to submit 3D primitives to be processed by the 3D pipeline. The commands and associated parameters passed to the graphics processor via the 3D primitive 2732 commands are forwarded to the vertex fetch function in the graphics pipeline. The vertex fetch function uses the 3D primitive 2732 command data to generate vertex data structures. The vertex data structures are stored in one or more return buffers. In some examples, the 3D primitive 2732 commands are used to perform vertex operations on the 3D primitives via a vertex shader. To process the vertex shader, the 3D pipeline 2722 dispatches shader execution threads to the graphics processor execution units.

[0360] In some examples, the 3D pipeline 2722 is triggered via the execution of 2734 commands or events. In some examples, a register write triggers command execution. In some examples, execution is triggered via a "go" or "kick" command in a command sequence. In some examples, command execution uses pipeline synchronization commands to trigger a command sequence flush through the graphics pipeline. The 3D pipeline will perform geometric processing on the 3D primitives. Once the operations are complete, the resulting geometric objects are rasterized, and the pixel engine colors the resulting pixels. For those operations, additional commands may also be included to control pixel coloring and pixel backend operations.

[0361] In some examples, when performing media operations, the graphics processor command sequence 2710 follows the media pipeline 2724 path. Generally, the specific uses and ways of programming the media pipeline 2724 depend on the media or compute operations to be performed. During media decoding, specific media decoding operations may be migrated to the media pipeline. In some examples, the media pipeline may also be bypassed, and media decoding may be performed entirely or partially using the resources provided by one or more general-purpose processing cores. In some examples, the media pipeline also includes elements for general-purpose graphics processing unit (GPGPU) operations, where the graphics processor is used to perform SIMD vector operations using compute shader programs that are not explicitly related to the rendering of graphics primitives.

[0362] In some examples, the media pipeline 2724 is configured in a manner similar to the 3D pipeline 2722. The set of commands for configuring the media pipeline state 2740 is dispatched or placed into the command queue before the media object commands 2742. In some examples, the commands for the media pipeline state 2740 may include data for configuring the media pipeline elements that will be used to process the media object. This includes data for configuring video decoding and video encoding logic within the media pipeline, such as the encoding or decoding format. In some examples, the commands for the media pipeline state 2740 also support the use of one or more pointers to "indirect" state elements that point to a batch of state settings.

[0363] In some examples, the media object command 2742 supplies a pointer to a media object for processing by a media pipeline. The media object includes a memory buffer that contains video data to be processed. In some examples, all media pipeline states must be valid before the media object command 2742 is issued. Once the pipeline state is configured and the media object command 2742 is queued, the media pipeline 2724 is triggered via an execute command 2744 or an equivalent execution event (e.g., a register write). The output from the media pipeline 2724 can then be post-processed by operations provided by the 3D pipeline 2722 or the media pipeline 2724. In some examples, GPGPU operations are configured and executed in a manner similar to media operations.

[0364] Program code can be applied to input information to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microprocessor, or any combination thereof.

[0365] The program code can be implemented in a high-level procedural programming language or an object-oriented programming language in order to communicate with the processing system. If desired, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described herein are not limited to the scope of any particular programming language. In any case, the language can be a compiled language or an interpreted language.

[0366] Examples of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of such implementations. An example can be implemented as a computer program or program code executing on a programmable system that includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0367] Such machine-readable storage media can include, but are not limited to, non-transitory tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as: hard disks; any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), rewritable compact disks (CD-RW), and magneto-optical disks); semiconductor devices such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM); phase change memory (PCM); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.

[0368] Accordingly, examples also include non-transitory tangible machine-readable media that contain instructions or contain design data, such as a Hardware Description Language (HDL), which defines the structures, circuits, devices, processors, and / or system features described herein. Such examples may also be referred to as program products.

[0369] Emulation (including binary translation, code morphing, etc.).

[0370] In some cases, an instruction converter can be used to convert instructions from a source instruction set architecture to a target instruction set architecture. For example, the instruction converter can translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions into one or more other instructions to be processed by the core. The instruction converter can be implemented in software, hardware, firmware, or a combination thereof. The instruction converter can be on the processor, outside the processor, or partially on the processor and partially outside the processor.

[0371] Figure 28FIG. 0 is a block diagram illustrating the use of a software instruction converter according to an example, the software instruction converter being operative to convert binary instructions in a source ISA into binary instructions in a target ISA. In the illustrated example, the instruction converter is a software instruction converter, but alternatively, the instruction converter may be implemented in software, firmware, hardware, or various combinations thereof. Figure 28 FIG. Figure 28 shows that a program in a high-level language 2802 can be compiled using a first ISA compiler 2804 to generate first ISA binary code 2806 that can be natively executed by a processor 2816 having at least one first ISA core. The processor 2816 having at least one first ISA core represents any such processor that can perform substantially the same functions as an Intel processor having at least one first ISA core by compatibly executing or otherwise processing (1) a substantial portion of the first ISA or (2) a target code version of an application or other software targeted to run on an Intel processor having at least one first ISA core, so as to achieve substantially the same results as an Intel processor having at least one first ISA core. The first ISA compiler 2804 represents a compiler operable to generate first ISA binary code 2806 (e.g., target code) that can be executed on the processor 2816 having at least one first ISA core with or without additional linking processing. Similarly, FIG. shows that a program in a high-level language 2802 can be compiled using an alternative ISA compiler 2808 to generate alternative ISA binary code 2810 that can be natively executed by a processor 2814 that does not have a first ISA core. The converted code need not be the same as the alternative ISA binary code 2810; however, the converted code will perform the general operations and be composed of instructions from the alternative ISA. Thus, the instruction converter 2812 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have a first ISA processor or core to execute the first ISA binary code 2806 through emulation, simulation, or any other process. Figure 28 FIG. Figure 28 shows that a program in a high-level language 2802 can be compiled using an alternative ISA compiler 2808 to generate alternative ISA binary code 2810 that can be natively executed by a processor 2814 that does not have a first ISA core. The converted code need not be the same as the alternative ISA binary code 2810; however, the converted code will perform the general operations and be composed of instructions from the alternative ISA. Thus, the instruction converter 2812 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device that does not have a first ISA processor or core to execute the first ISA binary code 2806 through emulation, simulation, or any other process.

[0372] IP Core Implementations

[0373] One or more aspects of at least some examples may be implemented by representative code stored on a machine-readable medium that represents and / or defines logic within an integrated circuit, such as a processor. For example, the machine-readable medium may include instructions representing various logic within the processor. When read by a machine, the instructions may cause the machine to fabricate logic for performing the techniques described herein. Such representations (referred to as “IP cores”) are reusable units of logic for an integrated circuit and may be stored on a tangible machine-readable medium as a hardware model describing the organization of the integrated circuit. The hardware model may be supplied to various customers or fabrication facilities that load the hardware model on fabrication machines for manufacturing the integrated circuit. The integrated circuit may be fabricated such that the circuit performs operations described in association with any of the examples described herein.

[0374] Figure 29 FIG. 2900 is a block diagram illustrating an IP core development system 2900 that may be used to fabricate an integrated circuit to perform operations. The IP core development system 2900 may be used to generate a modular, reusable design that may be incorporated into a larger design or used to build an entire integrated circuit (e.g., a system-on-a-chip integrated circuit). A design facility 2930 is capable of generating a software simulation 2910 of an IP core design in a high-level programming language (e.g., C / C++). The software simulation 2910 may be used to design, test, and verify the behavior of the IP core using a simulation model 2912. The simulation model 2912 may include functional simulation, behavioral simulation, and / or timing simulation. A register transfer level (RTL) design 2915 may then be created or synthesized from the simulation model 2915. The RTL design 2915 is an abstraction of the behavior of an integrated circuit that models the flow of digital signals between hardware registers, including associated logic performed using the modeled digital signals. In addition to the RTL design 2915, lower-level designs at the logic level or transistor level may also be created, designed, or synthesized. Thus, the specific details of the initial design and simulation may vary.

[0375] The RTL design 2915 or an equivalent can be further synthesized by a design facility into a hardware model 2920, which can be in a hardware description language (HDL) or some other representation of physical design data. The HDL can be further simulated or tested to verify the IP core design. A non-volatile memory 2940 (e.g., a hard disk, flash memory, or any non-volatile storage medium) can be used to store the IP core design for delivery to a third-party manufacturing facility 2965. Alternatively, the IP core design can be transmitted via a wired connection 2950 or a wireless connection 2960 (e.g., via the Internet). The manufacturing facility 2965 can then manufacture an integrated circuit that is at least partially based on the IP core design. The manufactured integrated circuit can be configured to perform operations in accordance with at least some of the examples described herein. Apparatus and method for attack-resistant encryption and decryption

[0376] Fault-injection attack (FIA) is becoming an increasingly common physical attack on the keys used by AES hardware accelerators. FIA reduces the key search space by using differential fault analysis (DFA) on ciphertext corrupted by laser / voltage / clock pulses. FIA on a conventional AES-256 engine typically targets the last three rounds of iterations to cause bit flips at intermediate circuit nodes using a guided laser pulse and voltage / clock glitches. Due to the avalanche effect of AES encryption, errors injected during the upstream AES rounds (0 - 10) cascade quickly, making first-order DFA impractical for realistic attacks. However, errors injected during the computations of rounds 13, 12, and 11 propagate down the AES logic cone, probabilistically corrupting 1, 4, and 16 ciphertext bytes respectively. Current AES attack countermeasures for error injection incur significant silicon and power overhead, limiting the adoption of such countermeasures.

[0377] Some existing solutions employ arithmetic and parity checkers in the AES data path to check for errors injected into the non-linear and linear parts of the AES, respectively. The arithmetic checker uses an inverse checker based on composite field arithmetic to detect errors injected into the non-linear inverse block. The parity checker detects errors such as those to the linear part, such as key addition, mix columns, and affine transformation. However, this solution causes a significant (e.g., 40%) increase in area and power. It is also a major challenge for the design team to integrate its memory encryption engine (e.g., multi-key total memory encryption (MKTME) engine) with significant pressure on the area and power budget. In addition, the current solution is not scalable for future memory technologies where the memory bandwidth requirements are increasing.

[0378] The last three rounds of AES operations are the target of the FIA attack to induce bit flips, thus generating exploitable error patterns in the output ciphertext. To provide error injection resistance, embodiments of the present invention use the round data path to perform redundant time-interleaved calculations on any one, any two, or all three of the last three rounds, and compare the outputs of the "regular" and "redundant" rounds. Any mismatch between the redundant and regular round outputs indicates the presence of an error. Specifically, due to the avalanche effect of AES encryption, errors injected during the upstream AES rounds (0-10) cascade quickly, making first-order DFA impractical for realistic attacks. However, errors injected during the calculations of rounds 13, 12, and 11 propagate down the AES logic cone, probabilistically corrupting 1, 4, and 16 ciphertext bytes, respectively. Differential cryptanalysis of the error ciphertext obtained by laser FIA on an unprotected AES engine can significantly reduce the key search space to 1-40b, thus motivating FIA countermeasures to protect the last 3 vulnerable AES rounds.

[0379] Embodiments of the present invention include an AES cryptographic circuit system that uses different isomorphic GF(2 4 ) 2 composite fields, while implementing the AES data path to maximize the spatial and temporal differences between regular and redundant calculations. The differences in AES arithmetic coupled with a random byte stream generate different error profiles for errors injected at the same node. Redundant AES round calculations and a reconfigurable byte data stream enable real-time detection of corrupted ciphertext, with an error coverage of up to 99.98%, while limiting the area overhead to 12%. Compared with existing solutions, this reduces the area / power overhead to 1 / 3.3. A. Example processor implementation

[0380] Figure 30FIG. illustrates an example processor 3001 on which embodiments of the present invention may be implemented. Each of the plurality of cores 102A-102B includes instruction fetch circuitry, instruction decoding circuitry, execution circuitry, and various registers 110 described above with respect to Figure 1 the core 102 illustrated in. The plurality of cores 102A-102B are coupled to a shared cache 3012 (e.g., an L3 cache) and to memory 3020 via a memory controller 3016.

[0381] In some examples, memory access (e.g., store or load) requests to memory 3020 generated by cores 102A-102B (e.g., by an address generation unit (AGU) of the execution circuitry) may be serviced by caches within cores 102A-102B and / or the shared cache 3012. Additionally or alternatively (e.g., for cache misses), the memory access requests may be serviced by memory 3020. Memory access requests generated by cores 102A-102B may be load or store operations. A load operation reads data from memory 3020 into a cache of the processor (e.g., cache 3012), and a store operation writes data to memory 3020. In some examples, the memory controller circuitry 3016 includes a direct memory access engine 3017, e.g., for performing accesses to memory 3020. The memory may be volatile memory (e.g., DRAM), non-volatile or persistent memory (e.g., non-volatile DIMM or non-volatile DRAM), and / or auxiliary (e.g., external) memory (e.g., not directly accessible by the processor). In some examples, the memory controller circuitry 3017 is used to perform compression and / or decompression of data, e.g., where multiple bits and / or bytes that are repeated in a data line are removed to allow compression based on the repetition (e.g., repetition-based compression / decompression). Various other compression techniques may also be used.

[0382] In some implementations, the cryptographic circuitry 3014 may receive a memory access request (e.g., a load or store operation) from one or more of its cores 102A-102B, the memory access request including an address and data to be encrypted (e.g., plaintext). The memory access request may be associated with a corresponding key (e.g., a key assigned to the hardware / software entity responsible for the request). For a store operation, the cryptographic circuitry 3014 may encrypt the data using the key to generate ciphertext (encrypted data), which is then stored in memory 3020. For a load operation, the cryptographic circuitry 3014 may read the requested ciphertext from a specified address in memory 3020 and decrypt the ciphertext using the key (or a different key).

[0383] In some embodiments, multiple cores 102A-102B use a cryptographic circuit system 3014 to perform cryptographic operations as described herein. Although illustrated in Figure 30 as being adjacent to the memory controller 3016 (e.g., coupled between the memory controller 3016 and the shared cache 3012 and / or between levels of the cache hierarchy), the cryptographic circuit system 3014 may be integrated into the memory controller 3016 or distributed across locations within the memory / cache subsystem.

[0384] The mode register 3114 may be programmed with mode control bits to configure the corresponding cryptographic circuit system 3014 in a particular operating mode. In some embodiments, programming of the mode register 3114 is restricted to trusted software components. Thus, an application or virtual machine must request a configuration change via the virtual machine monitor and / or via firmware executing on the secure processor. In some embodiments, the mode register 3114 may configure the FIA-resistant AES circuit system 3051 to operate in accordance with the FIA-resistant AES techniques described herein (e.g., using different isomorphic GF(2 4 ) 2 composite fields).

[0385] As an overview, the Advanced Encryption Standard (AES) is based on the Rijndael block cipher, where multiple transformations are performed using a secret key (e.g., the MKTME key) to transform intelligible data (plaintext) into an encrypted format (ciphertext). The transformations may include key expansion, substitution box (S-box), ShiftRows, MixColumns, and AddRoundKey. The key size indicates the number of transformation rounds for encrypting the plaintext. Each round consists of a number of processing steps, including one that depends on the encryption key. Multiple inverse rounds are applied to convert the ciphertext back to plaintext using the encryption key.

[0386] Cryptographic keys of lengths 128, 192, and 256 bits may be used to generate 128-bit data blocks. AES implementations transform plaintext to ciphertext or ciphertext to plaintext in 10, 12, or 14 consecutive rounds, depending on the length of the key.

[0387] In the key expansion operation, the key schedule transforms a key of size 128, 192, or 256 bits into 10, 12, or 14 round keys of 128 bits respectively. The round keys can be values derived from a cryptographic key, which is used to process plaintext in rounds of 128-bit blocks. For example, the round keys can be added to the encryption state (e.g., in an AES state register) using an Exclusive-OR (XOR) operation.

[0388] The S-box includes a non-linear byte substitution table for processing the state. For example, in a 128-bit (16-byte) input of a round, each byte is replaced by another byte according to the byte substitution table. The S-box is computationally intensive and is the most complex transformation in AES. In the shift rows transformation, the last three rows of the state are cyclically shifted by different offsets (e.g., row 0 is shifted by zero bytes, row 1 is shifted by one byte, row 2 is shifted by two bytes, and row 3 is shifted by three bytes). In the mix columns operation, the columns of the state are mixed independently of other columns to produce new columns. The mix columns function takes four bytes as input and outputs four bytes, where each input byte affects all four output bytes. For example, the state can be a 16-byte state divided into four columns. For example, the state input of the mix columns operation can be the output of the S-box operation shifted according to the shift rows transformation.

[0389] In some examples, additional processor components, such as a network interface circuitry (NIC) 3032, can rely on the cryptographic circuitry 3014 to encrypt and decrypt data in the memory 3020. Alternatively or additionally, these components can include their own integrated cryptographic circuitry for performing at least some of the FIA-resistant AES operations described herein. B. Embodiments of FIA-Resistant EAS Circuitry and Methods

[0390] The FIA-resistant AES circuitry 3051 can be implemented in hardware (e.g., as an ASIC) or using a combination of hardware and firmware / software. For example, some of the functional components described below can be implemented in circuitry, while other functional components can be implemented by executing firmware / software (e.g., executed by a controller integrated into the AES circuitry 3051).

[0391] In one embodiment, the FIA-resistant AES circuitry 3051 is configured to perform 14 iterative rounds of a single-cycle data path using 16-byte slices implemented as having a pair of isomorphic composite fields GF(2 4 ) 2 represented. Refer to Figure 31, in one embodiment, the round data path is divided into two groups, each group containing eight S-box modules and two mix column modules. Specifically, eight Sbox modules and two mix column modules are implemented with a first isomorphic composite field representation GF b (2 4 ) 2 , operate on the input bytes [15:12, 7:4] (as indicated by the patterned shading), and another eight Sbox modules and two mix column modules are implemented with another isomorphic composite field representation GF b (2 4 ) 2 , operate on the input bytes [11:8, 3:0] (as indicated by the non-shaded).

[0392] The AES circuit system 3051 resistant to FIA also performs up to three redundant iterations time-interleaved with the last three AES rounds. It should be noted that these specific details are not necessary to comply with the fundamental principles of the present invention. For example, for any non-zero value N, N redundant iterations can be performed for N AES rounds.

[0393] In some embodiments, rounds 0-10 operate in such a mode that at the end of each clock cycle, the intermediate output bytes are written into the 128b state register 3110. Additionally, in one embodiment, the output of round 10 is cross-mapped through the cross-mapping circuit system 3150 (GF a →GF b , GF b →GF a ), and word mixing is performed through the word mixing circuit system 3152 (127:96 ←→ 95:64, 63:32 ←→ 31:0), and is written into the redundant state register 3111.

[0394] During rounds 11-13, the data stream alternates between normal and redundant modes by alternately processing the input bytes from the state register 3110 or the redundant state register 3111 and writing the output back to the corresponding source. The redundant rounds (11r-13r) repeat the operations calculated by their regular counterparts (rounds 11-13), but are performed on the isomorphic cross-mapped, mixed version of the regular round inputs. Therefore, errors injected at the same circuit system nodes during regular and redundant rounds result in very different error profiles at the round output registers.

[0395] In some embodiments, before the comparison circuit system 3128 compares the redundant state output with the corresponding result from the AES state register 3110, the inverse cross-mapping circuit system 3160 and the mixing circuit system 3162 convert the redundant state output back to the original GF(2 4) 2 Indicates that, as indicated, when a mismatch is detected, the comparison circuitry 3128 generates an "error" signal indicating that there is a corrupted bit in the output ciphertext.

[0396] Figure 32 A cyclic diagram illustrating an embodiment of the AES data path. The redundant round operations (round 11r 3211B, round 12r 3212B, and round 13r 3213B) are interleaved with their regular corresponding round operations 3211A, 3212A, and 3213A, respectively, and errors are checked in subsequent cycles, as indicated at 3251, 3252, and 3253, respectively. The anti-FIA engine 3051 can be configured (e.g., via the mode register 3114) to check any one / any two / all three of the final rounds, resulting in a total latency of 15 / 16 / 17 cycles, respectively, to generate the AES-256 ciphertext 3190.

[0397] The choice of the composite field polynomial significantly affects the circuit implementation of the Sbox and the MixColumn data path groups 3101A - 3101B. Figure 33 Illustrates an example logic gate that implements a selection from 2880 isomorphic GF(2 4 ) 2 Indicates the basis field polynomials (x a , GF b ) selected for (GF 4 +x 3 +1, x 4 +x 3 +x 2 +x + 1) and the extension field polynomials (x 2 +x + 8, x 2 +3x + A) to maximize the separation of the error propagation characteristics for the same injected errors.

[0398] FIA simulations show that there is no separation (Hamming distance = 0) between the correct and incorrect ciphertexts from redundant computations on the homogeneous GFa data path, while when the same nodes are injected with errors, the use of GF a and GF b byte slices produces uniformly dissimilar ciphertexts (mean Hamming distance = 64).

[0399] In some embodiments, side-channel resistance is incorporated into the FIA-resistant AES circuitry 3051 by masking internal circuit nodes during normal and redundant round iterations to disrupt the correlation between the power / EM signature and the key. For example, random masks can be generated by a on-die pseudo-random number generator (RNG) and added to the input plaintext prior to the add-round-key operation in round 0. The presence of mask components in the downstream circuitry nodes disrupts the fidelity of the attacker's statistical power model, thus preventing accurate correlation (e.g., via differential power analysis).

[0400] Figure 34 Illustrative masked GF(2 4 ) inverse circuitry 3401 example. Mask compensation circuitry 3405 uses GF a and GF b data paths combined with word mixing during time-interleaved redundant rounds to provide protection against FIAs using randomly generated mask components (m2, m4, and m8). In some embodiments, the randomly generated mask components are also cross-mapped and word mixed between GF a and GF b during redundant computations. Depending on the computation mode, mask compensation circuit 3405 operates on 128b inputs from either a mask register or a redundant mask register. Similar to the AES round data path, the choice of GF a and GF b polynomials affects the circuit implementation of the mask compensation circuit, generating different error profiles for injected errors at masked circuit nodes.

[0401] Illustrative GF a (2 4 ) 2 and GF b (2 4 ) 2 composite field representations for different polynomial components of the data path are shown in two rows 3421 and 3422. Injected error 3430 is shown as a bit flip on mask bit m0[3], which propagates eight errors 3431 (row 3421) across 2-to-the-power-of (^2), 4-to-the-power-of (^4), and 8-to-the-power-of (^8) outputs in the GF a implementation, while the GF b data path (row 3422) propagates six downstream errors 3431, triggering a mismatch when the normal and redundant round outputs are compared in subsequent cycles.

[0402] Refer to Figure 35, when testing the working embodiment of the AES circuit system resistant to FIA with FIA attacks, the results show that the mean total error coverage from under-voltage attacks is 99.98%, indicating that the error injection attack resistance is increased to 5000 times compared with the existing AES circuit system implementation.

[0403] In Figure 36 FIG. illustrates a method according to an embodiment of the present invention. The method may be executed on various architectures described herein, but is not limited to any particular processor or system architecture.

[0404] At 3601, multiple encryption or decryption rounds are performed to generate a first output, and at 3602, redundant encryption / decryption rounds corresponding to one or more of the encryption / decryption rounds are performed to generate a second output. At 3603, the first output is compared with the second output. If a mismatch is detected at 3604, an error is generated at 3605. If no mismatch is detected, at 3606, the encrypted data is stored in a memory (for encryption) or the decrypted data is loaded from the memory (for decryption).

[0405] Embodiments of the present invention may include the steps described above. These steps may be embodied as machine-executable instructions that may be used to cause a general-purpose or special-purpose processor to execute these steps. Alternatively, these steps may be performed by specific hardware components that include hardwired logic for performing these steps, or by any combination of programmed computer components and custom hardware components.

[0406] Examples

[0407] The following are example implementations of different embodiments of the present invention.

[0408] Example 1. A processor, comprising: execution circuitry for executing instructions and generating memory access requests, the memory access requests including load requests for reading data from a memory and store requests for storing data in the memory; and cryptographic circuitry for performing multiple rounds of encryption or decryption to encrypt or decrypt data respectively, the cryptographic circuitry for performing one or more redundant rounds for corresponding ones of the multiple rounds, the one or more redundant rounds including spatial or temporal differences relative to the corresponding one or more rounds; the cryptographic circuitry for generating an error when a mismatch is detected between the output of the redundant round and the output of the corresponding round.

[0409] Example 2. The processor of Example 1, wherein one or more of the redundant rounds are time-interleaved with corresponding ones of the rounds.

[0410] Example 3. The processor according to Example 1 or 2, wherein the cryptographic circuitry is configured to modify an input to a corresponding one or more rounds to generate a corresponding input to one or more redundant rounds.

[0411] Example 4. The processor according to any one of Examples 1-3, wherein the input to a corresponding one or more rounds is to be isomorphically cross-mapped and mixed to generate a corresponding input to one or more redundant rounds.

[0412] Example 5. The processor according to any one of claims 1-4, wherein the cryptographic circuitry comprises: a first status register for storing an encryption or decryption status for a plurality of rounds; and a second status register for storing a redundant encryption or decryption status for one or more redundant rounds.

[0413] Example 6. The processor according to any one of Examples 1-5, wherein the cryptographic circuitry further comprises: circuitry for converting the redundant encryption or decryption status to a converted status representation corresponding to the encryption or decryption status.

[0414] Example 7. The processor according to any one of Examples 1-6, wherein the cryptographic circuitry further comprises: comparison circuitry for comparing the converted status representation with the encryption or decryption status to detect a mismatch.

[0415] Example 8. The processor according to any one of claims 1-7, wherein the cryptographic circuitry comprises: a first N Sbox circuits and a first M mix column circuits, the first N Sbox circuits and the first M mix column circuits being configured to be implemented with a first isomorphic composite field representation that operates on a first set of input bytes in each of a plurality of rounds; and a second N Sbox circuits and a second M mix column circuits, the second N Sbox circuits and the second M mix column circuits being implemented with a second isomorphic composite field representation that operates on a second set of input bytes in each of a plurality of rounds.

[0416] Example 9. A method comprising: performing cryptographic rounds to generate a first output; performing redundant cryptographic rounds corresponding to one or more of the cryptographic rounds to generate corresponding second outputs; comparing each of the second outputs with the corresponding first output to detect any mismatches; storing the encrypted data or loading the decrypted data corresponding to the cryptographic rounds in the absence of detected mismatches; and generating an error or error condition in the presence of detected mismatches.

[0417] Example 10. The method according to Example 9, wherein one or more of the redundant cryptographic rounds are time-interleaved with corresponding ones of the cryptographic rounds.

[0418] Example 11. The method according to Example 9 or 10, further comprising: modifying an input to one or more cryptographic rounds to generate a corresponding input to one or more redundant cryptographic rounds.

[0419] Example 12. The method according to any one of Examples 9-11, wherein the input to the corresponding one or more cryptographic rounds is to be isomorphically cross-mapped and mixed to generate a corresponding input to one or more redundant cryptographic rounds.

[0420] Example 13. The method according to any one of Examples 9-13, further comprising: storing an encryption or decryption state of a plurality of cryptographic rounds in a first state register; and storing a redundant encryption or decryption state of one or more redundant cryptographic rounds in a second state register.

[0421] Example 14. The method according to any one of Examples 9-13, further comprising: converting the redundant encryption or decryption state to a converted state representation corresponding to the encryption or decryption state.

[0422] Example 15. The method according to any one of Examples 9-14, further comprising: comparing the converted state representation with the encryption or decryption state to detect a mismatch.

[0423] Example 16. A machine-readable medium having program code stored thereon, which when executed by a machine causes the machine to perform operations, the operations including: performing cryptographic rounds to generate a first output; performing redundant cryptographic rounds corresponding to one or more of the cryptographic rounds to generate corresponding second outputs; comparing each of the second outputs with the corresponding first output to detect any mismatches; storing encrypted data or loading decrypted data corresponding to the cryptographic rounds in the absence of detected mismatches; and generating an error or error condition in the presence of detected mismatches.

[0424] Example 17. The machine-readable medium according to Example 16, wherein one or more of the redundant cryptographic rounds are to be time-interleaved with corresponding ones of the cryptographic rounds.

[0425] Example 18. The machine-readable medium according to Example 16 or 17, further comprising: modifying an input to one or more cryptographic rounds to generate a corresponding input to one or more redundant cryptographic rounds.

[0426] Example 19. The machine-readable medium according to any one of Examples 16-18, wherein the input to the corresponding one or more cryptographic rounds is to be isomorphically cross-mapped and mixed to generate a corresponding input to one or more redundant cryptographic rounds.

[0427] Example 20. The machine-readable medium according to any one of Examples 16-19, further comprising program code for causing the following operations: storing an encryption or decryption state of a plurality of cipher rounds in a first status register; and storing a redundant encryption or decryption state of one or more redundant cipher rounds in a second status register.

[0428] Example 21. The machine-readable medium according to any one of Examples 16-20, further comprising program code for causing the following operations: converting a redundant encryption or decryption state into a converted state representation corresponding to the encryption or decryption state.

[0429] Example 22. The machine-readable medium according to any one of Examples 16-21, further comprising program code for causing the following operations: comparing the converted state representation with the encryption or decryption state to detect a mismatch.

[0430] As described herein, an instruction may refer to a specific configuration of hardware such as an application specific integrated circuit (ASIC) that is configured to perform certain operations or has a predefined function or software instructions stored in a memory embodied in a non-transitory computer-readable medium. Thus, the techniques shown in the figures may be implemented using code and data stored on and executed on one or more electronic devices (e.g., a terminal station, a network element, etc.). Such electronic devices use computer machine-readable media (internally and / or via a network with other electronic devices) to store and transmit the code and data, the computer machine-readable media such as non-transitory computer machine-readable storage media (e.g., a magnetic disk; an optical disk; a random access memory; a read only memory; a flash device; a phase change memory) and transitory computer machine-readable communication media (e.g., electrical, optical, acoustic, or other forms of propagated signals - such as, a carrier wave, an infrared signal, a digital signal, etc.). Additionally, such electronic devices typically include a collection of one or more processors coupled to one or more other components such as one or more storage devices (non-transitory machine-readable storage media), user input / output devices (e.g., a keyboard, a touch screen, and / or a display), and a network connection. The coupling of the collection of processors to the other components is typically through one or more buses and bridges (also known as bus controllers). The storage devices and signals carrying network traffic represent one or more machine-readable storage media and machine-readable communication media, respectively. Thus, the storage devices of a given electronic device typically store code and / or data for execution on the collection of one or more processors of that electronic device. Of course, different combinations of software, firmware, and / or hardware may be used to implement one or more portions of the embodiments of the present invention. Throughout this detailed description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of the present invention. However, it will be apparent to one of ordinary skill in the art that the present invention may be practiced without some of these specific details. In some instances, well-known structures and functions are not described in detail so as not to obscure the subject matter of the present invention. Accordingly, the scope and spirit of the present invention should be determined according to the appended claims.

Claims

1. A processor, comprising: an execution circuit system for executing instructions and generating memory access requests, wherein the memory access requests include load requests to read data from the memory and storage requests to store data in the memory; as well as cryptographic circuitry for performing a plurality of rounds of encryption or decryption to encrypt or decrypt the data, respectively, the cryptographic circuitry for performing one or more redundant rounds for corresponding one or more rounds of the plurality of rounds, the one or more redundant rounds comprising a spatial or temporal difference relative to the corresponding one or more rounds; The cryptographic circuitry is configured to generate an error upon detecting a mismatch between an output of a redundant round and an output of a corresponding round.

2. The processor according to claim 1, wherein: The one or more redundant rounds are to be time-interleaved with the corresponding one or more rounds.

3. The processor according to claim 1 or 2, wherein: The cryptographic circuitry is for modifying inputs to the corresponding one or more rounds to generate corresponding inputs to the one or more redundant rounds.

4. The processor according to claim 3, wherein: The inputs to the corresponding one or more rounds are to be homogeneously cross-mapped and mixed to generate the corresponding inputs to the one or more redundant rounds.

5. The processor according to any one of claims 1 to 4, wherein: The cryptographic circuit system comprises: A first status register, used to store the encryption or decryption status of the multiple rounds; and The second state register is used to store the redundant encryption or decryption state of the one or more redundant rounds.

6. The processor according to claim 5, wherein: The cryptographic circuit system further comprises: Circuitry for converting the redundant encryption or decryption state into a converted state representation corresponding to the encryption or decryption state.

7. The processor according to claim 6, wherein: The cryptographic circuit system further comprises: Comparison circuitry is provided for comparing the converted state representation with the encrypted or decrypted state to detect the mismatch.

8. The processor according to any one of claims 1 to 7, wherein: The cryptographic circuit system comprises: first N Sbox circuits and first M hybrid column circuits, the first N Sbox circuits and the first M hybrid column circuits being configured to be implemented with a first homogeneous composite field that operates on a first set of input bytes in each of the plurality of rounds; and The second N Sbox circuits and the second M hybrid column circuits are configured to be implemented with a second homogeneous composite field, the second homogeneous composite field operating on a second set of input bytes in each of the multiple rounds.

9. A method comprising: performing a cryptographic round to generate a first output; performing redundant cryptographic rounds corresponding to one or more of the cryptographic rounds to generate corresponding second outputs; comparing each of the second outputs to a corresponding first output to detect any mismatch; In the event that no mismatch is detected, storing the encrypted data or loading the decrypted data corresponding to the cryptographic round; as well as In the event that a mismatch is detected, an error or fault condition is generated.

10. The method according to claim 9, wherein: One or more redundant cryptographic rounds are to be time interleaved with corresponding one or more of the cryptographic rounds.

11. The method according to claim 9 or 10, further comprising: Inputs to the one or more cryptographic rounds are modified to generate corresponding inputs to one or more redundant cryptographic rounds.

12. The method according to claim 11, wherein: The inputs to the corresponding one or more cryptographic rounds are to be isomorphically cross-mapped and mixed to generate the corresponding inputs to the one or more redundant cryptographic rounds.

13. The method according to any one of claims 9 to 12, further comprising: storing the encryption or decryption status of the plurality of cryptographic rounds in a first status register; as well as Redundant encryption or decryption states of one or more redundant cryptographic rounds are stored in a second state register.

14. The method according to claim 13, further comprising: The redundant encryption or decryption state is converted to a converted state representation corresponding to the encryption or decryption state.

15. The method according to claim 14, further comprising: The converted state representation is compared to the encrypted or decrypted state to detect the mismatch.

16. A machine-readable medium having program code stored thereon, wherein when the program code is executed by a machine, the machine performs operations comprising: performing a cryptographic round to generate a first output; performing redundant cryptographic rounds corresponding to one or more of the cryptographic rounds to generate corresponding second outputs; comparing each of the second outputs to a corresponding first output to detect any mismatch; In the event that no mismatch is detected, storing the encrypted data or loading the decrypted data corresponding to the cryptographic round; as well as In the event that a mismatch is detected, an error or fault condition is generated.

17. The machine-readable medium of claim 16, wherein: One or more redundant cryptographic rounds are to be time interleaved with corresponding one or more of the cryptographic rounds.

18. The machine-readable medium of claim 16 or 17, further comprising: Inputs to the one or more cryptographic rounds are modified to generate corresponding inputs to one or more redundant cryptographic rounds.

19. The machine-readable medium of claim 18, wherein: The inputs to the corresponding one or more cryptographic rounds are to be isomorphically cross-mapped and mixed to generate the corresponding inputs to the one or more redundant cryptographic rounds.

20. The machine-readable medium of claim 18 or 19, further comprising program code for causing the following operations: storing the encryption or decryption status of the plurality of cryptographic rounds in a first status register; and The redundant encryption or decryption states of the one or more redundant cryptographic rounds are stored in a second state register.