Fair use of multiple context shared cryptographic hardware
Patent Information
- Application Number
- CN202211013295.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-10
- Filing Date
- 2022-08-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-08-23
Smart Images

Figure CN116781246B_ABST
Abstract
Description
Technical Field
[0001] At least one embodiment relates to processing resources for performing and facilitating the transfer of confidential data. For example, at least one embodiment relates to hardware circuitry for the fair use of multiple context-sharing cryptographic hardware. Background Technology
[0002] Accelerator circuitry includes Direct Memory Transfer (DMA) circuitry to access system memory independently of the Central Processing Unit (CPU). DMA circuitry can also be used for memory-to-memory copying or data movement within or between memory. When data needs to be protected, DMA circuitry can implement cryptographic circuitry to encrypt and decrypt data being copied from and to secure memory. Some cryptographic algorithms use sequential operations that require sequential analysis of the data. These sequential operations present challenges for multiple clients sharing cryptographic circuitry, such as when accelerator circuitry in a data center is shared among multiple users. Some implementations limit transfer sizes to support fairness in arbitration across uses. This is not ideal because some transfer sizes are very large, while others are relatively small. Alternatively, some implementations create authentication tags for each block in the encrypted data stream and use separate initialization vectors (IVs). However, this increases the memory footprint of the cryptographic hardware. Attached Figure Description
[0003] Figure 1 It is a block diagram of a computing system having an accelerator circuit according to at least some embodiments, the accelerator circuit including a replication engine that supports fairness between multiple users or between multiple data streams belonging to a single user in a single application;
[0004] Figure 2 It is a block diagram of a replication engine of an accelerator circuit according to at least some embodiments;
[0005] Figure 3 It is a functional diagram of a flow that schedules a single data transfer to be performed by a DMA circuit, according to at least some embodiments;
[0006] Figure 4 The illustrations depict a first application having multiple push buffers and a second application having a single push buffer, according to at least some embodiments.
[0007] Figures 5A-5B The processing flow of a DMA engine during multiple time slices is illustrated according to at least some embodiments;
[0008] Figure 6 It is a flowchart of encryption operations for partial transmission according to at least some embodiments;
[0009] Figure 7 It is a flowchart of a decryption operation for partial transmission according to at least some embodiments;
[0010] Figure 8 This is a flowchart of a method for scheduling partial transmission to support fairness among multiple applications, according to at least some embodiments; and
[0011] Figure 9 It is a block diagram of a computing system with an accelerator according to at least some embodiments, the accelerator including a replication engine that supports fairness among multiple users or multiple data streams of the computing system. Detailed Implementation
[0012] As mentioned above, DMA circuits can be used for memory-to-memory copying or data movement within memory and can include cryptographic hardware for data protection. Sequential encryption algorithms using sequential operations present some challenges for multiple clients sharing the encryption hardware. In particular, the Advanced Encryption Standard Galois Counter Mode (AES-GCM) is an authenticated encryption algorithm that performs both encryption and authentication of the data stream. Hardware implementation of AES-GCM circuitry is very expensive because each 16 bytes to be encrypted simultaneously requires a 128-bit multiplier. AES-GCM is a sequential operation that requires sequential analysis of data to compute the GHASH function. A single AES key K is used to encrypt the data and obtain authenticated data. The component used by GCM to generate the message authentication code is called GHASH. If multiple users attempt to utilize the AES-GCM hardware engine, because the block counter, initialization vector (IV), key (KEY), and GHASH require state tracking, one user's operation is serialized and completed before another user's operation is serialized and completed. This does not guarantee any fairness between users, as one user's transfer may be much larger than another user's transfer. Furthermore, if a single user or application attempts to use the AES-GCM hardware engine for multiple cryptographic streams belonging to that user or application, the operations of one cryptographic stream are serialized and completed before the operations of another cryptographic stream are serialized and completed because the block counter, IV, key, and GHASH require state tracking. This does not guarantee any fairness among multiple cryptographic streams in an application.
[0013] Several aspects and embodiments of this disclosure address these and other challenges by providing scheduler circuitry that splits data transfers into a set of partial transfers (or portions) (e.g., 8KB), where each partial transfer has a fixed size and is required to be completed before a context switch to another application launches. A replication engine (CE) can sequentially execute a set of partial transfers for a first application over a time period (e.g., until a time slice timeout occurs). The CE stores one or more pieces of data (e.g., hash keys, block counters, etc.) for encryption or decryption in the application's secure memory, calculated based on the last partial transfer (e.g., the last partial transfer completed before the time slice timeout). The IV value remains unchanged throughout the single replication, and a counter is appended to the IV and increments once for each specified block. For example, the IV can be 96 bits, and the counter can be a 32-bit counter that increments once every 16-byte block. When the application's data transfer is resumed by the CE (e.g., for a subsequent time slice), one or more pieces of data for encryption or decryption are retrieved and used. The CE then sequentially executes the remaining partial transfers over a second time period using the retrieved values (e.g., until another time slice timeout occurs). Once all partial transmissions are complete, the CE stores or outputs the authentication tag computed in the final partial transmission of that group. In this way, the accelerator circuitry supports fairness across multiple context-sharing cryptographic hardware. Accelerator circuitry can guarantee fairness across multiple context-sharing cryptographic hardware in certain situations. Accelerator circuitry can be a graphics processing unit (GPU), deep learning accelerator (DLA) circuitry, intelligent processing unit (IPU), neural processing unit (NPU), tensor processing unit (TPU), neural network processor (NNP), data processing unit (DPU), vision processing unit (VPU), application-specific integrated circuit (ASIC), or field-programmable gate array (FPGA). Accelerator circuitry can address the computational needs of the neural network inference phase by providing building blocks that accelerate core deep learning operations. For example, deep learning accelerators can be used to accelerate different neural networks, such as convolutional neural networks (CNN), recurrent neural networks (RNN), fully connected neural networks, and so on.
[0014] Accelerator circuits can be scheduled by a host central processing unit (CPU) coupled to the accelerator circuit. Alternatively, accelerator circuits can be scheduled locally by firmware to ensure minimal latency. Accelerator circuits can be used in different types of layers in these neural networks, such as fixed-function engines for convolution, activation functions, pooling, batch normalization, etc. It should be noted that from an algorithmic perspective, neural networks can be specified using a set of layers (referred to as “raw layers” in this paper), such as bias and batch normalization. These raw layers can be compiled or transformed into another set of layers (referred to as “hardware layers” in this paper), where each hardware layer is used as a basic element for scheduling to execute on the accelerator circuit. The mapping between raw layers and hardware layers can be m:n, where m is the number of raw layers and n is the number of hardware layers. For example, in a neural network, raw layer bias, batch normalization, and local response normalization (LRN), such as rectified linear units (ReLU), can be compiled into a single hardware layer. In this case, m:n is 3:1. Each hardware layer can be represented by basic hardware instructions from the accelerator circuitry for performing operations, and each layer can communicate with another layer via a memory interface. For example, a first fixed-function engine in the DLA circuitry can execute the first layer (which receives input tensors), perform operations on the input tensors to generate an output tensor, and store the output tensor in system memory, such as dynamic random-access memory (DRAM) coupled to the accelerator. A second fixed-function engine can execute the second layer, which receives the output tensor from the first layer as a second input tensor from memory, performs operations on the second input tensor to generate a second output tensor, and stores the second output tensor in DRAM. Each communication introduces tensor read and tensor write operations in the memory interface.
[0015] Therefore, several aspects of this disclosure allow encryption hardware to be shared among multiple users while supporting fairness among users, despite varying transmission sizes. Several aspects of this disclosure do not require transmissions to be software-split and pre-encrypted individually, which would generate multiple authentication tags for each split. Several aspects of this disclosure support Quality of Service (QoS) across multiple users sharing the same encryption hardware, regardless of the size of individual transmissions. Several aspects of this disclosure allow encryption hardware to be shared among multiple data streams while supporting fairness among streams, despite varying transmission sizes. Several aspects of this disclosure support QoS across multiple data streams sharing the same encryption hardware, regardless of the size of individual transmissions. In some cases, QoS can be guaranteed across multiple data streams.
[0016] Figure 1This is a block diagram of a computing system 100 having an accelerator circuit 102 according to at least some embodiments. The accelerator circuit 102 includes a replication engine 120 that supports fairness between multiple users of the accelerator circuit 102, or between multiple data streams belonging to a single user within a single application. The computing system 100 is considered a headless system, wherein cell-by-cell management of the accelerator circuit 102 occurs on a main system processor CPU 104. The accelerator circuit 102 includes an interrupt interface 106, a configuration space bus (CSB) interface 108, a main data bus interface 110 (data backbone interface (DBBIF)), a secondary data bus interface 112, and a replication engine (CE) 120, as described in more detail below. The CPU 104 and the accelerator circuit 102 are coupled to system memory 114 (e.g., DRAM). The accelerator circuit 102 is coupled to system memory 114 via the main data bus interface 110. The accelerator circuit 102 may be coupled to secondary memory 116, such as video memory (DRAM and / or SRAM), via the secondary data bus interface 112. CSB interface 108 can be a control channel interface that implements register files (e.g., configuration registers) and interrupt interfaces. In at least one embodiment, CSB interface 108 is a synchronous, low-bandwidth, low-power, 32-bit control bus designed for CPU 104 to access configuration registers in accelerator circuitry 102. The interrupt interface can be a 1-bit level-driven interrupt. An interrupt line can be asserted when a task completes or an error occurs.
[0017] The accelerator circuit 102 may also include a memory interface block that interfaces with the memory using one or more bus interfaces. In at least one embodiment, the memory interface block uses a main data bus interface 110 connected to system memory 114. System memory 114 may include DRAM. The main data bus interface 110 may be shared with the CPU and input / output (I / O) peripherals. In at least one embodiment, the main data bus interface 110 is a data backbone (DBB) interface connecting the accelerator circuit 102 and other memory subsystems. The DBB interface is a configurable data bus that can specify different address sizes, different data sizes, and issue requests of different sizes. In at least one embodiment, the DBB interface uses an interface protocol, such as AXI (Advanced Extensible Interface) or other similar protocols. In at least one embodiment, the memory interface block uses an auxiliary data bus interface 112 connected to an auxiliary memory 116 dedicated to the accelerator circuit 102. The auxiliary memory 116 may include DRAM. The auxiliary memory 116 may be video memory. The accelerator circuit 102 may also include a memory interface connected to a higher bandwidth memory dedicated to the accelerator circuit 102. The memory can be an on-chip SRAM used to provide higher throughput and lower access latency.
[0018] For example, during inference, a typical process begins with a management processor (microcontroller or CPU) coupled to accelerator circuitry 102 sending hardware layer configuration and activation commands. If data dependencies do not preclude this, multiple hardware layers can be sent to different engines and activated simultaneously (e.g., if the input of another layer does not depend on the output of the previous layer). In at least one embodiment, each engine may have a double buffer for its configuration register, which allows the configuration of a second layer to begin processing when the active layer completes. Once a hardware engine has completed its active task, accelerator circuitry 102 can interrupt the management processor to report completion, and the management processor can restart the process. This command-execution-interrupt flow repeats until the inference of the entire network is complete. In at least one embodiment, the interrupt interface may signal that replication is complete. In another embodiment, semaphore release (typically a flag written to system memory that the CPU thread is polling) can be used to let the software know that the workload has been completed.
[0019] Figure 1 The computing system 100 represents a more cost-sensitive system for the cell-by-cell management of the accelerator circuitry 102 than a computing system with a dedicated controller or coprocessor. The computing system 100 can be considered a small-scale system model. Small-scale system models can be used for cost-sensitive connected Internet of Things (IoT) devices, artificial intelligence (AI), and task-oriented automation systems where cost, area, and power are primary drivers. Cost, area, and power savings can be achieved through the configurable resources of the accelerator circuitry 102. Neural network models can be pre-compiled and their performance can be optimized, allowing for larger models to be reduced in terms of load complexity. In turn, the reduced load complexity allows DLA implementations to scale down, where models consume less storage space and require less time for system software to load and process. In at least one embodiment, the computing system 100 can execute one task at a time. Alternatively, the computing system 100 can execute multiple tasks simultaneously. For the computing system 100, context switching does not cause the CPU 104 to be overloaded by servicing numerous interrupts from the accelerator circuitry 102. This eliminates the need for an additional microcontroller, and CPU 104 performs memory allocation and other subsystem management operations. As described herein, accelerator circuit 102 includes a replication engine 120 that supports fairness among multiple users of accelerator circuit 102. Further details of replication engine 120 will be referenced below. Figure 2 Describe it.
[0020] Figure 2This is a block diagram of a replication engine 120 of accelerator circuitry according to at least some embodiments. The replication engine 120 includes a hardware scheduler circuitry 202 (labeled ESCED of the engine scheduler) and a direct memory access (DMA) circuitry 204. The hardware scheduler circuitry 202 is coupled to auxiliary memory 206 and the DMA circuitry 204. The DMA circuitry 204 is coupled to a memory management unit (MMU), and the MMU 230 is coupled to system memory (…). Figure 2 (Not shown in the image). MMU 230 can provide routing functionality between system memory and accelerator circuitry. MMU 230 can provide paths for all engines on the accelerator (including replication engine 120) to access any location in memory (e.g., video memory, system memory, etc.). MMU 230 can perform access checks to allow only authorized access via the interface. MMU 230 can restrict and report unauthorized access. DMA circuitry 204 includes encryption circuitry that implements certified encryption algorithms to encrypt data retrieved from secure memory or decrypt received data to be stored in secure memory. In at least one embodiment, as... Figure 2 As shown, the DMA circuit 204 includes a logical copy engine (LCE) 208 and a physical copy engine (PCE) 210. The DMA circuit 204 may include multiple PCEs coupled to the LCE 208. The LCE 208 may include a secure memory 212 that can store encrypted IVs, block counters, and hash keys in each channel key slot. Each channel key slot, which can be assigned to an application, can store the application's context in a designated slot within the secure memory 212. The LCE 208 may also include a secure private interface 214 for receiving configuration information, encryption and decryption keys, an IV random number generator, and secure SRAM programming from a secure hub or other secure circuitry managing private keys. In at least one embodiment, the encryption and decryption keys and the IV are generated by a random number generator from a secure processor on the GPU.
[0021] In at least one embodiment, such as Figure 2As shown, PCE 210 includes front-end circuitry 216, a read pipeline 218, a write pipeline 220, and encryption circuitry 222 that provides secure data transfer for applications requiring confidentiality. In at least one embodiment, encryption circuitry 222 is an AES-GCM hardware engine implementing AES256-GCM cryptography. In at least one embodiment, encryption circuitry 222 is an AES-GCM circuit. Alternatively, encryption circuitry 222 may implement encryption algorithms in other sequences, where the underlying encryption hardware is shared among multiple users (e.g., multiple applications in a time-sliced manner). For example, encryption hardware may be shared by multiple users in a cloud infrastructure. For another example, encryption hardware can be used in a virtualization environment where a hypervisor allows the underlying hardware to support multiple guest virtual machines (VMs) by virtually sharing its resources (including accelerator circuitry 100). In another embodiment, DMA circuitry 204 includes LCE 208, a first PCE 210, and a second PCE. A first PCE 210 is coupled to an LCE 208 and includes an encryption circuit 222, a first read pipeline 218, and a write pipeline 220. A second PCE is coupled to an LCE 208 and includes a second encryption circuit, a second read pipeline, a second write pipeline, and a second front-end circuit. In at least one embodiment, the second encryption circuit is a second AES-GCM circuit. Each additional PCE may include front-end circuitry, a read pipeline, a write pipeline, and encryption circuitry.
[0022] In at least one embodiment, the copy engine 120 can encrypt data related to the data transfer. To encrypt the data transfer, a context with a valid SRAM index pointing to a slot in secure memory 212 allocated to the application is loaded onto LCE 208. The KEY indicated in the slot of secure memory 212 is loaded onto encryption circuitry 222 (AES hardware engine). The first IV used is SRAM.IV+1, which is an incrementing IV stored in LCE 208. PCE 210 generates a memory request (read / write). PCE 210 reads plaintext data from a first region of memory (Calculated Protected Region (CPR)), encrypts the plaintext data with the KEY and IV, and adds it to an authentication tag (AT or AuthTag). During encryption operations, PCE 210 reads from protected memory (e.g., video memory), internally encrypts the data using encryption circuitry 222, and writes the encrypted data to an unprotected region (e.g., system memory or video memory). In at least one embodiment, PCE 210 writes encrypted data to a second region of memory (Non-Computationally Protected Region (NonCPR)). At the end of replication (or the last replication split in the time slice), PCE 210 writes the used IV to the second region of memory (NonCPR) and writes the computed authentication tag to the second region of memory (NonCPR). When interacting with the MMU, the request may carry a region identifier. The region identifier indicates the location where the memory region must be a CPR or a Non-Computationally Protected Region (NonCPR). Replication engine 120 may interact with the MMU to obtain the address of each region. The region identifier is specified by replication engine 120 when making an MMU translation request because the MMU tracks the CPR and NonCPR attributes of the memory region. If the region identifier specified by replication engine 120 does not match the attributes of the target memory location, the MMU will block access and return an error (e.g., MMU_NACK) to replication engine 120. The CPR is the first memory region containing the decrypted data. The CPR may be a memory sandbox, accessible only to selected clients and inaccessible to any malicious actors. NonCPR is any memory area other than CPR. NonCPR is untrusted because it can be accessed by malicious actors. The replication engine 120 ensures that data movement from NonCPR to CPR must follow a decryption path; that is, nonCPR needs to contain encrypted data that only the replication engine 120 with the correct key can understand. Similarly, the replication engine 120 ensures that any data movement from CPR to NonCPR traverses an encrypted path. Encrypted data in NonCPR is accessible to malicious actors but cannot be tampered with because malicious actors do not have the cryptographic key to understand the encrypted data.The replication engine 120 can write authentication tags to NonCPR, so users can detect the damage from malicious actors when decrypting.
[0023] In at least one embodiment, the copy engine 120 can decrypt data related to the data transfer. To decrypt the data transfer, a context is loaded (CTX LOAD) on LCE 208, the context having a valid SRAM index pointing to a slot in secure memory 212 allocated to the application. The KEY indicated in the slot of secure memory 212 is loaded onto encryption circuitry 222 (AES hardware engine). The first IV used is IB.IV+1, which is the IV tracked and incremented in hardware scheduler circuitry 202 and passed to LCE 208. PCE 210 reads the expected authentication tag from memory, reads the password data from the second memory region (NonCPR) of memory, decrypts the password data with the KEY and IV, and adds it to the authentication tag. During the decryption operation, PCE 210 reads from unprotected memory (e.g., in system memory or video memory), internally decrypts the data using encryption circuitry 222, and writes the decrypted data to a protected region (e.g., CPR). In at least one embodiment, PCE 210 writes plaintext data to the first memory region (CPR). During the final copy split, the PCE210 reads the authentication tag from the authentication tag address provided in the method and compares the calculated authentication tag with the provided authentication tag. If the values match, the operation succeeds. If there is no match, the PCE210 raises a fatal interrupt, no semaphore release occurs, and channel recovery is required. Channel recovery (also known as robust channel recovery or RC recovery) is a mechanism used by the resource manager or GPU PF driver to mark all pending jobs on the engine as invalid by indicating an error in each working channel. The engine is then reset. Channel errors are used by the resource manager (or GPU PF driver) to let the software layer (e.g., CUDA) know that the job is not yet complete.
[0024] In at least one embodiment, the IV is 96 bits and contrasts with two components, including a 64-bit channel counter with a unique identifier for each channel and a 32-bit message counter that starts at zero and is incremented with each encryption / decryption at the start of channel (SOC). The 96-bit RNG mask is a key mask stored in Secure PRI. The copy IV (COPY_IV) is RNGXOR[CHANNEL_CTR,++MSG_CTR]. The copy engine detects that the IV has exceeded the maximum number of copies by checking whether the MESSAG_CTR+1 value used in the COPY_IV construction is zero. The copy engine 120 keeps track of the encrypted IV used in each encrypted copy and performs pre-incrementing and save-and-restore operations based on SRAM. The encrypted IV is passed to the encryption circuit 222 after an XOR operation with the RNG in the decryption IV method after each copy. The IV stored in SRAM is reflected based on the completion of the copy. The replication engine 120 can have multiple encrypted copies visible to the PCE and maintain two counters, including the IVs that should be sent on network-encrypted copies and the last completed copy. In context saving (CTXT_SAVE), the IV from the last copy is saved to SRAM. The IV used for decryption is stored in the instance block and passed to the replication engine 120 via the decryption IV method during decryption copying. If MESSAGE_CTR = 0, the replication engine 120 can detect overflows and interrupts. The replication engine 120 can perform an XOR operation with the correct RNG before passing the decrypted IV from the LCE to the front-end circuitry 216.
[0025] In at least one embodiment, the replication engine 120 includes a secure private interface 214. The secure private interface 214 is accessible to security software to provide secure configurations or keys and to query the interrupt status of encryption and decryption. The replication engine 120 can connect as a client to the security hub 224, allowing the dedicated on-chip security processor (SEC2) 226 and the GPU system processor (GSP) 228 to access these secure private interfaces 214, but not to BAR0. The GSP 228 can be used to offload GPU initialization and management tasks. The SEC2226 manages encryption keys and other security information used by the accelerator circuitry 100.
[0026] In at least one embodiment, the secure memory 212 is a secure SRAM with N entries (e.g., 512 entries), each entry having a valid bit. Each entry has lower and higher 128-bit components. The lower component may include a first encrypted IV counter, a second encrypted IV counter, an IV channel identifier, one or more key indices, preemption information, and a block counter. The higher component may include a first authentication tag, a second authentication tag, a third authentication tag, and a fourth authentication tag. The secure SRAM can be programmed via registers through the secure private interface 214. The SRAM can support read, write, and invalidation functions. When the lower 128 bits of a 256-bit entry are programmed by the SEC2226 / GSP 228, the SRAM index can be marked as valid. An attempt to read an invalid SRAM entry will return 0x0 in the data register. In the event of a fatal error, the state of the SRAM cannot be guaranteed to be valid. The replication engine 120 can automatically invalidate the SRAM index in the event of a fatal error, thus requiring software to reprogram the SRAM index.
[0027] During operation in at least one embodiment, hardware scheduler circuitry 202 receives a first descriptor of a first data transfer of a first size associated with a first application. The first size may represent the full copy size of the first data transfer. The first descriptor specifies a first index in secure memory 212 corresponding to the first application. Hardware scheduler circuitry 202 splits the first data transfer into a first set of portions. Each portion should be less than or equal to the first size and needs to be completed before context switching to another application. For example, each of the multiple portions may be 8KB. Alternatively, other sizes may be used. These portions may be chunks, partial transfers, or partial copies, which together constitute a full copy of the first data transfer. During a first time period, DMA circuitry 204 sequentially executes a first subset of the first set of portions using a first IV and a first encryption key associated with the first application. In at least one embodiment, replication engine 120 executes each split portion (i.e., chunk) consecutively on PCE 210 upon completion.
[0028] At the end of the first time period, DMA circuit 204 stores a first calculated hash key and a first block counter in secure memory 212 at a designated first index. In at least one embodiment, replication engine 120 checks the calculated values of the partial hash key and block counter in secure memory 212 at the index corresponding to the loaded context. The partial hash key, also known as the sub-hash key (H), is an intermediate value ultimately used to calculate the authentication tag. For example, when the context's time slice expires, secure memory 212 contains the current block counter and the partial hash key calculated from the point of the last split. The current block counter and the partial hash key can be retained in secure memory 212 for subsequent time slices to complete the transfer. DMA circuit 204 can then be used for another application. During a second time period following the first time period, DMA circuit 204 uses the first calculated hash key and the first block counter stored in secure memory 212 at a designated first index to sequentially execute a second subset of the first set of parts. That is, when the original user's time slice is reloaded after arbitration among multiple users, the block counter and partial hash key can be recovered from secure memory 212 to complete the partial copy execution. At the end of the second time period, DMA circuit 204 stores the first authentication tag associated with the first data transfer. It is assumed that the first data transfer was completed within the second time period. If the first data transfer was not completed during the second time period, DMA circuit 204 continues execution of the remainder of the first group in subsequent time periods.
[0029] Hardware scheduler circuit 202 can receive a second descriptor for a second data transfer of a second size associated with a second application. The second descriptor specifies a second index in secure memory 212 corresponding to the second application. Hardware scheduler circuit 202 splits the second data transfer into a second set of portions. Each portion should be less than or equal to the second size and needs to be completed before context switching to another application. During a third time period, DMA circuit 204 sequentially executes a first subset of the second set of portions using a second IV and a second encryption key associated with the second application. At the end of the third time period, DMA circuit 204 stores a second calculated hash key and a second block counter in secure memory 212 at a specified second index. During a fourth time period following the third time period, DMA circuit 204 sequentially executes a second subset of the second set of portions using the second calculated hash key and the second block counter stored in secure memory 212 at the specified second index. At the end of the fourth time period, DMA circuit 204 stores a second authentication tag associated with the second data transfer. It is assumed that the second data transfer is completed within the fourth time period. If the second data transfer is not completed during the second time period, the DMA circuit 204 continues to execute the remainder of the second group in subsequent time periods. In one embodiment, the first size and the second size are different. Due to the difference in size, the hardware scheduler circuit 202 can independently guarantee fairness of QoS requirements between the first application and the second application. Alternatively, the hardware scheduler circuit 202 can independently guarantee fairness of QoS requirements between multiple data streams of the same application or belonging to the same user.
[0030] In at least one embodiment, the first descriptor is an encrypted descriptor. LCE 208 retrieves the first IV from secure memory 212, for example, via secure private interface 214, and retrieves the first encryption key from secure memory storing the key. Encryption circuit 222 generates a first block cipher using the first encryption key, the first IV, and a first value of the first block counter. Encryption circuit 222 generates a second block cipher using the first encryption key, the first IV, and a second value of the first block counter. Encryption circuit 222 generates a first portion of ciphertext using a first portion of plaintext and the second block cipher. Encryption circuit 222 uses the first portion of ciphertext and zero blocks (or a second temporary value) to compute a first value of a first computed hash key. Encryption circuit 222 generates a third block cipher using the first encryption key, the first IV, and a third value of the first block counter. Encryption circuit 222 generates a second portion of ciphertext using a second portion of plaintext and the third block cipher. Encryption circuit 222 uses the second portion of ciphertext and the first value of the first computed hash key to compute a second value of the first computed hash key. Encryption circuit 222 uses the last value of the first block cipher and the first calculated hash key to generate the first authentication tag.
[0031] In at least one embodiment, the first descriptor is a decryption operation descriptor. LCE 208 retrieves the first IV and the first decryption key. Encryption circuit 222 generates a first block cipher using the first encryption key, the first IV, and a first value of the first block counter. Encryption circuit 222 generates a second block cipher using the first encryption key, the first IV, and a second value of the first block counter. Encryption circuit 222 generates a first portion of plaintext using a first portion of the ciphertext and the second block cipher. Encryption circuit 222 calculates a first value of a first calculated hash key using the first portion of the ciphertext, a zero block (or a second temporary value), and the first portion of the ciphertext. Encryption circuit 222 generates a third block cipher using the first encryption key, the first IV, and a third value of the first block counter. Encryption circuit 222 generates a second portion of plaintext using the second portion of the ciphertext and the third block cipher. Encryption circuit 222 calculates a second value of the first calculated hash key using the second portion of the ciphertext and the first value of the first calculated hash key. Encryption circuit 222 generates a first authentication tag using the first block cipher and the last value of the first calculated hash key.
[0032] Figure 3 This is a functional diagram of a process 300, according to at least some embodiments, scheduling a single data transfer performed by DMA circuitry. A push buffer 302 allocated to the application is stored in auxiliary memory 206. Figure 3(Not shown in the image). Push buffer 302 may include multiple data transfers 304, 306 (also known as single-copy DMA) and a semaphore acquisition mechanism 308. The push buffer includes specifications of the operations that the GPU context will perform for a specific client. The push buffer is stored in memory. Software can place the semaphore acquisition mechanism 308 at the end of push buffer 302 and have the engine release the semaphore. Push buffer 302 may also store indices assigned to locations where the IV, KEY, and partial authentication tag are stored in secure memory 212 during encryption or decryption.
[0033] During the application's time slice, hardware scheduler circuitry 202 (ESCED) receives a first application descriptor from push buffer 302 for a first data transfer 304. Hardware scheduler circuitry 202 includes a copy splitter 310 that splits the first data transfer 304 (single copy DMA) into a set of partial transfers 312. Each partial transfer 312 has a fixed size (e.g., 8KB) smaller than the size of the first data transfer (e.g., 1GB). Each partial transfer 312 is required to complete before a context switch to another application begins. Each partial transfer 312 (e.g., 8KB copy) contains a binary descriptor represented by one or more methods. LCE 208 receives the partial transfers 312 from hardware scheduler circuitry 202, and LCE 208 schedules a subset of the partial transfers 312 to execute on PCE 210 during the time slice. In some cases, hardware scheduler circuitry 202 sends only a subset of the partial transfers 312 to LCE 208 for execution during the time slice. PCE 210 executes a subset of partial transfers 312 sequentially using the first context until a time slice timeout occurs in the application. In response to the time slice timeout, PCE 210 stores the current value of the first hash key calculated based on the last partial transfer completed before the time slice timeout and the current value of the first block counter in secure memory. It should be noted that... Figure 3 Only data transfers for a single application are shown. The hardware scheduler circuit 202 can receive data transfers from other push buffers corresponding to other applications, for example... Figure 4 As shown.
[0034] Figure 4 The illustration shows a first application with multiple push buffers and a second application with a single push buffer, according to at least some embodiments. The first application 402 may be allocated multiple push buffers 404, 406, and 408. Push buffers 404, 406, and 408 may be stored in auxiliary storage 206. Figure 4(Not shown in the image). Each push buffer 404, 406, 408 may include multiple data transfers (also known as single-copy DMA) and semaphore acquisition mechanisms. The second application 410 may be allocated push buffer 412. Push buffer 412 may be stored in auxiliary memory 206 (…). Figure 4 (Not shown in the image). Push buffer 412 may include multiple data transfer and semaphore acquisition mechanisms.
[0035] In at least one embodiment, the DMA buffer may store a first push buffer 404 for a first application 402 and a second push buffer 412 for a second application 410. The first push buffer 404 stores the specifications of the operations to be performed for the first application 402, as well as a first IV, a first block counter, and a first encryption key identifier stored in secure memory 212. The second push buffer 412 stores the specifications of the operations to be performed for the second application 410, as well as a second IV, a second block counter, and a second encryption key identifier stored in secure memory 212. In a further embodiment, a third push buffer 406 for the second application 410 is stored in the DMA buffer. Each of the storage of the first push buffer 404, the second push buffer 412, and the third push buffer 406 is acquired by a semaphore at the end of the corresponding push buffer released by the DMA circuitry.
[0036] Figures 5A-5B A processing flow 500 of a DMA engine 502 during multiple time slices, according to at least some embodiments, is illustrated. For example... Figure 5A As shown, during a first time period 506 (e.g., a first time slice), the first M portions 501 of a first data transfer are received by the DMA engine 502. The DMA engine 502 uses a first IV / KEY 503 to execute the first M portions 501 of the first data transfer. The first IV / KEY 503 can be loaded from secure memory 504. The DMA engine 502 executes the first M portions 501 until a context switch 518. At the context switch 518, the DMA engine 502 stores the current value of the first hash key calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 518) and the current value of the first block counter 505 in secure memory 504. A context switch 518 may occur when the time slice expires. The first time slice may be a first time period 506 allocated to a first application to access the DMA engine 502. The second time slice may be a second time period 508 allocated to a second application.
[0037] like Figure 5AAs shown, during the second time period 508 (e.g., the second time slice), the first M portions 507 of the second data transfer are received by the DMA engine 502. The DMA engine 502 uses a second IV / KEY 509 to execute the first M portions 507 of the second data transfer. The second IV / KEY 509 can be loaded from secure memory 504. The DMA engine 502 executes the first M portions 507 until a context switch 520 occurs. At the context switch 520, the DMA engine 502 stores the current value of the second hash key and the current value of the second block counter 511 calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 520) in secure memory 504. A context switch 520 may occur when the time slice expires. Processing flow 500 can perform similar operations for additional applications (if any) during additional time periods.
[0038] like Figure 5A As shown, during the third time period 510 (e.g., the third time slice), the second M portions 513 of the first data transfer are received by the DMA engine 502. The DMA engine 502 retrieves the current value of the first hash key and the current value of the first block counter 505 from the secure memory 504, calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 518). The DMA engine 502 uses the first IV / KEY 503 and the current values of the first hash key and the first block counter 505 retrieved from the secure memory 504 to execute the second M portions 513 of the first data transfer. The DMA engine 502 executes the second M portions 513 until a context switch 522 occurs. At the context switch 522, the DMA engine 502 stores the current value of the second hash key and the current value of the second block counter 519, calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 522), in the secure memory 504.
[0039] like Figure 5AAs shown, during the fourth time period 512 (e.g., the fourth time slice), the middle M portions 517 of the second data transfer are received by the DMA engine 502. The DMA engine 502 retrieves the current value of the second hash key and the current value of the second block counter 511 from the secure memory 504, calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 520). The DMA engine 502 uses the second IV / KEY 509 and the current values of the second hash key and the second block counter 511 retrieved from the secure memory 504 to execute the second M portions 517 of the second data transfer. The DMA engine 502 executes the second M portions 517 until a context switch 524 occurs. At the context switch 524, the DMA engine 502 stores the current value of the second hash key and the current value of the second block counter 519, calculated based on the last portion of the transfer completed before the time slice expires (e.g., context switch 522), in the secure memory 504. Processing flow 500 can perform similar operations for additional applications (if any) during additional time periods.
[0040] like Figure 5B As shown, during time period X 514 (e.g., the Xth time slice), the last M portions 521 (or fewer than M portions) of the first data transfer are received by DMA engine 502. DMA engine 502 retrieves the current value of the first hash key and the current value of the first block counter 523 from secure memory 504, calculated based on the last portion transfer completed before the last time slice expires. DMA engine 502 uses the first IV / KEY 503 and the current values of the first hash key and the first block counter 523 retrieved from secure memory 504 to execute the last M portions 521 of the first data transfer. DMA engine 502 executes the last M portions 521 until a context switch 528 occurs. At context switch 528, DMA engine 502 calculates and stores the first authentication tag 525 in secure memory 504. At this point, the first data transfer is complete.
[0041] like Figure 5BAs shown, during the X+1 time period 516 (e.g., X+1 time slice), the additional M portions 527 of the second data transfer are received by the DMA engine 502. The DMA engine 502 retrieves the current value of the second hash key and the current value of the second block counter 529 from the secure memory 504, calculated based on the last portion transfer completed before the last time slice expires. The DMA engine 502 uses the second IV / KEY 509 and the current values of the second hash key and the second block counter 529 retrieved from the secure memory 504 to perform the additional M portions 527 of the second data transfer. The DMA engine 502 performs the additional M portions 527 until a context switch 530. At the context switch 530, the DMA engine 502 stores the current value of the second hash key and the current value of the second block counter 531, calculated based on the last portion transfer completed before the time slice expires (e.g., context switch 530), in the secure memory 504. Processing flow 500 can perform similar operations for additional applications (if any) during additional time periods. It should be noted that the first data transmission of the first application is complete, but the second data transmission of the second application is not yet complete.
[0042] like Figure 5B As shown, during a Y time period 534 (e.g., a Y time slice), the last M portions 533 (or fewer than M portions) of the second data transfer are received by the DMA engine 502. The DMA engine 502 retrieves the current value of the second hash key and the current value of the second block counter 535 from the security memory 504, calculated based on the last portion transfer completed before the last time slice timeout. The DMA engine 502 uses the second IV / KEY 509 and the current values of the second hash key and the second block counter 535 retrieved from the security memory 504 to execute the last M portions 533 of the second data transfer. The DMA engine 502 executes the last M portions 533 until a context switch 536 or when these portions have been completed. At the context switch 536 or when these portions are completed, the DMA engine 502 calculates and stores the second authentication tag 537. At this point, the second data transfer is complete. It should be noted that... Figures 5A-5B This demonstrates how two applications can share the DMA engine 502 in a time-slice manner while maintaining fairness between context switches. Alternatively, the DMA engine 502 can be shared by multiple data streams in a time-slice manner while maintaining fairness between context switches.
[0043] Figure 6This is a flowchart of an encryption operation 600 for partial data transmission according to at least some embodiments. Encryption operation 600 is a simplified AES-GCM operation, illustrating a context switch before the complete data transmission is finished. For AES-GCM, blocks are sequentially numbered using a block counter (32'1). The value of the block counter is combined with a first IV (96'IV) (block 602) and encrypted with an AES block cipher to obtain a first result (block 604). Specifically, the first IV and the first value of the first block counter are encrypted with a first block cipher using a first encryption key to obtain the first result. The block counter is incremented by the IV (96'IV) (32'2) (block 606) and encrypted with an AES block cipher to obtain a second result (block 608). Specifically, the first IV and the second value of the first block counter are encrypted with a second block cipher using a first encryption key to obtain the second result. The second result and the first plaintext 610 are combined (e.g., XOR'd) (block 612) to obtain the first ciphertext 614. In block 616, the first ciphertext 614 is combined with zero block 615 to obtain the first value 618 of the first calculated hash key. The first value 618 is a partial authentication tag used for the first data transmission. The first value 618 can be stored in secure memory 660 before context switching 620. If no context switching occurs at this time, encryption operation 600 continues. It should be noted that fairness schemes can be applied to other values of the IV and counter, such as a 64-bit IV, etc.
[0044] The block counter increments by the IV (96'IV) (32'3) (block 622) and is encrypted with an AES block cipher to obtain a third result (block 624). Specifically, the third value of the first IV and the first block counter is encrypted with a second block cipher using the first encryption key to obtain a third result. The third result and the second plaintext 626 are combined (e.g., XORed) (block 628) to obtain the second ciphertext 630. The second ciphertext 630 is combined with the first value 618 of the hash key of the first calculation stored in the secure memory 660 (block 632) to obtain the second value 634 of the hash key used for the first calculation. The second value 634 is a partial authentication tag used for the first data transmission. The second value 634 may be stored in the secure memory 660 before the context switch 636. If no context switch occurs at this time, the encryption operation 600 continues.
[0045] At the end of the data transmission, the block counter is incremented by IV (96'IV) (32'N) (block 637) and encrypted with an AES block cipher to obtain a fourth result (block 638). Specifically, the first IV and the nth value of the first block counter are encrypted with the Nth block cipher using the first encryption key to obtain the fourth result. The fourth result and the Nth plaintext 640 are combined (e.g., XOR'd) (block 642) to obtain the Nth ciphertext 644. The Nth ciphertext 644 is combined with the Nth value of the first calculated hash key stored in the secure memory 660 (block 646) to obtain the Nth value 648 of the first calculated hash key. Since this is the last block of the data transmission, the Nth value 648 is combined with a ciphertext to obtain a fifth result 652. The fifth result 652 is combined with the first result from block 604 to obtain the first authentication tag 654.
[0046] Figure 7 This is a flowchart of a decryption operation 700 for partial data transmission according to at least some embodiments. The decryption operation 700 is a simplified AES-GCM operation, illustrating a context switch before the complete data transmission is finished. For AES-GCM, blocks are sequentially numbered using a block counter (32'1). The value of the block counter is combined with a first IV (96'IV) (block 702) and encrypted with an AES block cipher to obtain a first result (block 704). Specifically, the first IV and the first value of the first block counter are encrypted with a first block cipher using a first encryption key to obtain the first result. The block counter increments with the IV (96'IV) (32'2) (block 706) and is encrypted with an AES block cipher to obtain a second result (block 708). Specifically, the first IV and the second value of the first block counter are encrypted with a second block cipher using a first encryption key to obtain the second result. The second result and the first ciphertext 710 are combined (e.g., XOR'd) (block 612) to obtain the first plaintext 714. In block 716, the first ciphertext 710 and zero block 715 are combined to obtain a first value 718 for the hash key used in the first calculation. The first value 718 is a partial authentication tag used for the first data transmission. The first value 718 may be stored in secure memory 760 before a context switch 720 occurs. If no context switch occurs at this time, the encryption operation 700 continues.
[0047] The block counter increments by the IV (96'IV) (32'3) (block 722) and is encrypted with an AES block cipher to obtain a third result (block 724). Specifically, the third value of the first IV and the first block counter is encrypted with a second block cipher using the first encryption key to obtain a third result. The third result and the second ciphertext 726 are combined (e.g., XORed) (block 728) to obtain the second plaintext 730. The second ciphertext 726 is combined with the first value 718 of the first calculated hash key stored in the secure memory 760 (block 732) to obtain the second value 734 of the first calculated hash key. The second value 734 is a partial authentication tag used for the first data transmission. The second value 734 may be stored in the secure memory 760 before the context switch 736. If no context switch occurs at this time, the encryption operation 600 continues.
[0048] At the end of the data transmission, the block counter is incremented by the IV (96'IV) (32'N) (block 737) and encrypted with an AES block cipher to obtain a fourth result (block 738). Specifically, the first IV and the nth value of the first block counter are encrypted with the Nth block cipher using the first encryption key to obtain the fourth result. The fourth result and the Nth ciphertext 740 are combined (e.g., XOR'd) (block 742) to obtain the Nth plaintext 744. The Nth ciphertext 740 is combined with the Nth value of the first calculated hash key stored in secure memory 760 (block 746) to obtain the Nth value 748 of the first calculated hash key. Since this is the last block of the data transmission, the Nth value 748 is combined with a ciphertext to obtain a fifth result 752. The fifth result 752 is combined with the first result from block 704 to obtain a first authentication tag 754. The authentication tag can be compared with the expected authentication tag. If a match occurs, the operation is successful. If no match occurs, an error is detected as described herein.
[0049] Figure 8 This is a flowchart of a method 800 for scheduling portion delivery to support fairness among multiple applications, according to at least some embodiments. Method 800 can be executed by processing logic including hardware, software, firmware, or any combination thereof. In at least one embodiment, method 800 is performed by… Figure 1 The accelerator circuit 102 performs the operation. In at least one embodiment, method 800 is performed by... Figure 1 The replication engine 120 executes the method. In at least one embodiment, the method 800 is performed by... Figure 2 The hardware scheduler circuit 202 executes the method. In at least one embodiment, the method 800 is performed by... Figure 2 The hardware scheduler circuit 202 and DMA circuit 204 are executed.
[0050] Reference Figure 8Method 800 begins with the processing logic receiving a first descriptor (block 802) for a first data transfer of a first size associated with a first application. The first descriptor specifies a first index in secure memory corresponding to the first application. The processing logic splits the first data transfer into a first set of portions (block 804). Each portion has a size smaller than the first size, and one portion must be completed before a context switch to another application, but not all portions must be completed before a context switch. During a first period, the processing logic sequentially executes a first subset of the first set of portions using a first IV and a first encryption key associated with the first application via an authenticated encryption algorithm (block 806). At the end of the first period, the processing logic stores a first computed hash key and a first block counter in secure memory at a specified first index (block 808). During a second period following the first period, the processing logic sequentially executes a second subset of the first set of portions using the first computed hash key and the first block counter stored in secure memory at a specified first index (block 810). At the end of the second period, the processing logic stores a first authentication tag associated with the first data transfer (block 812), and then method 800 ends.
[0051] In a further embodiment, the processing logic receives a second descriptor for a second data transfer of a second size associated with a second application. The second descriptor specifies a second index in secure memory corresponding to the second application. The processing logic splits the second data transfer into a second set of portions. Each portion is smaller than the second size and needs to be executed before the context is switched to another application. During a third time period, the processing logic sequentially executes a first subset of the second set of portions using a second IV and a second encryption key associated with the second application. At the end of the third time period, the processing logic stores a second calculated hash key and a second block counter in secure memory at a specified second index. During a fourth time period following the third time period, the processing logic sequentially executes a second subset of the second set of portions using the second calculated hash key and the second block counter stored in secure memory at the specified second index. At the end of the fourth time period, the processing logic stores a second authentication tag associated with the second data transfer.
[0052] In one embodiment, the first descriptor is an encryption operation descriptor. In this embodiment, the processing logic retrieves a first IV and a first encryption key from secure memory. The processing logic encrypts the first IV and a first value of a first block counter using a first block cipher of the first encryption key to obtain a first result. The processing logic encrypts the first IV and a second value of the first block counter using a second block cipher of the first encryption key to obtain a second result. The processing logic combines the second result with first plaintext to obtain a first ciphertext. The processing logic combines the first ciphertext with a zero block to obtain a first value of a first calculated hash key. The first value is a partial authentication tag for the first data transmission. In a second time period, the processing logic retrieves the first IV, the current value of the first block counter, and the current value of the first calculated hash key. The processing logic encrypts the first IV and the current value of the first block counter using a third block cipher of the first encryption key to obtain a third result. The processing logic combines the third result with second plaintext to obtain a second ciphertext. The processing logic combines the second ciphertext with the current value of the first calculated hash key to obtain a fourth result. The processing logic combines the fourth result with a piece of ciphertext to obtain a fifth result. The processing logic combines the fifth result with the first result to obtain the first authentication label.
[0053] In another embodiment, the first descriptor is a decryption operation descriptor. In this embodiment, the processing logic retrieves a first IV and a first encryption key from secure memory. The processing logic encrypts the first IV and a first value of a first block counter using a first block cipher of the first encryption key to obtain a first result. The processing logic encrypts the first IV and a second value of the first block counter using a second block cipher of the first encryption key to obtain a second result. The processing logic combines the second result with a first ciphertext to obtain a first plaintext. The processing logic combines the first ciphertext, the first plaintext, and a zero block to obtain a first value of a first calculated hash key. The first value is a partial authentication tag for the first data transmission. During a second time period, the processing logic retrieves the first IV, the current value of the first block counter, and the current value of the first calculated hash key. The processing logic encrypts the first IV and the current value of the first block counter using a third block cipher of the first encryption key to obtain a third result. The processing logic combines the third result with a second ciphertext to obtain a second plaintext. The processing logic combines the second plaintext with the current value of the first calculated hash key to obtain a fourth result. The processing logic combines the fourth result with a piece of ciphertext to obtain a fifth result. The processing logic combines the fifth result with the first result to obtain the first authentication label.
[0054] In another embodiment, the processing logic receives a first descriptor of a first data transfer of a first size associated with a first data stream, the first descriptor specifying a first index in secure memory corresponding to the first data stream. The processing logic splits the first data transfer into a first set of portions. Each portion has a size smaller than the first size, and one portion must be completed before a context switch to another data stream, but not all portions must be completed before a context switch. During a first time period, the processing logic sequentially executes a first subset of the first set of portions using a first IV and a first encryption key associated with the first data stream via an authenticated encryption algorithm. At the end of the first time period, the processing logic stores a first computed hash key and a first block counter in secure memory at a specified first index. During a second time period following the first time period, the processing logic sequentially executes a second subset of the first set of portions using the first computed hash key and the first block counter stored in secure memory at a specified first index. At the end of the second time period, the processing logic stores a first authentication tag associated with the first data transfer.
[0055] In a further embodiment, the processing logic receives a second descriptor of a second data transfer of a second size associated with a second data stream. The second descriptor specifies a second index in secure memory corresponding to the second data stream. The processing logic splits the second data transfer into a second set of portions. Each portion is smaller than the second size and needs to be executed before the context is switched to another data stream. During a third time period, the processing logic sequentially executes a first subset of the second set of portions using a second IV and a second encryption key associated with the second data stream. At the end of the third time period, the processing logic stores a second calculated hash key and a second block counter in secure memory at a specified second index. During a fourth time period following the third time period, the processing logic sequentially executes a second subset of the second set of portions using the second calculated hash key and the second block counter stored in secure memory at the specified second index. At the end of the fourth time period, the processing logic stores a second authentication tag associated with the second data transfer.
[0056] Figure 9This is a block diagram of a computing system 900 with an accelerator according to at least some embodiments, the accelerator including a replication engine supporting fairness among multiple users or multiple data streams of the computing system. The computing system 900 is considered a leading system, wherein a main system processor CPU 104 delegates high-interrupt-frequency tasks to an accompanying microcontroller 904 coupled to accelerator circuitry 102. Except that the computing system 900 includes the accompanying microcontroller 904, the computing system 900 is similar to computing system 100, as indicated by similar reference numerals. The computing system 900 can be considered a larger system characterized by the addition of a dedicated control coprocessor and may include high-bandwidth SRAM to support accelerator circuitry 102.
[0057] In some cases, Figure 9 Larger models are used when higher performance and versatility are required. Performance-oriented systems can perform inference on many different network topologies; therefore, they maintain a high degree of flexibility. Additionally, these systems can perform many tasks simultaneously, rather than serializing inference operations, so that inference operations do not consume too much processing power on CPU 104. Accelerator circuitry 102 may include a memory interface coupled to a dedicated high-bandwidth SRAM to address these needs. The SRAM can be used by accelerator circuitry 102 as a cache. The SRAM can also be used by other high-performance computer vision-related components on the system to further reduce communication to main system memory 114 (e.g., DRAM). Accelerator circuitry 102 enables an interface with microcontroller 904 (or a dedicated control coprocessor) to limit the interrupt load on CPU 104. In at least one embodiment, microcontroller 904 may be a RISC-V-based PicoRV32 processor, an ARM Cortex-M or Cortex-R processor, or other microcontroller designs. Using the dedicated coprocessor (microcontroller 904), the main processor (CPU 104) can handle some of the tasks associated with managing accelerator circuitry 102. For example, while the hardware scheduler circuit is responsible for scheduling and fine-grained programming of the accelerator circuit 102, the microcontroller 904 or CPU 104 can still handle some coarse-grained scheduling of the accelerator circuit 102, input-output memory management (IOMMU) mapping of memory accesses (as needed), memory allocation of input data and fixed-weight arrays on the accelerator circuit 102, and synchronization between other system components and tasks running on the accelerator circuit 102.
[0058] The techniques disclosed herein can be incorporated into any processor capable of processing neural networks, such as a central processing unit (CPU), GPU, deep learning accelerator (DLA) circuitry, intelligent processing unit (IPU), neural processing unit (NPU), tensor processing unit (TPU), neural network processor (NNP), data processing unit (DPU), vision processing unit (VPU), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), etc. Such processors can be incorporated into personal computers (e.g., laptops), data centers, Internet of Things (IoT) devices, handheld devices (e.g., smartphones), vehicles, robots, voice-controlled devices, or any other device that performs inference, training, or any other processing on neural networks. Such processors can be used in virtualization systems, enabling an operating system running in a virtual machine on the system to utilize the processor.
[0059] As an example, processors incorporating the techniques disclosed herein can be used to process one or more neural networks in a machine to identify, classify, manipulate, process, operate, modify, or navigate physical objects in the real world. For instance, such processors can be used in autonomous vehicles (e.g., cars, motorcycles, helicopters, drones, flying devices, ships, submarines, delivery robots, etc.) to enable the vehicles to move in the real world. Additionally, such processors can be used in robots in factories to select components and assemble them into assemblies.
[0060] As an example, a processor incorporating the techniques disclosed herein can be used to process one or more neural networks to identify one or more features in an image or to alter, generate, or compress the image. For example, such a processor can be employed to enhance images rendered using rasterization, ray tracing (e.g., using NVIDIA RTX), and / or other rendering techniques. In another example, such a processor can be employed to reduce the amount of image data transmitted from a rendering device to a display device over a network (e.g., the Internet, mobile telecommunications networks, Wi-Fi networks, and any other wired or wireless network system). Such transmission can be utilized to stream image data from servers or data centers in the cloud to user devices (e.g., personal computers, video game consoles, smartphones, other mobile devices, etc.) to enhance services such as streaming images (e.g., NVIDIA GeForce Now (GFN), Google Stadia, etc.).
[0061] As an example, processors incorporating the techniques disclosed herein can be used to process one or more neural networks for any other type of application capable of utilizing neural networks. Such applications may involve translating languages, recognizing and removing sounds from audio, detecting anomalies or defects in the production of goods and services, monitoring living and non-living things, medical diagnosis, decision-making, and so on.
[0062] Other variations are within the spirit of this disclosure. Therefore, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrated embodiments are shown in the accompanying drawings and described in detail above. However, it should be understood that this disclosure is not intended to be limited to one or more specific forms disclosed, but rather is intended to cover all modifications, alternative constructions, and equivalents falling within the spirit and scope of this disclosure as defined in the appended claims.
[0063] Unless otherwise stated herein or obviously contradicted by the context, the terms “a” and “an” and “described”, and similar references used in the context of describing the disclosed embodiments (especially in the context of the claims) should be interpreted as encompassing both singular and plural forms, and should not be construed as definitions of the terms. Unless otherwise stated, the terms “comprising,” “having,” “including,” and “containing” should be interpreted as open-ended terms (meaning “including, but not limited to”). The term “connection,” when unmodified and referring to a physical connection, should be interpreted as partially or wholly contained within, attached to, or joined together, even if something is in between. Unless otherwise stated herein, statements of numerical ranges herein are intended only as a simplified representation of each individual value falling within that range. Each individual value is incorporated into the specification as if it were individually stated herein. In at least one embodiment, unless otherwise stated or contradicted by the context, the use of the terms “set” (e.g., “set of items”) or “subset” will be interpreted as a non-empty collection comprising one or more members. Furthermore, unless otherwise stated or contradicted by the context, the term "subset" for a corresponding set does not necessarily mean a proper subset of the corresponding set, but rather that a subset can be equal to the corresponding set.
[0064] Unless otherwise explicitly stated or clearly contradicted by the context, connective phrases such as “at least one of A, B, and C” or “at least one of A, B, and C” are understood in the context to generally refer to items, terms, etc., which can be A or B or C, or any non-empty subset of the set A, B, and C. For example, in an exemplary example of a set with three members, the connective phrases “at least one of A, B, and C” and “at least one of A, B, and C” refer to any of the following sets: {A}, {B}, {C}, {A, B}, {A, C}, {B, C}, {A, B, C}. Therefore, such connective language is generally not intended to imply that some embodiments require the presence of at least one of A, at least one of B, and at least one of C. Additionally, unless otherwise stated or contradicted by the context, the term “multiple” indicates a plural state (e.g., “multiple items” means multiple items). In at least one embodiment, the number of multiple items is at least two, but may be more when explicitly stated or indicated by the context. Furthermore, unless otherwise stated or clearly understood from the context, the phrase “based on” means “at least partially based on” rather than “based on only”.
[0065] Unless otherwise stated herein or clearly contradicted by the context, the operations of the processes described herein may be performed in any suitable order. In at least one embodiment, processes such as those described herein (or variations thereof and / or combinations thereof) are executed under the control of one or more computer systems configured with executable instructions and are implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that are executed jointly on one or more processors via hardware or a combination thereof. In at least one embodiment, the code is stored on a computer-readable storage medium, for example, in the form of a computer program, which in at least one embodiment includes a plurality of instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable storage medium that excludes transient signals (e.g., propagating transient electrical or electromagnetic transmissions) but includes non-transitory data storage circuitry (e.g., buffers, caches, and queues). In at least one embodiment, code (e.g., executable code or source code) is stored on one or more non-transitory computer-readable storage media (or other memory for storing executable instructions) on which executable instructions are stored, which, when executed by one or more processors of a computer system (i.e., as a result of execution), cause the computer system to perform the operations described herein. In at least one embodiment, the set of non-transitory computer-readable storage media includes multiple non-transitory computer-readable storage media, and one or more individual non-transitory storage media lack the complete code, while the multiple non-transitory computer-readable storage media collectively store the complete code. In at least one embodiment, the executable instructions are executed such that different instructions are executed by different processors; for example, the non-transitory computer-readable storage media stores the instructions, and the main central processing unit (“CPU”) executes some of the instructions, while the graphics processing unit (“GPU”) and / or data processing unit (“DPU”) (possibly together with the GPU) executes the other instructions. In at least one embodiment, different components of the computer system have separate processors, and the different processors execute different subsets of the instructions.
[0066] Therefore, in at least one embodiment, the computer system is configured to implement one or more services that perform the operations of the processes described herein, either individually or collectively, and such a computer system is configured with suitable hardware and / or software to enable the implementation of the operations. Furthermore, the computer system implementing at least one embodiment of this disclosure is a single device, and in another embodiment it is a distributed computer system comprising multiple devices operating in different ways, such that the distributed computer system performs the operations described herein, and that a single device does not perform all the operations.
[0067] The use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate embodiments of this disclosure and does not constitute a limitation on the scope of this disclosure unless otherwise required. No language in the specification should be construed as indicating that any unclaimed element is essential to the practice of the disclosure.
[0068] All references cited in this article, including publications, patent applications and patents, are incorporated herein by reference as if each reference were individually and specifically indicated to be incorporated herein by reference and the entire contents of which are described herein.
[0069] The terms “coupled” and “connected”, and their derivatives, may be used in the specification and claims. It should be understood that these terms may not be intended to be synonyms with each other. Rather, in certain examples, “connected” or “coupled” may be used to indicate that two or more elements are in direct or indirect physical or electrical contact with each other. “Coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.
[0070] Unless otherwise expressly stated, it will be understood that throughout this specification, terms such as “processing,” “computing,” “determining,” etc., refer to the actions and / or processes of a computer or computing system or similar electronic computing device that process and / or convert data represented as physical quantities (e.g., electrons) in the registers and / or memory of the computing system into other data represented as physical quantities in the memory, registers, or other such information storage, transmission, or display devices of the computing system.
[0071] Similarly, the term "processor" can refer to any device or part of memory that processes electronic data from registers and / or memory and converts that electronic data into other electronic data that can be stored in registers and / or memory. As a non-limiting example, a "processor" can be a CPU or a GPU. A "computing platform" can include one or more processors. As used herein, a "software" process can include, for example, software and / or hardware entities that perform work over time, such as tasks, threads, and intelligent agents. Likewise, each process can refer to multiple processes that execute instructions sequentially or intermittently, sequentially, or in parallel. In at least one embodiment, the terms "system" and "method" are used interchangeably herein, provided that a system can embody one or more methods, and a method can be considered a system.
[0072] In this document, reference may be made to obtaining, acquiring, receiving, or inputting analog or digital data into a subsystem, computer system, or computer-implemented machine. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog and digital data can be accomplished in various ways, such as by receiving data as a parameter to a function call or a call to an application programming interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data via a serial or parallel interface. In at least one embodiment, the process of obtaining, acquiring, receiving, or inputting analog or digital data can be accomplished by transmitting data from a providing entity to an acquiring entity via a computer network. In at least one embodiment, reference may also be made to providing, outputting, transmitting, sending, or presenting analog or digital data. In various examples, the process of providing, outputting, transmitting, sending, or presenting analog or digital data can be implemented by transmitting data as an input or output parameter to a function call, an application programming interface, or an inter-process communication mechanism.
[0073] While the description herein illustrates exemplary embodiments of the described technologies, other architectures may be used to implement the described functionality and are intended to fall within the scope of this disclosure. Furthermore, although specific assignments of responsibilities have been defined above for descriptive purposes, various functions and responsibilities may be assigned and divided in different ways depending on the circumstances.
[0074] Furthermore, although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter claimed in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms for implementing the claims.
Claims
1. An accelerator circuit, comprising: Secure storage; Dispatcher circuit; as well as A direct memory access (DMA) circuit is coupled to the scheduler circuit and the secure memory, wherein the DMA circuit includes encryption circuitry that implements a certified encryption algorithm to encrypt data retrieved from the secure memory or decrypt received data to be stored in the secure memory. The scheduler circuit is used for: A first descriptor is received from a first application for a first data transfer of a first size, the first descriptor specifying a first index in the secure memory corresponding to the first application; as well as The first data transmission is split into a first group of parts, each part being less than or equal to the first size and needing to be executed before the context is switched to another application; The DMA circuit described therein is used for: During the first time period, a first subset of the first group portion is executed sequentially using a first initialization vector IV and a first encryption key associated with the first application, and a first calculated hash key and a first block counter are stored in the secure memory at a designated first index; as well as During a second period following the first period, a second subset of the first set of portions is executed sequentially using the hash key of the first computation stored in the secure memory at a designated first index and the first block counter, and the first authentication tag associated with the first data transfer is stored.
2. The accelerator circuit according to claim 1, wherein: The scheduler circuit is further used for: A second descriptor is received from a second application for a second data transfer of a second size, the second descriptor specifying a second index in the secure memory corresponding to the second application; as well as The second data transfer is split into a second set of parts, each part being less than or equal to the second size, and this process needs to be completed before the context is switched to another application; The DMA circuit is further used for: During the third time period, a first subset of the second group of parts is executed sequentially using the second IV and the second encryption key associated with the second application; At the end of the third time period, the second calculated hash key and the second block counter are stored in the secure memory at the designated second index; During the fourth period following the third period, a second subset of the second set of parts is executed sequentially using the hash key of the second computation stored in the secure memory at the designated second index and the second block counter; as well as At the end of the fourth time period, a second authentication tag associated with the second data transmission is stored.
3. The accelerator circuit according to claim 1, wherein, The certified encryption algorithm is Advanced Encryption Standard Galois Counter Mode (AES-GCM).
4. The accelerator circuit according to claim 1, wherein, The first descriptor is an encryption operation descriptor, wherein the DMA circuit is further used for: Retrieve the first IV and the first encryption key from the secure memory; The first result is obtained by encrypting the first IV and the first value of the first block counter using the first block password of the first encryption key; The second result is obtained by encrypting the second value of the first IV and the first block counter using the second block cipher of the first encryption key; The second result is combined with the first plaintext to obtain the first ciphertext; The first ciphertext is combined with a zero block to obtain a first value of the hash key calculated first, wherein the first value is a partial authentication tag used for the first data transmission; Retrieve the current value of the first IV, the first block counter, and the current value of the first calculated hash key; A third result is obtained by encrypting the current value of the first IV and the first block counter using the third block cipher of the first encryption key; The third result is combined with the second plaintext to obtain the second ciphertext; The second ciphertext is combined with the current value of the first calculated hash key to obtain the fourth result; The fourth result is combined with a piece of ciphertext to obtain the fifth result; as well as The fifth result is combined with the first result to obtain the first authentication label.
5. The accelerator circuit according to claim 1, wherein, The first descriptor is a decryption operation descriptor, wherein the DMA circuitry is further used for: Retrieve the first IV and the first encryption key from the secure memory; The first result is obtained by encrypting the first IV and the first value of the first block counter using the first block password of the first encryption key; The second result is obtained by encrypting the second value of the first IV and the first block counter using the second block cipher of the first encryption key; The second result is combined with the first ciphertext to obtain the first plaintext; The first ciphertext and the zero block are combined to obtain a first value of the hash key calculated first, wherein the first value is a partial authentication tag used for the first data transmission; Retrieve the current value of the first IV, the first block counter, and the current value of the first calculated hash key; A third result is obtained by encrypting the current value of the first IV and the first block counter using the third block cipher of the first encryption key; The third result is combined with the second ciphertext to obtain the second plaintext; The second plaintext is combined with the current value of the hash key calculated in the first step to obtain the fourth result; The fourth result is combined with a piece of ciphertext to obtain the fifth result; as well as The fifth result is combined with the first result to obtain the first authentication label.
6. The accelerator circuit of claim 2, wherein the first size and the second size are different, and wherein the scheduler circuit is configured to guarantee fairness of QoS requirements across the first application and the second application, regardless of the first size and the second size.
7. The accelerator circuit according to claim 1, wherein, The accelerator circuit is a graphics processing unit (GPU), a deep learning accelerator (DLA) circuit, an intelligent processing unit (IPU), a neural processing unit (NPU), a tensor processing unit (TPU), a neural network processor (NNP), a data processing unit (DPU), a vision processing unit (VPU), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
8. The accelerator circuit of claim 2 further includes a DMA buffer for storing a first push buffer for a first application and a second push buffer for a second application, wherein the first push buffer stores a specification of an operation to be performed for the first application and a first index, and wherein the second push buffer stores a specification of an operation to be performed for the second application and a second index.
9. The accelerator circuit according to claim 8, wherein, The DMA buffer is used to store a third push buffer for the second application, wherein the second push buffer stores a semaphore at the end of the second push buffer, and the semaphore will be released by the DMA circuit.
10. The accelerator circuit according to claim 1, wherein, The DMA circuit includes: Logic Computing Engine (LCE); A first physical computing engine (PCE), coupled to the LCE and including the encryption circuitry, a first read pipeline, and a first write pipeline; and A second PCE is coupled to the LCE and includes a second encryption circuit, a second read pipeline, and a second write pipeline.
11. A data transmission method, comprising: The accelerator circuit receives a first descriptor of a first data transfer of a first size associated with a first application, the first descriptor specifying a first index in secure memory corresponding to the first application; The accelerator circuitry splits the first data transmission into a first set of parts, each part being less than or equal to the first size and needing to be executed before the context is switched to another application; During the first time period, the accelerator circuit sequentially executes a first subset of the first group of parts using a first initialization vector IV and a first encryption key associated with the first application through a certified encryption algorithm, and the accelerator circuit stores a first calculated hash key and a first block counter in the secure memory at a designated first index. as well as During a second period following the first period, the accelerator circuit sequentially executes a second subset of the first set of portions using the hash key of the first computation stored in the secure memory at a designated first index and the first block counter, and the accelerator circuit stores the first authentication tag associated with the first data transmission.
12. The method of claim 11, further comprising: The accelerator circuit receives a second descriptor for a second data transfer of a second size associated with a second application, the second descriptor specifying a second index in the secure memory corresponding to the second application; The accelerator circuit splits the second data transmission into a second set of parts, each part being less than or equal to the second size, and needs to complete execution before the context is switched to another application; During the third time period, the accelerator circuit sequentially executes a first subset of the second set of parts using a second IV and a second encryption key associated with the second application; At the end of the third time period, the accelerator circuit stores the second calculated hash key and the second block counter at a designated second index in the secure memory; During the fourth period following the third period, the accelerator circuit sequentially executes a second subset of the second set of parts using the hash key of the second computation stored in the secure memory at a designated second index and the second block counter; as well as At the end of the fourth time period, the accelerator circuit stores the second authentication tag associated with the second data transmission.
13. The method of claim 11, wherein the certified encryption algorithm is Advanced Encryption Standard Galois Counter Mode (AES-GCM).
14. The method of claim 11, wherein the first descriptor is an encryption operation descriptor, and wherein the method further comprises: Retrieve the first IV and the first encryption key from the secure memory; The first result is obtained by encrypting the first IV and the first value of the first block counter using the first block password of the first encryption key; The second result is obtained by encrypting the second value of the first IV and the first block counter using the second block cipher of the first encryption key; The second result is combined with the first plaintext to obtain the first ciphertext; The first ciphertext is combined with a zero block to obtain a first value of the hash key calculated first, wherein the first value is a partial authentication tag used for the first data transmission; Retrieve the current value of the first IV, the first block counter, and the current value of the first calculated hash key; A third result is obtained by encrypting the current value of the first IV and the first block counter using the third block cipher of the first encryption key; The third result is combined with the second plaintext to obtain the second ciphertext; The second ciphertext is combined with the current value of the first calculated hash key to obtain the fourth result; The fourth result is combined with a piece of ciphertext to obtain the fifth result; as well as The fifth result is combined with the first result to obtain the first authentication label.
15. The method of claim 11, wherein the first descriptor is a decryption operation descriptor, and the method further comprises: Retrieve the first IV and the first encryption key from the secure memory; The first result is obtained by encrypting the first IV and the first value of the first block counter using the first block password of the first encryption key; The second result is obtained by encrypting the second value of the first IV and the first block counter using the second block cipher of the first encryption key; The second result is combined with the first ciphertext to obtain the first plaintext; The first ciphertext and the zero block are combined to obtain a first value of the hash key calculated first, wherein the first value is a partial authentication tag used for the first data transmission; Retrieve the current value of the first IV, the first block counter, and the current value of the first calculated hash key; A third result is obtained by encrypting the current value of the first IV and the first block counter using the third block cipher of the first encryption key; The third result is combined with the second ciphertext to obtain the second plaintext; The second ciphertext is combined with the current value of the hash key calculated in the first step to obtain the fourth result; The fourth result is combined with a piece of ciphertext to obtain the fifth result; as well as The fifth result is combined with the first result to obtain the first authentication label.
16. An accelerator circuit, comprising: The replication engine CE includes Advanced Encryption Standard Galois Counter Mode (AES-GCM) hardware, the CE being configured to perform encryption and authentication on multiple applications, wherein the CE includes secure memory for storing a first context including a first encryption key, a first initialization vector IV, a first hash key, and a first block counter associated with a first application among the multiple applications, and a second context including a second encryption key, a second IV, a second hash key, and a second block counter associated with a second application among the multiple applications; The engine scheduler coupled to the CE, wherein: The engine scheduler is used for: Receive an encryption operation descriptor or a decryption operation descriptor for a first data transfer of a specified size for the first application; as well as The first data transfer is split into a set of partial transfers, each with a fixed size less than or equal to a specified size, and must be completed before the context is switched to another application; The replication engine CE is used for: The set of partial deliveries are executed sequentially using the first context until the time slice of the first application times out; and In response to the time slice timeout of the first application, the current value of the calculated first hash key and the current value of the first block counter, which were transmitted based on the last part completed before the time slice timeout, are stored in the secure memory.
17. The accelerator circuit according to claim 16, wherein, The CE is also used for: Retrieve the current value of the first hash key and the current value of the first block counter from the secure memory; The remaining portion of the transfer of the first set of partial transfers is executed sequentially using the current value of the first hash key and the current value of the first block counter until the second time slice of the first application times out; as well as In response to the second time slice timeout of the first application, the second current value of the first hash key and the second current value of the first block counter, calculated based on the last part of the transmission completed before the second time slice timeout, are stored in the secure memory.
18. The accelerator circuit of claim 16, wherein the CE is further configured to: Retrieve the current value of the first hash key and the current value of the first block counter from the secure memory; The remaining portion of the set of partial transfers is sequentially performed using the current value of the first hash key and the current value of the first block counter until the second time slice of the first application times out; and In response to the second time slice timeout of the first application, the authentication tag calculated in the last part of the set of partial transmissions is stored.
19. The accelerator circuit of claim 16, wherein the accelerator circuit is a graphics processing unit (GPU), a deep learning accelerator (DLA) circuit, an intelligent processing unit (IPU), a neural processing unit (NPU), a tensor processing unit (TPU), a neural network processor (NNP), a data processing unit (DPU), a vision processing unit (VPU), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
20. The accelerator circuit of claim 16, wherein the CE comprises: Logic Computing Engine (LCE); as well as A first physics computing engine (PCE) is coupled to the LCE and includes a first AES-GCM circuit, a first read pipeline, and a first write pipeline. as well as A second PCE is coupled to the LCE and includes a second AES-GCM circuit, a second read pipeline, and a second write pipeline.
Citation Information
Patent Citations
Data decryption processing method and device and data encryption processing method and device
CN110636081A
Method and system for high-speed processing IPSec security protocol packets
US20020188839A1