Model encryption on ai accelerators

By encrypting AI models with a symmetric key and decrypting them using a private key fused into the AI accelerator, the techniques address the challenge of ensuring confidentiality and integrity of AI models in cloud-based environments, effectively reducing the risk of unauthorized access.

WO2025096745A1PCT designated stage expired Publication Date: 2025-05-08MTS IP HLDG LTD +4

Patent Information

Application Number
PCT/US2024/053838
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-14
Filing Date
2024-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing technologies face challenges in ensuring the confidentiality and integrity of artificial intelligence (AI) models, particularly when these models are executed on AI accelerators in cloud-based environments, where they are vulnerable to unauthorized access and data breaches.

Method used

The described techniques involve encrypting AI models with a symmetric key generated and owned by the model owner, and then decrypting the model using a private key fused into the AI accelerator, thereby ensuring that the decryption process occurs within a trusted execution environment, such as a GPU.

Benefits of technology

This approach effectively protects the confidentiality and integrity of AI models by ensuring that only the AI accelerator, which is securely fused with the private key, can decrypt and access the model, thereby reducing exposure to unauthorized access and data breaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024053838_08052025_PF_FP_ABST
    Figure US2024053838_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A device may determine at least one public key and at least one private key associated with at least one accelerator in the cloud environment. A device may encrypt, using the at least one public key, at least one symmetric encryption key. A device may transmit the encrypted symmetric encryption key to the accelerator. A device may decrypt at the accelerator, the encrypted symmetric encryption key using the at least one private key. A device may encrypt a dataset using the at least one decrypted symmetric encryption key, to generate an encrypted dataset. A device may determine, by a server, the at least one accelerator based at least in part on the at least one public key and the at least one private key. A device may load, by the server, the encrypted dataset into the at least one accelerator.
Need to check novelty before this filing date? Find Prior Art

Description

MODEL ENCRYPTION ON Al ACCELERATORSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This Application claims the benefit of U.S. Provisional Patent Application 63 / 596,129, filed November 3, 2023, and entitled “MODEL ENCRYPTION ON Al ACCELERATORS,” and U.S. Provisional Patent Application 63 / 660,093, filed June 14, 2024, and entitled “MODEL ENCRYPTION ON Al ACCELERATORS,” which are hereby incorporated by reference in their entireties. In cases where this application and a document incorporated by reference conflict, this application controls.FIELD OF INVENTION

[0002] The invention relates generally to model encryption and decryption and, more specifically, to model encryption and decryption on artificial intelligence (Al) accelerators.BACKGROUND

[0003] Privacy-preserving machine learning serves as an important tool in ensuring customer data privacy. In particular, machine learning and artificial models carrying out tasks rely on a trusted execution environment and expect to run the models with confidentiality and integrity in environments, including cloud-based environments. Artificial Intelligence (Al) accelerators serve to improve performance when executing machine learning tasks. As such, privacy of Al accelerators plays an integral role in ensuring efficient privacy-preserving machine learning.SUMMARY

[0004] The confidentiality and integrity of a trained model is of utmost importance. Techniques described herein provide one or more solutions, including through model confidentiality, by encrypting the model with a symmetric key that is generated and owned by the model owner. A symmetric decryption key reaches the graphics processing unit (GPU) and allows the GPU to decrypt the model for loading in memory and carrying out runtime protections. In this way, confidentiality and integrity of the model is preserved. A trained model is protected at runtime in the GPU memory from adversaries (either local orremote) attempting to access a model encryption key (MEK) or the model itself. Described techniques provide a solution through model encryption with a symmetric key, with the GPU and / or Al accelerators, decrypting the model, thereby providing run-time protection to ensure the confidentiality and integrity of the model.

[0005] Some of the described techniques avoid relying on the CPU or other external devices to decrypt the model before loading the model to the Al accelerator. A private key provided or managed by the customer is fused into the Al accelerator, thereby allowing the Al accelerator to decrypt the model. Thus, the Al accelerator is responsible for decrypting the model.

[0006] Some of the described techniques use private keys or secrets (the secrets used to derive a key), inside the Al accelerator, through for example, a soft fuse, a hard fuse, protected memory or registers. This generates a different key for every Al accelerator, or one key associated with a set of Al accelerators.

[0007] Some of the described techniques use private symmetric keys in accelerators, fusing the private symmetric keys into an Al accelerator, allowing the Al accelerator to decrypt a model, such as customer loaded models.

[0008] In some embodiments, the customer can wrap a private symmetric key for the model, encrypt it and provide the model to a host device. The host device loads the model to the Al accelerator, with the Al accelerator being fused with the private symmetric key associated with the model. The Al accelerator further decrypts the model before loading the model into the system memory. The decrypted model is accessible by Al accelerators for use, including for Al inferencing. As the host receives an encrypted model, with decryption occurring with the fusing of the private symmetric keys with the Al accelerator, the host is unable to access the decrypted model, thereby protecting the data integrity of the model and limiting exposure to outside modules including but not limited to cloud storage, service providers, and external attackers.

[0009] In some embodiments, the GPU generates a set of asymmetric encryptions keys (i.e. public and private keys), as part of the Device Identity Composition Engine (DICE) with the device identity engine deriving unique and per- attestation keys.

[0010] In some embodiments, encryption keys are accessed from one or more Al accelerators, and further furnished for a standard attestation process. The model is encryptedby the customer with public keys and transmitted to a service provider. Server devices within the service provider loads the encrypted model into the one or more Al accelerators, with the one or more Al accelerators decrypting the at least one model to be placed in memory or storage.

[0011] In some embodiments, attempts to access the decrypted model from memory are verified by the at least one or more Al accelerators using range checks or other standard mechanisms.

[0012] In some embodiments, the model is encrypted by the customer using a private key, also referred to as the MEK. The MEK is the public part of an asymmetric encryption key pair and is derived by the last firmware layer in the DICE architecture. The model and the MEK are transferred to the host, with the MEK never being released by the GPU and is used to unwrap the MEK.

[0013] In some embodiments, a public encryption key is exported by the GPU through an attestation manifest for use with a MEK. In this process, the GPU providing the attestation manifest will be able to access the MEK using the asymmetric private decryption key for unwrapping the MEK.

[0014] In some embodiments, the GPU launches a Trusted Execution Environment (TEE) before unwrapping the MEK. The GPU launches the TEE in which the memory used by the processor cores is encrypted by a TEE Session key so that the unwrapped MEK is neither inaccessible nor modifiable.

[0015] In some embodiments, the MEK and encrypted model in host memory / storage is transferred to the GPU. Both the MEK and encrypted model are sent over the PCIe link by the CPU to the GPU. Since both the MEK and encrypted model are traversing the PCIe bus in encrypted form, any unauthorized access to the MEK and encrypted model by malicious endpoints on the same root complex is reduced or eliminated.

[0016] In some embodiments, once the model is transferred to the GPU, the Al accelerator cores enter a special execution mode serving as a TEE, where GPU microcode (i.e., immutable RTL code) generates an ephemeral session key to encrypt all transactions flowing outside from the GPU cores, including any shared registers, caches, TLBs, DRAM, etc.

[0017] In some embodiments, the MEK is loaded in memory (now double encrypted by the TEE key) and unwrapped. The model can at this point be decrypted and loaded in TEEmemory. The GPU is able to generate output tokens and transfer them in plain text or encrypted back to the host. When the workload is done, the GPU executes a tear down command (i.e., microcode) in which the ephemeral session key of the TEE is deleted and / or revoked, turning all GPU memory into garbage (MEK, Model, Tokens, etc.).

[0018] Some techniques described herein relate to a method for encrypting data in a cloud environment, the method including: determining at least one public key and at least one private key associated with at least one accelerator in the cloud environment; encrypting, using the public key, at least one symmetric encryption key; transmitting the encrypted symmetric encryption key to the accelerator; decrypting at the accelerator, the encrypted symmetric encryption key using the at least one private key; encrypting a dataset using the at least one decrypted symmetric encryption key, to generate an encrypted dataset; determining, by a server, the at least one accelerator based at least in part on the at least one public key and the at least one private key; and loading, by the server, the encrypted dataset into the at least one accelerator.

[0019] Some techniques described herein relate to a method, wherein the accelerator executes machine learning tasks and Al functions,

[0020] Some techniques described herein relate to a method, wherein the at least one public key is determined based at least in part on an attestation process.

[0021] Some techniques described herein relate to a method, further including: storing the at least one private key in a memory location of the at least one accelerator.

[0022] Some techniques described herein relate to a method, wherein the dataset includes at least one Al model.

[0023] Some techniques described herein relate to a method, further including wrapping the private key using at least a portion of the symmetric key pair.

[0024] Some techniques described herein relate to a method, further including implementing, using at least the at least one Al model, at least one machine learning task.

[0025] Some techniques described herein relate to a method, wherein the at least one machine learning task includes inferencing.

[0026] Some techniques described herein relate to a method for decrypting data, the method including: generating, for a GPU, at least one asymmetric encryption key pair, the at least oneasymmetric encryption key pair including a public key and a private key; determining, based at least in part on a verification, that the public key is associated with at least one Al accelerator; encrypting, using at least the public key a model to generate an encrypted model; determining, by a host, the at least one accelerator based at least in part on the private key and the public key; loading, by the host, the encrypted model into the at least one accelerator; and decrypting the encrypted model to generate a decrypted model for storage in the at least one accelerator.

[0027] Some techniques described herein relate to a method, wherein the GPU and the at least one accelerator are part of a trusted execution environment.

[0028] Some techniques described herein relate to a method for decrypting data, the method including: determining, based at least in part on a verification, at least one public key associated with a trusted execution environment, wherein the trusted execution environment includes at least one trusted computing component; encrypting, using the at least one public key, at least one model to generate at least one encrypted model; transmitting, the at least one encrypted model to a server; loading, in the trusted execution environment, the at least one encrypted model; and decrypting the at least one encrypted model in the trusted execution environment, to generate at least one decrypted model.

[0029] Some techniques described herein relate to a method, wherein the trusted execution environment is included in a GPU.

[0030] Some techniques described herein relate to a method, wherein a memory of the GPU stores the at least one decrypted model, wherein the memory is within a trust boundary associated with the trusted execution environment.

[0031] Some techniques described herein relate to a method for encrypting a model, the method including: encrypting, using at least one model encryption key, the model; transmitting, to a host device, the model and the at least one model encryption key; generating, by a device core layer, an asymmetric encryption key pair including a private portion and a public portion; wrapping the at least one model encryption key, using at least the public portion of the asymmetric encryption key pair; and storing the private portion of the asymmetric encryption key pair and the at least one model encryption key in a GPU, wherein the GPU unwraps the at least one model encryption key using the private portion of the asymmetric encryption key pair.

[0032] Some techniques described herein relate to a method, wherein the GPU is part of a trusted execution environment.

[0033] Some techniques described herein relate to a method, further including: decrypting the model to generate a decrypted model; and storing the decrypted model in a memory of the GPU, wherein the memory is within a trust boundary associated with the trusted execution environment.

[0034] Some techniques described herein relate to a non-transitory machine-readable medium having instructions stored thereon, which when executed by a processor, cause the processor to perform operations, the operations including: encrypting, using at least one model encryption key, a model; transmitting, through at least one bus, the model and the at least one model encryption key, to a host device; generating, by a device core layer, an asymmetric encryption key including a private portion and a public portion; wrapping the at least one model encryption key using at least the public portion; and storing the private portion and the at least one model encryption key in a GPU, wherein the GPU unwraps the at least one model encryption key using the private portion.

[0035] Some techniques described herein relate to a non-transitory machine -readable medium, wherein the GPU is part of a trusted execution environment.

[0036] Some techniques described herein relate to a system including one or more circuits, the one or more circuits including at least one CPU, at least one GPU, and at least one communication bus, the one or more circuits configured to perform operations including: transmitting, through the at least one communication bus, at least one model encryption key and an encrypted model to the at least one GPU; executing, using respective one or more accelerator cores of the at least one GPU, a trusted execution environment (TEE), wherein the at least one GPU generates an ephemeral session key for encrypting the at least one model encryption key; initiating a session in the TEE, wherein the encrypted model is decrypted during the session to generate a decrypted model; executing, during the session, the decrypted model in the TEE, to implement at least one machine learning task; and terminating, using the ephemeral session key by the at least one GPU, the session.

[0037] Some techniques described herein relate to an encryption device for model protection, the encryption device including one or more circuits configured to perform operations including: generating, by a GPU, at least one asymmetric encryption key; deriving, by a device identity engine, at least one GPU public key; encrypting during runtime, using the at least one GPU public key, a model encryption key; decrypting, using a private key associated with the at least one GPU public key, the model encryption key; storing a decrypted model in a memorylocation in the GPU; executing a request from a host device for access to the decrypted model stored in the memory location; denying the request from the host device for access to the decrypted model; and executing, using the decrypted model, an inference task in the GPU.

[0038] All combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are part of the inventive subject matter disclosed herein. The terminology used herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.BRIEF DESCRIPTIONS OF THE DRAWINGS

[0039] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0040] FIG. 1, reproduced below, shows an embodiment in which model encryption and decryption are performed using a system for Al model protection.

[0041] FIG. 2 is a flowchart of an encryption method for Al model protection according to an embodiment of the present invention.

[0042] FIG. 3 is a flowchart of a verification and decryption method for Al model protection according to an embodiment of the present invention.DETAILED DESCRIPTION

[0043] The following description and drawings are illustrative of the invention and are not to be construed as limiting the invention. The techniques disclosed herein relate generally to model encryption and decryption and, more specifically, to model encryption and decryption on artificial intelligence (Al) accelerators. The techniques disclosed herein avoid relying on the CPU or other external devices to decrypt the model before loading the model to the Alaccelerator. A private key provided or managed by the customer is fused into the Al accelerator, thereby allowing the Al accelerator to decrypt the model.

[0044] Referring to FIG. 1, a model encryption system 100 comprises at least three parties: an attested key service 101, a customer (or model owner) 102, a software or service provider running on a server 103 and Al accelerators 104. The attested key service 101 provides key identifiers, including public keys for attesting the identity of the accelerators. The software or service provider 103 is running on a server (i.e. host) which may be any kind of servers or a cluster of servers, such as web or cloud servers, application servers, backend servers or a combination thereof. Further, the server may be a cloud server or a server of a data center that provides a variety of cloud services to clients, such as, for example, cloud storage, cloud computing services, machine learning training services, data mining services, etc. For example, the software or service provider 103 may send or transmit an instruction (e.g. artificial intelligence (Al) training, inference instruction, etc.) for execution on a server. In response to the instruction, the server communicates with Al accelerators 104 to carry out execution of the instruction. Al accelerators 104 serve to carry out, for example, instructions related to machine learning, providing a faster mechanism for carrying out computation intensive tasks. The Al accelerators 104 may include dedicated processors to carry out such functionality.

[0045] Although the accelerators are described as Al accelerators, the accelerators 104 can represent accelerators including cryptography accelerators, compression accelerators, graphics accelerators, artificial intelligence and inference engines, smart network interface controllers (SmartNICs), and other custom or special-purpose circuitry implemented using field- programmable gate arrays (FPGAs), application- specific integrated circuits (ASICs), or other types of programmable or fixed function integrated circuits. The embodiments described herein refer to GPU, which serve as special purpose components representing the Al accelerators 104 and can be characterized as hardware accelerators. The special-purpose components (i.e. GPU) or accelerators 104 may be implemented using any suitable type and / or combination of circuitry and / or logic. The trusted execution environment can include at least one GPU, or some combination of CPUs, special purpose computing components such as accelerators 104 or GPUs and / or other processing components.

[0046] The model encryption system 100 depicts model encryption and decryption using customer private keys that are fused with Al accelerators 104. Model encryption system 100 includes the attested key service 101, the model owner (customer) 102 and the software orservice provider 103. The private keys associated with the Al accelerators 104 are inaccessible to the attested key service 101, the model owner 102 and the software 103. The Al accelerators 104 serve to provide authenticated encryption of the model data. The Al accelerators receives encrypted model data and further uses a key (i.e. private key) to decrypt, thus verifying the authenticity of the data. A customer looking to protect their model data, must ask for the public keys of the Al accelerators 104, which are furnished by the attested key service 101 after an attestation process. The model is encrypted by the customer with the public keys. The model is then sent to the service provider 103. The servers of the service provider load the encrypted model into the Al accelerators 104. The model is then decrypted by the Al accelerators 104 and placed in the private memories of the accelerators. Any attempts to access the model in the memory are protected by the Al accelerator, such as by using range checks or other mechanisms to protect the model from access.

[0047] FIG. 2 illustrates a flowchart for an example method 200 for an encryption method for Al model protection according to an embodiment of the present invention. In act 210, a public key and a private key are determined for at least one accelerator or a set of accelerators 104 (e.g. Al accelerators). The at least one public key is provided by the attested key service 101 after a standard attestation or verification process. In act 220, a symmetric encryption key is encrypted using the public key. In act 230, the encrypted symmetric encrypted key is transmitted to the accelerator 104. The public keys associated with at least one Al accelerator, or a set of Al accelerators 104 are provided by the attested key service 101 after a standard attestation or verification process. In act 240, the encrypted symmetric encryption key is decrypted at the accelerator, using the at least one private key. A dataset (i.e. the Al model) is encrypted using the at least one decrypted symmetric encryption key, to generate an encrypted dataset. The server 103 (or servers), which can be part of a service provider, including but not limited to, cloud environments, then loads the encrypted dataset into the at least one accelerator or a set of accelerators 104 for decryption. The encrypted dataset is then decrypted in the accelerator 104 and placed in the memory of the accelerator 104. Any attempts to access the dataset in the memory of the accelerator 104 are protected by the accelerator 104 taking measures using range checks or other standard mechanisms. In act 250, a dataset is encrypted using the at least one decrypted symmetric encryption key, to generate an encrypted dataset. In act 260, server determines the accelerator based at least in part on the at least one public key and the at least one private key. In act 270, the encrypted dataset is loaded by the server ontothe at least one accelerator 104. Thus, method 200 and the encryption method ensure customer data privacy by protecting the data. The dataset (e.g. Al model) is encrypted, and the encrypted dataset is assigned to an accelerator 104 for decryption and usage. In some embodiments, the private keys are wrapped with the encrypted dataset, and then provided to a host device.

[0048] FIG. 3 is a flowchart of a decryption method 300 for model protection according to an embodiment of the present invention. In act 310, a model encryption key (MEK) is determined, the MEK used for encrypting a model. The MEK is used for encrypting the Al model. Model encryption and decryption uses a customer private key which are fused with the accelerators 104. This way, accelerators 104 are used, instead of a CPU or other vulnerable devices for decrypting the model (i.e. Al model) for inference use. Each key can be associated with an accelerator or a set of accelerators 104. This allows for scenarios in which the model when decrypted is not visible or accessible to a host device, host memory, cloud software, the service provider 103 and external attackers. In act 320, the MEK and the encrypted model are transmitted to a host. As such, the model is encrypted while in the host device and can be decrypted with private keys that reside in an accelerator 104. In these cases, the host does not have access to the decrypted model, limiting exposure to cloud software, service providers, and external attackers to have access. In act 330, at least one asymmetric encryption key pair is generated. This asymmetric encryption key pair is generated at least by a GPU and includes at least a public key and a private key. In act 340, the private key is generated. This private key is associated with at least one Al accelerator 104. In act 350, at least one public key is determined to be associated with at least one Al accelerator 104. In act 360, the model is encrypted using the at least one public key. The encrypted model is transmitted to a server and an accelerator (i.e. Al accelerator 104) determined by the server 103, based at least in part on the private key and the public key. The server 103 loads the encrypted model into the at least one accelerator 104. In act 370, the encrypted model is decrypted to obtain a model for storage in the at least one accelerator 104. The private keys reside inside the accelerator hardware through soft or hard fuses. Alternatively, the private keys may reside in a protected memory or register. In an embodiment, a different key is provided for every accelerator; alternatively, the same key may be provided for a group of accelerators 104. As an alternative to using a private key, in some embodiments a secret is used to derive a key. In some embodiments, the private key is a symmetric key that is generated and owned by the customer or model owner(e.g., in the customer’s or model owner’s hardware security model (HSM)). The private secret is used to wrap the symmetric model key.

[0049] In some embodiments, the model is encrypted before being provided to the host. The model can only be decrypted with one or more private keys that resides in the Al accelerator. The host does not have access to the decrypted model. In this way, exposure and access of the model (e.g., to cloud software, service providers, and external attackers) is reduced or eliminated.

[0050] Some disclosed techniques include: using one or more private symmetric keys (or a secret used to derive a key) in an Al accelerator, fusing the private key(s) into the Al accelerator that allows the accelerator to decrypt customer loaded models. Customers can wrap their private key(s) for their model, encrypt the model, and provide the model to the host. The host then can load the model to the Al accelerator. The Al accelerator decrypts the model before loading the model into the memory of the Al accelerator. The Al accelerator is configured to perform model decryption because the Al accelerator is fused with the private key(s) from the customer. Once decrypted, the model (e.g., in plain text) can be consumed by the Al accelerator for usage, such as for Al inferencing.

[0051] Some disclosed techniques operate in processor systems, including multi-processor, hardware execution environments, software execution environments, CPUs, edge devices, machine learning applications, GPUs and Al Accelerators. In some embodiments, the processor systems include at least one CPU, at least one Al accelerator, at least one GPU, memory, at least one I / O device and at least one Trusted Execution Environment (TEE). A TEE is an environment in which tasks are executed and are expected to have high levels of trust in that surrounding environment and the capacity to avoid and / or ignore threats which may impact the environment. A TEE ensures that code and data are isolated and protected from unauthorized users, attackers, from outside environments and even within the surrounding environment, including but not limited to components such as the operating system. The environment may face security issues, vulnerabilities and thus relies on a TEE to safely execute the tasks in that environment. The TEE ensures that there is a higher level of trust for execution of code and / or tasks with respect to validity, isolation, and access of items accessible in this space, in comparison to an environment that is not a TEE. This serves to ensure that components executed within this TEE are also trustworthy, including but not limited to theoperating system and software applications. In some embodiments, the TEE includes at least one GPU and / or at least one Al accelerator.

[0052] In some disclosed techniques, the symmetric decryption key is provided to the GPU in a secure way, and the GPU decrypts the model, loaded in memory and provides runtime protections around confidentiality and integrity of the model. As the only trusted entity is the GPU, the GPU generates a set of asymmetric (public and private) encryption keys as part of the device identity engine where unique and per-device attestation keys are also derived (e.g., Device Identity Composition Engine (DICE)). The GPU exports the public encryption key as part of the attestation manifest, and the client then uses it to wrap the Customer Model Encryption Key (MEK). Only the GPU which provides the attestation manifest is able to unwrap the MEK (using the corresponding asymmetric private decryption key). Before unwrapping the MEK, the GPU launches a Trusted Execution Environment (TEE) in which all the memory used by the cores is encrypted by a TEE session key so that even when unwrapped, no one can see the MEK in plaintext or access or modify it. With the MEK, the GPU can then decrypt the model as it is read from the CPU memory / storage (and transferred in encrypted form via PCIe down to the GPU) and load it in GPU memory. In this embodiment, the model is never accessible in plain text outside the GPU TEE.Security Considerations

[0053] In various embodiments, the systems and methods utilize one or more of the following security considerations.

[0054] First, the GPU is preferably configured to generate asymmetric encryption keys rooted in hardware to support MEK wrapping. The customer encrypts the MEK with a GPU Public Key derived from the device-unique identity engine (DICE). The GPU decrypts the MEK using its Private Key counterpart and does not release this key.

[0055] Second, only the intended GPU can decrypt the Customer Encryption Key. The Customer Encryption Key was wrapped with a per-device Public Key. No other GPU is able to impersonate the attested GPU and gain access to the MEK by having the same Private Key that decrypts the MEK or by based on the attestation report from the intended GPU.

[0056] Third, the GPU is preferably configured to cryptographically demonstrate the model was loaded without any modification. This can be done with a signed report of the modelmeasurement / hash. The GPU should also prevent any modification of memory where the loaded model resides until the workload is torn down or until next reset.

[0057] Fourth, during runtime, the model is encrypted in memory and access is protected by the GPU. No entity local to the host or remote (network) is able to infer or extract the model or parts of it in plaintext.Conclusion

[0058] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.

[0059] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.

[0060] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0061] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0062] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0063] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0064] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of Aand B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements) etc.

[0065] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

CLAIMS1. A method for encrypting data in a cloud environment, the method comprising: determining at least one public key and at least one private key associated with at least one accelerator in the cloud environment; encrypting, using the at least one public key, at least one symmetric encryption key to generate at least one encrypted symmetric encryption key; transmitting the at least one encrypted symmetric encryption key to the at least one accelerator; decrypting, at the at least one accelerator, the at least one encrypted symmetric encryption key using the at least one private key, to generate at least one decrypted symmetric encryption key; encrypting a dataset using the at least one decrypted symmetric encryption key, to generate an encrypted dataset; determining, by a server, the at least one accelerator based at least in part on the at least one public key and the at least one private key; and loading, by the server, the encrypted dataset into the at least one accelerator.

2. The method of claim 1, wherein the at least one accelerator executes machine learning tasks and Al functions.

3. The method of claim 1, wherein the at least one public key is determined based at least in part on an attestation process.

4. The method of claim 2, further comprising: storing, the at least one private key in a memory location of the at least one accelerator.

5. The method of claim 1, wherein the dataset includes at least one Al model.

6. The method of claim 1, wherein the at least one private key is fused on the at least one accelerator.

7. The method of claim 5, further comprising implementing, using at least the at least one Al model, at least one machine learning task.

8. The method of claim 7, wherein the at least one machine learning task comprises inferencing.

9. A method for decrypting data, the method comprising: generating, for a GPU, at least one asymmetric encryption key pair, the at least one asymmetric encryption key pair including a public key and a private key; determining, based at least in part on a verification, that the public key is associated with at least one accelerator; encrypting, using at least the public key a model to generate an encrypted model; determining, by a host, the at least one accelerator based at least in part on the private key and the public key; loading, by the host, the encrypted model into the at least one accelerator; and decrypting the encrypted model to generate a decrypted model for storage in the at least one accelerator.

10. The method of claim 9, wherein the GPU and the at least one accelerator are part of a trusted execution environment.

11. A method for decrypting data, the method comprising: determining, based at least in part on a verification, at least one public key associated with a trusted execution environment, wherein the trusted execution environment includes at least one trusted computing component; encrypting, using the at least one public key, at least one model to generate at least one encrypted model; transmitting, the at least one encrypted model to a server; loading, in the trusted execution environment, the at least one encrypted model; and decrypting the at least one encrypted model in the trusted execution environment, to generate at least one decrypted model.

12. The method of claim 11, wherein the trusted execution environment is included in a GPU.

13. The method of claim 12, wherein a memory of the GPU stores the at least one decrypted model, wherein the memory is within a trust boundary associated with the trusted execution environment.

14. A method for encrypting a model, the method comprising: encrypting, using at least one model encryption key, the model; transmitting, to a host device, the model and the at least one model encryption key; generating, by a device core layer, an asymmetric encryption key pair comprising a private portion and a public portion; wrapping the at least one model encryption key, using at least the public portion of the asymmetric encryption key pair; and storing the private portion of the asymmetric encryption key pair and the at least one model encryption key in a GPU, wherein the GPU unwraps the at least one model encryption key using the private portion of the asymmetric encryption key pair.

15. The method of claim 14, wherein the GPU is part of a trusted execution environment.

16. The method of claim 15, further comprising: decrypting the model to generate a decrypted model; and storing the decrypted model in a memory of the GPU, wherein the memory is within a trust boundary associated with the trusted execution environment.

17. A non-transitory machine-readable medium having instructions stored thereon, which when executed by a processor, cause the processor to perform operations, the operations comprising: encrypting, using at least one model encryption key, a model; transmitting, through at least one bus, the model and the at least one model encryption key, to a host device; generating, by a device core layer, an asymmetric encryption key comprising a private portion and a public portion; wrapping the at least one model encryption key using at least the public portion; and storing the private portion and the at least one model encryption key in a GPU, wherein the GPU unwraps the at least one model encryption key using the private portion.

18. The non-transitory machine-readable medium of claim 17, wherein the GPU is part of a trusted execution environment.

19. A system comprising one or more circuits, the one or more circuits comprising at least one CPU, at least one GPU, and at least one communication bus, the one or more circuits configured to perform operations comprising: transmitting, through the at least one communication bus, at least one model encryption key and an encrypted model to the at least one GPU; executing, using respective one or more accelerator cores of the at least one GPU, a trusted execution environment (TEE), wherein the at least one GPU generates an ephemeral session key for encrypting the at least one model encryption key; initiating a session in the TEE, wherein the encrypted model is decrypted during the session to generate a decrypted model; executing, during the session, the decrypted model in the TEE, to implement at least one machine learning task; and terminating, using the ephemeral session key by the at least one GPU, the session.

20. An encryption device for model protection, the encryption device including one or more circuits configured to perform operations comprising: generating, by a GPU, at least one asymmetric encryption key; deriving, by a device identity engine, at least one GPU public key; encrypting during runtime, using the at least one GPU public key, a model encryption key; decrypting, using a private key associated with the at least one GPU public key, the model encryption key; storing a decrypted model in a memory location in the GPU; executing a request from a host device for access to the decrypted model stored in the memory location; denying the request from the host device for access to the decrypted model; and executing, using the decrypted model, an inference task in the GPU.

Citation Information

Patent Citations

  • Confidential verification of FPGA code

    US20180270068A1

  • Processing of reduction and broadcast operations on large datasets with mutli-dimensional hardware accelerators

    US20220292399A1

  • System and method for transferring data from non-volatile memory to a process accelerator

    US20220413732A1

  • Storage device including storage controller and operating method

    US20230135891A1

Cited By

  • AI model edge device binding protection method based on hardware security module

    CN120658440A