Chipset, system and method for confidential computing
The chipset offloads decryption and encryption to generic processors, addressing inefficiencies in confidential AI/ML computing by enhancing confidentiality and efficiency through secure in-model operations, suitable for various platforms without additional hardware.
Patent Information
- Application Number
- PCT/EP2024/050806
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-24
AI Technical Summary
State-of-the-art solutions for confidential AI/ML computing lack secure contractual agreements between mutually distrusting parties, are platform-specific, and rely on costly cryptographic operations, making them inefficient for secure data and model protection.
A chipset with on-chip generic processors and application-specific processors offloads decryption and encryption operations to the generic processors, enabling confidential computing without additional hardware, using authenticated key exchanges and in-model cryptographic operations.
Enhances confidentiality and efficiency by freeing up application-specific processors for AI/ML tasks, ensuring secure data and model protection without additional hardware or software stacks, and preventing micro-architectural leakage.
Smart Images

Figure EP2024050806_24072025_PF_FP_ABST
Abstract
Description
[0001] CHIPSET, SYSTEM AND METHOD FOR CONFIDENTIAL COMPUTING
[0002] TECHNICAL FIELD
[0003] The present disclosure relates to the field of computer technology. For instance, the disclosure relates to a chipset, system, and method for confidential computing.
[0004] BACKGROUND
[0005] Traditionally, confidential computing is a collection of cryptographic protocols that are held as the gold standard for computing data without revealing any information about the data without help from a trusted third party. There are many cryptographic primitives to realize confidential computing: Yao's garbled circuit, oblivious transfer, zero-knowledge proof, etc. Eventhough these primitives provide provable security guarantees, realizing them even on the modem platform is prohibitively expensive due to the requirement on computation, memory, and amount of message exchange. Trusted execution environments (TEEs) strike a balance by realizing confidential computing by having a trade-off in trust assumption. Silicon root of trust (S-RoT) may be used to integrate security directly into the hardware level.
[0006] TEEs such as Intel SGX, AMD SEV, ARM TrustZone, RISC-V keystone, etc., drastically reduce the trusted computing base (TCB) and provide security to applications, known as enclaves, without having to trust the operating system and hypervisor. Thus, the attack surface is reduced by eliminating two of the largest sources of vulnerabilities for a system. TEEs use isolation primitives provided by an OS CPU (a CPU running operating system (OS)) to exclude all software, but a single target application (commonly known as an enclave) from the software trusted computing base (TCB). Only the OS CPU is part of the hardware TCB, while the remaining hardware in the system is considered untrusted, as can be seen in the figure above. Even memory is not included in the TCB and can only be used in conjunction with memory encryption and integrity protection. Such a trust model makes the TEEs ideal candidates for the aforementioned trusted path applications where the software stack is attacker-controlled.
[0007] Apart from the isolation of enclaves from the untrusted OS, another crucial security primitive of TEEs is remote attestation. A remote verifier can ensure that a (legitimate) platform is running an enclave with the correct code. In a simplified form, at the end of a remote attestation process, the platform generally proves to a remote verifier that it is running the proper enclave with proper firmware. It is noted that there exist some differences between remote attestation mechanisms in different TEE technologies. However, the remote attestation mechanisms across different platforms are comparable. In Intel SGX, the attestation report also contains CPU version number (SVN) that denotes the CPU microcode version, and provides the remote verifier with the public key of the enclave. This public key could be later used by the remote verifier to establish a secure channel to provision secrets to the enclave.
[0008] Machine learning (ML) or artificial intelligence (Al) recently has become a significant workload on both consumer product and high-performance computing (HPC) servers. Typically, these workloads are executing on-domain specific accelerators such as Al accelerators, GPUs, FPGAs, etc. Confidential computing and confidential AI / ML are a major offering from top cloud providers in order to offer hardware-enforced security against compromised software stack and minimize the trust assumption.
[0009] SUMMARY
[0010] There are various scenarios for confidential AI / ML training and inference, which are given below as examples:
[0011] Outsourced training: This is a scenario where the model and data provider do not have sufficient computing resources to execute the model. Therefore, the model and data provider offload their model and data to an untrusted cloud provider. The objective of this scenario is to train the model without leaking the model to the cloud provider / data owner and keep the data confidential from the cloud provider and the model owner. Al as a service: This is the first-party cloud scenario where the cloud provider renders the model and associated software stack for inference as well as the computing resource. Therefore the cloud provider is also the model owner. The data provider sends confidential inference data to the cloud for inference and gets back the inference result. The objective of this scenario is to keep the inference data confidential from the cloud provider.
[0012] Confidential training and fine-tuning: Here the model provider sends the base model along with the training / fine- tuning framework. The data provider has additional data to fine-tune the model with. This fine-tuning could be done either at the data provider’s computing resources or the cloud if the computing requirement is too high. After the fine-tuning, the data provider gets the fine-tuned model. The objective of this scenario is to keep the base model confidential from the data owner and the cloud provider and keep the additional data confidential from the model provider and the cloud.
[0013] For a cloud Al service provider, it is conventional to use a solution based on standard trusted execution environments (TEEs) such as intel SGX or AMD SEV to create a secure channel from clients to TEE enclaves. All the computations are executed inside a TEE enclave to ensure that the integrity and confidentiality of data and ML model is preserved from the untrusted operating system and hypervisor.
[0014] However, state-of-the-art solutions do not provide a secure contractual agreement between a model provider and data provider while maintaining the privacy of both the data and the model. Further, the state-of-the-art solutions are specific to certain platforms (CPU and Al accelerator) and may be dependent on specific hardware implementations.
[0015] There are three parties that are mutually distrusting: the model provider; the data provider; the infrastructure provider. The model provider may desire to send an encrypted model to the Al infrastructure. The data provider may desire to send encrypted data to the Al infrastructure provider. It can be assumed that all the software stack, i.e., the OS, hypervisor, and other applications are untrusted. All the hardware, except the CPU cores where the application is deployed and specific Al accelerator, are untrusted. There is a CPU TEE where the user can deploy workload. Both the CPU and the Al accelerator have their individual root of trust to carry out attestation, key exchange, and establish a secure channel.
[0016] In view of the above-mentioned problems and disadvantages, this disclosure aims to improve confidential computing for ALML tasks. For instance, an objective of this disclosure may be to protect the model and the data confidential from the untrusted software stack and cloud provider, where the data and model providers are mutually distrusting parties.
[0017] A first aspect of this disclosure provides a chipset for confidential computing. The chipset comprises one or more on-chip generic processors, and a plurality of application-specific processors for machine learning.
[0018] The one or more generic processors are configured to obtain encrypted input data of a machine learning model, decrypt the encrypted input data using a first secret to obtain decrypted input data, and provide the decrypted input data to the applicationspecific processors.
[0019] The application-specific processors are configured to obtain the decrypted input data, and provide the decrypted input data to the machine learning model for inferencing to obtain an output.
[0020] Optionally, the chipset may be at least a part of an application-specific hardware adapted to accelerate Al and ML applications. For instance, the chipset may be at least a part of an Al accelerator or neural processing unit (NPU).
[0021] The one or more generic processors are on-chip, such that they are adapted to assist the application-specific processors. The one or more on-chip generic processors comprised in the chipset are not adapted to host an operating system or the like. Instead, the one or more on-chip generic processors may be adapted to, among other purposes, perform data decryption and / or encryption during model inferencing. For instance, the one or more on-chip generic processors may be built based on ARM architecture or reduced instruction set computing (RISC) architecture.
[0022] In an implementation form of the first aspect, the one or more generic processors may be configured to encrypt the output using the first secret to obtain an encrypted output.
[0023] In a further implementation form of the first aspect, the machine learning model comprises an input layer, a plurality of hidden layers, and an output layer. The application-specific processors may be configured to perform machine learning operations at the input layer, the hidden layers, and the output layer. The one or more generic processors may be configured to perform the decryption of the input data between operations of the application-specific processors at the input layer and at the hidden layers. Optionally, when the encryption of the output is performed, the one or more generic processors may be configured to perform the encryption of the output between between operations of the application-specific processors at the hidden layers and at the output layer.
[0024] In a further implementation form of the first aspect, the chipset may further comprise a memory unit connected with the one or more generic processors and the application-specific processors. The machine learning model comprises an input layer, a plurality of hidden layers, and an output layer. The application-specific processors may be configured to store the encrypted input data in the memory unit after feeding the encrypted input data to the input layer. The one or more generic processors may be configured to decrypt the stored encrypted input data, and store the decrypted input data in the memory unit. The applicationspecific processors may be configured to obtain the decrypted input data from the memory unit, feed the decrypted input data to the hidden layers to obtain the output from the last layer of the hidden layers, and store the output in the memory unit.
[0025] In a further implementation form of the first aspect, the one or more generic processors may be configured to obtain the output from the memory unit, encrypt the output, and store the encrypted output in the memory unit.
[0026] In a further implementation form of the first aspect, the one or more generic processors may be configured to perform a first authenticated key exchange with a provider device of (or, which provides) the input data to obtain the first secret.
[0027] Optionally, the first authenticated key exchange may be a Diffie-Hellman key exchange.
[0028] In a further implementation form of the first aspect, the one or more generic processors may be configured to use Advanced Encryption Standard with Galois / Counter Mode (AES-GCM) to decrypt the encrypted input data. Optionally, when the output is encrypted, the AES-GCM may be used by the one or more generic processors.
[0029] In a further implementation form of the first aspect, the machine learning model comprises a set of encrypted weights. The one or more generic processors may be configured to decrypt the encrypted weights using a second secret, and provide the decrypted weights to the application-specific processors. The application-specific processors may be configured to perform machine learning inference using the decrypted weights.
[0030] In a further implementation form of the first aspect, the one or more generic processors may be further configured to validate the decrypted based on Secure Hash Algorithm 3 (SHA-3). In a further implementation form of the first aspect, the one or more generic processors may be configured to perform a second authenticated key exchange with an owner device of (or, which provides / owns) the machine learning model to obtain the second secret.
[0031] Optionally, the second authenticated key exchange may be a Diffie-Hellman key exchange.
[0032] In a further implementation form of the first aspect, after obtaining the output, the one or more generic processors may be configured to delete the input data and / or the output and / or the machine learning model.
[0033] Optionally, memory and / or cache that are used to store the input data and / or the output may be zero-filled. In this way, micro- architectural leakage can be prevented.
[0034] A second aspect of this disclosure provides a system (or an apparatus) comprising one or more chipsets according to the first aspect or any implementation form thereof.
[0035] Optionally, the system (or the apparatus) may be an Al accelerator.
[0036] A third aspect of this disclosure provides a method for confidential computing. The method is performed by a chipset comprising one or more on-chip generic processors, and a plurality of application-specific processors for machine learning. The method comprises the following steps: obtaining, by the one or more generic processors, encrypted input data of a machine learning model; decrypting, by the one or more generic processors, the encrypted input data using a first secret to obtain decrypted input data; providing, by the one or more generic processors, the decrypted input data to the application-specific processors; obtaining, by the application-specific processors, the decrypted input data; and providing, by the application-specific processors, the decrypted input data to the machine learning model for inferencing to obtain an output.
[0037] The method of the third aspect may share the same optional features and advantages as the chipset of the first aspect.
[0038] A fourth aspect of this disclosure provides a computer program comprising instructions which, when the program is executed by a computer, cause the computer to cany out the method according to the third aspect.
[0039] A fifth aspect of the present disclosure provides a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to cany out the method according to the third aspect.
[0040] It has to be noted that all elements, units, and means described in the present disclosure could be implemented in software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity, which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. BRIEF DESCRIPTION OF DRAWINGS
[0041] The above-described aspects and implementation forms will be explained in the following description in relation to the enclosed drawings, in which
[0042] FIG. 1 shows an example of a chipset according to this disclosure;
[0043] FIG. 2 shows a further example of a chipset according to this disclosure;
[0044] FIG. 3 shows a schematic illustration of a machine learning model;
[0045] FIG. 4 shows an example of protecting a machine learning model according to this disclosure; and
[0046] FIG. 5 shows an application scenario of this disclosure;
[0047] FIG. 6 shows an example of key exchanges between a model provider and an Al infrastructure provider, and between a model user and the Al infrastructure provider; and
[0048] FIG. 7 shows an application scenario of this disclosure.
[0049] DETAILED DESCRIPTION OF EMBODIMENTS
[0050] In FIGs. 1-7 below, corresponding elements may share the same features and function likewise.
[0051] FIG. 1 shows an example of a chipset 100 according to this disclosure. The chipset 100 comprises at least one on-chip generic processor 110, and a plurality of application-specific processors 120.
[0052] The chipset 100 may be at least a part of an Al accelerator. The Al accelerator may be an application-specific hardware adapted to accelerate Al and ML applications. For instance, The Al accelerator may be a Neural Processing Unit (NPU), Tensor Processing Unit (TPU), or a graphics processing unit (GPU). The plurality of application-specific processors 120 may be referred to as a collection of processing resources that are specifically adapted to perform AI / ML-related operations (e.g., vector / matrix computation, convolution, and the like). For instance, the plurality of application-specific processors 120 may comprise components like parallel Arithmetic and Logical Units (ALUs), caches, floating-point units or vector processors, registers, and memory to store thread information. The at least on-chip generic processor 110 is adapted to assist the applicationspecific processors 120. For instance, the on-chip generic processors 110 may be configured to perform some operations that are not suitable to be executed by the application-specific processors 120. In this disclosure, the at least one on-chip generic processor 110 may also be referred to as a generic core 110, and the plurality of application-specific processors 120 may also be referred to as Al cores 120. For instance, the generic core 110 may be an ARM core 110, and an Al core 120 may be a DaVinci core (when an Al accelerator comprising the chipset is based on Huawei Ascend DaVinci Architecture). It is also noted that the Al core may be referred to some other names when different architectures are used. For instance, the Al core may also be a Compute Unified Device Architecture (CUD A) core, or a Graphics Core Next (GCN) compute unit (CU).
[0053] The generic core 110 is configured to obtain encrypted input data of an ML model, decrypt the encrypted input data using a first secret to obtain decrypted input data, and provide the decrypted input data to the Al cores 120.
[0054] These operations may be performed during an inference phase. During the inference phase, the Al cores 120 are configured to receive the encrypted input data using an input layer of the ML model. Then, the generic core 110 is configured to perform inlayer (or in-model) decryption of the encrypted input data and provide the decrypted input data to the Al cores 120. That is, the decryption of the input data is performed between the input layer and hidden layers of the ML model. After obtaining the decrypted input data, the Al cores 120 are configured to provide the decrypted input data to the ML model (e.g., the hidden layers of the ML model) for inferencing to obtain an output. Optionally, AES-GCM may be used to decrypt the encrypted input data. AES-GCM offers authenticated decryption, which means that apart from the decryption, it offers message authentication code (MAC) generation and verification that offer data integrity and authenticity.
[0055] By performing the inlayer decryption, the ML model can be self-sufficient to protect itself and the input data. By offloading the decryption operation from the Al cores 120 to the generic core 110, the heterogenous nature of the chipset for an Al accelerator is leveraged. This solution does not require additional hardware and software stack to implement the crypto and security-related operations.
[0056] By using the generic core 110 to perform the decryption, crypto and security-related operations can be offloaded from Al cores 120. In this way, Al compute resources can be freed up to maximize utilization and efficiency by leveraging the heterogeneous nature of the chipset for Al accelerators such as NPU and GPU.
[0057] Optionally, the generic core 110 may be configured to encrypt the output by using the first secret to obtain an encrypted output. Optionally, AES-GCM may be used to encrypt the output. AES-GCM offers authenticated encryption, which means that apart from the encryption, it offers message authentication code (MAC) generation and verification that offer data integrity and authenticity.
[0058] It is noted that the encryption of the output is also an inlayer (or in-model) encryption. That is, the encryption of the output is performed between the hidden layers and the output layer.
[0059] The first secret is shared between the chipset and an owner device that provides the input data. The owner device may be configured to decrypt the encrypted output using the first secret. In this way, the confidentiality can be further enhanced.
[0060] The generic core 110 and the Al cores 120 may be connected via a bus connection. For example, the generic core 110 and the Al cores 120 may be connected via Peripheral Component Interconnect Express (PCI-e), or CHIE interconnect.
[0061] It is noted that the chipset 100 shown in FIG. 1 is schematically shown for illustration purposes only. Other components are not shown for the sake of simplicity.
[0062] FIG. 2 shows a further example of a chipset 100 according to this disclosure. Based on the chipset 100 in FIG. 1, the chipset 100 may further comprise a memory unit 210 shown in FIG. 2. For instance, the memory unit 210 may be a High Bandwidth Memory (HBM).
[0063] The memory unit 210, the generic core 110, and the Al cores 120 may be connected via bus connection, e.g., PCI-e, or CHIE interconnect.
[0064] During inferencing, the Al cores 120 may be configured to store the encrypted input data in the memory unit 210 after feeding the encrypted input data to the input layer. The generic core 110 is then configured to decrypt the stored encrypted input data, and store the decrypted input data in the memory unit 210. The Al cores 120 is then configured to obtain the decrypted input data from the memory unit 210, provide the decrypted input data to the hidden layers for inferencing, obtain the output from the last layer of the hidden layers, and store the output in the memory unit 210.
[0065] The encryption of the output is optional and can be configurable. If the encryption of the output is enabled, the generic core 110 is configured to encrypt the output and store the encrypted output in the memory unit 210. Then, the Al cores 120 are configured to obtain the encrypted output from the memory unit 210 and provide the encrypted output to the output layer for final processing. If the encryption of the output is disabled, then the Al cores 120 are configured to provide the unencrypted output to the output layer.
[0066] FIG. 3 shows a schematic illustration of a machine learning model. As depicted in FIG. 3, on the left-hand side of the arrow, a conventional ML model is shown, which comprises an input layer, multiple hidden layers, and an output layer. On the right- hand side of the arrow, an improved ML model according to this disclosure is shown for confidential computing (inferencing). Apart from the input layer 310, hidden layers 330, and the output layer 350, the improved ML model further comprises a decryption layer 320 between the input layer 310 and the hidden layers 330, and an encryption layer 340 between the hidden layers 330 and the output layer 350. In the decryption layer 320, decryption of input data is performed. In the encryption layer 340, encryption of output is performed. ML-related operations at the input layer 310, hidden layers 330, and the output layer 350 are performed by the Al cores 120. Decryption at the decryption layer 320 and / or encryption at the encryption layer 340 are performed by the generic core 110.
[0067] As an example, AES-GSM operators may be inserted between the input layer 310 and the hidden layers 330, and between the hidden layers 330 and the output layer 350, respectively, which are configured to perform decryption and encryption, respectively. The AES-GSM operators are invoked by the generic core 110.
[0068] It is noted that the decryption and encryption operations are configurable and can be switched on / off depending on various needs. That is, either one, or both of the decryption and encryption operations can be activated.
[0069] When both the decryption layer 320 and the encryption layer 340 are activated, it can be seen that the initial input data is encrypted (even when the inferencing starts), and the final output is also encrypted (before the inferencing finishes). The decrypted data is processed during the ML model inferencing at the hidden layers 330. In this way, confidentiality is improved for the ML model inferencing.
[0070] The chipset is normally seen as a TEE environment. By placing the decryption and encryption into the ML model, the difficulty of data hacking is increased. Therefore, the confidentiality of the input data and / or output data during ML model inferencing can be further improved.
[0071] FIG. 4 shows an example of protecting ML model according to this disclosure. On the left-hand side of FIG. 4, a structure of an ML model is shown, which comprises information fields such as header, input variables, layers, weights, and binary fields. The weights comprise model parameters. Typically, the model parameters are generated through extensive training and fine- tuning. A model owner provides this kind of information to an infrastructure provider such that the ML model can be instantiated for inferencing. At the same time, the model owner may desire to protect the ML model. Because letting any untrusted third party have access to the weights may harm the interest of the model owner. One solution is to send the model weights encrypted to the Al cores, and the Al cores can decrypt the model weights before execution. However, the model weights are often very large and thus, require extensive computation to decrypt them. This disclosure provides a solution for protecting the ML model during inferencing. Similar to the cryptographic operation on the data, an in-model encryption is used to encrypt the model weights. The encrypted model weights is loaded by the Al cores 120. During the execution (inferencing), the generic core 110 is configured to obtain a pointer to the encrypted model weights, and perform decryption on the encrypted model weights to obtain decrypted model weights. Then, the generic core 110 is configured to return the decrypted model weights to the Al cores 120 for further execution. Optionally, AES-GCM may be used to perform the encryption and decryption of the model weights. Optionally, SHA-3 may be used for hashing the model weights, such that the integrity of the model can be ensured after the decryption. For instance, the hash of the model weights after the model is deployed can be sent to both the model owner and input data provider to ensure that the integrity of the model is preserved. This solution requires no additional hardware modification and therefore, can be implemented on any existing chipset of an Al accelerator.
[0072] Optionally, the generic core 110 and the Al cores 120 may be further configured to delete the input data and / or the output. For instance, the generic core 110 and the Al cores 120 may be further configured to flash cache and zero-fill memory that are used during the model inferencing. When the model is no longer used, the Al cores may also be configured to delete the model weights (e.g., by zero-filling a corresponding memory area that stores the model weights). In this way, micro-architectural leakage to the attacker may be prevented.
[0073] FIG. 5 shows an application scenario of this disclosure. As depicted in FIG. 5, the decryption of the input data, encryption of the output data, decryption of the model weights, and cleanup can be individually activated, or may be combined to achieve different levels of confidentiality.
[0074] FIG. 6 shows an example of key exchanges between a model provider and an Al infrastructure provider, and between a model user and the Al infrastructure provider. The Al infrastructure provider may use one or more Al accelerators to host AI / ML models and provide services to users.
[0075] For obtaining the first secret that is used to decrypt the input data, the generic core 110 may be configured to perform a first authenticated key exchange 610 with a provider device that provides the input data (also referred to as an input provider). Optionally, the first authenticated key exchange may be a Diffie-Hellman key exchange.
[0076] When the encryption of the output is enabled, the chipset 100 may be configured to provide the encrypted output 630 to the input provider.
[0077] For obtaining the second secret that is used to decrypt the model weights, the generic core 110 may be configured to perform a second authenticated key exchange 620 with a provider device that provides the model (also referred to as a model provider). Optionally, the second authenticated key exchange may be a Diffie-Hellman key exchange.
[0078] FIG. 7 shows a method 700 for confidential computing of this disclosure. The method 700 is performed by a chipset comprising one or more on-chip generic processors 110, and a plurality of application-specific processors 120 for machine learning. The method comprises the following steps:
[0079] Step 701 : obtaining, by the one or more generic processors 110, encrypted input data of a machine learning model;
[0080] Step 702: decrypting, by the one or more generic processors 110, the encrypted input data using a first secret to obtain decrypted input data;
[0081] Step 703 : providing, by the one or more generic processors 110, the decrypted input data to the application-specific processors; Step 704: obtaining, by the application-specific processors 120, the decrypted input data; and
[0082] Step 705: providing, by the application-specific processors 120, the decrypted input data to the machine learning model for inferencing to obtain an output.
[0083] It is noted that before step 701 , the encrypted input data is already loaded into the input layer of the machine learning model, which is executed by the application-specific processors 120. The one or more generic processors 110 is configured to obtain the encrypted input data as an output of the input layer, decrypt the encrypted input data, and provide the decrypted input data to the hidden layers (e.g., the first layer of the hidden layers).
[0084] The method of FIG. 7 may share the corresponding features mentioned above with respect to FIG. 1-6, which are not repeated herein.
[0085] In summary, the present disclosure provides a solution for confidential computing where an ML model is used for inferencing. For instance, a computing system / device (e.g., Al accelerator) is proposed to establish confidential AI / ML infrastructure to perform the operation with encrypted data and models without depending on dedicated cryptographic hardware. An example of a workflow may be as follows:
[0086] 1. A model provider and a service provider that employs at least one Al accelerator (comprising the chipset mentioned above) perform an authenticated Diffie-Hellman key exchange to establish a shared secret (e.g., as illustrated in FIG. 6).
[0087] 2. Similarly, an input data owner and the Al accelerator establish a shared secret by executing authenticated Diffie- Hellman (e.g., as illustrated in FIG. 6).
[0088] 3. The model owner adds additional layers such as model decryption (and optional verification), input data decryption (and optional verification), result encryption (and optional tag calculation), and data clean-up to the original model (e.g., as illustrated in FIG. 5).
[0089] 4. The model owner encrypts the model weights with the shared secret between the model owner and the Al accelerator and sends them to the Al accelerator.
[0090] 5. The input data owner encrypts the input data with the shared secret between the input data owner and the Al accelerator.
[0091] 6. The Al accelerator loads the model and starts executing (inferencing). The Al accelerator decrypts the model weights with the shared key between the model provider and the Al accelerator. Optionally, the Al accelerator may verify the decrypted model weights between the model owner and the Al accelerator.
[0092] 7. Upon receiving the encrypted input data at the input layer of the model, the Al accelerator decrypts the encrypted input data and may optionally verify the decrypted data between the input data provider and the Al accelerator.
[0093] 8. The Al accelerator continues executing the ML model with the decrypted model weights and the input data.
[0094] 9. After all the hidden layers of the ML model are executed, the added layer encrypts the output results with the shared secret between the input data provider and the Al accelerator.
[0095] 10. After the model execution is done, the hardware executes the next added layer, i.e., it cleans up all the microarchitectural state and memories associated with the generic core 110 and the Al cores 120 such as flushing the caches, and zero-filling the memory.
[0096] An application scenario of this disclosure may be confidential model inferencing. For instance, a data provider (e.g., a hospital providing patient data) may need to use an AI / ML model to analyze sensitive data. An AI / ML model provider may provide a trained AI / ML model to an infrastructure provider such that the trained AI / ML model can be executed. The infrastructure provider may also be a cloud Al service provider that offers access to the data provider for using the AI / ML model. The data provider can provide the encrypted input data to the infrastructure provider for inferencing. The data exchange between the data provider and the infrastructure provider, and between the model provider and the infrastructure provider are encrypted. Moreover, during the inferencing, the input and output of the model are also encrypted. Therefore, model and data security can be ensured.
[0097] According to this disclosure, crypto and security -related operations is offloaded to the generic core 110, such that the Al cores 120 may focus on executing Al-related operations. Moreover, in-model operators are used for executing crypto and security- related operations (e.g., decryption and encryption). In this way, the model itself is self-sufficient to protect itself and the data. No additional hardware and software stack is needed to implement the crypto and security-related operations. Further, by cleaning up memory and cache, it can be ensured that no model and data leakage after the model inferencing is over. By encrypting the model parameters and only decrypting them during model inferencing, it can be ensured that the model parameters stay encrypted outside the chipset 100 (of the Al accelerator) and is only decrypted inside the chipset 100. Therefore, model parameter leakage can be prevented.
[0098] The solution of this disclosure can be applied to any Al accelerator that comprises at least one generic processor and a plurality of Al processors. No additional hardware modification is needed.
[0099] It is noted that the elements in the present disclosure may comprise processing circuitry configured to perform, conduct or initiate the various operations of the elements described herein, respectively. The processing circuitry may comprise hardware and software. The hardware may comprise analog circuitry or digital circuitry, or both analog and digital circuitry. The digital circuitry may comprise components such as application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), digital signal processors (DSPs), or multi-purpose processors. Optionally, the processing circuitry comprises one or more processors and a non-transitory memory connected to the one or more processors. The non-transitory memory may carry executable program code which, when executed by the one or more processors, causes the device to perform, conduct or initiate the operations or methods described herein, respectively.
[0100] The present disclosure has been described in conjunction with various aspects as examples as well as implementations. However, other variations can be understood and effected by those persons skilled in the art and practicing the claimed subject matter, from the studies of the drawings, this disclosure and the independent claims. In the claims as well as in the description the word “comprising” does not exclude other elements or steps and the indefinite article “a” or “an” does not exclude a plurality. A single element or another unit may fulfill the functions of several entities or items recited in the claims. The mere fact that certain measures are recited in the mutual different dependent claims does not indicate that a combination of these measures cannot be used in an advantageous implementation.
Claims
CLAIMS1. A chipset (100) for confidential computing, the chipset (100) comprising one or more on-chip generic processors (110), and a plurality of application-specific processors (120) for machine learning, wherein the one or more generic processors (110) are configured to: obtain encrypted input data of a machine learning model, decrypt the encrypted input data using a first secret to obtain decrypted input data, and provide the decrypted input data to the application-specific processors (120); wherein the application-specific processors (120) are configured to: obtain the decrypted input data; and provide the decrypted input data to the machine learning model for inferencing to obtain an output.
2. The chipset (100) according to claim 1, wherein the one or more generic processors (110) are configured to encrypt the output using the first secret to obtain an encrypted output.
3. The chipset (100) according to claim 1 or 2, wherein the machine learning model comprises an input layer (310), a plurality of hidden layers (330), and an output layer (350), wherein the application-specific processors (120) are configured to perform machine learning operations at the input layer (310), the hidden layers (330), and the output layer (350); and the one or more generic processors (110) are configured to perform the decryption of the input data between the operations of the input layer (310) and of the hidden layers (330), and perform the encryption of the output between the operations of the hidden layers (330) and the output layer (350).
4. The chipset (100) according to claim 1 or 2, wherein the chipset (100) further comprises a memory unit (210) connected with the one or more generic processors (110) and the application-specific processors (120), and the machine learning model comprises an input layer (310), a plurality of hidden layers (330), and an output layer (350), wherein the application-specific processors (120) are configured to store the encrypted input data in the memory unit (210) after feeding the encrypted input data to the input layer (310); the one or more generic processors (110) are configured to decrypt the stored encrypted input data, and store the decrypted input data in the memory unit (210); and the application-specific processors (120) are configured to obtain the decrypted input data from the memory unit (210), feed the decrypted input data to the hidden layers (330) to obtain the output from the last layer of the hidden layers (330), and store the output in the memory unit (210).
5. The chipset (100) according to claim 4, wherein the one or more generic processors (110) are configured to obtain the output from the memory unit (210), encrypt the output, and store the encrypted output in the memory unit (210).
6. The chipset (100) according to any one of claims 1 to 5, wherein the one or more generic processors (110) are configured to perform a first authenticated key exchange (610) with a provider device of the input data to obtain the first secret.
7. The chipset (100) according to any one of claims 1 to 6, wherein the decryption of the encrypted input data is based on Advanced Encryption Standard with Galois / CounterMode, AES-GCM.
8. The chipset (100) according to any one of claims 1 to 7, wherein the machine learning model comprises a set of encrypted weights; the one or more generic processors (110) are configured to decrypt the encrypted weights using a second secret, and provide the decrypted weights to the application-specific processors (120); and the application-specific processors (120) are configured to perform machine learning inference using the decrypted weights.
9. The chipset (100) according to claim 8, wherein the one or more generic processors (110) are further configured to validate the decrypted weights based on Secure Hash Algorithm 3, SHA-3.
10. The chipset (100) according to claim 9, wherein the one or more generic processors (110) are configured to perform a second authenticated key exchange (620) with an owner device of the machine learning model to obtain the second secret.
11. The chipset (100) according to any one of claims 1 to 10, wherein after obtaining the output, the one or more generic processors (110) are configured to delete the input data and / or the output and / or the machine learning model.
12. A system comprising one or more chipsets (100) according to any one of claims 1 to 11.
13. A method (700) for confidential computing executed by a chipset comprising one or more on-chip generic processors, and a plurality of application-specific processors for machine learning, wherein the method comprises: obtaining (701), by the one or more generic processors, encrypted input data of a machine learning model; decrypting (702), by the one or more generic processors, the encrypted input data using a first secret to obtain decrypted input data; providing (703), by the one or more generic processors, the decrypted input data to the application-specific processors; obtaining (704), by the application-specific processors, the decrypted input data; and providing (705), by the application-specific processors, the decrypted input data to the machine learning model for inferencing to obtain an output.
14. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to perform the method according to claims 13.
15. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method according to claim 13.
Citation Information
Patent Citations
Method and system for securing neural network models
US20220327222A1