Protection of data and ai models executed on hardware accelerators
The integration of a hardware cipher chip and TPM on GPUs encrypts and decrypts AI models securely, addressing the lack of protection in current technologies and enabling secure distribution and use of AI models.
Patent Information
- Application Number
- PCT/US2024/044176
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-28
- Publication Date
- 2026-03-05
AI Technical Summary
Current technologies lack effective mechanisms to protect artificial intelligence (AI) models stored on graphics processing units (GPUs) from unauthorized access and duplication, relying primarily on legal terms and inadequate encryption methods that can be easily bypassed.
Implementing a hardware cipher chip between the GPU RAM and cache to encrypt and decrypt AI models and data, using asymmetric and symmetric encryption methods, and integrating a trusted platform module (TPM) to manage cryptographic keys, ensuring secure data transfer and storage on GPUs.
Provides robust protection against unauthorized use and theft of AI models by encrypting data in transit and at rest, enabling secure business models for selling AI models to customers, even on less sophisticated hardware.
Smart Images

Figure US2024044176_05032026_PF_FP_ABST
Abstract
Description
Docket No. 202408872PROTECTION OF DATA AND Al MODELS EXECUTED ON HARDWARE ACCELERATORSBACKGROUND
[0001] Artificial Intelligence (Al) models generated with deep learning frameworks (e.g., PyTorch and the like) are highly valuable assets. Creating a large language model for generative Al from scratch is a significant cost on account of the electrical energy consumed during data preparation and model training, rent for the required GPU hardware, costs of scarce and highly trained Al experts, and the like. Protection is lacking for such assets (e.g., Al models) after they are distributed to an end user and loaded into their target Al computing hardware (e.g., NVIDIA GPUs). In some cases, protection for such assets is limited to binding their usage with legal terms, but malicious or adversary market participants cannot be prevented from ignoring such legal terms.BRIEF SUMMARY
[0002] Embodiments of the invention address and overcome one or more of the described- herein shortcomings by providing methods, systems, and apparatuses for protecting computing assets (e.g., Al models) from unauthorized use or theft. In particular, for example, embodiments define technical solutions for protecting artificial intelligence (Al) models and other assets used on a graphics processing unit (GPU) from use by unauthorized parties.
[0003] In an example aspect, a protection computing system includes a central processing system comprising a central processing unit, system random access memory (RAM), and a system bus coupled to the central processing unit (CPU) and the system RAM. The protection computing system further includes a graphics processing system coupled to the central processing system. The graphics processing system is configured to execute artificial intelligence (Al) models. The graphics computing system includes a graphics processing unit (GPU), a GPU RAM, and a GPU cache memory communicatively coupled to the GPU and the GPU RAM. The graphics processing system further includes a hardware cipher chip configured to encrypt and decrypt the Al models and data associated with the Al models. In various examples, the central processing system further comprises a trusted platform module (TPM) configured to provide a cipher key to the hardware cipher chip. The cipher key can used to encrypt and decrypt the Al models or data associated with the Al models.Docket No. 202408872
[0004] In an example, the hardware cipher chip is located between the GPU RAM and the GPU cache, such that the hardware cipher chip is configured to: encrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU cache to the GPU RAM; and decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU RAM to the GPU cache. For example, the central processing system can be further configured to provide a public key to the hardware cipher chip, wherein the public key is configured to encrypt the Al models and data associated with the Al models, so as generate asymmetric encrypted data. The system RAM can transfer the asymmetric encrypted data to the GPU RAM via direct memory access or via a direct channel between the CPU and the GPU RAM. The hardware cipher chip can secure a private key configured to decrypt the asymmetric encrypted data in the GPU RAM, so as to define a decryption and so as to generate asymmetric decrypted data. The hardware cipher chip can further secure a symmetric key used to, after the decryption, re-encrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the GPU RAM. In another example, the central processing system is further configured to inject a symmetric key into the GPU RAM, wherein the symmetric key is configured to encrypt and decrypt the Al models and data associated with the Al models, so as generate symmetric encrypted data in the GPU RAM. In yet another example, the hardware cipher chip can be located between the system RAM and the GPU RAM, such that the hardware cipher chip is configured to decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the system RAM to the GPU RAM.
[0005] In another example aspect, a protection computing system can perform various operations to protect Al models or data associated with Al models from being maliciously copied or hacked. The protection computing system can define a central processing unit (CPU) side and a graphics processing unit (GPU) side coupled to the CPU side. The protection computing system can, responsive to a request for an artificial intelligence (Al) model, receive a request for a public key from a GPU of the GPU side. The system can receive the Al model that is encrypted with the public key, so as to define an encrypted model. The system can then place the encrypted model in a random access memory (RAM) of the GPU side. The GPU side can include a hardware cipher chip that encrypts the Al model or data associated with the Al model when the Al model or the data associated with the Al model is transferred from the GPU cache to the RAM of the GPU side. The hardware cipher chip can decrypt the Al model or data associated with the Al model before transferring the Al model or the data associated with the Al model from the RAM of the GPU side to the GPU cache. In some cases, the hardware cipher chip uses a private key to decrypt the encrypted model in theDocket No. 202408872RAM of the GPU side, so as to define a decryption and so as to generate asymmetric decrypted data. The hardware cipher chip can use a symmetric key to, after the decryption, reencrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the RAM of the GPU side.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0006] The foregoing and other aspects of the present invention are best understood from the following detailed description when read in connection with the accompanying drawings. For the purpose of illustrating the invention, there is shown in the drawings embodiments that are presently preferred, it being understood, however, that the invention is not limited to the specific instrumentalities disclosed. Included in the drawings are the following Figures:
[0007] FIG. 1 is a block diagram of an example protection system configured to protect models and data running on graphics processing units (GPUs) from being copied, in accordance with an example embodiment.
[0008] FIG. 2 is a block diagram of another example protection system configured to protect models and data running on GPUs from being copied, in accordance with an example embodiment.
[0009] FIG. 3 is a flow diagram that illustrates operations that can be performed by the protection system shown in FIG. 2, in accordance with an example embodiment.
[0010] FIG. 4 is a block diagram of another example protection system configured to protect models and data running on GPUs from being copied, in accordance with another example embodiment.
[0011] FIG. 5 is a flow diagram that illustrates operations that can be performed by the protection system shown in FIG. 4, in accordance with an example embodiment.
[0012] FIG. 6 is a block diagram of another example protection system configured to protect models and data running on GPUs from being copied, wherein the protection uses a trusted software driver, in accordance with another example embodiment.
[0013] FIG. 7 is a flow diagram that illustrates operations that can be performed by the protection system shown in FIG. 6, in accordance with an example embodiment.
[0001] FIG. 8 is a block diagram of another example protection system configured to protect models and data running on GPUs from being copied, in accordance with another example embodiment.Docket No. 202408872
[0015] FIG. 9 illustrates a computing environment within which embodiments of the disclosure may be implemented.DETAILED DESCRIPTION
[0016] As an initial matter, it is recognized herein that current approaches to graphics processing unit (GPU) security support, for instance from Nvidia, rely on trusted container technology. For example, Nvidia’s current trusted container technology is said to ensure that only trusted containers can access a given GPU, and that only containers with access to an authorization key can read / write the data in the GPU RAM. Nvidia also offers Nvidia RDMA and GPUDirectStorage technologies that send data from a remote device to the Nvidia GPU. It is recognized herein, however, that currently technologies do not protect the copyright of data (e.g., models) that are transferred, but rather only purport to protect data from access.
[0017] In contrast, embodiments described herein protect copyrighted data that is stored in GPU memory, in particular Al models stored in a GPU for example, from access by any software or hardware other than the related GPU with the corresponding key. Furthermore, embodiments can be implemented in edge devices with cheaper GPUs as well as high-end GPUs used in cloud environments.
[0018] By way of further background, it is recognized herein that current approaches to protecting a model from unintended use is to not sell it, but instead make its functionality available via a cloud-based service. An example of this approach is ChatGPT, which is currently only available as a service, and which is executed on OpenAI’s own servers. In other approaches, such as Industrial Edge, platforms are offered in which Al applications and highly specialized Al models can be bought and sold so as to be executed on customer-operated hardware. Such platforms can apply some restrictions to try to keep applications “inside” their system, such that applications cannot immediately be extracted and run locally on unofficial or unlicensed devices. It is recognized herein, however, that distribution of Al models that are loaded into applications (e.g., Industrial Edge applications) for execution (e.g., Al SDK pipeline packages deployed to an Al Inference Server) are often not protected against illegal or unintended duplication of the models. Consequently, legal terms might be the only mechanism to protect the models, as technical mechanisms for protection are lacking. It is further recognized herein that even if current frameworks offered protection mechanisms such as encrypting the data that represents the model, such data needs to be decrypted before being loaded on the GPU, thereby placing the model into the GPU memory from where a malicious user might extract the model with little effort.Docket No. 202408872
[0019] Embodiments described herein define hardware modifications to GPUs, such that data (e.g., models) can be sent to a given GPU in an encrypted manner. Additionally, data (e.g., models) that are sent back from the GPU can be encrypted. Furthermore, various embodiments employ symmetric encryption so as to expedite processing as compared to asymmetric private / public key encryption.
[0020] Referring initially to FIGs. 1 , 2, 4, 6, and 8, an example protection computing system 100 can include a central processing system (or CPU side) 102 and a graphics computer or processing system (or GPU side) 104 that is communicatively coupled to the CPU side 102. The CPU side 102 can define an industrial computer system configured to monitor or control various manufacturing or automated operations, though it will be understood that the computer system 102 can be configured to perform alternative or additional operations, for instance various operations using Al models or data, and all such operations are contemplated as being within the scope of this disclosure. The central processing system 102 can define a central processing unit (CPU) 108 coupled to a system bus 110 that is coupled to system random access memory (RAM) 112. The system RAM 112 may define computer readable storage media in the form of volatile memory. The system RAM 112 may include other dynamic storage device(s) (e.g., dynamic RAM, static RAM, and synchronous DRAM). The GPU side 104 can define a graphics processing unit (GPU) 1 14 coupled to a GPU cache RAM or GPU cache 1 16. The GPU cache RAM 116 can be coupled to a GPU bus 118 that is coupled to GPU RAM 120, such that the GPU 114 is coupled to the GPU RAM 120 via the GPU bus 1 18.
[0021] Referring now to FIG. 1 , the CPU side 102 can define a trusted platform module (TPM) or trusted processing unit (TPU) 122 configured to secure hardware through integrated cryptographic keys. The TPM 122 can be communicatively coupled to the system bus 110. The GPU side 104 can define a hardware (HW) cipher chip 101 configured to encrypt data. For example, the HW cipher 101 can define a cryptoprocessor configured to encrypt data. Still referring to FIG. 1 , in an example, the HW cipher 101 is coupled to the GPU cache RAM 116 and the GPU bus 118, so as to be located between the GPU bus 1 18 and the GPU cache RAM 116. The HW cipher chip 101 can perform hardware acceleration for transferring data from the GPU RAM 120 to the GPU cache RAM 116. When data is copied back from the cache RAM 116 to the GPU RAM 220, the HW cipher 101 can accelerate the encryption of the clear text data that is transferred.
[0022] In various examples, the hardware cipher 101 requires a key for encryption and decryption, for instance a cipher key 106. In an example, the key 106 is provided from the TPM 122 to the hardware cipher 101 using (via) the system bus 110. In another example, the key 106 can be transferred via a direct and encrypted channel from the TPM 122 to the HWDocket No. 202408872 cipher 101 , thereby providing increased protection. In yet another (less secure) example, the CPU 108 can provide the key 106 to the hardware cipher 101 via the bus 1 10, for instance when the CPU side 102 does not include the TPM 122. Thus, the TPM 122 can be configured to provide one or more keys to the hardware cipher chip 101 , and the keys can be used to encrypt or decrypt the Al models and the data associated with the Al models.
[0023] In various examples, when the key 106 is placed in the HW cipher chip 101 , the key 106 cannot be read from the HW cipher chip 101 , thereby defining a secure storage of the key 106. In various examples, the GPU cache 1 16 is closely integrated with the GPU 114, such that reading the content of the cache RAM 1 16 is impossible or near impossible for all but the most sophisticated (e.g., government-backed with unlimited budget) hackers.
[0024] Referring now to FIGs. 2 and 3, in some examples, the GPU 114 is equipped with a hard-coded private / public key pair (e.g., public key 201 and private key 203) and a symmetric key 205 for fast encryption and decryption. For example, the public key 201 , private key 203, and symmetric key 205 can stored in the HW cipher chip 101. In an example, the private key 203 and the symmetric key 205 are protected and cannot be accessed or read, and the public key 201 is made freely available. For example, the private key 203 can be hard-coded into the GPU RAM 120 such that the private key 201 does not change. In an example, the symmetric key 205 can be created dynamically, such that the symmetric key 205 can change (e.g., be different) each time. At 302, a user or customer 301 can request or purchase an Al model from a vendor system 303, for instance from a cloud-based store of the vendor 303. Unless otherwise specified, Al model (or model) or data can be used interchangeably herein, without limitation, and an Al model can refer to any machine learning (ML) model or large language model (LLM) executed on the GPU 114. For example, the Al model might be trained to perform image classification, facial recognition, or industrial anomaly detection, though it will be understood that Al models or applications can vary as desired, and all such Al models and applications are contemplated as being within the scope of this disclosure.
[0025] With continuing reference to FIGs. 2 and 3, when the user 301 or customer buys or otherwise requests the Al model, the vendor 303 (e.g., via the cloud) might send a request to the computer system 102 of the user 301 , for the public key 201 of the GPU system 104 (at 304). At 306, the computer system 102 can send the request for the public key 201 to the GPU system 104. At 308, in response to the request, the GPU system or GPU side 104 can send the public key 201 to the CPU system 102. At 310, the computer system 102 can return the public key 201 to the vendor system 303, in particular a cloud server of the vendor system 303. At 312, the vendor system 303 can use the public key 201 to encrypt the model and distribute the encrypted model to the customer 301 . In particular, at 314, the vendor systemDocket No. 202408872303 can send the encrypted model to the central processing system 102. At 316, the CPU system or side 102 can send the encrypted model to the GPU system or side 104, so that the encrypted model is put into the GPU RAM 120. In particular, in an example, the system RAM 112 can receive the encrypted model and can transfer the encrypted model to the GPU RAM 120 via direct memory access (DMA) between the system RAM 112 and the GPU RAM 120. Alternatively, the CPU 108 can receive the encrypted model and directly transfer the encrypted model to the GPU RAM 120.
[0026] In various examples, the public key 201 is signed with a trusted certificate to ensure authenticity, such that the seller (vendor system 303) can authenticate the GPU system 104. It is recognized here that, in some cases in which the public key 201 is not signed, a malicious user might send an arbitrary public key 201 , which might enable the user to decrypt the corresponding model.
[0027] Still referring to FIGs. 2 and 3, at 318, the GPU system 104, in particular the HW cipher 101 , can use the private key 203 to decrypt the encrypted model in the GPU RAM 120 and immediately encrypt the model again using the symmetric key 205, thereby transforming the model from asymmetric encrypted data to symmetric encrypted data while performing high speed encryption / decryption. In some cases, the symmetric key 205 is hard-coded. In other examples, the symmetric key can be generated randomly. At 320, the GPU system 104 can use the model or data from the model. In particular, for example, the GPU 114 can read data from the model from the GPU RAM 120 into the cache memory 1 16, via the HW cipher chip 101 that can use the symmetric key 205 to decrypt the data before placing the data into the cache 1 16. Similarly, when data is copied back from the cache 116 to the GPU RAM 120, the cipher chip 101 can encrypt the data before is copied to the GPU RAM 120, for protection.
[0028] Referring now to FIG. 4, in accordance with another embodiment, the GPU system 104, in particular the HW cipher chip 101 , can be equipped with a symmetric key 401 that can be overwritten or shadowed, for instance by a key that is injected into the HW cipher chip 101 by a trusted driver application. Without being bound by theory, such an embodiment might be implemented based on the value of the model that is being protected.
[0029] Referring also to FIG. 5, at 302, as previously described (common reference numbers in different drawings refer to substantially the same or similar features), the user or customer 301 can request or purchase an Al model from the vendor system 303, for instance from a cloud-based store of the vendor 303. At 502, the vendor system 303 can encrypt the model with the symmetric key 401 . At 504, the vendor system 303 can send the encrypted model to the CPU system 102 of the user 301. At 316, the CPU system 104 can send the encrypted model to the GPU system or side 104, so that the encrypted model is put into the GPU RAMDocket No. 202408872 120. At 506, on a separate and secure path, for instance via email or letter, the vendor system 303 can send the symmetric key to the user 301 , so that the user 301 can receive (use) the symmetric key 401 at the CPU system 102 (at 508). In particular, for example, at 510, the CPU system 102 can inject the symmetric key 401 into the GPU system 104, so as to update the encryption key implemented on the GPU system 104. Thus, at 512, the GPU system 104 can use the model or data from the model. In particular, for example, the GPU 114 can read data from the model from the GPU RAM 120 into the cache memory 116, via the HW cipher chip 101 that can use the symmetric key 401 to decrypt the data before placing the data into the cache 1 16. Similarly, when data is copied back from the cache 116 to the GPU RAM 120, the cipher chip 101 can encrypt the data before is copied to the GPU RAM 120, for protection. In accordance with the example illustrated in FIGs. 4 and 5, the user 301 might be trusted to secure the model and key 401 , so as to not share them. Thus, with reference to FIGs. 2 and 4, the hardware cipher chip 101 can be located between the GPU RAM 120 and the GPU cache 116, such that the hardware cipher chip 101 is configured to encrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU cache 116 to the GPU RAM 120; and decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU RAM 120 to the GPU cache 116.
[0030] Referring now to FIGs. 6 and 7, in accordance with another embodiment, a trusted driver software of the user 201 can generate a symmetric key (at 700), such that hardware is not relied on to support the security of the model. At 702, the user 301 request or purchase an Al model from the vendor system 303, and the user 301 can provide the symmetric key to the vendor system 303. At 502, the vendor system 303 can encrypt the model with this symmetric key. At 704, the vendor system 303 can provide the encrypted model and key to the CPU system 102. At 705, a trusted driver installed on the customer’s system (CPU system 104) can copy the data into the GPU RAM 120 and decrypt it on the way (at 706) with the symmetric key.
[0031] After 706, the data is in the GPU RAM 120 in an unencrypted state. It is recognized herein that a hacker, in some cases, might find ways to read the GPU RAM 120 and extract the model in this way. The risk can be mitigated, however, by having the model only partially decrypted in the GPU RAM 120 such that the model is decrypted on the fly (at 706). Furthermore, non-experienced users might find it difficult to extract the model, and an advantage of this embodiment is that it does not require specialized hardware for maintaining the data encrypted in memory. Instead, a dedicated software driver is used to implement theDocket No. 202408872 protection as the last step before unencrypting the model and loading it onto the GPU system 104 (at 705).
[0032] Referring now to FIG. 8, in accordance with another embodiment, the hardware cipher chip 101 is placed between the GPU RAM 120 and the system RAM 1 12. If data is encrypted, the cipher chip 101 can decrypt the data before placing it into the GPU RAM 120. Alternatively, or additionally, referring again to FIG. 3, after 316, in accordance with the example illustrated in FIG. 8, the GPU 104 can decrypt the model in the GPU RAM 120 using the private key 203, rather than re-encrypting the model at 318 as shown in the example in FIG. 3. The model can then be used, at 320. With respect to the connection between the bus 1 10 and the public key 201 , the CPU system 102, in some cases, only reads the public key 201 to transport it the vendor system 303 of the associated Al model. The model can then be put encrypted into the GPU RAM 120 (at 316), and decrypted by the GPU 114 using the private key 203. It is recognized herein that this example might provide sufficient protection because a hacker would be forced to expend significant effort and would be required to have a deep knowledge of the infrastructure of the GPU system 104 to monitor the GPU RAM 120 with external hardware.
[0033] In accordance with various embodiments described herein, Al model assets are protected with technical means, such that various business models involving selling models to customers are enabled in a secure manner.
[0034] Thus, as described herein, a protection computing system can include a central processing system comprising a central processing unit, system random access memory (RAM), and a system bus coupled to the central processing unit (CPU) and the system RAM. The protection computing system further can include a graphics processing system coupled to the central processing system. The graphics processing system is configured to execute artificial intelligence (Al) models. The graphics computing system includes a graphics processing unit (GPU), a GPU RAM, and a GPU cache memory communicatively coupled to the GPU and the GPU RAM. The graphics processing system further includes a hardware cipher chip configured to encrypt and decrypt the Al models and data associated with the Al models. In various examples, the central processing system further comprises a trusted platform module (TPM) configured to provide a cipher key to the hardware cipher chip. The cipher key can used to encrypt and decrypt the Al models or data associated with the Al models.
[0035] In an example, the hardware cipher chip is located between the GPU RAM and the GPU cache, such that the hardware cipher chip is configured to: encrypt the Al models and data associated with the Al models before transferring the Al models and data associated withDocket No. 202408872 the Al models from the GPU cache to the GPU RAM; and decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU RAM to the GPU cache. For example, the central processing system can be further configured to provide a public key to the hardware cipher chip, wherein the public key is configured to encrypt the Al models and data associated with the Al models, so as generate asymmetric encrypted data. The system RAM can transfer the asymmetric encrypted data to the GPU RAM via direct memory access or via a direct channel between the CPU and the GPU RAM. The hardware cipher chip can secure a private key configured to decrypt the asymmetric encrypted data in the GPU RAM, so as to define a decryption and so as to generate asymmetric decrypted data. The hardware cipher chip can further secure a symmetric key used to, after the decryption, re-encrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the GPU RAM. In another example, the central processing system is further configured to inject a symmetric key into the GPU RAM, wherein the symmetric key is configured to encrypt and decrypt the Al models and data associated with the Al models, so as generate symmetric encrypted data in the GPU RAM. In yet another example, the hardware cipher chip can be located between the system RAM and the GPU RAM, such that the hardware cipher chip is configured to decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the system RAM to the GPU RAM.
[0036] In another example aspect, a protection computing system can perform various operations to protect Al models or data associated with Al models from being maliciously copied or hacked. The protection computing system can define a central processing unit (CPU) side and a graphics processing unit (GPU) side coupled to the CPU side. The protection computing system can, responsive to a request for an artificial intelligence (Al) model, receive a request for a public key from a GPU of the GPU side. The system can receive the Al model that is encrypted with the public key, so as to define an encrypted model. The system can then place the encrypted model in a random access memory (RAM) of the GPU side. The GPU side can include a hardware cipher chip that encrypts the Al model or data associated with the Al model when the Al model or the data associated with the Al model is transferred from the GPU cache to the RAM of the GPU side. The hardware cipher chip can decrypt the Al model or data associated with the Al model before transferring the Al model or the data associated with the Al model from the RAM of the GPU side to the GPU cache. In some cases, the hardware cipher chip uses a private key to decrypt the encrypted model in the RAM of the GPU side, so as to define a decryption and so as to generate asymmetric decrypted data. The hardware cipher chip can use a symmetric key to, after the decryption, reDocket No. 202408872 encrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the RAM of the GPU side.
[0037] FIG. 9 illustrates an example of a computing environment within which embodiments of the present disclosure may be implemented. A computing environment 900 includes a computer system 910 that may include a communication mechanism such as a system bus 921 or other communication mechanism for communicating information within the computer system 910. The computer system 910 further includes one or more processors 920 coupled with the system bus 921 for processing the information. The CPU system 102 and the GPU system 104 may include, or be coupled to, the one or more processors 920.
[0038] The processors 920 may include one or more central processing units (CPUs), graphics processing units (GPUs), or any other processor known in the art. More generally, a processor as described herein is a device for executing machine-readable instructions stored on a computer readable medium, for performing tasks and may comprise any one or combination of, hardware and firmware. A processor may also comprise memory storing machine-readable instructions executable for performing tasks. A processor acts upon information by manipulating, analyzing, modifying, converting or transmitting information for use by an executable procedure or an information device, and / or by routing the information to an output device. A processor may use or comprise the capabilities of a computer, controller or microprocessor, for example, and be conditioned using executable instructions to perform special purpose functions not performed by a general purpose computer. A processor may include any type of suitable processing unit including, but not limited to, a central processing unit, a microprocessor, a Reduced Instruction Set Computer (RISC) microprocessor, a Complex Instruction Set Computer (CISC) microprocessor, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a System-on-a- Chip (SoC), a digital signal processor (DSP), and so forth. Further, the processor(s) 920 may have any suitable microarchitecture design that includes any number of constituent components such as, for example, registers, multiplexers, arithmetic logic units, cache controllers for controlling read / write operations to cache memory, branch predictors, or the like. The microarchitecture design of the processor may be capable of supporting any of a variety of instruction sets. A processor may be coupled (electrically and / or as comprising executable components) with any other processor enabling interaction and / or communication therebetween. A user interface processor or generator is a known element comprising electronic circuitry or software or a combination of both for generating display images or portions thereof. A user interface comprises one or more display images enabling user interaction with a processor or other device.Docket No. 202408872
[0039] The system bus 921 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may permit exchange of information (e.g., data (including computer-executable code), signaling, etc.) between various components of the computer system 910. The system bus 921 may include, without limitation, a memory bus or a memory controller, a peripheral bus, an accelerated graphics port, and so forth. The system bus 821 may be associated with any suitable bus architecture including, without limitation, an Industry Standard Architecture (ISA), a Micro Channel Architecture (MCA), an Enhanced ISA (EISA), a Video Electronics Standards Association (VESA) architecture, an Accelerated Graphics Port (AGP) architecture, a Peripheral Component Interconnects (PCI) architecture, a PCI-Express architecture, a Personal Computer Memory Card International Association (PCMCIA) architecture, a Universal Serial Bus (USB) architecture, and so forth.
[0040] Continuing with reference to FIG. 9, the computer system 910 may also include a system memory 930 coupled to the system bus 921 for storing information and instructions to be executed by processors 920. The system memory 930 may include computer readable storage media in the form of volatile and / or nonvolatile memory, such as read only memory (ROM) 931 and / or random access memory (RAM) 932. The RAM 932 may include other dynamic storage device(s) (e.g., dynamic RAM, static RAM, and synchronous DRAM). The ROM 931 may include other static storage device(s) (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). In addition, the system memory 930 may be used for storing temporary variables or other intermediate information during the execution of instructions by the processors 920. A basic input / output system 933 (BIOS) containing the basic routines that help to transfer information between elements within computer system 910, such as during start-up, may be stored in the ROM 931 . RAM 932 may contain data and / or program modules that are immediately accessible to and / or presently being operated on by the processors 920. System memory 930 may additionally include, for example, operating system 934, application programs 935, and other program modules 936. Application programs 935 may also include a user portal for development of the application program, allowing input parameters to be entered and modified as necessary.
[0041] The operating system 934 may be loaded into the memory 930 and may provide an interface between other application software executing on the computer system 910 and hardware resources of the computer system 910. More specifically, the operating system 934 may include a set of computer-executable instructions for managing hardware resources of the computer system 910 and for providing common services to other application programs (e.g., managing memory allocation among various application programs). In certain example embodiments, the operating system 934 may control execution of one or more of the programDocket No. 202408872 modules depicted as being stored in the data storage 940. The operating system 934 may include any operating system now known or which may be developed in the future including, but not limited to, any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
[0042] The computer system 910 may also include a disk / media controller 943 coupled to the system bus 921 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 941 and / or a removable media drive 942 (e.g., floppy disk drive, compact disc drive, tape drive, flash drive, and / or solid state drive). Storage devices 940 may be added to the computer system 910 using an appropriate device interface (e.g., a small computer system interface (SCSI), integrated device electronics (IDE), Universal Serial Bus (USB), or FireWire). Storage devices 941 , 942 may be external to the computer system 910.
[0043] The computer system 910 may also include a field device interface 965 coupled to the system bus 921 to control a field device 966, such as a device used in a production line. The computer system 910 may include a user input interface or GUI 961 , which may comprise one or more input devices, such as a keyboard, touchscreen, tablet and / or a pointing device, for interacting with a computer user and providing information to the processors 920.
[0044] The computer system 910 may perform a portion or all of the processing steps of embodiments of the invention in response to the processors 920 executing one or more sequences of one or more instructions contained in a memory, such as the system memory 930. Such instructions may be read into the system memory 930 from another computer readable medium of storage 940, such as the magnetic hard disk 941 or the removable media drive 942. The magnetic hard disk 941 (or solid state drive) and / or removable media drive 942 may contain one or more data stores and data files used by embodiments of the present disclosure. The data store 940 may include, but are not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed data stores in which data is stored on more than one node of a computer network, peer-to-peer network data stores, or the like. The data stores may store various types of data such as, for example, skill data, sensor data, or any other data generated in accordance with the embodiments of the disclosure. Data store contents and data files may be encrypted to improve security. The processors 920 may also be employed in a multi-processing arrangement to execute the one or more sequences of instructions contained in system memory 930. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and software.Docket No. 202408872
[0045] As stated above, the computer system 910 may include at least one computer readable medium or memory for holding instructions programmed according to embodiments of the invention and for containing data structures, tables, records, or other data described herein. The term “computer readable medium” as used herein refers to any medium that participates in providing instructions to the processors 920 for execution. A computer readable medium may take many forms including, but not limited to, non-transitory, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical disks, solid state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disk 941 or removable media drive 942. Non-limiting examples of volatile media include dynamic memory, such as system memory 930. Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up the system bus 921 . Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0046] Computer readable medium instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0047] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchartDocket No. 202408872 illustrations and / or block diagrams, may be implemented by computer readable medium instructions.
[0048] The computing environment 900 may further include the computer system 910 operating in a networked environment using logical connections to one or more remote computers, such as remote computing device 980. The network interface 970 may enable communication, for example, with other remote devices 980 or systems and / or the storage devices 941 , 942 via the network 971 . Remote computing device 980 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above relative to computer system 910. When used in a networking environment, computer system 910 may include modem 972 for establishing communications over a network 971 , such as the Internet. Modem 972 may be connected to system bus 921 via user network interface 970, or via another appropriate mechanism.
[0049] Network 971 may be any network or system generally known in the art, including the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection or series of connections, a cellular telephone network, or any other network or medium capable of facilitating communication between computer system 910 and other computers (e.g., remote computing device 980). The network 971 may be wired, wireless or a combination thereof. Wired connections may be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection generally known in the art. Wireless connections may be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular networks, satellite or any other wireless connection methodology generally known in the art. Additionally, several networks may work alone or in communication with each other to facilitate communication in the network 971 .
[0050] It should be appreciated that the program modules, applications, computerexecutable instructions, code, or the like depicted in FIG. 9 as being stored in the system memory 930 are merely illustrative and not exhaustive and that processing described as being supported by any particular module may alternatively be distributed across multiple modules or performed by a different module. In addition, various program module(s), script(s), plug-in(s), Application Programming Interface(s) (API(s)), or any other suitable computer-executable code hosted locally on the computer system 910, the remote device 980, and / or hosted on other computing device(s) accessible via one or more of the network(s) 971 , may be provided to support functionality provided by the program modules, applications, or computer-executable code depicted in FIG. 9 and / or additional or alternate functionality. Further, functionality may be modularized differently such that processing described as being supported collectively byDocket No. 202408872 the collection of program modules depicted in FIG. 9 may be performed by a fewer or greater number of modules, or functionality described as being supported by any particular module may be supported, at least in part, by another module. In addition, program modules that support the functionality described herein may form part of one or more applications executable across any number of systems or devices in accordance with any suitable computing model such as, for example, a client-server model, a peer-to-peer model, and so forth. In addition, any of the functionality described as being supported by any of the program modules depicted in FIG. 9 may be implemented, at least partially, in hardware and / or firmware across any number of devices.
[0051] It should further be appreciated that the computer system 910 may include alternate and / or additional hardware, software, or firmware components beyond those described or depicted without departing from the scope of the disclosure. More particularly, it should be appreciated that software, firmware, or hardware components depicted as forming part of the computer system 910 are merely illustrative and that some components may not be present or additional components may be provided in various embodiments. While various illustrative program modules have been depicted and described as software modules stored in system memory 930, it should be appreciated that functionality described as being supported by the program modules may be enabled by any combination of hardware, software, and / or firmware. It should further be appreciated that each of the above-mentioned modules may, in various embodiments, represent a logical partitioning of supported functionality. This logical partitioning is depicted for ease of explanation of the functionality and may not be representative of the structure of software, hardware, and / or firmware for implementing the functionality. Accordingly, it should be appreciated that functionality described as being provided by a particular module may, in various embodiments, be provided at least in part by one or more other modules. Further, one or more depicted modules may not be present in certain embodiments, while in other embodiments, additional modules not depicted may be present and may support at least a portion of the described functionality and / or additional functionality. Moreover, while certain modules may be depicted and described as sub-modules of another module, in certain embodiments, such modules may be provided as independent modules or as sub-modules of other modules.
[0052] Although specific embodiments of the disclosure have been described, one of ordinary skill in the art will recognize that numerous other modifications and alternative embodiments are within the scope of the disclosure. For example, any of the functionality and / or processing capabilities described with respect to a particular device or component may be performed by any other device or component. Further, while various illustrativeDocket No. 202408872 implementations and architectures have been described in accordance with embodiments of the disclosure, one of ordinary skill in the art will appreciate that numerous other modifications to the illustrative implementations and architectures described herein are also within the scope of this disclosure. In addition, it should be appreciated that any operation, element, component, data, or the like described herein as being based on another operation, element, component, data, or the like can be additionally based on one or more other operations, elements, components, data, or the like. Accordingly, the phrase “based on,” or variants thereof, should be interpreted as “based at least in part on.”
[0053] Although embodiments have been described in language specific to structural features and / or methodological acts, it is to be understood that the disclosure is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as illustrative forms of implementing the embodiments. Conditional language, such as, among others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments could include, while other embodiments do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, and / or steps are included or are to be performed in any particular embodiment.
[0054] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
Claims
Docket No. 202408872CLAIMSWhat is claimed is:1 . A protection computing system, comprising: a central processing system comprising a central processing unit, system random access memory (RAM), and a system bus coupled to the central processing unit (CPU) and the system RAM; a graphics processing system coupled to the central processing system, the graphics processing system configured to execute artificial intelligence (Al) models, the graphics processing system comprising a graphics processing unit (GPU), a GPU RAM, and a GPU cache memory communicatively coupled to the GPU and the GPU RAM, wherein the graphics processing system further comprises a hardware cipher chip configured to encrypt and decrypt the Al models and data associated with the Al models.
2. The protection computing system as recited in claim 1 , wherein the central processing system further comprises a trusted platform module (TPM) configured to provide a cipher key to the hardware cipher chip, the cipher key used to encrypt and decrypt the Al models and data associated with the Al models.
3. The protection computing system as recited in claim 1 , wherein the hardware cipher chip is located between the GPU RAM and the GPU cache, such that the hardware cipher chip is configured to: encrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU cache to the GPU RAM; and decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the GPU RAM to the GPU cache.
4. The protection computing system as recited in any one of the preceding claims, wherein the central processing system is further configured to provide a public key to the hardware cipher chip, the public key configured to encrypt the Al models and data associated with the Al models, so as generate asymmetric encrypted data.
5. The protection computing system as recited in claim 4, wherein the system RAM can transfer the asymmetric encrypted data to the GPU RAM via direct memory access or via a direct channel between the CPU and the GPU RAM.Docket No. 2024088726. The protection computing system as recited in claim 5, wherein the hardware cipher chip secures a private key configured to decrypt the asymmetric encrypted data in the GPU RAM, so as to define a decryption and so as to generate asymmetric decrypted data.
7. The protection computing system as recited in claim 6, wherein the hardware cipher chip further secures a symmetric key used to, after the decryption, re-encrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the GPU RAM.
8. The protection computing system as recited in any of claims 1 to 3, wherein the central processing system is further configured to inject a symmetric key into the GPU RAM, the symmetric key configured to encrypt and decrypt the Al models and data associated with the Al models, so as generate symmetric encrypted data in the GPU RAM.
9. The protection computing system as recited in claim 1 , wherein the hardware cipher chip is located between the system RAM and the GPU RAM, such that the hardware cipher chip is configured to decrypt the Al models and data associated with the Al models before transferring the Al models and data associated with the Al models from the system RAM to the GPU RAM.
10. A method performed by a protection computing system that comprises a central processing unit (CPU) side and a graphics processing unit (GPU) side coupled to the CPU side, the method comprising: responsive to a request for an artificial intelligence (Al) model, receiving a request for a public key from a GPU of the GPU side; receiving the Al model that is encrypted with the public key, so as to define an encrypted model; and placing the encrypted model in a random access memory (RAM) of the GPU side.11 . The method as recited in claim 10, wherein the GPU side further comprises a GPU cache and hardware cipher chip between the GPU cache and the RAM of the GPU side, the method further comprising: the hardware cipher chip encrypting the Al model or data associated with the Al model when the Al model or the data associated with the Al model is transferred from the GPU cache to the RAM of the GPU side.Docket No. 20240887212. The method as recited in claim 10, wherein the GPU side further comprises a GPU cache and hardware cipher chip between the GPU cache and the RAM of the GPU side, the method further comprising: the hardware cipher chip decrypting the Al model or data associated with the Al model before transferring the Al model or the data associated with the Al model from the RAM of the GPU side to the GPU cache.
13. The protection system as recited in claim 10, the method further comprising: the hardware cipher text using a private key to decrypt the encrypted model in the RAM of the GPU side, so as to define a decryption and so as to generate asymmetric decrypted data.
14. The method as recited in claim 13, the method further comprising: the hardware cipher chip using a symmetric key to, after the decryption, re-encrypt the asymmetric decrypted data, so as to generate symmetric encrypted data in the RAM of the GPU side.
15. The method as recited in claim 10, the method further comprising: receiving the encrypted model at a RAM of the CPU side, such that the placing of the encrypted model in the RAM of the GPU side is performed via direct memory access between the RAM of the CPU side and the RAM of the GPU side.
Citation Information
Patent Citations
Method and system for providing explanation for output generated by an artificial intelligence model
US20200313849A1
Method for implanting a watermark in a trained artificial intelligence model for a data processing accelerator
US20210109790A1
Secure computing device
US20220060455A1
System and method of securing ai model
US20240275583A1
Method for using artificial intelligence model and related apparatus
WO2022261878A1