Secure Re-encryption of Homomorphic Encryption Data

By using an obfuscation module to transform ciphertext streams and employing a hardware security module for decryption and re-encryption, the system addresses noise accumulation and privacy concerns in MLaaS FHE systems, enabling deeper model training and secure data protection.

JP7698384B2Active Publication Date: 2025-06-25INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023530628
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-20
Filing Date
2021-11-05
Publication Date
2025-06-25
Estimated Expiration
2041-11-05

AI Technical Summary

Technical Problem

Existing machine learning as a service (MLaaS) systems using fully homomorphic encryption (FHE) face challenges with noise accumulation in encrypted data during model training, which limits the scope of machine learning and can be computationally costly, and there is a risk of adversaries identifying processing details through communication links.

Method used

The system employs an obfuscation module (OM) to modify ciphertext streams between a service provider and a hardware security module (HSM), adding noise-reducing transformations like random delays, frequency changes, and random selection to obscure processing details, while using the HSM for decrypting and re-encrypting intermediate results with a client's private key.

Benefits of technology

This approach reduces noise in encrypted data, allowing for deeper model training without decrypting client data, maintains privacy, and prevents adversaries from deducing processing details, thus enhancing model security and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698384000001
    Figure 0007698384000001
  • Figure 0007698384000002
    Figure 0007698384000002
  • Figure 0007698384000003
    Figure 0007698384000003
Patent Text Reader

Abstract

A method for securely re-encrypting homomorphically encrypted data by receiving fully homomorphic encryption (FHE) information from a client device, training a machine learning model using the FHE information, generating an FHE ciphertext, applying a first transformation to the FHE ciphertext, generating an obfuscated FHE ciphertext, transmitting the obfuscated FHE ciphertext to a secure device, receiving a re-encrypted version of the obfuscated FHE ciphertext from the secure device, applying a second transformation to the re-encrypted version of the obfuscated FHE ciphertext, generating a de-obfuscated re-encrypted FHE ciphertext, determining FHE ML model parameters according to the de-obfuscated re-encrypted ciphertext, and transmitting the FHE ML model parameters to the client device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the protection of data protected using fully homomorphic encryption (FHE). In particular, the present disclosure relates to obfuscating FHE (fully homomorphic encryption) data passed between a service provider and a hardware security module (HSM) provider.

Background Art

[0002] Machine learning as a service (MLaaS) enables a client to provide a training data set to a cloud or other network-based service provider for model training. The service provider receives the training data set, develops a model using the training data, and passes the model parameters of the trained model to the client.

[0003] By using FHE for the encryption of training data and the decryption of model parameters, the service provider can progress the training process without accessing the training data set or knowing details about the training data and related model parameters. Fully homomorphic encryption provides encrypted data that can be used for model training and model parameter determination without decrypting the data. Training a model with encrypted data generates encrypted model parameters, which are passed to the user and decrypted using FHE. FHE enables the user to access the parameters of the trained model from the service provider without disclosing the training data or the parameters of the final model to the service provider.

Summary of the Invention

[0004] The following presents an overview for providing a basic understanding of one or more embodiments of the present disclosure. This overview is not intended to identify key elements or to delineate any scope of particular embodiments or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a prelude to the more detailed description that is presented later. In one or more embodiments described herein, a device, system, computer-implemented method, apparatus, or computer program product, or a combination thereof enables secure re-encryption of homomorphic encrypted data.

[0005] Aspects of the present invention relate to methods, systems, and computer-readable media for securely re-encrypting homomorphic encrypted data by receiving fully homomorphic encryption (FHE) information from a client device, training a machine learning model using the FHE information, generating an FHE ciphertext, applying a first transformation to the FHE ciphertext, generating an obfuscated FHE ciphertext, transmitting the obfuscated FHE ciphertext to a secure device and receiving a re-encrypted version of the obfuscated FHE ciphertext from the secure device, applying a second transformation to the re-encrypted version of the obfuscated FHE ciphertext, generating a de-obfuscated re-encrypted FHE ciphertext, determining FHE ML model parameters according to the de-obfuscated re-encrypted ciphertext, and transmitting the FHE ML model parameters to the client device.

[0006] The above and other objects, features, and advantages of the present disclosure will become more apparent through a more detailed description of some embodiments of the present disclosure in the accompanying drawings, where like references generally refer to like components in embodiments of the present disclosure.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0008] Some embodiments will be described in more detail with reference to the accompanying drawings in which embodiments of the present disclosure are illustrated. However, the present disclosure can be implemented in various ways and should not be construed as limited to the embodiments disclosed herein.

[0009] In an embodiment, one or more components of the system can adopt hardware and / or software to essentially solve highly technical problems (for example, receiving information from a client device, applying a first transformation to the information to generate obfuscated information by the first transformation, sending the obfuscated information to a security device, receiving the obfuscated information converted by the security device, applying a second transformation to the converted obfuscated information to obtain de-obfuscated converted information, and sending the de-obfuscated converted information to the client device, etc.). These solutions are not abstract. For example, due to the FHE encryption of data and model parameters, the secure transmission of FHE data over a network, the reception of FHE data, and the processing power required to facilitate the decryption of FHE data, they cannot be executed as a series of mental acts by humans. Furthermore, some of the processes executed can be executed by a dedicated computer for performing defined tasks related to the use of FHE. For example, a dedicated computer can be employed to perform tasks related to securely transmitting FHE data over a network, etc.

[0010] Any mathematical operation on FHE data generates noise in the data. The noise accumulates according to the depth of the operation (the number of sequential operations), and at high levels, there is a possibility of decryption errors. Efforts to reduce the noise of FHE data include bootstrapping that uses an encrypted secret key to reduce the data's noise. Bootstrapping can be computationally expensive to implement on a large scale. There is also a method where the service provider sends the FHE data to a hardware-assisted module, which securely decrypts, re-encrypts, and then sends back the FHE data. The latter approach has mainly two limitations. First, an adversary monitoring the communication link between the service provider and the hardware security can identify the types of calculations being performed by the service provider. Second, the hardware-assisted module may be able to decrypt the calculations being performed by the service provider. The present disclosure provides a method for obfuscating the data exchanged between the service provider and the hardware-assisted module to overcome the above two limitations.

[0011] Clients that require a trained machine learning model can request the model's construction from a service provider. The service provider can provide the model's architecture and the computing resources necessary for training the model. The client provides the training dataset. The client may want to protect the data in the training set while using the service provider for model construction and training. Through homomorphic encryption (HE), logical circuits can be evaluated using encrypted data, and in this example, a machine learning model can be trained. Since the HE data is encrypted before it reaches the service provider, the service provider cannot access the underlying data values. Training can be performed using the HE data without access to the relevant decryption key or the ability to decrypt the data. At each stage of the training process, noise is introduced into the encrypted data, so using HE for training may be limited. This noise can limit the scope of machine learning using HE. As model learning progresses using HE data, the noise in the data increases. As the level of noise increases, it may become impossible to use the encrypted data beyond the system's ability to decode the noisy signal. One solution is to use secret key encryption to periodically re-encrypt the HE ciphertext. This process is called "bootstrapping" and can reduce the noise in the ciphertext. This re-encryption / bootstrapping process is not available in all systems and may be computationally costly to implement. The disclosed embodiments enable the exchange of FHE data via a network with an HSM service provider, reducing noise through data decryption and re-encryption, while maintaining the privacy and security of the data from entities such as malicious entities that access the HSM service provider or network communication traffic.

[0012] In an embodiment, FHE data is passed from a client to a cloud or other network service provider. The service provider utilizes the FHE data for evaluating a logic circuit or in training a machine learning model such as a neural network model. The FHE data is encrypted by the client using a public-private key pair of a public key or a private key. The client does not pass the private key to the service provider. Training of the model or evaluation of the circuit can obtain intermediate results from the FHE input data. The intermediate results contain noise due to the basic encryption of the data. In this embodiment, the method passes the intermediate results from the service provider's system to an external entity such as a hardware security module (HSM). In this embodiment, the method blindly adds the intermediate results additively to enhance the security of the encrypted data. The HSM modifies the received ciphertext. The HSM decrypts and then re-encrypts the ciphertext containing the intermediate results received from the service provider's system. The client provides the private key to the HSM to decrypt and then re-encrypt the intermediate results from the service provider. The client can provide the private key to the HSM using a communication channel different from the channel used for the FHE data.

[0013] By decrypting and re-encrypting the intermediate results using the private key, the noise in the intermediate results can be reduced or removed. By passing the re-encrypted intermediate results to the service provider, model learning and circuit evaluation can be continued using the less noisy results. As a result, noise does not accumulate during training, enabling the training of a model structure with more layers. By training a model using the FHE training data, encrypted model parameters are generated. After training is completed, the encrypted model parameters are passed to the client and decrypted using the private key. The service provider can only "see" the encrypted data and the encrypted model parameters.

[0014] In this embodiment, the exchange of intermediate results from the service provider to the HSM is performed at regular intervals related to the model architecture. The ciphertext containing the intermediate results may be passed to the HSM after the calculation for each layer of the network or after the noise of the results has exhausted the multiplication depth of the system. The intervals and number of ciphertexts are related to the nature of the model's architecture structure and may serve as a means to obtain this information. For example, the number of model layers, the number of nodes per layer, the type of activation, and the training strategy can be distinguished from the analysis of the communication between the service provider's system and the HSM.

[0015] In the embodiment, the privacy protection of the model is enhanced by adding an obfuscation module (OM) between the service provider's system and the HSM. In this embodiment, the OM obfuscates the ciphertext passed between the system and the HSM and reduces the accuracy of any information obtained from the analysis of the HSM communication traffic of the system.

[0016] By applying obfuscation to the information, the disclosed embodiments enable the communication between the CAM and the HAM without the concern that the details of the CAM's processing can be identified from the messaging traffic. Obfuscation hides the details available in other ways when passing ciphertexts related to reducing the noise of the FHE data during the training of the machine learning model by the CAM.

[0017] In the embodiment, the OM adds additional ciphertexts to the ciphertext containing the intermediate results. The OM sends the additional ciphertexts to the HSM at random intervals, changing the regular pattern associated with the ciphertext containing the actual intermediate results.

[0018] In the embodiment, the OM changes the timing of the ciphertexts, delaying the transmission of some or all of the ciphertexts and changing the pattern of the original stream of ciphertexts received from the service provider's system. In the embodiment, the OM changes the frequency of the ciphertexts, shifting some or all of the ciphertexts along the time axis and removing all or part of the periodic aspect of the ciphertext stream.

[0019] In an embodiment, the ciphertext consists of a number of slots, and each slot contains a value related to an intermediate value. In this embodiment, the OM changes the order of the slots of the ciphertext before sending the text to the HSM. When receiving a re-encrypted version of the ciphertext, the OM rearranges the slot values of the ciphertext back to the original order before sending the re-encrypted ciphertext back to the service provider's system.

[0020] In an embodiment, the OM randomly selects ciphertexts for re-encryption from the entire set of ciphertexts of the service provider system. In this embodiment, the OM can randomly select ciphertexts using any known randomization function, including random selection based on the depth of the remaining ciphertexts (how many more ciphertext operations can be performed before the noise of the encrypted data exceeds the decryption noise threshold). When the depth of the remaining ciphertexts is high, the probability of selection is usually lower. As the depth of the remaining ciphertexts decreases, the selection probability of the ciphertexts will increase. As noise accumulates, the depth of the ciphertexts decreases until the point of ciphertext depth depletion where further productive evaluation or development of the machine learning model is no longer possible due to the noise. As the depth of the ciphertexts decreases, the probability of random selection increases, and the level of noise that rises before depletion can be reduced. When the level of noise indicated by the high level of the remaining ciphertexts is low, the need to reduce noise is low, and the selection probability of the ciphertexts is low.

[0021] In an embodiment, the OM determines an optimal trade-off between ensuring model privacy and the computing resource cost and process delay associated with adding obfuscation. In this embodiment, the method determines the ratio of the original ciphertext stream changed by obfuscation to the computational cost of the added ciphertexts, and evaluates the total time related to the delay of the ciphertexts against the time related to passing information without obfuscation.

[0022] FIG. 1 is a schematic diagram showing exemplary network resources related to the practice of the disclosed invention. The present invention may be practiced on any processor of the disclosed elements that process instruction streams. As shown in the figure, a networked Hardware Security Module (HSM) 110 is connected to the server subsystem 102 and is connected to the client device 104 either wirelessly or via a wired connection. The HSM 110 is connected to the server subsystem 102 of the service provider either wired or wirelessly via the OM 112 and the network 114. Also, the client device 104 can be connected to the server subsystem 102 either wired or wirelessly via the network 114. The client device 104, the HSM 110, and the OM 112 are configured with a Fully Homomorphic Encryption (FHE) security program (not shown) along with computing resources (processor, memory, network communication hardware) sufficient to execute the program. As shown in FIG. 1, the server subsystem 102 includes a server computer 150. FIG. 1 shows a block diagram of the components of the server computer 150 within the networked computer system 1000 according to an embodiment of the present invention. It should be understood that FIG. 1 is only an example of one embodiment and does not imply any limitation with respect to the environments in which different embodiments may be implemented. Many modifications can be made to the depicted environment. In an embodiment, (not shown), the OM 112 is present within the server subsystem 102. In this embodiment, all communication between the server subsystem 102 and the HSM 110 passes through the OM 112.

[0023] Server computer 150 can include a processor 154, a memory 158, a persistent storage 170, a communication unit 152, an input / output (I / O) interface 156, and a communication fabric 140. The communication fabric 140 provides communication among the cache 162, the memory 158, the persistent storage 170, the communication unit 152, and the input / output (I / O) interface 156. The communication fabric 140 can be implemented in any architecture designed to pass data or control information or combinations thereof among processors (such as microprocessors, communication and network processors, etc.), system memory, peripheral devices, and any other hardware components within the system. For example, the communication fabric 140 can be implemented with one or more buses.

[0024] The memory 158 and the persistent storage 170 are computer-readable storage media. In this embodiment, the memory 158 includes a random access memory (RAM) 160. Generally, the memory 158 can include any suitable volatile or non-volatile computer-readable storage media. The cache 162 is a high-speed memory that improves the performance of the processor 154 by holding recently accessed data and data near recently accessed data from the memory 158.

[0025] The program instructions and data used to practice embodiments of the present invention, for example, the FHE security program 175, are stored in the persistent storage 170 for execution or access or both by one or more of the respective processors 154 of the server computer 150 via the cache 162. In this embodiment, the persistent storage 170 includes a magnetic hard disk drive. Alternatively, or in addition to the magnetic hard disk drive, the persistent storage 170 may include a solid state hard drive, a semiconductor storage device, a read only memory (ROM), an erasable programmable read only memory (EPROM), a flash memory, or any other computer readable storage medium capable of storing program instructions or digital information.

[0026] Also, the medium used by the persistent storage 170 may be removable. For example, a removable hard disk may be used for the persistent storage 170. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer onto another computer readable storage medium that is also part of the persistent storage 170.

[0027] The communication unit 152 provides communication in these examples with the resources of the client computing device 104 and with other data processing systems or devices including 110. In these embodiments, the communication unit 152 includes one or more network interface cards. The communication unit 152 may provide communication through the use of either or both physical and wireless communication links. Software distribution programs, and other programs and data used in the practice of the present invention, may be downloaded through the communication unit 152 to the persistent storage 170 of the server computer 150.

[0028] The I / O interface 156 enables the input and output of data with other devices that can be connected to the server computer 150. For example, the I / O interface 156 can provide a connection to an external device 190 such as a keyboard, keypad, touch screen, microphone, digital camera, or other suitable input device or a combination thereof. The external device 190 can also include, for example, a portable computer-readable storage medium such as a thumb drive, portable optical or magnetic disk, and memory card. Software and data used to practice embodiments of the present invention, for example, the FHE security program 175 on the server computer 150, can be stored on such a portable computer-readable storage medium and loaded onto the persistent storage 170 via the I / O interface 156. The I / O interface 156 is also connected to a display 180.

[0029] The display 180 provides a mechanism for displaying data to the user and can be, for example, a computer monitor. The display 180 can also function as a touch screen such as that of a tablet computer.

[0030] In an embodiment, the client device 104 applies encryption to a training dataset intended for use in training a machine learning model. The encryption involves the use of a secret key or a public key for homomorphically encrypting the dataset. The homomorphically encrypted dataset is passed from the client device 104 to the server subsystem 102 via the network 114. The server subsystem 102 may be local and under the control of the same entity as the client device 104, or in some cases, the server subsystem 102 may be cloud resources provided for use by a cloud service provider for the development and training of a machine learning model. The server subsystem uses the provided data for evaluation or training without decrypting the data, so the overall process uses fully homomorphic encryption (FHE). The server subsystem 102 receives the FHE data from the client device 104 and uses the FHE data to evaluate a specified circuit or train a specified machine learning model. During the evaluation or training, the server subsystem generates intermediate values or ciphertexts related to the evaluation or training. Since these are ciphertexts of FHE data, they contain noise. Without efforts to reduce the noise, the noise will increase and decryption errors will occur in the final result.

[0031] In some systems, the server subsystem 102 passes the encrypted intermediate results to the HSM 110 as a stream of ciphertexts, each containing one or more intermediate results per ciphertext. The HSM 110 uses the secret key received from the client 104 to decrypt each ciphertext and then re-encrypts it. This decrypt-then-re-encrypt process reduces or removes the noise from the intermediate results carried in the ciphertexts. The HSM 110 returns the re-encrypted intermediate results to the server subsystem 102 as a stream of ciphertexts corresponding to the original stream.

[0032] In such a system, the nature of the architecture of the machine learning model may be determined using information leaked into the ciphertext stream. Information about the model, such as the number of model layers, the number of nodes per layer, the details of the CAM calculation, the type of activation, and the training strategy, can be determined from the analysis of the ciphertext stream.

[0033] In the embodiment shown in FIG. 1, the method adds an obfuscation module (OM) 112 to the entire networked system. OM 112 exists between the server subsystem 102 and the HSM 110, and all communication traffic between the server subsystem 102 and the HSM 110 passes through OM 112. In this embodiment, OM 112 modifies the ciphertext stream transmitted from the server subsystem 102 to the HSM 110. This modification obfuscates architecture details such as the number of model layers, the number of nodes per layer, the type of activation, and the training strategy, which would otherwise be identifiable from the messaging stream. OM 112 de-obfuscates the message passing from the HSM 110 to the server subsystem 102, providing the server subsystem with a ciphertext with reduced noise. In an embodiment, OM 112 exists as an integral part of the server subsystem 102, and the communication between the server subsystem 102 and the HSM 110 flows from the server subsystem 102 to OM 112, then to the HSM 110 via the network 114, and back along the same routing.

[0034] After receiving the de-obfuscated ciphertext, the method proceeds with the training of the ML model. Multiple epochs or iterations of training may occur. A set of multiple intermediate ciphertexts may be passed to the HSM 110 via OM 112 and may be returned to the server subsystem 102. After the completion of the training of the ML model, the method passes a set of FHE ML model parameters to the client device 104. The client device 104 decrypts the encryption of the FHE ML model parameters using the client's private key, enabling the client to utilize the ML model trained to analyze new data.

[0035] Figure 2 provides flowchart 200, which shows exemplary operations related to the implementation of the present disclosure. After the program starts, at block 210, the method of FHE security program 175 operating via obfuscation module (OM) receives information from a client device such as a computing agent module (CAM). In an embodiment, the information includes a ciphertext from the CAM that includes FHE data intermediate results, and the CAM includes a system processor of a cloud-based service provider that trains a machine learning model for a remote client. The received information includes data encrypted using the private or public key of the CAM's client.

[0036] At block 220, the method of FHE security program 175 operates via OM and applies a first transformation to the information received from the CAM. The applied transformation can supplement the information by adding generated spurious information, such as "fake" ciphertext, to the received information. When the fake ciphertext is added, the pattern of the ciphertext in the downstream communication from the OM changes. This change obfuscates any information identifiable in the original stream of information received from the CAM, including the number of model layers, the number of nodes per layer, the type and training strategy of the activation, and the details of the CAM's calculations. The applied transformation can include additively blinding the received information by adding data to each received ciphertext such that the information passed downstream is no longer the same as the information received from the CAM.

[0037] In an embodiment, the applied transformation introduces a random delay in the timing or frequency at which the received ciphertext or other information is passed along to downstream elements. In this embodiment, the delay can include delaying the timing of all or part of the original stream of ciphertext, or changing the timing of all or part of the original ciphertext while adding fake ciphertext to the stream of received information passed to downstream elements. In an embodiment, the method can change the frequency of the ciphertext by shifting some or all of the received ciphertext to disrupt any regular frequency patterns associated with the received ciphertext.

[0038] In an embodiment, the applied transformation randomly selects the received ciphertext and transmits only the selected ciphertext for downstream processing such as decryption and re-encryption by the HSM. In this embodiment, the transformation may randomly select the ciphertext using a random number generator to determine the next ciphertext for downstream processing. For example, from the first selected ciphertext, the transformation determines a random number indicating the number of ciphertexts to skip before the next ciphertext selected for downstream processing. Selecting only a portion of the entire set of ciphertexts for downstream processing changes the original pattern of the ciphertext message, improving the security of the FHE, but causing the entire stream of ciphertexts used by the CAM for circuit evaluation or machine learning model training to be more noisy. Since only a portion of the stream is passed to the HSM for decryption and re-encryption, the entire stream contains more noise. In this embodiment, the method can select the ciphertext according to the depth of the remaining ciphertext.

[0039] In an embodiment, the method balances the computational cost associated with various obfuscation techniques and the degree to which each technique or combination of techniques changes the original stream of information provided by the CAM.

[0040] In block 230, the method of the FHE security program 175 transmits the converted information from the OM to a downstream entity such as a hardware assistance module (HAM) like a hardware security module (HSM). The information transmitted to the HAM / HSM has been changed by the OM and is no longer the same as the stream of information received by the OM from the CAM.

[0041] The HAM / HSM converts the information received from the OM. In an embodiment, the HSM decrypts the information received from the OM using an encrypted secret key directly provided to the HSM by a client or a client device, and then re-encrypts the information. The steps of decrypting and re-encrypting the information reduce, or remove, the noise introduced by the original encryption applied to the original data by the client before transmitting the data from the client to the CAM for processing. By changing the noise level of the intermediate result, the CAM can use FHE from the client to train a more extensive and deeper machine learning model. Although the reduction of noise allows for the training of a deeper model, the CAM does not have the ability to decrypt the client's training data or view the client's unencrypted data. Since the stream of information passed from the CAM to the HAM is obfuscated, the HAM does not have the ability to derive meaningful information about the circuit under evaluation or the model being trained from the provided stream of information. After decrypting and re-encrypting the information received from the OM, the HAM transmits the re-encrypted information to the OM.

[0042] In block 240, the method of the FHE security program 175 has the OM receive the obfuscated information converted from the HAM / HSM. The original obfuscated information has been decrypted and re-encrypted by the HSM to remove the noise accumulated from the information.

[0043] In block 250, the OM applies the inverse of the original obfuscation transformation to the received information. For example, the OM removes the fake ciphertext from the received information and removes any additive blinding from the received ciphertext.

[0044] In block 260, the method of the FHE security program 175 sends the decrypted and transformed information back to the CAM. As an example, the OM decrypts the stream of re-encrypted ciphertext received from the HSM, removes the previously added fake ciphertext, and also removes any additive blinding artifacts. Then, the OM sends the decrypted stream of ciphertext to the CAM for further use in circuit evaluation or training of a machine learning model. In an embodiment, the method uses the received re-encrypted ciphertext when determining the parameters of the ML model. By using FHE data and encrypted ciphertext, FHE ML model parameters that cannot be read by the MLaaS provider can be generated.

[0045] In block 270, the method pulls back the FHE ML model parameters to the client device. In this embodiment, the client device decrypts the received FHE ML model parameters. Then, the client can use the decrypted ML model parameters to analyze new data.

[0046] Figure 3 provides a schematic diagram 300 of a networked computing environment according to an embodiment of the present invention. Schematic diagram 300 shows an environment for a scenario where a client owns a training dataset 311 and desires an ML model having ML model parameters 315 for analyzing data. As shown in the figure, the training data 311 becomes fully homomorphic encrypted (FHE) training data at 312 and is passed from the client device 310 to the MLaaS provider server 320. The client device 310 encrypts the training data at 312 using either the client secret key or the client public key from the client secret-public key pair to generate the FHE training data. The MLaaS server 320 trains the ML model 322 using the FHE training data. The progress of the training of the ML model results in the generation of intermediate results in the form of encrypted ciphertexts. These ciphertexts contain noise resulting from the use of the FHE training data. The method passes the encrypted ciphertexts through an obfuscation module (OM) 330 to change the ciphertext stream and hide information that can be otherwise identified about the ML model, including the number of model layers, the number of nodes per layer, the type of activation, and the training strategy. The OM 330 can add fake ciphertexts, change the frequency of the ciphertexts, randomly select only a portion of the overall ciphertext stream, or otherwise change the pattern of the ciphertext information.

[0047] As shown in the figure, the method passes the obfuscated ciphertexts from the OM 330 to an external HSM 340. The client device 310 provides the client secret key 318 of the client secret-public key pair associated with encrypting the FHE training data 312 to the HSM 340. The HSM decrypts the received ciphertext of the intermediate result using the received secret key 318 and then re-encrypts it. The decryption and subsequent re-encryption of the ciphertext reduce the noise of the ciphertext that enables further training of the ML model by the MLaaS server 320.

[0048] HSM340 returns the re-encrypted ciphertext back to the MLaaS provider server 320 via OM330. OM330 removes the obfuscation previously applied from the stream and performs the de-obfuscation of the stream of re-encrypted ciphertext.

[0049] The MLaaS server 320 utilizes the re-encrypted ciphertext to further train the ML model 322. The MLaaS server 320 can pass multiple iterations of the intermediate result ciphertext to HSM340 via OM330 and receive the corresponding re-encrypted ciphertext in return. The iterations can continue until the ML model is fully trained by the MLaaS server 320. The fully trained ML model includes a set of FHE ML model parameters that define the ML model. The FHE ML model parameters are encrypted and are not read by the MLaaS provider. The MLaaS server 320 passes the FHE ML model parameters to the client device 310. The client device 310 decrypts the FHE ML model parameters at 314 using the client private key. The client device 310 owns the decrypted ML model parameters 315 and can construct and use the ML model to analyze new data.

[0050] Figure 4 shows the communication messaging timeline between the CAM and the HSM without data obfuscation at 410 and with data obfuscation as described above at 420. As shown in the figure, the regular pattern and spacing of the messages 415 communicated along timeline 410 without the disclosed obfuscation were eliminated by obfuscating the information stream by adding an OM between the CAM and the HSM. The stream of obfuscated messages 425 arranged along timeline 420 no longer contains the details of the ML model architecture that have become apparent.

[0051] The disclosed embodiments provide for the use of MLaaS remote resources available from an edge cloud or a cloud resource provider. These resources enable an MLaaS consumer to access an ML architecture and associated computing environment resources as needed.

[0052] This disclosure includes a detailed description of cloud computing, but it should be understood that the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in combination with any other type of computing environment now known or later developed.

[0053] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services), where the resources can be rapidly provisioned and released with minimal management effort or service provider interaction. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0054] The characteristics are as follows.

[0055] On-demand self-service: A cloud consumer can unilaterally provision computing capabilities such as server time and network storage automatically as needed, without the need for human interaction with the service provider.

[0056] Broad network access: Computing capabilities are available over a network and can be accessed via standard mechanisms, thereby facilitating use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, PDAs).

[0057] Resource pooling: The provider's computing resources are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically assigned and re-assigned according to demand. In general, consumers have a sense of location independence since they do not manage or know the exact location of the provided resources. However, consumers may be able to specify location at a higher level of abstraction (e.g., country, state, data center).

[0058] Rapid elasticity: Computing capabilities can be provisioned quickly and elastically, allowing for some cases to scale out automatically and immediately, and to be released quickly and scale in immediately. To the consumer, the available computing capabilities for provisioning often appear limitless and can be purchased in any quantity at any time.

[0059] Measured service: Cloud systems leverage measurement capabilities at an appropriate level of abstraction for the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource use. Resource usage can be monitored, controlled, and reported to provide transparency to both the provider and consumer of the utilized service.

[0060] The service model is as follows.

[0061] Machine Learning as a Service (MLaaS): It is a function provided for consumers to use the provider's machine learning model architecture to train an ML model using the consumers' training data and then utilize the trained model. Consumers pass the data to the provider and receive the output from the trained model. Consumers do not manage the underlying cloud resources used for the model.

[0062] Software as a Service (SaaS): The function provided to consumers is that they can use the provider's application running on the cloud infrastructure. The application can be accessed from various client devices via a client interface such as a web browser (e.g., webmail). Consumers do not manage or control the underlying cloud infrastructure, including the network, server, operating system, storage, and even individual application functions. However, this does not apply to limited settings of user-specific application configurations.

[0063] Platform as a Service (PaaS): The function provided to consumers is to deploy the applications created or obtained by consumers to the cloud infrastructure using the programming languages and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including the network, server, operating system, and storage, but can control the deployed applications and, in some cases, also control the configuration of the hosting environment.

[0064] Infrastructure as a Service (IaaS): The functionality provided to consumers is to prepare processor, storage, network, and other basic computing resources that enable consumers to deploy and run any software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but can control the operating system, storage, and deployed applications, and in some cases, can partially control some network components (such as host firewalls).

[0065] The deployment models are as follows.

[0066] Private cloud: This cloud infrastructure is operated exclusively for a specific organization. This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.

[0067] Community cloud: This cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (such as mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by the organization or a third party and can exist on-premises or off-premises.

[0068] Public cloud: This cloud infrastructure is provided to an unspecified number of people or large industry groups and is owned by an organization that sells cloud services.

[0069] Hybrid Cloud: This cloud infrastructure combines two or more cloud models (private, community, or public). While retaining the entities specific to each model, they are bound by standard or individual technologies to achieve data and application portability (e.g., cloud bursting for load distribution between clouds).

[0070] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.

[0071] Here, FIG. 5 illustrates an exemplary cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10. In contrast, local computer devices used by cloud consumers (e.g., PDA or mobile phone 54A, desktop computer 54B, laptop computer 54C, or automotive computer system 54N or combinations thereof, etc.) can communicate. The nodes 10 can communicate with each other. The nodes 10 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, the private, community, public, or hybrid clouds described above or combinations thereof. Thereby, the cloud computing environment 50 can provide infrastructure, platform, software, or combinations thereof as a service, and cloud consumers do not need to maintain resources on local computer devices. Note that the types of computer devices 54A - N shown in FIG. 5 are merely exemplary, and it should be understood that the computing nodes 10 and the cloud computing environment 50 can communicate with any type of electronic device via any type of network or network addressable connection (e.g., using a web browser) or both.

[0072] Here, FIG. 6 shows a set of functional abstraction layers provided by the cloud computing environment 50 (FIG. 5). It should be understood in advance that the components, layers, and functions shown in FIG. 6 are merely exemplary, and the embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided.

[0073] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, a server 62 based on a reduced instruction set computer (RISC) architecture, server 63, blade server 64, storage device 65, as well as network and network components 66. In some embodiments, the software components include network application server software 67 and database software 68.

[0074] The virtualization layer 70 provides an abstraction layer. From this layer, for example, the following virtual entities can be provided: virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and operating systems 74, and virtual client 75.

[0075] As an example, the management layer 80 can provide the following functions. Resource provisioning 81 enables the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 enables cost tracking when resources are utilized within a cloud computing environment, and billing or invoicing for the consumption of these resources. As an example, these resources may include licenses for application software. Security enables not only the protection of data and other resources but also the identification and authentication of cloud consumers and tasks. The user portal 83 provides access to the cloud computing environment for consumers and system administrators. Service level management 84 enables the allocation and management of cloud computing resources so that the required service level is met. Planning and fulfillment of service quality assurance (SLA) 85 enables the advance arrangement and procurement of cloud computing resources expected to be needed in the future according to the SLA.

[0076] The workload layer 90 provides examples of functions available in a cloud computing environment. Examples of workloads and functions that can be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, delivery of virtual classroom education 93, data analysis processing 94, transaction processing 95, and the FHE security program 175.

[0077] The present invention can be a system, method, computer program product, or a combination thereof integrated at any possible level of technical detail. The present invention can be beneficially implemented in any single or parallel system that processes instruction streams. The computer program product may include a computer-readable storage medium storing computer-readable program instructions for causing a processor to execute aspects of the present invention.

[0078] The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. A more specific example of the computer-readable storage medium can include a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM (or flash memory), an SRAM, a CD-ROM, a DVD, a memory stick, a floppy disk, a punch card, a mechanically encoded device that records instructions in a raised structure within a groove, and suitable combinations thereof. As used herein, a computer-readable storage device should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0079] The computer-readable program instructions described in this specification can be downloaded from a computer-readable storage medium to respective computer devices / processing devices. Alternatively, they can be downloaded via a network (such as the Internet, LAN, WAN, or wireless network, or a combination thereof) to an external computer or external storage device. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computer device / processing device receives the computer-readable program instructions from the network and transfers them for storage in a computer-readable storage medium in each computer device / processing device.

[0080] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer as a stand-alone software package, or partially on the user's computer. Alternatively, it may be executed partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit for the purpose of implementing aspects of the present invention.

[0081] Embodiments of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block in the flowchart illustrations and / or block diagrams, and combinations of multiple blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0082] The above computer-readable program instructions may be provided to a computer, or a processor of other programmable data processing apparatus, to produce a machine. Thereby, these instructions, executed via the processor of such computer or other programmable data processing apparatus, create means for performing the functions / operations specified in one or more blocks in a flowchart and / or block diagram and / or both. The above computer-readable program instructions may further be stored in a computer-readable storage medium that can be instructed to function in a particular manner with respect to a computer, programmable data processing apparatus, or other device or combinations thereof. Thereby, the computer-readable storage medium in which the instructions are stored constitutes a product that includes instructions for performing the modes of functions / operations specified in one or more blocks in a flowchart and / or block diagram and / or both.

[0083] Alternatively, a computer-executable process may be generated by loading the computer-readable program instructions into a computer, other programmable apparatus, or other device and causing a series of operational steps to be executed on the computer, other programmable apparatus, or other device. Thereby, the instructions executed on the computer, other programmable apparatus, or other device perform the functions / operations specified in one or more blocks in a flowchart and / or block diagram and / or both.

[0084] The flowcharts and block diagrams in the drawings of the present disclosure illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for performing a particular logical function. In some other implementations, the functions shown within a block may be executed in an order different from the order shown in each figure. For example, two consecutive blocks shown may actually be executed simultaneously or substantially simultaneously, depending on the relevant functions, or may be executed in the reverse order in some cases. It should be noted that each block in a block diagram or flowchart or both, and combinations of multiple blocks in a block diagram or flowchart or both, can be executed by a dedicated hardware-based system that performs a particular function or operation, or executes a combination of dedicated hardware and computer instructions.

[0085] References to "one embodiment", "an embodiment", "exemplary embodiment", etc. in this specification indicate that the described embodiment may include a particular feature, structure, or characteristic, but not all embodiments necessarily include the particular feature, structure, or characteristic. Further, such phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in relation to an embodiment, it is submitted that it is within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in relation to other embodiments, whether or not explicitly described.

[0086] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit the present invention. The terms used in this specification are for the sole purpose of describing particular embodiments and are not intended to be limiting. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. As used herein, the term "comprises" or "comprising" or both thereof specify the presence of the stated feature, integer, step, operation, element, or component or combination thereof, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element components, or groups thereof or combinations thereof.

[0087] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limiting. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the present invention. The terms used herein were chosen in order to best explain the principles of the embodiments, the practical application or technical improvement of technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. Receiving fully homomorphic encryption (FHE) information from a client device; Training a machine learning (ML) model using the FHE information to generate an FHE ciphertext; Applying a first transformation to the FHE ciphertext to generate an obfuscated FHE ciphertext; Transmitting the obfuscated FHE ciphertext to a secure device; Receiving a re-encrypted version of the obfuscated FHE ciphertext from the secure device; Applying a second transformation to the re-encrypted version of the obfuscated FHE ciphertext to generate a de-obfuscated re-encrypted FHE ciphertext; Training the ML model using the re-encrypted FHE ciphertext to generate FHE ML model parameters; Transmitting the FHE ML model parameters to the client device. A computer-implemented method comprising the steps of

2. The computer-implemented method according to claim 1, wherein the FHE information includes data encrypted using a private key of the client device or a public key of the client device.

3. The computer-implemented method according to claim 1, wherein the re-encrypted FHE ciphertext includes the obfuscated FHE ciphertext re-encrypted using a private key of the client device.

4. The computer-implemented method according to claim 1, wherein the secure device includes a hardware security module (HSM).

5. The computer-implemented method according to claim 1, wherein applying the first transformation obfuscates the frequency pattern of the FHE ciphertext.

6. The computer-implemented method according to claim 1, wherein the first transformation is randomly applied to the FHE ciphertext.

7. The computer-implemented method according to claim 1, wherein applying the first transformation adds spurious data to the FHE information.

8. A computer program for protecting homomorphic encryption data, the computer program causing a computer to Receive fully homomorphic encryption (FHE) information from a client device; Train a machine learning (ML) model using the FHE information to generate an FHE ciphertext; Apply a first transformation to the FHE ciphertext to generate an obfuscated FHE ciphertext; Transmit the obfuscated FHE ciphertext to a secure device; Receive a re-encrypted version of the obfuscated FHE ciphertext from the secure device; Apply a second transformation to the re-encrypted version of the obfuscated FHE ciphertext to generate a de-obfuscated re-encrypted FHE ciphertext, and Train the ML model using the re-encrypted FHE ciphertext to generate FHE ML model parameters, and A computer program for realizing a function of transmitting the FHE ML model parameters to the client device. **Claim 9** The computer program according to claim 8, wherein the FHE information includes data encrypted using a private key of a client device or a public key of the client device. **Claim 10** The computer program according to claim 8, wherein the re-encrypted FHE ciphertext includes an obfuscated FHE ciphertext re-encrypted using a private key of a client device. **Claim 11** The computer program according to claim 8, wherein the secure device includes a hardware security module (HSM). **Claim 12** Applying the first transformation obfuscates the frequency aspect of the FHE ciphertext, according to the computer program of claim 8. **Claim 13** The first transformation is randomly applied to the FHE ciphertext, according to the computer program of claim 8. **Claim 14** Applying the first transformation adds spurious data to the FHE information, according to the computer program of claim 8. **Claim 15** A computer system for protecting homomorphic encrypted data, the computer system comprising One or more computer processors, and One or more computer-readable storage devices, and Program instructions stored on the one or more computer-readable storage devices for execution by the one or more computer processors, the stored program instructions comprising Program instructions for receiving fully homomorphic encryption (FHE) information from a client device, and Program instructions for training a machine learning (ML) model using the FHE information to generate an FHE ciphertext, and Program instructions for applying a first transformation to the FHE ciphertext to generate an obfuscated FHE ciphertext, and Program instructions for transmitting the obfuscated FHE ciphertext to a secure device, and Program instructions for receiving a re-encrypted version of the obfuscated FHE ciphertext from the secure device, and Program instructions for applying a second transformation to the re-encryption version of the obfuscated FHE ciphertext to generate a de-obfuscated re-encrypted FHE ciphertext, Program instructions for training the ML model using the re-encrypted FHE ciphertext to generate FHE ML model parameters, Program instructions for transmitting the FHE ML model parameters to the client device, comprising a computer system.

16. The computer system according to claim 15, wherein the FHE information includes data encrypted using a private key of the client device or a public key of the client device.

17. The computer system according to claim 15, wherein the re-encrypted FHE ciphertext includes an obfuscated FHE ciphertext re-encrypted using a private key of the client device.

18. The computer system according to claim 15, wherein the secure device includes a hardware security module (HSM).

19. The computer system according to claim 15, wherein applying the first transformation obfuscates the frequency aspect of the FHE ciphertext.

20. The computer system according to claim 15, wherein the first transformation is randomly applied to the FHE ciphertext.

Citation Information

Patent Citations

  • Information processing method and information processing system

    JP2019168590A

  • Method for confidential execution of a program operating on data encrypted by a homomorphic encryption

    US20170244553A1

  • Searching Over Encrypted Model and Encrypted Data Using Secure Single-and Multi-Party Learning Based on Encrypted Data

    US20200366459A1