Method, neural network, computer program (neural network training using homomorphic encryption)

JP7926943B2Active Publication Date: 2026-09-30INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023044209
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-21
Filing Date
2023-03-20
Publication Date
2026-09-30
Estimated Expiration
2043-03-20

Smart Images

  • Figure 0007926943000004
    Figure 0007926943000004
  • Figure 0007926943000005
    Figure 0007926943000005
  • Figure 0007926943000006
    Figure 0007926943000006
Patent Text Reader

Abstract

To provide a neural network training optimization method using homomorphically encrypted elements and a dropout algorithm for regularization, a neural network, and a computer program.SOLUTION: The method includes receiving, via a neural network, a training dataset containing samples that are encrypted using homomorphic encryption. The method also includes determining a packing formation and selecting a dropout technique during training of the neural network based on the packing technique. The method further includes starting with a first packing formation from the training dataset, and inputting the first packing formation in an iterative or recursive manner into the neural network using the selected dropout technique, with a next packing formation from the training dataset acting as an initial input that is applied to the neural network for a next iteration, until a stopping metric is produced by the neural network.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to neural network training using homomorphically encrypted elements, and more specifically relates to optimized training of neural networks using homomorphically encrypted elements and a dropout algorithm for regularization. [Background Art]

[0002] Neural networks, or artificial neural networks, are a subset of machine learning comprising an input layer, one or more hidden layers, and an output layer. Each node, or artificial neuron, connects to another and has associated weights and a threshold. If the output of any individual node exceeds a specified threshold, the node is activated and sends data to the next layer of the network. Various types of neural networks exist, such as modular neural networks, recurrent neural networks, generative adversarial networks, deep neural networks, spiking neural networks, feedforward neural networks, and physical neural networks. [Summary of the Invention] [Problem to be Solved by the Invention]

[0003] Homomorphic encryption is a form of encryption that allows a user to perform computations on encrypted data without first decrypting the data. The resulting computation remains in encrypted form, which, when decrypted, yields an output that is the same as or similar to the output produced if the operation had been performed on unencrypted data. This allows data to be encrypted and outsourced to other environments such as commercial cloud environments for processing while remaining in an encrypted state. [Means for Solving the Problem]

[0004] Embodiments of the present disclosure include a method for optimizing the training of a neural network using homomorphically encrypted elements having a dropout algorithm for regularization. The method includes receiving a training dataset containing samples encrypted using homomorphic encryption via input to the neural network, and determining a packing configuration used by the homomorphic encryption to pack the samples in the training dataset. The method also includes selecting a dropout technique during training of the neural network based on the packing technique. The method further includes starting with a first packing configuration from the training dataset and inputting the first packing configuration into the neural network iteratively or recursively using the selected dropout technique, such that the next packing configuration from the training dataset acts as an initial input applied to the neural network for the next iteration, until a stopping metric is generated by the neural network.

[0005] Additional embodiments of the present disclosure include a computer program product for optimizing the training of a neural network using homomorphically encrypted elements having a dropout algorithm for regularization, one or more computer-readable storage media, and program instructions stored in the one or more computer-readable storage media, wherein the program instructions are executable by a processor and cause the processor to receive a training dataset containing samples encrypted using homomorphic encryption via input to a neural network, and to determine a packing configuration used by the homomorphic encryption used to pack the samples in the training dataset. The computer program product also includes instructions causing the processor to select a dropout technique during training of the neural network based on the packing technique. The computer program product also includes instructions causing the processor to start with a first packing configuration from the training dataset, and to input the first packing configuration into the neural network iteratively or recursively using the selected dropout technique, with the next packing configuration from the training dataset acting as an initial input applied to the neural network for the next iteration, until a stop metric is generated by the neural network.

[0006] Further embodiments of the present disclosure include a neural network trained using homomorphically encrypted elements having a dropout algorithm for regularization. The neural network may be implemented in a system including memory, a processor, and local data storage in which computer executable code is stored. The computer executable code includes program instructions that can be executed by the processor to cause the processor to perform the above method. This summary is not intended to illustrate any implementation, embodiment, or combination thereof of the present disclosure. [Brief explanation of the drawing]

[0007] These and other features, aspects and advantages of the embodiments of this disclosure will be better understood with respect to the following description, the appended claims and the appended drawings.

[0008] [Figure 1] This is a block diagram illustrating the operation of the key operating elements for training a neural network, and is used in one or more embodiments of the present disclosure.

[0009] [Figure 2] This is a block diagram showing different groups of neurons for dropout on a neural network, used by one or more embodiments of the present disclosure.

[0010] [Figure 3A] This is a block diagram illustrating different dropout techniques for inputting batch inputs into a neural network, as used by one or more embodiments of the present disclosure. [Figure 3B] This is a block diagram illustrating different dropout techniques for inputting batch inputs into a neural network, as used by one or more embodiments of the present disclosure.

[0011] [Figure 4] This is a flowchart illustrating the process of training a neural network performed according to embodiments of the present disclosure.

[0012] [Figure 5] This is a top block diagram showing an exemplary computer system that may be used in any implementation of one or more of the methods, tools, modules, and related functions described herein that may be implemented.

[0013] [Figure 6]1 illustrates a cloud computing environment according to an embodiment of the present disclosure.

[0014] [Figure 7] 1 illustrates an abstraction model layer according to an embodiment of the present disclosure.

[0015] While the present disclosure is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the disclosure is not intended to be limited to the particular forms disclosed. On the contrary, the disclosure is intended to cover all modifications, equivalents, and alternatives falling within the scope of the present disclosure. In the accompanying drawings, like reference numerals are used to designate like parts. DETAILED DESCRIPTION OF EMBODIMENTS

[0016] The present disclosure relates to neural network training using homomorphically encrypted elements, and more specifically, to optimized training of neural networks using homomorphically encrypted elements and a dropout algorithm for regularization. While the present disclosure is not necessarily limited to such applications, various aspects of the disclosure may be understood through discussion of various examples using this context.

[0017] Homomorphic encryption enables any computation to be performed on encrypted data (ciphertext) without decryption. Given a public key pk, a secret key sk, an encryption function Math. and a decryption function σ(), the homomorphic encryption operation Math. can be defined when there exists another operation ×, such that Math. where χ1 and χ2 are plaintext, each of which encodes a vector consisting of a plurality of integers or fixed points. Approximate homomorphic encryption can also be used, as a result of which the result is similarly and easily decrypted, and a certain amount of noise is taken into account. It should be noted that each homomorphic encryption operation brings a certain amount of noise to the ciphertext. Decryption failure can occur when accumulated noise increases beyond the noise budget.

[0018] Homomorphic encryption can be applied and used as an input to a neural network that can be regarded as a privacy-preserving neural network. An artificial neural network, or neural network, consists of an input layer of neurons (or nodes, units), a hidden layer of a plurality of neurons, and a final layer of output neurons. Neural networks interconnect neurons that use mathematical or computational models for information processing. Neural networks include, but are not limited to, modular neural networks, recurrent neural networks, generative adversarial networks, deep neural networks, spiking neural networks, feedforward neural networks, and physical neural networks.

[0019] Neural networks have tens of thousands of parameters, and some such networks can have millions of parameters. With such a large number of parameters, neural networks can be flexible and can be adapted to a variety of complex data sets. However, a drawback of this complexity is that neural networks can tend to overfit to their training sets. There are several techniques that can avoid overfitting to the training set. For example, a technique called early stopping is introduced, which can interrupt training when the performance of the neural network on a validation set begins to degrade. Another regularization technique is l1 and l2 regularization that constrains the connection weights of the neural network. Dropout is also a common regularization technique used to prevent overfitting.

[0020] In dropout, at every training stage, every neuron (including input neurons but excluding output neurons) can be temporarily "dropped out," meaning that it will be completely ignored during that training stage but may become active in the next stage.

[0021] Encrypting individual values ​​for each input neuron can result in very large ciphertexts, which can be inconvenient from a user's perspective and, as a direct consequence, may require high bandwidth. To mitigate this problem, multiple values ​​can be "packed" into a single ciphertext. Using multiple elements packed into a single ciphertext, a ciphertext mask can be used to perform an operation on a specific element within that ciphertext. For example, masking can be applied to a ciphertext with three elements so that only the third element is used.

[0022] However, the limitations of using homomorphic encryption inputs for neural networks remain because training these models becomes computationally intensive due to dropout and masking increases the multiplication depth of these neural networks. Multiplication depth, as referred to herein, refers to the longest number of consecutive multiplication operations that can be performed before a bootstrap operation is required. Consecutive operations can mean a series of multiplication operations before being interrupted by an addition operation. The maximum number of multiplication operations in a pass can be the number between inputs and outputs, or between bootstrap operations. Furthermore, current machine learning techniques using homomorphic encryption are limited to simple, non-standard types of machine learning models. These non-standard machine learning models have not proven efficient and accurate when dealing with more practical and complex datasets.

[0023] Embodiments of the present disclosure may overcome the above and other problems by providing dropout methodologies including single neurons per batch dropout, groups of neurons per sample dropout, and groups of neurons per batch dropout technique. These methodologies may reduce the multiplication depth of circuits operating for fully homomorphic encryption by omitting some plaintext masking. The dropout methodologies provide multiple packing and dropout options that can be selected based on the performance and accuracy of the model.

[0024] More specifically, embodiments of the present disclosure provide a technique for training a neural network using homomorphic cryptographic inputs that remove or reduce masking requirements, at least in part based on packing or dropout techniques used to train the neural network. For example, training samples may be packed into ciphertext that passes through the neural network. Neurons are also packed together within the ciphertext based on the packing of the training samples. When dropout selects neurons to be ignored that would require masking, rather than neurons that are simply ignored, an entire block of ciphertext is ignored.

[0025] For example, but not limited to, a neural network is trained using homomorphic encrypted inputs. The training dataset is homomorphically encrypted using a predetermined packing configuration. This packing configuration may involve each sample being in its own ciphertext, or multiple samples being packed into a ciphertext to form a batch. Note that other packing configurations may be implemented and applied to the training dataset. Based on the packing configuration, a dropout technique is selected, and neurons in the neural network are packed into neuron groupings corresponding to that dropout technique. These neuron groupings may be based on a method that avoids the need to mask data when neurons are dropped during training. Once established, the neural network may input a first packing configuration from the training dataset and, in an iterative manner, begin training the neural network using the next packing configuration from the training dataset, which acts as the initial input to the neural network for the next iteration. This process continues until a stopping metric is achieved. The stopping metric could be, for example, when the neural network achieves a predetermined accuracy during validation, or when accuracy no longer improves after the training cycle.

[0026] In some embodiments, the neurons of the neural network are packed into neuron groups. These neuron groups may group neurons by layer, resulting in a neuron group containing only neurons within the same layer. For example, for illustrative purposes only, a layer may contain four neurons, with one neuron group containing the first and second neurons, and another neuron group containing the third and fourth neurons. Neurons in previous or subsequent layers cannot be included in any group. In some embodiments, the neuron groups are based on the packing configuration used by homomorphic encryption used to pack samples into a training dataset. For example, a training sample may be converted into a ciphertext having four features. Based on that packing, a neuron group may consist of four neurons.

[0027] In some embodiments, the packing configuration is based on the dropout technique implemented on the neural network. For example, the neural network may utilize a dropout technique that drops a group of neurons instead of a single neuron. When a neuron is selected for dropping, the group of neurons to which that neuron belongs is dropped entirely. Dropping entire groups of neurons eliminates the need for masking, thereby reducing the multiplication depth of the circuit manipulating the fully homomorphic encrypted input.

[0028] In some embodiments, the packing configuration of the input (e.g., a training dataset) packs multiple samples within a ciphertext. The ciphertext may contain at least two samples that can improve the optimization of the neural network. The number of samples per ciphertext may be based on the noise generated by the neural network when it is input. For example, the ciphertext may have a sample threshold, in which case the noise generated by the neural network may be excessively large.

[0029] Figure 1 is a block diagram of a neural network training environment 100 for optimized training of a neural network using homomorphically encrypted elements and a dropout algorithm for regularization. The neural network training environment 100 includes a training dataset 110, a homomorphic encryption module 120, a packing configuration 130, and a neural network 140. The neural network includes neurons, synapses, and neuron groups 148-1, 148-2, 148-3, 148-4, 148-5, 148-6, 148-7, and 148-8. For the purposes of this specification, the exemplary embodiments will be assumed to be implemented as part of a machine learning training methodology. However, this is only one possible implementation and is not intended to limit the disclosure. Other implementations in which machine learning techniques are utilized may also be used without departing from the spirit and scope of this disclosure.

[0030] The training dataset 110 is a set of data from the neural network training environment 100 used as training data for the neural network 140. The training dataset 110 contains a set of samples, each containing a label with one or more features. In some embodiments, the training dataset 110 is divided into a training set, a validation set, and a test set. The validation set may be a subset of the training dataset 110 used to validate the pseudo-labeled dataset generated by the neural network 140. The test set may also be a subset of the training dataset 110 used to test the neural network 140 after training and validation.

[0031] The training dataset 110 may be used, for example, for image classification, in which case the training data may include images with known classifications. The training dataset 110 may also include data that should be handled carefully, personal, confidential, or a combination thereof. Therefore, it is necessary to process such information in an encrypted form by encrypting such data and training the neural network 140 with it. Thus, during decryption, the neural network 140 can operate accurately on encrypted information.

[0032] The homomorphic encryption module 120 is a component of the neural network training environment 100 configured to homomorphically encrypt the training dataset 110. The homomorphic encryption module 120 provides an encryption system for encrypting the training dataset 110, so that computations can be performed on the encrypted data without decryption, thereby enabling secure outsourced computations. In some embodiments, the homomorphic encryption module 120 encrypts multiple samples into a single packed ciphertext known as a batch. Single instruction multiple data (SIMD) techniques may be used to perform operations on those values ​​in parallel. In some embodiments, each sample is packed into a ciphertext. A packing configuration 130 is generated, which is the ciphertext produced by the homomorphic encryption module 120 when encrypting the training dataset 110. The packing configuration 130 may describe how the data is packed, whether in batch format, in the form of one ciphertext per sample, or in some other packing configuration.

[0033] In addition, the homomorphic encryption module 120 may apply multiple types of encryption schemes that can perform different classes of computations on the encrypted data. These schemes include, but are not limited to, partial homomorphic encryption, part-of-the-order homomorphic encryption, leveled full homomorphic encryption, and full homomorphic encryption.

[0034] The neural network 140 is a component of the neural network training environment 100, which is trained based on the input training dataset 110. The neural network 140 may include multiple neurons (nodes) arranged in various layers. These neurons form adjacent layers, and they include connections or edges between them. Each connection between neurons may have associated weights that can assist the neural network 140 in evaluating the received input. The neural network 140 also includes neuron groups 148-1, 148-2, 148-3, 148-4, 148-5, 148-6, 148-7, 148-N (collectively, "neuron group 148"), where N is a variable integer representing any number of possible neuron groups 148.

[0035] The input layer 142 is a layer of the neural network 140 configured to provide information to the neural network 140. The input layer 142 may receive inputs such as encrypted samples from the packing of the dataset 110 in the packing configuration 130 and feed them into the neural network 140. In some embodiments, the input layer 142 may input additional information into the neural network 140. For example, the input layer 142 may input the packing configuration 130 in the form of batches having multiple samples in a single ciphertext, where each input neuron inputs features from the sample.

[0036] The hidden layer 144 is a layer of the neural network 142 configured to perform computations and transfer information from one layer to another. The hidden layer 144 may contain a set of hidden neurons to form the hidden layer 144. While Figure 1 shows only three layers, each with four neurons, it will be understood that the neural network 140 may contain multiple hidden layers, each with multiple neurons, depending on the configuration of the neural network 140.

[0037] Overfitting and underfitting of the input data can be addressed and mitigated through dropout techniques. Low bias may result in the neural network 140 overfitting the data, while high bias may result in the neural network 140 underfitting the data. Overfitting occurs when the neural network 140 learns its training data well but cannot generalize beyond that training data. Underfitting occurs when the neural network 140 is unable to generate accurate predictions for the training data or enable data.

[0038] Dropout, also commonly referred to as dilution or dropconnect, is a regularization technique used to reduce overfitting in neural networks. This dropout technique involves omitting neurons (both hidden and visible) during the training process of the neural network. In addition, the weights associated with synapses can be reduced or "thinned" independently of or in conjunction with the omission of neurons.

[0039] The process by which neurons are dropped or driven to zero can be achieved by setting their weights to zero, ignoring the neuron, or negate the neuron's computation by any other means that does not affect the final result or create a new, unique case. In some embodiments, when a neuron is selected for dropout, the entire neuron group 148 to which that neuron belongs will be dropped.

[0040] Neuron group 148 is a group of neurons within a layer packed into a single ciphertext. As shown in Figure 1, each layer contains four neurons, each having neuron group 148 with two neurons per group. However, it should be noted that other configurations are possible. For example, a layer can have any number of neurons, having various combinations of neuron groups with that layer. A layer may contain neuron groups of only one neuron, and other neuron groups of several neurons.

[0041] The configuration of neuron groups can be configured to optimize the performance of the neural network 140. If, during testing, a neuron group of a particular size performs such that the neural network achieves a higher performance metric, then the group can be selected for the entire neural network 140. In some embodiments, neuron groups 148 are selected based on a packing configuration 130 of the input data. For example, if the packing configuration includes a sample having four features, then neuron group 148 can be packed with four neurons as each feature is input to the neurons during testing. Neuron groups 148 can also be packed and dropped to avoid the need to mask data, thereby reducing the multiplication depth of the neural network 140. The output layer 146 is a layer of the neural network 140 configured to transfer information from the neural network 140 to an external destination.

[0042] During the training stage, the neural network 140 learns the optimal weights for each neuron. The optimal configuration can then be applied to test data. Exemplary applications of such a neural network 140 include image classification, object recognition, speech recognition, or data that may be sensitive, private, privileged, personal, confidential, or a combination thereof.

[0043] It should be noted that Figure 1 is intended to show typical key components of an exemplary neural network training environment 100. In some embodiments, individual components may have greater or less complexity than those shown in Figure 1, and there may be components other than or in addition to those shown in Figure 1, and the number, type, and configuration of such components may vary.

[0044] Figure 2 is a block diagram showing a neural network configuration 200 having a neural network with different groups of neurons in each layer, according to an embodiment of the present disclosure. The neural network 200 includes layers 220, 230, 240, and 250.

[0045] Layer 220 contains four neurons in two neuron groups. Each neuron group has two neurons present in layer 220. Layer 230 also contains four neurons; however, these neurons are grouped individually. The neuron groups in layer 230 can be seen as single neuron groups or as a layer without any neuron groups. Layer 240 contains groups that include all of the neurons within that layer, while layer 250 contains different neuron group sizes within its single layer 250. Layer 250 contains a neuron group of two neurons and two other neuron groups, each having a single neuron.

[0046] The neuron group configurations in Figure 2 are for illustrative purposes only and are used solely to illustrate various ways in which the neural network 200 may be composed of different neuron groups. In some embodiments, the neural network 200 includes neuron groups of comparable size, while in some other embodiments, the neural network 200 includes neuron groups that are sized differently. It should be noted that the neuron group configuration is determined based on the performance and accuracy of the neural network, and multiple neuron group configurations may be tested to achieve the most optimal approach in terms of performance and dropout.

[0047] Figures 3A and 3B are block diagrams showing a neural network 310 having a batch packing configuration as input using dropout technique. First, in Figure 3A, the neural network 310 is input 320 in a homomorphic cryptographic batch packing configuration. This batch packing configuration includes batches 320-1, 320-2, 320-3, and 320-4 (collectively, "batch 320"). Batch 320 may contain features from multiple samples within the same ciphertext packing configuration. For example, batch 320-1 may contain features of the same type from three different samples. Batch 320 can be input to the neural network 310 simultaneously or sequentially to improve performance and utility.

[0048] Once the batch packing configuration 320 is input to the neural network and training occurs, neurons are selected for dropout, as indicated by the neurons along with the diagonal pattern. In addition, synapses connected to the selected neurons are effectively turned off, as indicated by the dashed lines connecting them to the selected neurons. Thus, in some embodiments, the homomorphic cryptographic batch packing configuration 320 can be input to a neural network 310 in which one neuron is selected for dropout.

[0049] Again in Figure 3B, we see a neural network 310 receiving input 320 in a homomorphic encryption batch packing configuration. This batch packing configuration includes batches 320-1, 320-2, 320-3, and 320-4 (collectively, “batch 320”). Batch 320 may contain features from multiple samples within the same ciphertext packing configuration. For example, batch 320-1 may contain features of the same type from three different samples. Batch 320 can be input to the neural network 310 simultaneously or sequentially to improve performance and utility.

[0050] Once the batch packing configuration 320 is input to the neural network and training occurs, neurons are selected for dropout, as indicated by the neurons along with the diagonal pattern. However, in this case, the neuron group 330 in which the dropped-out neurons reside is also dropped. Therefore, both neurons in neuron group 330 are dropped for its training iterations. In addition, the synapses connected to both of these neurons are effectively turned off, as indicated by the dashed lines connecting them to the selected neurons. Thus, in some embodiments, the homomorphic cryptographic batch packing configuration 320 may be input to a neural network 310 in which one neuron is selected for dropout, causing its neuron group 330 to be similarly dropped for its training iterations.

[0051] Therefore, the exemplary embodiment provides a mechanism for optimizing the training of a neural network using homomorphic encrypted data by utilizing a dropout technique that can reduce the multiplicative depth of the neural network. In addition, the mechanism of the exemplary embodiment may work in conjunction with other neural network training techniques or other computing systems or combinations thereof to perform operations that utilize machine learning models and techniques to optimize the performance of neural networks, more specifically, to optimize a neural network inputting homomorphic encrypted data.

[0052] Figure 4 is a flowchart illustrating a process 400 for training a neural network using homomorphic encryption elements and a dropout algorithm for regularization, according to embodiments of the present disclosure. As shown in Figure 4, the process 400 begins with receiving a training dataset 110 encrypted using homomorphic encryption 120 via input to the neural network 140. This is shown in step 410. The training dataset may consist of a set of samples, each containing a label with one or more features. In some embodiments, the training dataset 110 is divided into a training set, a validation set, and a test set. The validation set may be a subset of the training dataset 110 used in validating a pseudo-labeled dataset generated by the neural network 140.

[0053] A homomorphic encryption packing configuration 130 for the training dataset 110 is determined. This is shown in step 420. The packing configuration 130 may describe how the data is packed into ciphertext, whether in batch format, one ciphertext per sample format, or some other packing configuration. Once the packing configuration is determined, a dropout technique is selected based on the type of packing configuration 130 used to homomorphically encrypt the training dataset 110. This is shown in step 430. The dropout technique organizes the neurons in the neural network 140 into neuron groups 148 so that when a neuron is dropped, a mask may not be required when the neuron is dropped for that training iteration. For example, the packing configuration 130 packs the samples and features such that the neuron group 148 has four neurons per group. When four neurons per group are used, in this particular case, then a mask does not need to be applied when the entire neuron group is dropped. The configuration of the neuron groups can also be configured to optimize the performance of the neural network 140.

[0054] In some embodiments, the packing configuration is selected based on the dropout technique and the neuron groups 148 of the neural network 140. For example, an optimized neuron group 148 for the neural network 140 is two neurons per neuron group 148. Thus, the packing configuration 130 can pack the sample in a way that accommodates the dropout technique and the neuron groups 148 and avoids the need to use a mask during dropout.

[0055] The neural network is trained using a homomorphic encryption dataset 110, as shown in step 440. Once a packing configuration 130 is determined and a dropout technique is set for the configured neuron groups, the neural network 140 may start with a first packing configuration 130 from the training dataset 110 and input that configuration 130 into the neural network 140. In an iterative or recursive manner, training continues using the selected dropout technique applied to account for regularization by inputting the next packing configuration 130 until the neural network generates a stopping metric. The stopping metric may interrupt training when the neural network's performance against the validation set begins to degrade or when the neural network 140 achieves a predetermined level of accuracy.

[0056] Referring here to Figure 5, a top block diagram (e.g., a neural network training environment 100) of an exemplary computer system 500 can be used in the implementation of one or more of the methods, tools, and modules described herein, and any associated functions, according to embodiments of the present disclosure (e.g., using one or more processor circuits or computer processors of a computer). In some embodiments, the main components of the computer system 500 may include one or more processors 502, memory 504, terminal interface 512, I / O (input / output) device interface 514, storage interface 516, and network interface 518, all of which may be connected to communicate directly or indirectly via a memory bus 503, an I / O bus 508, and an I / O bus interface 510 for intercomponent communication.

[0057] The computer system 500 may include one or more general-purpose programmable central processing units (CPUs) 502-1, 502-2, 502-3, and 502-N, which are collectively referred to herein as processors 502 (e.g., graphics processing units, physical processing units, application-specific integrated circuits, field-programmable gate arrays). In some embodiments, the computer system 500 may include multiple processors, which is typical for relatively large systems; however, in other embodiments, the computer system 500 may alternatively be a single-CPU system. Each processor 502 may execute instructions stored in memory 504 and may include one or more levels of onboard cache.

[0058] Memory 504 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 522 (for example, processing within memory) or cache memory 524. Computer system 500 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example only, storage system 526 may be provided for reading from and writing to a non-removable non-volatile magnetic medium, such as a “hard drive”. Not shown, a magnetic disk drive may be provided for reading from and writing to a removable non-volatile magnetic disk (for example, a “floppy disk”), or an optical disk drive may be provided for reading from and writing to a removable non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical media. In addition, memory 504 may include flash memory (for example, a flash memory stick drive or flash drive). Memory devices may be connected to the memory bus 503 by one or more data medium interfaces. The memory 504 may include at least one program product having a set of program modules (e.g., at least one) configured to perform functions of various embodiments.

[0059] Although the memory bus 503 is shown in Figure 5 as a single-bus structure providing a direct communication path between the processor 502, memory 504, and I / O bus interface 510, the memory bus 503 may, in some embodiments, include multiple different buses or communication paths, which can be configured in any variety of forms, such as hierarchical, star or web configurations, point-to-point links, multiple hierarchical buses, parallel and redundant paths, or any other suitable type of configuration. Furthermore, while the I / O bus interface 510 and I / O bus 508 are shown as single, respective units, the computer system 500 may, in some embodiments, include multiple I / O bus interface units, multiple I / O buses, or both. Additionally, while multiple I / O interface units are shown to isolate the I / O bus 508 from various communication paths running to various I / O devices, in other embodiments, some or all of such I / O devices may be directly connected to one or more system I / O buses.

[0060] In some embodiments, the computer system 500 may be a multi-user mainframe computer system, a single-user system, or a server computer or similar device that has little or no direct user interface but receives requests from other computer systems (clients). Furthermore, in some embodiments, the computer system 500 may be implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, a network switch or router, or any other suitable type of electronic device.

[0061] It should be noted that Figure 5 is intended to show the main representative components of an exemplary computer system 500. However, in some embodiments, the individual components may be more or less complex than those shown in Figure 5, and there may be components other than or added to those shown in Figure 5, and the number, type, and configuration of such components may vary.

[0062] One or more programs / utilities 528, each having at least one set of program modules 530 (e.g., a neural network training environment 100), can be stored in memory 504. A program / utility 528 may include a hypervisor (also called a virtual machine monitor), one or more operating systems, one or more application programs, other program modules, and program data. Each of the operating systems, one or more application programs, other program modules, and program data, or any combination thereof, may include the implementation of a networking environment. A program 528 or program module 530 or any combination thereof generally performs functions or methods of various embodiments.

[0063] While this disclosure includes a detailed description of cloud computing, it should be understood that implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of this disclosure can be implemented in conjunction with any other type of computing environment that is currently known or may be developed in the future.

[0064] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal administrative effort or minimal interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.

[0065] The characteristics are as follows:

[0066] On-demand self-service: Cloud consumers can automatically and unilaterally provision computing power, such as server time and network storage, as needed, without the need for human interaction with the service provider.

[0067] Broad network access: Functionality is available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0068] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model where different physical and virtual resources are dynamically allocated and reallocated according to demand. Consumers generally do not have control or knowledge of the exact location of the resources provided, but there is location independence in that they may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0069] Rapid Adaptability: Features can be provisioned quickly and adaptively, sometimes automatically, to scale out rapidly, and to release quickly and scale in rapidly. To consumers, the capacity available for provisioning often appears unlimited and can be purchased in any quantity at any time.

[0070] Measurement Services: Cloud systems automatically control and optimize resource usage by leveraging measurement capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and users.

[0071] The service model is as follows:

[0072] Software as a Service (SaaS): The capabilities provided to consumers utilize the provider's applications running on cloud infrastructure. These applications are accessible from various client devices through thin client interfaces such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0073] Platform as a Service (PaaS): The functionality offered to consumers is the ability to deploy applications they have created or acquired, written using programming languages ​​and tools supported by the provider, onto cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they do control the deployed applications and, in some cases, the configuration of the application hosting environment.

[0074] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provisioning of processing, storage, networking, and other fundamental computing resources, allowing consumers to deploy and run any software, which may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do control the operating system, storage, and deployed applications, and in some cases, have limited control over selected networking components (e.g., host firewalls).

[0075] The deployment model is as follows:

[0076] Private Cloud: A cloud infrastructure operated for a single organization only. A private cloud may be managed by that organization or a third party, and may reside on-premises or off-premises.

[0077] Community Cloud: A cloud infrastructure shared by several organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by these organizations or third parties and may reside on-premises or off-premises.

[0078] Public cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services.

[0079] Hybrid Cloud: Cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain separate entities but are linked together by standard or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0080] Cloud computing environments are service-oriented and focus on statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure that includes a network of interconnected neurons.

[0081] Referring here to Figure 6, an exemplary cloud computing environment 600 is shown. As shown, the cloud computing environment 600 includes one or more cloud computing nodes 610 to which local computing devices used by cloud consumers, such as a personal digital assistant (PDA) or mobile phone 620-1, a desktop computer 620-2, a laptop computer 620-3, or an automotive computer system 620-4, or a combination thereof, can communicate. The nodes 610 may communicate with each other. They may be physically or virtually grouped within one or more networks, such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above (not shown). This makes it possible for the cloud computing environment 600 to provide infrastructure, platform, or software, or a combination thereof, as a service that does not require cloud consumers to maintain resources on their local computing devices for that purpose. It should be understood that the types of computing devices 620-1 to 620-4 shown in Figure 6 are for illustrative purposes only, and that the computing node 610 and the cloud computing environment 600 may communicate with any type of computerized device over any type of network and / or network-addressable connection (e.g., using a web browser).

[0082] Referring now to Figure 7, a set of functional abstraction layers 700 provided by the cloud computing environment 600 (Figure 6) is shown. It should be understood that the components, layers, and functionalities shown in Figure 7 are for illustrative purposes only, and embodiments of this disclosure are not limited thereto. As shown, the following layers and corresponding functionalities are provided:

[0083] The hardware and software layer 710 includes hardware and software components. Examples of hardware components include a mainframe 711, a RISC (Reduced Instruction Set Computer) architecture-based server 712, a server 713, a blade server 714, a storage device 715, and network and networking components 716. In some embodiments, the software components include network application server software 717 and database software 718.

[0084] The virtualization layer 720 provides an abstraction layer from which the following examples of virtual entities may be provided: a virtual server 721, virtual storage 722, a virtual network 723 including a virtual private network, a virtual application and operating system 724, and a virtual client 725.

[0085] In one example, the management layer 730 may provide the following functions: Resource provisioning 731 provides the dynamic procurement of computing and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 732 provides cost tracking as resources are used within the cloud computing environment and billing or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, as well as protection for data and other resources. The user portal 733 provides consumers and system administrators with access to the cloud computing environment. Service level management 734 provides the allocation and management of cloud computing resources to ensure that the required service levels are met. Service level agreement (SLA) planning and execution 735 provides the proactive preparation and procurement of cloud computing resources that are expected to be needed in the future in accordance with the SLA.

[0086] The workload layer 740 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include mapping and navigation 741, software development and lifecycle management 742 (e.g., neural network training environment 100), provision of virtual classroom education 743, data analysis processing 744, transaction processing 745, and analysis systems 746.

[0087] This disclosure may be a system, method, or computer program product or combination thereof at any possible level of technical detail of integration. A computer program product may include a computer-readable storage medium (or more mediums) having computer-readable program instructions for causing a processor to perform aspects of this disclosure.

[0088] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooves on which instructions are recorded, and any suitable combination thereof. When used herein, computer-readable storage media should not be interpreted as being transient signals in themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through fiber optic cables), or electrical signals transmitted through wires.

[0089] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device, or to an external computer or external storage device, via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computing / processing device receives computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium within each computing / processing device.

[0090] Computer-readable program instructions for performing the operations disclosed herein may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk®, C++, and procedural programming languages ​​such as the C programming language or similar programming languages. Computer-readable program instructions may run entirely on the user's computer, partially on the user's computer as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (for example, via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to perform an aspect of the present disclosure.

[0091] Aspects of the present disclosure will be described herein with reference to flowcharts or block diagrams or combinations thereof of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It will be understood that each block in a flowchart or block diagram or combination thereof, and any combination of blocks in a flowchart or block diagram or combination thereof, can be implemented by computer-readable program instructions.

[0092] These computer-readable program instructions can be provided to a computer processor or other programmable data processing device to generate a machine, and as a result, instructions executed via the computer processor or other programmable data processing device create means for implementing functions / operations specified in one or more blocks of a flowchart or block diagram or a combination thereof. Furthermore, these computer-readable program instructions may be stored in a computer-readable storage medium, and such instructions can instruct a computer, a programmable data processing device or other device, or a combination thereof, to function in a particular manner, thereby the computer-readable storage medium storing the instructions contains a product containing instructions that implement modes of functions / operations specified in one or more blocks of a flowchart or block diagram or a combination thereof.

[0093] Computer-readable program instructions can also be loaded onto a computer, other programmable device, or other device to perform a series of operational steps on that computer, other programmable device, or other device to generate a computer implementation process, and as a result, the instructions executed on the computer, other programmable device, or other device implement the functions / operations specified by one or more blocks in a flowchart or block diagram or a combination thereof.

[0094] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions containing one or more executable instructions that implement a specified logical function. In some alternative implementations, the functions described in a block may be performed in an order different from the order shown in the drawings. For example, two blocks shown consecutively may actually be achieved as a single step, or they may be executed simultaneously, substantially simultaneously, partially or entirely in overlapping time, and blocks may also be executed in reverse order depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, or combination thereof, and the combination of blocks in a block diagram or flowchart, or combination thereof, may be implemented by a dedicated hardware-based system that performs a specified function or operation, or performs a combination of dedicated hardware and computer instructions.

[0095] The terminology used herein is solely for the purpose of describing specific embodiments and is not intended to limit the various embodiments. As used herein, the singular forms “a,” “an,” and “the” are intended to include the plural form as well, unless the context explicitly indicates otherwise. It will be further understood that the terms “include” or “including,” when used herein, specify the presence of a described feature, integer, stage, operation, element, or component, or a combination thereof, but do not exclude the presence or addition of one or more other features, integers, stages, operations, elements, components, or groups thereof, or combinations thereof. In the preceding detailed description of exemplary embodiments among the various embodiments, references were made to the accompanying drawings (similar numbers representing similar elements) which form part of this specification and illustrate specific examples of embodiments in which the various embodiments may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice them, but other embodiments may be used, and logical, mechanical, electrical, and other variations may be made without departing from the scope of the various embodiments. In the preceding description, numerous specific details were provided to give a complete understanding of the various embodiments. However, the various embodiments may be carried out without these specific details. In other instances, well-known circuits, structures, and techniques are not shown in detail so as not to obscure the embodiments.

[0096] When different reference numbers have a common number followed by different characters (e.g., 100a, 100b, 100c) or a common number followed by different punctuation marks (e.g., 100-1, 100-2, or 100.1, 100.2), using only the reference symbol without characters or a following number (e.g., 100) may refer to a group of multiple elements, any subset of a group, or an exemplary specimen of a group as a whole.

[0097] Throughout this description, it should be understood that the term “mechanism” is used to refer to elements of the present invention that perform various operations and functions. As used herein, the term “mechanism” may be an implementation of a function or aspect of an exemplary embodiment in the form of an apparatus, procedure, or computer program product. In the case of a procedure, the procedure is implemented by one or more devices, apparatus, computer, or data processing system, etc. In the case of a computer program product, the logic embodied in or represented by computer code or instructions on the computer program product is executed by one or more hardware devices to implement a function or to perform an operation associated with a particular “mechanism.” Thus, the mechanisms described herein may be implemented as specialized hardware, software that runs on the hardware and thereby constitutes the hardware to implement a specialized function of the present invention that the hardware would not otherwise be able to perform, software instructions stored on a medium which the instructions are readily executable by the hardware and thereby specifically configure the hardware to perform the described function and the specific computer operations described herein, procedures or methods for performing the function, or any combination of the above.

[0098] Furthermore, references to “models” or “model” in this specification specifically refer to computer-running machine learning models. These models include algorithms and statistical models that computer systems use to perform specific tasks without using explicit instructions, but instead rely on patterns and inference. Machine learning algorithms build computer-running models based on known sample data as “training data” to make predictions or decisions without being explicitly programmed to perform the task in question. Examples of machine learning models include, but are not limited to, supervised machine learning models such as convolutional neural networks (CNNs) and deep neural networks (DNNs), as well as unsupervised machine learning models such as isolation forest models, one-class support vector machine (SVM) models, and local outlier factor models, and ensemble learning mechanisms such as random forest models.

[0099] Furthermore, the phrase "at least one of..." when used with a list of items means that one or more different combinations of the listed items may be used, and only one of each item in the list may be required. In other words, "at least one of..." means that any combination and any number of items can be used from the list, but not all of the items in the list are required. An item can be a specific object, thing, or category.

[0100] For example, though not limited to, “at least one of item A, item B, or item C” could include item A, item A and item B, or item B. This example could also include item A, item B, and item C, or item B and item C. Of course, any combination of these items is possible. In some examples, “at least one of” could be, for example, two of item A, one of item B, ten of item C, four of item B, seven of item C, or any other preferred combination.

[0101] Different examples of the word “embodiment” as used herein do not necessarily refer to the same embodiment, although they may. Any data and data structures illustrated or described herein are merely examples, and different amounts of data, data types, fields, number and types of fields, field names, number and types of rows, records, entries, or data organization may be used in other embodiments. Furthermore, any data may be combined with logic so that separate data structures are not required. Therefore, the prior detailed descriptions should not be taken in an restrictive sense.

[0102] The various embodiments described herein are presented for illustrative purposes only and are not intended to be exhaustive or limit the scope to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles of the embodiments, the practical applications of the technology found in the market or technical improvements thereto, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

[0103] While this disclosure has described specific embodiments, variations and modifications thereof are expected to be apparent to those skilled in the art. Therefore, the following claims are intended to be construed as encompassing all such variations and modifications that fall within the true spirit and scope of this disclosure.

[0104] The descriptions of the various embodiments of this disclosure are presented for illustrative purposes only and are not intended to be exhaustive or to limit the scope to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best describe the principles of the embodiments, their practical applications, or the technical improvements of the technology found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for training a neural network, wherein the method is The training dataset, which includes samples encrypted using homomorphic encryption, The step of receiving via input to the aforementioned neural network; A step of determining the packing configuration used by the homomorphic encryption used to pack the samples in the training dataset; The step of selecting a dropout technique during training of the neural network based on the packing configuration; and The process begins with a first packing configuration from the training dataset, and then iteratively or recursively inputs the first packing configuration to the neural network using the selected dropout technique, so that the next packing configuration from the training dataset acts as the initial input applied to the neural network for the next iteration, until a stopping metric is generated by the neural network. A method that includes [a certain feature].

2. The method according to claim 1, wherein the packing configuration is based on the dropout technique implemented on the neural network.

3. The method according to claim 1, wherein the neurons of the neural network are packed into groups of neurons.

4. The method according to claim 3, wherein the dropout technique causes a group of neurons, including neurons selected for drop, to be dropped.

5. The method according to claim 3, wherein the neuron group is packed in the same configuration as the packing configuration of the training dataset.

6. The method according to claim 3, wherein the neuron group is packed as ciphertext.

7. The method according to claim 1, wherein the packing configuration is in the form of a batch containing at least two training samples from the training dataset in a single ciphertext.

8. The method according to claim 7, wherein the dropout technique causes a group of neurons, including neurons selected for drop, to be dropped.

9. The method according to any one of claims 1 to 8, wherein the packing configuration is a homomorphic encrypted ciphertext.

10. A computer program comprising computer-readable instructions, wherein, when the computer-readable instructions are executed on a computing device, the computing device, The training dataset, which includes samples encrypted using homomorphic encryption, The procedure received via input to a neural network; A procedure for determining the packing configuration used by the homomorphic encryption used to pack the samples in the training dataset; A procedure for selecting a dropout technique during training of the neural network based on the packing configuration; and A procedure to start with a first packing configuration from the training dataset, and to input the first packing configuration into the neural network iteratively or recursively using the selected dropout technique, such that the next packing configuration from the training dataset acts as the initial input applied to the neural network for the next iteration, until a stopping metric is generated by the neural network. A computer program that executes something.

11. The computer program according to claim 10, wherein the packing configuration is based on the dropout technique implemented on the neural network.

12. The computer program according to claim 10, wherein the neurons of the neural network are packed into groups of neurons.

13. The computer program according to claim 12, wherein the dropout technique causes a group of neurons, including neurons selected for dropping, to be dropped.

14. The computer program according to claim 12, wherein the neuron groups are packed in the same configuration as the packing configuration of the training dataset.

15. The computer program according to any one of claims 10 to 14, wherein the packing configuration is in the form of a batch containing at least two training samples from the training dataset in a single ciphertext.

16. The computer program according to claim 15, wherein the dropout technique causes a group of neurons, including neurons selected for dropping, to be dropped.

Citation Information

Patent Citations

  • Fully homomorphic encryption deep learning reasoning method and system based on FPGA

    CN112699384A

  • Inference device and method

    JP2019046460A

  • Learning and inferring insights from encrypted data

    US20200019867A1

  • Training a neural network model

    US20200372344A1