Packing of Machine Learning Models Using Pruning and Reordering
The system addresses inefficiencies in HE by pruning and rearranging machine learning models based on importance, enhancing performance through optimized packing, resulting in improved latency, energy savings, and accuracy during inference.
Patent Information
- Application Number
- JP2024572631
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-05-31
- Publication Date
- 2025-07-30
AI Technical Summary
Existing methods for pruning machine learning models under homomorphic encryption (HE) fail to optimize latency, energy savings, and accuracy due to inefficient packing and pruning strategies, leading to suboptimal performance during inference.
A processor-based system for pruning and rearranging machine learning models based on neuron or weight importance, followed by packing, to reduce ciphertext calculations under specific constraints, enabling pruning-aware packing that improves inference performance.
The system achieves significant efficiency improvements, particularly for larger tiles, with up to a 50% increase in zero tiles and minimal accuracy loss, optimizing latency, energy, and memory usage during homomorphic encryption inference.
Smart Images

Figure 2025524387000001_ABST
Abstract
Description
Technical Field
[0001] This technique relates to machine learning models. More specifically, this technique relates to the execution of machine learning models under homomorphic encryption.
[0002] Homomorphic encryption (HE) enables the execution of operations on encrypted data. Such an encryption system can be used, for example, in a client-server scenario where a client desires that a server execute a function f(x). The client can provide x, and the function f can be obtained from a different source. HE enables the server to compute the function f(x) homomorphically without learning the specific values of the variable x. The client can then use the secret key to decrypt the result encrypted using the corresponding public key. In some schemes, multiple clients can provide multiple keys. For example, in a fully homomorphic encryption (FHE) scheme with multiple keys, all clients can have their own secret keys and provide the associated public keys used to encrypt the results to the server.
[0003] The HE operation can be executed using the single instruction multiple data (SIMD) paradigm, in which a message is divided into an array of values called slots. A single HE operation is applied to all these slots at once. In particular, a single ciphertext encrypts a fixed-size vector, and the homomorphic operations on the ciphertext are executed on a slot-by-slot basis for the elements of the plaintext vector. In the CKKS HE scheme, for example, up to thousands of encrypted values can be stored in a single encrypted message and processed at once. To take advantage of the SIMD feature, more than one input element can be packed and encrypted in all ciphertexts. The packing method can thus have a dramatic impact on latency, throughput, communication cost, and memory requirements. Therefore, a method of packing or grouping these values into encrypted messages can be used to improve performance. For example, a simple way to pack a plaintext matrix could be to pack it in row-major order until all slots of a given ciphertext are "filled", and then create a new ciphertext and repeat this. HELayers is an exemplary software development kit (SDK) that automates the packing process for data scientists. In particular, HELayers uses a special packing technique called tiled tensors. A tiled tensor is a data structure that packs a tensor within a fixed-size chunk called a tile. For example, a tensor can be a vector or a matrix. Since each tile can be encrypted into a single ciphertext and different elements of each tile are mapped to different slots of that ciphertext, this tiled tensor data structure with a fixed size naturally fits well with HE. In addition, tensors can also be used to implement various layers of neural networks.For example, in one solution, a 5-dimensional tensor denoted as C, X, Y, F, B is adopted, where C is the channel dimension encoding the input channels, X and Y are the dimensions of the width and height of the image, F is the filter dimension encoding different filters for each layer, and B is the batch dimension encoding different images to be classified. Additionally, the same tensor can be covered by different-shaped tiles of the same size. For example, a matrix can be simply covered by column vectors or row vectors, but as long as the number of elements in the tile matches the number of slots in the ciphertext, the matrix can also be covered by 2-dimensional tiles. Additionally, tile tensors allow other operations such as replicating elements along one or more dimensions. In some frameworks, it is also possible to easily switch between one tile shape and another, and to easily set the amount of replication along each dimension. In the following, the inventors use tile shapes to include the amount of replication along each of these dimensions. Different tile shapes can result in different performances. For example, one tile shape may require more memory but may have optimal execution time, while another shape may be optimal in memory but may require more time to execute. To find the best shape supported by those systems for a given purpose, in some ways, an optimizer that scans the shape configuration space and reports the detected best shape is used. In the context of pruning, packing the neurons and weights of a neural network into tiles causes the problem that pruning can only be done at the resolution of the entire tile (i.e., the ciphertext), and not at the resolution of a single neuron or weight.
[0004] FHE operations can be significantly more expensive compared to their plaintext counterparts. For example, FHE operations can have a computational cost that is about three to five orders of magnitude higher than the operations performed in plaintext. One optimization used in the plaintext neural network domain is the use of pruning by operation removal. Pruning improves latency by reducing the number of operations that must be executed, suppresses overfitting, and thus improves the accuracy of the deployed network. By pruning the network, zeros are introduced in the weights and / or activations, and as a result, computations involving these values can be skipped. Thus, pruning reduces latency and energy for inference execution. For example, in a simple weight pruning scheme, all weights with values below a certain threshold can be removed. Consequently, this reduces the number of operations that need to be formed during inference. Pruning complex models generally results in improved test accuracy through reduced variance, while excessive simplification can lead to underfitting and thus a decrease in accuracy. Some pruning techniques may be able to remove most of the weights with only a small decrease in accuracy. Additionally, by retraining the network after pruning, the error output at each neuron can be reduced by training, and thus the accuracy loss can be mitigated.
[0005] However, one issue regarding pruning in the context of HE - compliant inference is that latency or energy savings due to reduction in the number of operations do not necessarily correspond to the degree of pruning. For example, latency or energy savings do not necessarily have to correspond to the number of weights removed. Instead, the actual reduction in operations can also depend on the packing method. This is because the zeros introduced during pruning can be packed within the same ciphertext message along with other non - zero values, in which case the entire ciphertext has to be retained as is, and the number of operations to be performed on this ciphertext remains unchanged. Thus, in order to prune the entire message and enjoy the latency and energy benefits, it may be necessary that all the packed values within the ciphertext message are zero. One solution is to prune in groups that match the shape of the tiles encoded in the ciphertext message. However, this pruning method can lead to the removal of important weights, which can consequently result in a significant decrease in accuracy during inference. Furthermore, with these pruning methods, there is a possibility that the satisfaction of the optimality constraints involving latency, energy, and accuracy is not guaranteed. Summary of the Invention
[0006] According to the embodiments described in this specification, the system may include a processor for pruning a machine learning model based on the importance of neurons or weights. This processor may further rearrange and pack the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculation under selected constraints. Therefore, this processor may enable pruning-aware packing for machine learning models that improve performance during inference. Preferably, the processor is for performing pruning and packing in parallel. In this embodiment, pruning may be able to better improve the efficiency of packing. Optionally, the importance is based on the criticality of the neurons. In this embodiment, neurons that are not important for the accuracy of the model during inference may be pruned to improve efficiency. Optionally, the importance is based on the value of the weights. In this embodiment, weights that are not important for model accuracy may be flagged and thus ignored during inference. Optionally, the selected constraints include inference accuracy constraints. In this embodiment, a specific accuracy may be ensured during inference. Optionally, the selected constraints include memory constraints. In this embodiment, a specific memory usage may be ensured during inference. Optionally, the selected constraints include latency constraints. In this embodiment, a specific latency may be ensured during inference. Preferably, the step of pruning the machine learning model includes the step of removing operations from the machine learning model. In this embodiment, the efficiency of the machine learning model during inference may be improved. Preferably, the ciphertext calculation includes performing encrypted inference on the homomorphic of the pruned, rearranged, and packed machine learning model. In this embodiment, the encrypted inference on the homomorphic may have improved accuracy and performance.
[0007] According to another embodiment described herein, the method may comprise, via a processor, pruning a machine learning model based on the importance of neurons or weights. The method may further comprise, via the processor, rearranging and packing the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculations under selected constraints. Thus, the method may enable pruning-aware packing for machine learning models that improve performance during inference. Optionally, the method may further comprise performing homomorphically encrypted inference using the pruned, rearranged, and packed machine learning model. In this embodiment, the homomorphically encrypted inference may have improved accuracy and performance. Optionally, the step of pruning the machine learning model comprises pruning the weights of the machine learning model by setting to zero weights having values not exceeding a threshold. In this embodiment, the weights may be flagged and ignored during inference. Optionally, the step of pruning the machine learning model comprises pruning the neurons of the machine learning model. In this embodiment, the neurons may be removed and not used during training and inference. Optionally, the step of rearranging the machine learning model comprises using balanced clustering. In this embodiment, the maximum number of zero tiles can be discovered more efficiently. Optionally, the step of rearranging the machine learning model comprises alternately rearranging the rows and columns of the weight matrix corresponding to the weights between the layers of the machine learning model until convergence is detected. In this embodiment, the maximum number of zero tiles can be discovered. Optionally, the method comprises expanding the pruned and rearranged machine learning model and un-pruning the zero values within the packing shape with partially zero-valued pruning. In this embodiment, the accuracy lost during pruning can be recovered.Optionally, the method comprises simulating a pruned and packed machine learning model and obtaining latency scores and memory scores associated with a plurality of packing shapes and pruning thresholds, where the pruned, sorted, and packed machine learning model has a pruning threshold and a packing shape that minimize an objective function based on selected constraints. In this embodiment, any selected constraints can be used to ensure that the constraints are satisfied during inference.
[0008] According to another embodiment described herein, a computer program product for pruning and packing a machine learning model may comprise a computer-readable storage medium having program code embodied therewith. The program code is executable by a processor to cause the processor to prune a machine learning model based on the importance of neurons or weights. The program code may also cause the processor to rearrange and pack the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculation under selected constraints. Accordingly, the program code may enable pruning-aware packing for a machine learning model that improves performance during inference. Optionally, the program code may also cause the processor to set to zero weights having values that do not exceed a threshold. In this embodiment, zero weights may be flagged and need not be considered during training and inference. Optionally, the program code may also cause the processor to rearrange the machine learning model using heuristics. In this embodiment, zero tiles to be pruned can be increased more efficiently. Optionally, the program code may also cause the processor to further rearrange the machine learning model using balanced clustering. In this embodiment, zero tiles can be discovered more efficiently. Optionally, the program code may also cause the processor to alternately rearrange the rows and columns of a weight matrix corresponding to the weights between layers of the machine learning model until convergence is detected. In this embodiment, the maximum number of zero tiles can be discovered. Optionally, the program code may also cause the processor to retrain the pruned machine learning model and execute the pruned machine learning model to obtain an accuracy score for the pruned machine learning model associated with a particular pruning threshold. In this embodiment, the accuracy score may be used to select the best combination of pruning, rearrangement, and packing.Optionally, the program code may also cause the processor to simulate a pruned and packed machine learning model and obtain latency scores and memory scores associated with a plurality of packing shapes and pruning thresholds, where the pruned, reordered, and packed machine learning model has a pruning threshold and a packing shape that minimize an objective function based on the selected constraints. In this embodiment, the latency scores and memory scores may be used by an objective function calculator to select the best combination of pruning, reordering, and packing. Optionally, the program code may also cause the processor to perform homomorphically encrypted inferences using the pruned, reordered, and packed machine learning model. In this embodiment, homomorphically encrypted inferences may be performed more efficiently and accurately.
Brief Description of the Drawings
[0009]
Figure 1
[0010]
Figure 2
[0011]
Figure 3A
[0012]
Figure 3B
[0013]
Figure 4
[0014]
Figure 5
[0015]
Figure 6
[0016]
Figure 7
[0017]
Figure 8
[0018]
Figure 9
[0019]
Figure 10
[0020]
Figure 11
[0021]
Figure 12
[0022]
Figure 13
[0023]
Figure 14
[0024]
Figure 15
[0025]
Figure 16
[0026]
Figure 17
DETAILED DESCRIPTION OF THE INVENTION
[0027] According to embodiments of the present disclosure, a system comprises a processor for pruning a machine learning model based on neuron and weight importance. The processor is further for rearranging and packing the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculations under selected constraints. Accordingly, embodiments of the present disclosure provide a method of pruning-aware packing for machine learning model inference under homomorphic encryption (HE) that enjoys maximum performance benefits from the pruning stage without even a minimal drop in accuracy. In particular, in the case of experiments on autoencoder neural networks, significant efficiency improvements were recorded especially for larger tiles. In particular, in an exemplary iterative k-means rearrangement algorithm, the number of tiles having only zero elements increased from 40% to 50%, from 20% to 40%, and from 8% to 40% for tile sizes of 4x4, 8x8, and 16x16, respectively.
[0028] Referring now to FIG. 1, the block diagram illustrates an exemplary system for pruning, reordering, and packing a machine learning model. The exemplary system is generally referred to by reference numeral 100. FIG. 1 includes a computing device 102. For example, the computing device 102 can be a server. In some examples, the computing device 102 can be a node of a cloud computing service. The computing device 102 includes a network pruner 104, a network reorderer 106, a network packer 108, and an objective function evaluator 110. It is shown that the computing device 102 receives a machine learning model 112. For example, the machine learning model 112 can be any suitable machine learning model trained to perform HE operations. In various examples, the machine learning model 112 can be a neural network. For example, the machine learning model 112 can be a convolutional neural network (CNN), an autoencoder, or any other suitable machine learning model. In various examples, the machine learning model 112 may or may not be encrypted. The system 100 also includes selected constraints 114 shown to be received by the computing device 102. For example, the selected constraints 114 can be any suitable constraints such as inference accuracy constraints, memory constraints, latency constraints, amortized latency, power constraints, energy constraints, or any combination thereof. The system also comprises a pruned, reordered, and packed machine learning model 116 shown to be output by the computing device 102.
[0029] In the example of FIG. 1, computing device 102 receives a machine learning model 112 and a selected constraint 114, and outputs a pruned, reordered, and packed machine learning model 116 that satisfies the selected constraint 114. In various examples, the machine learning model 112 may or may not be encrypted. For example, the machine learning model 112 may be encrypted after being trained with proprietary information. Thus, the weights of the machine learning model 112 may be deployed in encrypted form to an untrusted computing device 102. The computing device 102 may thus learn the shape of the machine learning model 112, such as the number of layers and the number of parameters for each layer, without knowing any of the values of the parameters. In some examples, the activation inputs from the client may also be encrypted under HE. Additionally, if some operations are removed, the computing device 102 may also be able to learn which operations were removed. In this way, by keeping the model secret, the underlying proprietary information can be kept secret.
[0030] In some examples, the machine learning model 112 may not need to be encrypted. For example, the machine learning model 112 may be trained with publicly available data that is not subject to any constraints and thus may not need to be kept secret. In other examples, the machine learning model 112 may be encrypted. For example, the machine learning model 112 may be trained with data that is subject to access constraints.
[0031] Referring still to FIG. 1, the network pruning device 104 of the computing device 102 can prune the machine learning model 112. For example, the network pruning device 104 can prune the machine learning model 112 using any number or type of suitable pruning thresholds or parameters. In various examples, the threshold can be a value of a weight having a larger L1 norm compared to a certain fixed ratio of other weights, and is referred to herein as L1-based pruning. In some examples, any other suitable pruning parameter can be received. For example, the pruning parameter can include which of the weights or neurons to prune, and whether to use random pruning, global pruning, or local pruning. For example, in random pruning, neurons can be randomly pruned, or the weights of the model can be randomly set to zero. In global pruning, all layers are pruned at once. In local pruning, all layers are pruned according to other parameters. When using random pruning, these parameters may not have an effect. However, considering, for example, L1-based pruning, the parameters can have a strong effect. For example, if the processor prunes 50% of the network, only the initial layers can be pruned. In various examples, any of six different pruning configurations from the combination of these parameters {W,L}×{R,L1}×{W,N}, where G / R / {W,N} is the same as L / R / {W,N}, W refers to the weight of pruning, N refers to the pruning neuron, G refers to the global pruning method, L refers to the local pruning method, and L1 refers to L1-based pruning. In some examples, the network pruning device 104 is prune in this specification packA packing-based pruning configuration, also referred to as such, can be used. For example, the network pruning device 104 can first select the size of the packing shape. In the example of a tile tensor, the size of the packing shape can be the tile size. In various examples, the tile size can be 2x2, 4x8, or 8x8. The network pruning device 104 can then divide all matrices into tiles. The network pruning device 104 can calculate the minimum value, maximum value, or average value of their values for all tiles and prune the tile with the lowest result.
[0032] The network permutator 106 can permute the machine learning model 112 after the machine learning model has been pruned. For example, the network permutator 106 can use any suitable heuristic, such as a balanced clustering heuristic, to permute the machine learning model 112. In some examples, the network permutator 106 can use the k-means clustering heuristic as described in more detail in FIG. 5. In some examples, the network permutator 106 can alternately permute the rows and columns of the weight matrix corresponding to the weights between the layers of the machine learning model as described in more detail with respect to FIG. 6.
[0033] The network packer 108 can pack the pruned and permuted machine learning model using any suitable packing shape or size. For example, the network packer 108 can pack the pruned and permuted machine learning model using various different packing shapes and sizes.
[0034] In some examples, the network pruner 104, the network sorter 106, and the network packer 108 can generate a plurality of pruned, sorted, and packed machine learning models. In various examples, the objective function evaluator 110 can evaluate each combination of different pruning, sorting, and packing based on the objective function and one or more selected constraints 114. An exemplary algorithm for calculating an exemplary objective function is described with respect to FIGS. 3A and 3B.
[0035] In various examples, a resulting pruned, sorted, and packed machine learning model 116 is output and can be used for inference in a HE environment. An exemplary pruned, sorted, and packed machine learning model 116 used in this way is described with respect to FIG. 2.
[0036] As a specific example of a technique, the HELayers packing solution can be used with a CKKS SEAL implementation targeting 128-bit security. For training, GPUs can be installed in a cluster of server-class machines. In training, PyTorch version 1.11.0 accelerated with CUDA version 11.6 can be used. The network architecture used can be an autoencoder network architecture, in which case a square activation layer follows all fully connected (FC) layers. For example, square activation can be used instead of rectified linear unit (ReLU) activation to support non-interactive solutions that require HE-friendly networks. In various examples, more refined activations, such as more advanced or trainable activations, can be used additionally or alternatively to achieve better results. As an example, the autoencoder network can include an FC with 32 neurons or 64 neurons. In some examples, the autoencoder network can include multiple FC layers, such as three FC layers with sizes of 64 neurons, 32 neurons, and 64 neurons. The autoencoder network can be trained on the MNIST dataset, first released in 1998, which has 60,000 images of 28x28x1 = 768 pixels. Thus, the input size and decoder output size of the autoencoder can be 768. In some examples, the decoder can be fused with the encoder as an additional FC layer of the relevant size and trained together. The number of training epochs and retraining epochs can be set to {20,10}, {30,20}, or any other suitable value. A batch of 10 samples can be used, and the learning rate of the Adam optimizer can be set to 1e-3. In various examples, the loss function used can be any suitable loss function, such as the mean squared error (MSE) between the input image and the reconstructed image.In this example, the HELayers use data structures called CtileTensor and PtileTensor to hold tiled tensors of encrypted and unencrypted data respectively. These have an API called encode for encoding (packing) the data before encrypting it. Thus, in some examples, the HELayers can be adapted to automatically identify zero tiles by modifying different encoding functions and testing whether all of their elements are zero for all tiles. If a tile contains only zeros, the processor need not allocate it and may instead include a new flag indicating that this is a zero tile. In various examples, when considering binary addition and multiplication operations that take two inputs, if only one of the inputs has a flag set, the addition function can be modified to return the other object and the multiplication function can be modified to return a new null tile with this flag set. If both inputs are zero, the returned element can be a null tile.
[0037] It should be understood that the block diagram of FIG. 1 is not intended to show that system 100 should include all of the components shown in FIG. 1. Rather, system 100 may include fewer components, or additional components not shown in FIG. 1 (e.g., additional computing devices, or additional machine learning models, machine learning models that have been pruned, reordered, and packed, or additional processing such as extensions). For example, system 100 may additionally include a model extender to reduce neurons or weights that include zero values. For example, the model extender may perform an operation to reverse a pruning operation. In particular, the model extender may search for tiles that do not hold only zero values and unprune the zero elements within them. The unpruned weight elements can then be trained to improve model accuracy. In this way, the model extender may recover some of the model accuracy lost due to initial pruning. For example, if a tile has non-zero elements and is not reduced as a result, the system cannot ignore the tile during inference. Thus, the model extender may instead make maximum use of the elements to enhance the performance of the model during inference.
[0038] FIG. 2 is a block diagram showing an exemplary system for generating an encrypted result based on encrypted data using a machine learning model packed using pruning and reordering. Exemplary system 200 includes elements from FIG. 2, which are also referred to. For example, system 200 includes a pruned, reordered, and packed machine learning model 116. It is shown that the pruned, reordered, and packed machine learning model 116 of system 200 receives encrypted data 202 and outputs encrypted result 204. For example, encrypted data 202 can be any information to be classified, such as an image. In various examples, encrypted result 204 may include a classification of the input encrypted data 202. In some examples, encrypted result 204 may also include a confidence score for the classification.
[0039] As previously explained, one practical application of HE is for encrypted inference on neural networks running on the cloud. For example, system 200 can be used for the diagnosis of COVID-19 through the classification of X-ray images of patients in a hospital environment, in which case the encrypted X-ray images are securely transmitted to the cloud. In this example, computing device 102 can be a server that executes a machine learning model trained on different systems using proprietary data that is not publicly available, and thus, the network parameters can be encrypted to hide them from the server. The parameters can include weights or biases. In this example, without the server learning anything about the image or network parameter values, the client can obtain the encrypted classification result 204 from the server. The client can then decrypt the encrypted result 204 using a key. For example, the key can correspond to the key used to encrypt the encrypted X-ray image.
[0040] It should be understood that the block diagram of FIG. 2 is not intended to show that system 200 should include all of the components shown in FIG. 2. Rather, system 200 can include fewer components, or additional components not shown in FIG. 2 (e.g., additional data, or additional results, etc.). For example, a pruned, reordered, and packed machine learning model can alternatively be a pruned, reordered, extended, and packed machine learning model, or can be the product of any combination of these operations, as explained in FIG. 17 below.
[0041] Figures 3A and 3B are process flow diagrams of an exemplary process that can select a combination of network packing, pruning, and rearrangement based on an objective function. Process 300 can be implemented using any suitable computing device such as the computing device 1200 of FIG. 12 or the system 100 of FIG. 1. For example, the methods described below can be implemented by the computing device 102, the processor 1202, or the processor 1502 of FIGS. 1, 12, and 15.
[0042] FIG. 3A shows a process of pruning, rearranging, and packing a machine learning model for inference on a network using a two-dimensional tile tensor and a batch size of 1. In various examples, since the pruning algorithm generally takes the value of each parameter as input, blocks 302-336 are part of the pre-unfolding process and can be executed on a system having access to the plaintext network parameters. Additionally, in the example of FIG. 3A, pruning individual weights based on a threshold is considered. A general goal in the example of FIG. 3A can be to find the value of a pruning threshold (PRUNE_THRES) and the shape of the tile tensor to optimize the machine learning model for a given objective function. In various examples, the objective function can be based on any combination of accuracy, latency, memory requirements, and energy consumption, among any other suitable selected constraints.
[0043] In block 302, process 300 begins. In various examples, the processor can receive a trained model, a set of pruning thresholds, and a set of different tile shapes. For example, among other suitable values, the pruning threshold can be a value of "1" as in the examples of FIGS. 4 and 5 below. In various examples, the pruning threshold can be an L1-based pruning threshold.
[0044] In decision diamond 304, the processor determines whether each pruning threshold PRUNE_THRES within the received set of pruning thresholds THRES_ALL has been processed. If all pruning thresholds THRES_ALL have been processed, the process may continue at decision diamond 318. If not all pruning thresholds THRES_ALL have been processed, the process may continue at block 306.
[0045] At block 306, the processor sets the received trained model as the model to be processed. In some examples, the processor may process multiple models and thus may pick one out of a set of models provided to be pruned, sorted, and packed.
[0046] In decision diamond 308, the processor determines whether all layers within the model have been processed. If all layers within the model have been processed, the process may continue at block 312. If not all layers within the model have been processed, the process may continue at block 310.
[0047] At block 310, the processor prunes the selected layer of the model based on the selected pruning threshold PRUNE_THRES. For example, the processor may prune a machine learning model based on the threshold. In various examples, the processor may use weight pruning, neuron pruning, or a combination thereof.
[0048] At block 312, the processor retrains the pruned model. For example, the processor may retrain the pruned network to recover some of the accuracy loss resulting from the pruning.
[0049] In block 314, the processor executes the trained and pruned model to obtain an accuracy score for the trained and pruned model. For example, the updated accuracy score can be obtained by performing inference on a test set for the pruned and retrained machine learning model.
[0050] In block 316, the processor adds a combination of the pruning threshold PRUNE_THRES, the model, and the associated accuracy score for the model pruned using the pruning threshold PRUNE_THRES to the set of model records MODEL_RECS. For example, the set of model records MODEL_RECS can be stored in a file.
[0051] In decision diamond 318, the processor determines whether each combination of the pruning threshold PRUNE_THRES, the model, and the associated accuracy score for the model pruned using the pruning threshold PRUNE_THRES has been processed using sorting. If not, the process can continue at decision diamond 320. If so, the process can continue at decision diamond 330.
[0052] In decision diamond 320, the processor determines whether all of the different tile sizes and tile shapes received in TILE_SHAPES have been processed for a particular combination of the pruning threshold PRUNE_THRES, the model, and the associated accuracy score for the model pruned using the pruning threshold PRUNE_THRES. If so, the process may continue with additional combinations in decision diamond 318. If not, the process may continue with additional permutations in block 322. For example, the set TILE_SHAPES may be an independent set of integer tuples. As an example, the values of TILE_SHAPES may be {(2,2), (3,3)} for a set of two 2D tiles with length = 2, width = 2, and length = 3, width = 3, respectively.
[0053] In block 322, the processor permutes the model using a combination of tile shapes T1 and T2. For example, the processor may permute the model using any suitable heuristic such as an iterative clustering algorithm. In some examples, the heuristic may be a balanced clustering heuristic. In some examples, the processor may use the balanced k-means clustering technique as illustrated in FIG. 5.
[0054] In block 324, the processor packs the permuted model. For example, the processor may pack the weights and activation tensors into T1xT2 tiles. In various examples, the processor may also discard zero tiles.
[0055] In block 326, the processor simulates the packed and permuted model and generates associated latency values and memory values for the packed and permuted model. For example, these latency scores and memory scores can be calculated in response to receiving selected latency and memory constraints from a client. Thus, the processor can simulate the network and obtain estimated values of the latency requirements and memory requirements of the packed and permuted model.
[0056] In block 328, the processor adds to the record file a pruning threshold PRUNE_THRES, the model, and an associated accuracy score for the model pruned using the pruning threshold PRUNE_THRES, and a combination of the latency score and memory score for the model when packed and permuted using tile shapes T1 and T2. The record file can thus include rows corresponding to all combinations of different pruning thresholds, permutations, and tile tensor shapes.
[0057] In decision diamond 330, the processor determines whether each record in the record file has been processed to generate an objective function score. If not, the process can continue in block 332 to process additional records. If so, the process can continue in block 334.
[0058] In block 332, the processor calculates an objective function score for each record in the record file. For example, the objective function score can depend on the objective function and various selected optimization constraints. In the examples of FIGS. 3A and 3B, these optimization constraints include accuracy, latency, and memory constraints.
[0059] In block 334, the processor selects a record row from the record file associated with the lowest objective function score calculated in block 332. For example, the best record row can be selected according to the objective function and optimization constraints.
[0060] In block 336, the process ends. In some examples, the processor may output a pruned, sorted, and packed model using the selected record row of block 334.
[0061] The process flow diagram of FIG. 3A is not intended to indicate that the operations of process 300A should be executed in any particular order, or that all operations of process 300A should be included in all cases. For example, although tile tensor grouping is shown for illustrative purposes, any other suitable grouping may alternatively be used. In addition, process 300A of FIG. 3A considers pruning individual weights based on a threshold. However, other pruning methods such as pruning groups of weights that will be packed together in the same encrypted message, as well as techniques in the prior art such as pruning based on activation threshold, dynamic pruning, and weight splicing among other suitable pruning techniques, may be used. Further, process 300A may include any suitable number of additional operations. In some examples, process 300A uses an exhaustive search strategy to find the optimal point, but a local search strategy may alternatively be used. For example, an exhaustive search to find a permutation matrix with the maximum number of zero tiles may cost O(M!N!), which can be prohibitively expensive even for medium-sized weights. Thus, in some embodiments, process 300A may instead permute rows and columns heuristically to make the problem more tractable. For example, method 300A may use the exemplary k-means heuristic described in FIG. 5. In some examples, method 300A may further include expanding tiles with partial zero values to further improve accuracy and efficiency, as described in FIG. 3B.
[0062] Figure 3B is a process flow diagram of an exemplary process that can select a combination of network packing, pruning, permutation, and expansion based on an objective function. Process 300B can be implemented using any suitable computing device, such as the computing device 1200 of FIG. 12 or the system 100 of FIG. 1. For example, the methods described below can be implemented by the computing device 102, the processor 1202, or the processor 1502 of FIGS. 1, 12, and 15.
[0063] The process 300B of FIG. 3B includes elements that are similarly referenced in FIG. 3A. Additionally, in decision diamond 338, the processor determines whether all layers in the model have been further processed. If not, the process continues at block 340. If so, process 300B can continue at block 342.
[0064] At block 340, the processor performs an expansion operation. For example, in the expansion operation, any partially zero tiles within a layer of the machine learning model can be unpruned.
[0065] At block 342, the processor retrains the model. For example, the machine learning model can be retrained with the unpruned values, improving the accuracy of the resulting retrained model.
[0066] At block 344, the processor executes the machine learning model on a test set of data to generate an updated accuracy score for the retrained model. For example, due to additional weights that became available during training, the updated accuracy score can be higher. In various examples, the updated accuracy score can replace the accuracy score in the record file and be used instead of the previous accuracy score when calculating the objective function at block 332.
[0067] The process flow diagram of FIG. 3B is not intended to indicate that the operations of process 300B should be performed in any particular order or that all operations of process 300B should be included in all cases. For example, the example of FIG. 3B shows example P3E of FIG. 17, but in some examples, FIG. 3B may include pruning-based packing such as P4E of FIG. 17 or pruning-based semi-packing such as P5E of FIG. 17.
[0068] FIG. 4 is a diagram showing an exemplary process of pruning, reordering, and packing a weight matrix. The exemplary process 400 can be executed by any suitable processor such as the processor of computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0069] The process 400 of FIG. 4 includes an initial weight matrix 402. As an example, the numbered rows of the weight matrix 402 correspond to the first layer of the neural network, and the numbered columns of the weight matrix respond to the second layer of the neural network. FIG. 4 shows a simple example of a 4x8 weight matrix and a 2x2 tile packing shape. As shown in FIG. 4, the values of the weight matrix range from 0.1 to 1.9 and correspond to the weights between neurons in two layers. The process 400 includes a pruned weight matrix 404 in which values less than 1.0 have been pruned from the weight matrix 402.
[0070] Process 400 further shows a first pruned and packed weight matrix 406 where one of the eight tiles is a zero tile that contains all zero values. In FIG. 4, this zero-valued tile tensor that can be discarded is indicated by the bold solid contour line. The other non-zero tiles are indicated using dashed contour lines. Process 400 further includes a pruned, permuted weight matrix 408 to which the best permutation has been applied. For example, any number of different permutations may be performed and the best permutation may be selected and applied based on maximizing the zero tiles. In some examples, the best permutation may be selected using an objective function as described herein. Process 400 further includes a pruned, permuted, and packed weight matrix 410 in which four of the eight tiles are zero tiles indicated by the bold contour lines. The pruning 412 of the weight matrix 402 is indicated by the first arrow. The packing 414 of the pruned weight matrix 404 is indicated by the second arrow. The permutation 416 of the pruned weight matrix 404 is indicated by the third arrow. The packing 418 of the pruned, permuted weight matrix 408 is indicated by the fourth arrow. FIG. 4 further shows a neural network 420 corresponding to the weight matrix 402, a pruned neural network 422 with pruned weights corresponding to the zeros of the pruned weight matrix 402, and a pruned, permuted neural network 424 having the row and column order corresponding to the pruned, permuted weight matrix 408.
[0071] In the example of FIG. 4, the applied exemplary pruning threshold has a value of 1, and thus, in the pruned weight matrix 404, weights having a value < 1.0 are set to zero. As shown in block 406 of FIG. 4, if the weight tensor is packed as is at the stage shown in block 404, only one of the eight tile tensors contains all zeros and can thus be discarded by the processor. The remaining seven tile tensors in block 406 contain a mix of non-zeros in addition to zeros. Thus, these tile tensors cannot be discarded. To increase the number of zero tensors that can be discarded, the processor can, therefore, rearrange the rows and columns of the tensor before packing the rearranged tile tensors. For example, the processor can rearrange the rows and columns such that zero values are grouped together as much as possible. In various examples, the processor can use a rearrangement procedure to perform this regrouping, resulting in an optimal rearrangement 408. For example, any suitable rearrangement procedure can be used. In some examples, the rearrangement procedure used can be an alternating rearrangement process that rearranges the rows and columns of the weight matrix as described in FIG. 5.
[0072] By rearranging the rows and columns according to a balanced k-means rearrangement algorithm, the processor increases the number of zero tiles to a maximum of four, as shown in block 410. For example, the processor may use the balanced k-means rearrangement described in FIG. 5 below. Such an increase in zero tiles directly translates to a reduction in the execution time of the network when inference is performed. Further, the rearrangement of the rows and columns of the weight matrix 404 is equivalent to shuffling the neurons within one or more layers of the weight matrix 404 and thus does not affect the overall functionality of the neural network.
[0073] It should be understood that the block diagram of FIG. 4 is not intended to show that process 400 should include all of the components shown in FIG. 4. Rather, process 400 may include fewer components, or additional components not shown in FIG. 4 (e.g., additional layers, neurons, weights, tile shapes, dimensions, or additional permutations, etc.). In various examples, higher-dimensional tile tensors such as 2x2x256 tile tensors may alternatively be used. In some examples, a batch dimension may also be used. For example, the batch dimension may include the use of a subset of the original weight matrix for pruning purposes. For example, given a ciphertext encrypting a vector of 1024 elements, block 410 may need to prune all 1024 packed elements, which may result in a decrease in accuracy. Alternatively, the processor may instead assume an inference system that performs inferences on a batch of 256 samples at a time. In that case, the tile tensor will allocate 2x2 slots per sample in all ciphertexts. Thus, the processor may only need to prune 2x2 tiles from the weight matrix, which may be much more feasible to execute.
[0074] FIG. 5 is a diagram showing an exemplary process of permutation using a balanced variation of the k-means method. Exemplary process 500 may be executed by any suitable processor such as the processor of computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0075] In various examples, exemplary heuristics for rearrangement are based on the k-means clustering technique. More specifically, the example of FIG. 5 shows the use of balanced k-means. The process 500 of FIG. 5 includes a first weight matrix 502. For example, the first weight matrix 502 can be a pruned weight matrix having a tile tensor of T1xT2. A set of numbers for the rows and a set of numbers for the columns are used to indicate the initial ordering of the rows and columns, respectively. In iteration 0 504, the initial weight matrix 502 includes only one zero-tile tile tensor indicated by the thick line, where all the values of the tile tensor are zero.
[0076] In various examples, the rows of the pruned weight matrix 502 can be regarded as vectors, and the first iteration of the k-means method 506 can be applied to generate a new weight matrix 508 having rearranged rows to increase the number of zero tiles. In particular, the new weight matrix 508 includes two zero tiles indicated by the thick line. In addition, the new order of the rows is indicated by bold numbers. In particular, row 0 is shifted down two places so as to be placed between rows 2 and 3.
[0077] In the exemplary process 500 of FIG. 5, the new matrix 508 is then transposed 510 to generate a transposed matrix 512. At arrow 514, the processor can then apply a second iteration of the k-means method to the transposed matrix 512 to generate a second new matrix 516. The second new matrix 516 indicates the new ordering of the original columns as shown by the bold numbers.
[0078] At arrow 518, the processor can transpose the second new matrix 516 to generate a transposed second new matrix 520. The transposed second new matrix 520 includes four zero tiles as indicated by the thick outlined font.
[0079] In various examples, process 500 is repeated until convergence. For example, convergence can be achieved when no additional zero tiles result from the row and column permutations. As an example, if the processor prunes 400 elements, the tile size is 2x2 = 4 elements, and the processor detects 100 zero tiles, convergence can be assumed. However, alternatively, the processor may stop process 500 after a given threshold. For example, the processor may stop process 500 after 80% of the elements form zero tiles. In some examples, the distance function used can be the Hamming distance. For example, non-zero cells can be treated as having a value of "1". In various examples, the number of clusters used by the processor for k-means is equal to the number of tiles along a row or column, depending on the iteration being performed. For example, given an MxN matrix and t1xt2 tiles, the number of clusters in iteration i can be equal to ceil(M / t1) [when i is even] and ceil(N / t2) [when i is odd]. In this example, for an 8x16 matrix with 4x2 tiles, the number of clusters would thus be 2, 8, 2, 8, … etc.
[0080] It should be understood that the block diagram of FIG. 5 is not intended to show that system 500 should include all of the components shown in FIG. 5. Rather, system 500 may include fewer components or additional components not shown in FIG. 5 (e.g., additional layers, neurons, weights, tile shapes, or additional types or iterations of permutations, etc.). In various examples, process 500 may alternatively use higher-dimensional tile tensors. For example, in the process, a 2x2x256 tile tensor may be used. In various examples, the k-means clustering technique of process 500 may alternatively be replaced with any other suitable balanced clustering technique. For example, alternative balanced clustering techniques may include agglomerative clustering or graph partitioning techniques such as the normalized cut (Ncut) technique, which measures the sum of dissimilarities between different groups and the sum of similarities within a group when treating image segmentation as a graph partitioning problem, first described in 1997. Other clustering techniques that may be used with balanced variations include the Gaussian Mixture Model (GMM) and Density-Based Spatial Clustering of Applications with Noise (DBSCAN).
[0081] FIG. 6 is a diagram showing an exemplary process for permutation of weights for a multi-layer neural network. Exemplary process 600 may be implemented using any suitable computing device, such as computing device 1200 of FIG. 12 or system 100 of FIG. 1. For example, process 600 may be implemented by computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0082] The exemplary process 600 includes a first permutation 602 and a second permutation 604. As indicated by the two arrows 606, the processor may repeat permutation 602 and permutation 604 until convergence.
[0083] In various examples, for a single weight matrix for a two-layer network, the processor may permute the rows and columns independently. However, for deeper networks such as the neural network shown in FIG. 6, when permuting the rows or columns of a given layer, it may also affect the weights for the adjacent layers. In the example of FIG. 6, the exemplary neural network includes five layers labeled as A, B, C, D, and E. The set of weights between and connecting the various layers A, B, C, D, and E are labeled as W AB , W BC , W CD , and W DE respectively. In the example of FIG. 6, the transposes of W BC and W DE are labeled as
Number
Number
[0084] In this example, shuffling the neurons within layer B leads to permuting the rows of the transposed weight matrix
Number
Number
Number
[0085] In block 604, the processor can similarly rearrange the remaining set of layers along the columns. For example, the processor weights the matrix W AB For the columns of, W CD For the columns of, and the transposed weight matrix
Number
Number
Number
Number
[0086] In block 606, the process is repeated. For example, the processor may iterate through blocks 602 and 604 until convergence is achieved. In some examples, convergence may be detected based on the fact that no additional zero tiles result from the rearrangement. In various examples, convergence may be based on a preset maximum iteration count to handle oscillatory behavior. For example, if the algorithm oscillates between rearrangements using 41 zero tiles and 42 zero tiles, convergence may be detected after a preset number of oscillations. In some examples, the processor may detect convergence in response to determining that the count of zero tiles obtained in the last N iterations has not changed.
[0087] It should be understood that the diagram of FIG. 6 is not intended to show that process 600 should include all of the components shown in FIG. 6. Rather, process 600 may include fewer components or additional components not shown in FIG. 6 (e.g., additional layers, neurons, weights, tile shapes, or additional types or iterations of rearrangement, etc.).
[0088] FIG. 7 is a diagram showing an exemplary process of neuron pruning using rearrangement. Exemplary process 700 may be implemented using any suitable computing device such as computing device 1200 of FIG. 12 or system 100 of FIG. 1. For example, process 700 may be implemented by computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0089] The exemplary process 700 for neuron pruning in FIG. 7 is shown for a four-layer neural network having 6, 4, 8, and 4 neurons in layers A, B, C, and D, respectively. The neurons in layers A, B, C, and D are vectors X A , X B , X C , and X Dcan be described as. FIG. 7 also shows the associated weight matrices W AB , W CD , and the transposed weight matrices
Number
[0090] In block 702, the original 4-layer neural network includes all of its original weights. In block 704, after pruning 706 indicated by the arrow, as shown by the gray blocks, most of the original weights have been removed. For example, pruning 706 can be performed using any suitable pruning technique, such as by a pruning threshold. In the example of FIG. 7, the pruning threshold can be a neuron critical threshold. However, as shown by the bold blocks, only a total of 4 out of the 2x2 packings contain all zeros, and thus are considered zero tiles corresponding to neurons that can be discarded.
[0091] In block 708, after rearrangement 710, the number of zero tiles has increased to a total of 11 zero tiles that can be discarded. By discarding 11 instead of 2 encrypted messages, the resulting inference latency of the pruned, rearranged, and packed network can be significantly improved.
[0092] It should be understood that the diagram of FIG. 7 is not intended to show that process 700 should include all of the components shown in FIG. 7. Rather, process 700 may include fewer components, or additional components not shown in FIG. 7 (e.g., additional layers, neurons, weights, tile shapes, or additional types or iterations of permutations, etc.).
[0093] FIG. 8 is a diagram showing an exemplary process of weight pruning using permutation. Exemplary process 800 may be implemented using any suitable computing device such as computing device 1200 of FIG. 12 or system 100 of FIG. 1. For example, process 800 may be implemented by computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0094] The exemplary process 800 for weight pruning in FIG. 8 is similarly shown for a four-layer neural network having 6, 4, 8, and 4 neurons in each of layers A, B, C, and D, respectively. FIG. 8 also shows the associated weight matrices W AB , W CD , and the transposed weight matrix
Number
[0095] In the example of FIG. 8, at block 804, the weights corresponding to zero tiles are discarded via pruning 806, while the neurons themselves are maintained. For example, pruning an entire neuron can be too aggressive for a particular network. Thus, as shown in FIG. 8, the processor can instead alternatively prune only the weights more conservatively. However, in the example of weight pruning, the processor cannot simply drop the last few rows and columns as explained in FIG. 7. Instead, in the weight pruning example of FIG. 8, the processor tags each zero tile with a label that requests the server to skip any calculations that use this tile at block 808 after rearrangement 810.
[0096] FIG. 9 is a process flow diagram of an exemplary method for packing, pruning, and rearranging a machine learning model under selected constraints. Method 900 can be implemented using any suitable computing device such as the computing device 1200 of FIG. 12 or the system 100 of FIG. 1. For example, the methods described below can be implemented by the computing device 102, the processor 1202, or the processor 1502 of FIGS. 1, 12, and 15.
[0097] At block 902, the processor receives a trained machine learning model and selected constraints. For example, the trained machine learning model may or may not be encrypted. In some examples, the machine learning model can be a neural network such as a convolutional neural network. In various examples, the selected constraints can include an inference accuracy constraint, a memory constraint, a latency constraint, or any combination thereof.
[0098] At block 904, the processor prunes the trained machine learning model based on the importance of neurons and weights. For example, the processor may set weights having values that do not exceed a threshold to zero. In some examples, the processor prunes the weights of the machine learning model. For example, the processor may prune the weights by setting weights having values that do not exceed a threshold to zero and flagging specific packing shapes, such as tiles of weights, as zero tiles to be ignored.
[0099] At block 906, the processor rearranges and packs the remaining neurons and weights of the pruned machine learning model to reduce the amount of ciphertext computation under selected constraints. In various examples, the processor may use heuristics to rearrange the machine learning model. For example, the heuristic may be a balanced clustering heuristic. In some examples, the processor may rearrange the machine learning model by alternately rearranging the rows and columns of the weight matrix corresponding to the weights between the layers of the machine learning model until convergence is detected. In some examples, the processor may retrain the pruned machine learning model and execute the pruned machine learning model to obtain an accuracy score for the pruned machine learning model associated with a specific pruning threshold. In some examples, the processor may simulate the pruned and packed machine learning model to obtain latency scores and memory scores associated with multiple packing shapes and pruning thresholds. For example, the pruned, rearranged, and packed machine learning model may have a pruning threshold and a packing shape that minimize an objective function based on selected constraints. In various examples, the processor may prune and pack the machine learning model in parallel. For example, the processor may calculate a specific combination of pruning, rearrangement, and packing such that the maximum number of neurons or weights are pruned from the machine learning model and apply the combination to the machine learning model.
[0100] The process flow diagram of FIG. 9 is not intended to indicate that the operations of method 900 should be performed in any particular order or that all of the operations of method 900 should be included in all cases. Further, method 900 may include any suitable number of additional operations. For example, method 900 may further comprise performing homomorphically encrypted inferences using a pruned, reordered, and packed machine learning model. In some examples, the processor prunes neurons of the machine learning model. For example, the processor may remove the last empty columns and empty rows of each of the pruned and reordered weight matrices. In various examples, method 900 may also further comprise expanding a pruned, packed, and reordered machine learning model to utilize zero values within the packing shape.
[0101] FIG. 10 is a process flow diagram of an exemplary method that may select a combination of network packing, pruning, and reordering based on an objective function. Method 1000 may be implemented using any suitable computing device, such as computing device 1200 of FIG. 12 or system 100 of FIG. 1. For example, the methods described below may be implemented by computing device 102, processor 1202, or processor 1502 of FIGS. 1, 12, and 15.
[0102] At block 1002, the processor receives a trained machine learning model and an objective function. For example, the trained machine learning model may be encrypted or may not be encrypted. The objective function may include various constraints.
[0103] In block 1004, for each of the plurality of selected pruning techniques and parameters, the processor prunes a layer of the machine learning model using the selected pruning technique, retrains the pruned machine learning model, and executes the retrained machine learning model on a test set to generate an updated accuracy score. For example, each of the pruning techniques and parameters may be associated with a different updated accuracy score.
[0104] In block 1006, for each combination of the plurality of selected packing configurations and pruning techniques, the processor rearranges the pruned machine learning model to increase the number of zero-value packings, discards the zero-value packings, packs the rearranged and pruned machine learning model, and simulates the pruned and packed machine learning model to estimate metrics of interest. For example, the metrics of interest may include latency, memory usage, among other potential metrics of interest.
[0105] In block 1008, the processor calculates an objective function for each pruned and packed machine learning model corresponding to a particular combination of the selected packing configuration and pruning technique, based on the corresponding updated accuracy score and the metrics of interest.
[0106] In block 1010, the processor outputs the pruned and packed machine learning model having the lowest objective function. For example, a pruned and packed machine learning model that minimizes the objective function given a particular set of constraints may be output.
[0107] The process flow diagram of FIG. 10 is not intended to indicate that the operations of method 1000 should be performed in any particular order or that all of the operations of method 1000 should be included in all cases. Further, method 1000 may include any suitable number of additional operations. For example, method 1000 may further comprise performing homomorphically encrypted inferences using a pruned, permuted, and packed machine learning model. In some examples, method 1000 may expand a pruned, packed, and permuted machine learning model, undo pruning for tiles that do not have all zero values, and retrain these tiles to improve the inference accuracy of the network. For example, each of the unpruned tiles may have a complete set of values instead of having zero values from pruning.
[0108] FIG. 11 is a process flow diagram of an exemplary method that may generate encrypted results based on encrypted data using a machine learning model packed using pruning and permutation. Method 1100 may be implemented using any suitable computing device such as the computing device 1200 of FIG. 12 or the system 200 of FIG. 2. For example, the methods described below may be implemented by the pruned, permuted, and packed machine learning model 116 of FIG. 2.
[0109] In block 1102, the processor transmits encrypted data to a pruned, permuted, and packed machine learning model. For example, the encrypted data may include an encrypted image or any other type of data to be classified. In various examples, the pruned, permuted, and packed machine learning model may be pruned, permuted, and packed using techniques described herein, such as via methods 900 or 1000 of FIGS. 9 and 10 above.
[0110] In block 1104, the processor receives encrypted results from a pruned, sorted, and packed machine learning model. For example, the encrypted results can be classifications or images or other data.
[0111] The process flow diagram of FIG. 11 is not intended to indicate that the operations of method 1100 should be executed in any particular order or that all of the operations of method 1100 should be included in all cases. Further, method 1100 may include any suitable number of additional operations. For example, method 1100 may include a step of decrypting the encrypted results using a key corresponding to the key used to encrypt the encrypted data transmitted to the pruned, sorted, and packed machine learning model.
[0112] The present disclosure includes a detailed description regarding cloud computing, but it should be understood that the implementations of the teachings recited herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment, now known or hereafter developed.
[0113] Cloud computing is a service delivery model that enables convenient on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0114] The characteristics are as follows.
[0115] On-demand self-service: Cloud consumers can automatically and on-demand provision computing capabilities such as server time and network storage unilaterally without the need for human interaction with a service provider.
[0116] Broad network access: The capabilities are available over a network and accessed through standard mechanisms that promote use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0117] Resource pooling: The provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to demand. Consumers generally have no control or knowledge over the exact location of the provided resources, but location independence exists in that they may be able to specify location at a higher level of abstraction (e.g., country, state, or data center).
[0118] Rapid elasticity: Capabilities can be rapidly and elastically provisioned, in some cases automatically, scaling out quickly and also being released quickly and scaling in rapidly. To the consumer, the capabilities available for provisioning often appear limitless and can be purchased in any quantity at any point in time.
[0119] Measured service: Cloud systems automatically control and optimize resource use by leveraging measurement capabilities appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts) at some level of abstraction. Resource usage can be monitored, controlled, and reported, thereby providing transparency for both the provider and consumer of the utilized service.
[0120] The service model is as follows.
[0121] Software as a Service (SaaS): The ability provided as a service to the consumer is to use the provider's applications running on the cloud infrastructure. The applications are accessible from various client devices through a client interface such as a web browser (e.g., web-based email). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even the individual application capabilities, with the exception of limited user-specific application configurations.
[0122] Platform as a Service (PaaS): The ability provided as a service to the consumer is to deploy the consumer-created or -acquired applications, created using programming languages and tools supported by the provider, onto the cloud infrastructure. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but controls the deployed applications and, in some cases, the application hosting environment configuration.
[0123] Infrastructure as a Service (IaaS): The ability provided as a service to the consumer is to provision processing, storage, network, and other basic computing resources, and the consumer can deploy and run any software that may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but controls the operating system, storage, deployed applications, and, in some cases, selectively controls the selected networking components (e.g., host firewalls).
[0124] The deployment model is as follows.
[0125] Private cloud: The cloud infrastructure is operated solely for an organization. It may be managed by the organization or a third party and may exist on-premises or off-premises.
[0126] Community cloud: The cloud infrastructure is shared by multiple organizations and supports a specific community with common concerns (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by the organizations or a third party and may exist on-premises or off-premises.
[0127] Public cloud: The cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services.
[0128] Hybrid cloud: This cloud infrastructure is a composite of two or more clouds (private, community, or public) that remains a unique entity but is joined together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0129] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing, there is an infrastructure that includes a network of interconnected nodes.
[0130] FIG. 12 is a block diagram of an exemplary computing device that can pack, prune, and reorder a machine learning model under selected constraints. Computing device 1200 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, computing device 1200 can be a cloud computing node. Computing device 1200 can be described in the general context of computer system executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like that perform particular tasks or implement particular abstract data types. Computing device 1200 can be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media including memory storage devices.
[0131] Computing device 1200 can include a processor 1202 for executing stored instructions and a memory device 1204 for providing temporary memory space for the operation of the instructions during operation. The processor can be a single-core processor, a multi-core processor, a computing cluster, or any number of other configurations. Memory 1204 can include random access memory (RAM), read-only memory, flash memory, or any other suitable memory system.
[0132] Processor 1202 may be connected to an input / output (I / O) device interface 1208 adapted to connect computing device 1200 to one or more I / O devices 1210 through a system interconnect 1206 (e.g., PCI (Registered Trademark), PCI-Express (Registered Trademark), etc.). The I / O devices 1210 may include, for example, a keyboard and a pointing device, and the pointing device may include, among other things, a touchpad or a touch screen. The I / O devices 1210 may be a plurality of built-in components of the computing device 1200, or may be a plurality of devices externally connected to the computing device 1200.
[0133] Processor 1202 may also be linked to a display interface 1212 adapted to connect computing device 1200 to a display device 1214 through system interconnect 1206. The display device 1214 may include a display screen that is a built-in component of the computing device 1200. The display device 1214 may include, among other things, a computer monitor, a television, or a projector that is externally connected to the computing device 1200. Additionally, a network interface controller (NIC) 1216 may be adapted to connect the computing device 1200 to a network 1218 through the system interconnect 1206. In some embodiments, the NIC 1216 may transmit data using any suitable interface or protocol, such as, among other things, an Internet small computer system interface. The network 1218 may be, among other things, a cellular network, a wireless network, a wide area network (WAN), a local area network (LAN), or the Internet. An external computing device 1220 may be connected to the computing device 1200 through the network 1218. In some examples, the external computing device 1220 may be an external web server 1220. In some examples, the external computing device 1220 may be a cloud computing node.
[0134] Processor 1202 may also be linked through system interconnect 1206 to a storage device 1222 that may include a hard drive, an optical drive, a USB flash drive, an array of drives, or any combination thereof. In some examples, the storage device may include a model pruning module 1224, a model reordering module 1226, and a model packing module 1228. The model pruning module 1224 may receive a machine learning model and one or more selected constraints. For example, the selected constraints may include an inference accuracy constraint, a memory constraint, a latency constraint, an amortized latency, a power constraint, an energy constraint, or any combination thereof. The model pruning module 1224 may prune the machine learning model based on the importance of neurons and weights. For example, the importance may be based on the criticality of the neurons. The criticality of a neuron may be a measure of the resulting accuracy loss in response to removing a particular neuron. In some examples, the importance may be based on the value of the weights. For example, a pruning threshold may be used to set to zero weights having values that do not exceed the threshold. The model pruning module 1224 may remove operations from the machine learning model. In some examples, the operations may be associated with one or more neurons. The model reordering module 1226 and the model packing module 1228 may reorder and pack the remaining neurons and weights of the pruned machine learning model to reduce the amount of ciphertext calculations under the selected constraints. In various examples, the model reordering module 1226 may reorder the machine learning model using any suitable heuristic, such as a balanced clustering heuristic. For example, the balanced clustering heuristic may be a balanced k-means clustering. In some examples, the model reordering 1226 may reorder the machine learning model using an alternating row and column reordering. The model packing module 1228 may pack the machine learning model using any suitable packing method.The model packer module 1228 may use a packing method that reduces ciphertext calculations by maximizing the number of zero-valued packing shapes. In some examples, the model pruner module 1224 and the model packer module 1228 may prune and pack in parallel. For example, pruning and packing may be based on a combination of pruning, packing, and permutation determined using an objective function. The objective function evaluator 1230 may calculate the objective function for each of any number of combinations of packing methods, permutation techniques, and pruning thresholds or parameters.
[0135] It should be understood that the block diagram of FIG. 12 is not intended to show that computing device 1200 should include all of the components shown in FIG. 12. Rather, computing device 1200 may include fewer components, or additional components not shown in FIG. 12 (e.g., additional memory components, embedded controllers, modules, additional network interfaces, etc.). For example, computing device 1200 may also include a model expander that expands a pruned, packed, and reordered model, cancels pruning for tiles that do not have all zero values, and retrains these tiles to improve the inference accuracy of the network. In some examples, computing device 1200 may further include an execution module for performing encrypted inference execution in the homomorphic of a pruned, reordered, and packed machine learning model. Further, any of the functions of model pruning module 1224, model reordering module 1226, and model packing module 1228 may be implemented partially or fully within hardware and / or within processor 1202. For example, the functions may be implemented, inter alia, using logic implemented in an application specific integrated circuit, an embedded controller, or within the logic implemented within processor 1202. In some embodiments, the functions of model pruning module 1224, model reordering module 1226, and model packing module 1228 may be implemented using logic, where, when referred to herein, logic may include any suitable hardware (e.g., processor, inter alia), software (e.g., application, inter alia), firmware, or any suitable combination of hardware, software, and firmware.
[0136] Referring now to FIG. 13, an exemplary cloud computing environment 1300 is shown. As shown, the cloud computing environment 1300 includes one or more cloud computing nodes 1302 that can communicate with local computing devices used by cloud consumers, such as, for example, a personal digital assistant (PDA) or cellular phone 1304A, a desktop computer 1304B, a laptop computer 1304C, and / or an automotive computer system 1304N. The nodes 1302 can communicate with each other. They may be physically or virtually grouped within one or more networks, such as a private, community, public, or hybrid cloud as described above herein, or combinations thereof (not shown). Thereby, the cloud computing environment 1300 can provide infrastructure, platform, and / or software as services such that a cloud consumer need not maintain resources on a local computing device therefor. The types of computing devices 1304A - N shown in FIG. 13 are for illustrative purposes only, and it is understood that the computing nodes 1302 and the cloud computing environment 1300 can communicate with any type of computerized device via any type of network and / or network addressable connection (e.g., using a web browser).
[0137] Referring now to FIG. 14, a set of functional abstractions provided by the cloud computing environment 1300 (FIG. 13) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 14 are for illustrative purposes only and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided.
[0138] The hardware and software layer 1400 includes hardware components and software components. Examples of hardware components may include mainframe 1401, RISC (Reduced Instruction Set Computer) architecture-based server 1402; server 1403; blade server 1404; storage device 1405; and network and networking components 1406. In some embodiments, the software components include network application server software 1407 and database software 1408.
[0139] The virtualization layer 1410 provides an abstraction layer that can provide the following examples of virtual entities, namely, virtual server 1411; virtual storage 1412; virtual network 1413 including a virtual private network; virtual applications and operating systems 1414; and virtual client 1415.
[0140] In one example, the management layer 1420 can provide the functions described below. Resource provisioning 1421 provides for the dynamic procurement of computing resources and other resources utilized to execute tasks within a cloud computing environment. Metering and pricing 1422 provides for cost tracking when resources are utilized within a cloud computing environment, and for billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides identity verification for cloud consumers and tasks, and protection for data and other resources. User portal 1423 provides access to the cloud computing environment for consumers and system administrators. Service level management 1424 provides for cloud computing resource allocation and management such that the required service levels are met. Service level agreement (SLA) planning and fulfillment 1425 provides for the prior commitment and procurement of cloud computing resources where future requirements are expected to comply with the SLA.
[0141] The workload layer 1430 provides examples of functions that a cloud computing environment can utilize. Examples of workloads and functions that can be provided from this layer include mapping and navigation 1431; software development and lifecycle management 1432; virtual classroom education provision 1433; data analysis processing 1434; transaction processing 1435; and machine learning model optimization 1436.
[0142] The present invention can be a system, method, and / or computer program product integrated at any possible level of technical detail. The computer program product can include one (or more) computer-readable storage media having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0143] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes, hereinafter, namely, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a punch card, or a mechanically encoded device such as a raised structure in a groove in which instructions are recorded, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through an electric wire.
[0144] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium in each respective computing / processing device.
[0145] The computer-readable program instructions for carrying out the operations of the present invention may be written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages, and may be either code or object code. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit in order to carry out aspects of the present invention.
[0146] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the technique. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0147] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions for implementing the function / act specified in one or more blocks of the flowchart and / or block diagram.
[0148] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0149] Referring now to FIG. 15, there is shown a block diagram of an exemplary tangible non-transitory computer-readable medium 1500 that can pack, prune, and permute a machine learning model under selected constraints. The tangible non-transitory computer-readable medium 1500 can be accessed by a processor 1502 via a computer interconnect 1504. Further, the tangible non-transitory computer-readable medium 1500 can include code that instructs the processor 1502 to perform the operations of methods 900 and 1000 of FIGS. 9 and 10.
[0150] The various software components discussed herein can be stored on a tangible non-transitory computer-readable medium 1500 as shown in FIG. 15. For example, the model pruning module 1506 includes code for pruning a trained machine learning model based on the importance of neurons and weights. The model pruning module 1506 also includes code for setting to zero weights having values that do not exceed a threshold. In some examples, the model pruning module 1506 includes code. In some examples, the model pruning module 1506 includes code. The model permutation module 1508 includes code for permuting the remaining neurons and weights of the pruned machine learning model to reduce the amount of ciphertext calculations under selected constraints. The model permutation module 1508 further includes code for permuting the machine learning model using heuristics. For example, the model permutation module 1508 may include code for permuting the machine learning model using balanced clustering. In some examples, the model permutation module 1508 may include code for alternately permuting the rows and columns of the weight matrix corresponding to the weights between the layers of the machine learning model until convergence is detected. The model packing module 1510 includes code for packing the neurons and weights of the pruned machine learning model to reduce the amount of ciphertext calculations under selected constraints. The model packing module 1510 also includes code. The objective function evaluator module 1512 includes code for simulating the pruned and packed machine learning model and obtaining latency scores and memory scores associated with a plurality of packing shapes and pruning thresholds. For example, the objective function evaluator module 1512 includes code for detecting a pruning threshold and a packing shape that minimizes the objective function based on selected constraints.In some examples, the objective function evaluator module 1512 includes code for retraining the pruned machine learning model, executing the pruned machine learning model, and obtaining an accuracy score for the pruned machine learning model associated with a particular pruning threshold.
[0151] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or combinations of dedicated hardware and computer instructions. It should be understood that any number of additional software components not shown in FIG. 15 may be included within the tangible non-transitory computer-readable medium 1500 depending on the particular application. For example, the computer-readable medium 1500 may also include code for performing homomorphically encrypted inferences using a pruned, reordered, and packed machine learning model. In some examples, the computer-readable medium 1500 may also include code for expanding a pruned and reordered machine learning model and unpruning zero values within a pruned packing shape with partially zero values.
[0152] FIG. 16 is a diagram illustrating an exemplary process of weight pruning and rearrangement using an exemplary expansion. The exemplary process 1600 may be implemented using any suitable computing device such as the computing device 1200 of FIG. 12 or the system 100 of FIG. 1 with an optional model expander. For example, process 700 may be implemented by the computing device 102, the processor 1202, or the processor 1502 of FIGS. 1, 12, and 15.
[0153] In block 1602, the processor receives a trained neural network having layers A, B, C, D, and weights W AB , W BC , W CD . The transpose matrix of the weights W BC is labeled as
Number
Number
[0154] In block 1604, a plurality of weight values are pruned to zero via pruning 1606, resulting in two zero tiles that only contain zero values. As shown, in block 1604, the accuracy of the resulting neural network may decrease, but the efficiency is increased.
[0155] In block 1608, the order of the weight matrix is rearranged via the rearrangement operation 1610, increasing the number of zero tiles to a total of seven zero tiles. As shown in block 1608, there is no impact on the accuracy, but the efficiency is increased.
[0156] In block 1612, the accuracy of the neural network is increased via an extension operation 1614. In particular, in order to increase the accuracy of the neural network, the zero values of the tiles that are partially zero are utilized by an extend operation 1614. In particular, the extend operation 1614 can un-prune any zero value in a tile that is partially zero, such that the value can be used for training. Thus, block 1612 can recover most of the accuracy loss of block 1604.
[0157] It should be understood that the diagram of FIG. 16 is not intended to show that process 1600 should include all of the components shown in FIG. 16. Rather, process 1600 can include fewer components, or additional components not shown in FIG. 16 (e.g., additional layers, neurons, weights, tile shapes, or additional types or iterations of permutations, or final packing, etc.).
[0158] FIG. 17 is a diagram showing an exemplary set of different combinations of pruning, permutation, extension, and packing according to the embodiments described herein. An exemplary combination can be implemented by system 100 to generate a pruned, permuted, and packed learning model, or a pruned, permuted, extended, and packed machine learning model.
[0159] Figure 17 shows various combinations of training, pruning, reordering, expansion, retraining, and packing according to the techniques described herein. These different combinations are referred to by the acronyms P2, P2T, P3, P3E, P4, P4E, P5E, and P6. As shown in Figure 17, each of the combinations begins by training a machine learning model. For example, the machine learning model can be a neural network. In various examples, once the trained machine learning model is ready, the processor can prune the neurons or weights of the trained machine learning model based on some criterion. In all strategies except P2T, first, pruning is performed by one of the six pruning configurations discussed in FIG. 1 above. In the example of P2T, the initial pruning is packing-based pruning. Since P2T performs packing-based pruning and the inventors prune complete tiles, there is no need for the processor to perform additional steps such as reordering or expansion in P2T. In contrast, when performing non-packing-aware pruning, the pruned weights or neurons may not necessarily be arranged in a nice way, which will result in a wide cancellation of tile operations. Thus, the processor can apply additional operations such as reordering or expansion to improve the efficient use of tiles. As described above, the reordering operation may include reordering the rows and columns of the weight matrix after the pruning operation to aggregate zero elements together. The expansion operation reverses the pruning operation. For example, the expansion operation may include searching for tiles that do not hold only zero values and unpruning the zero elements within these tiles. Thus, in the examples of P3, P3E, P4, P4E, P5E, and P6, reordering is also performed next.
[0160] In the example of P4, instead of expanding the model as in P3, the processor may perform a second pruning-aware packing stage to reduce all incomplete zero tiles. In the examples of P5 and P6, after the first rearrangement stage, the processor arranges the tiles that are partially zero and performs semi-packing-aware-pruning method Prune semi-pack that prunes some, but not all, of the further elements within them. Thereafter, the processor may reapply the rearrangement algorithm. In this way, the processor may assist the rearrangement heuristic while maintaining the initially applied pruning configuration. After the second rearrangement stage, the processor may determine whether to expand the tiles or perform packing-aware pruning on them based on the number of zeros within the tiles. Finally, for all combinations, the processor may perform retraining of the machine learning model to increase accuracy and perform the final packing of the retrained machine learning model. In various examples, these integrated combinations of rearrangement, expansion, pruning, and packing provide various trade-offs between accuracy, performance, and memory consumption, and thus provide options for various use cases.
[0161] It should be understood that the figure of FIG. 17 is not intended to show that the set of processes 1700 should include all of the components shown in FIG. 17. Rather, process 1700 may include a fewer number of components, or additional components not shown in FIG. 17 (e.g., additional training, pruning, rearrangement, expansion, retraining, or packing, etc.). Accordingly, FIG. 17 is not intended to be an exhaustive list of the various combinations of operations described herein.
[0162] The descriptions of the various embodiments of this technique have been presented for illustrative purposes, but these are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best describe the principles of the embodiments, the practical application, or a technical improvement to the technology found in the marketplace, or to enable other skilled artisans to understand the embodiments disclosed herein.
Claims
1. Pruning a machine learning model based on the importance of neurons or weights; and Rearranging and packing the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculation under selected constraints A system comprising a processor for.
2. The system according to claim 1, wherein the processor is for performing pruning and packing in parallel.
3. The system according to claim 1, wherein the importance is based on the criticality of the neurons.
4. The system according to claim 1, wherein the importance is based on the value of the weights.
5. The system according to claim 1, wherein the selected constraints include an inference accuracy constraint.
6. The system according to claim 1, wherein the selected constraints include a memory constraint.
7. The system according to claim 1, wherein the selected constraints include a latency constraint.
8. The system according to claim 1, wherein the procedure for pruning the machine learning model has a procedure for deleting operations from the machine learning model.
9. The system according to claim 1, wherein the ciphertext calculation includes executing encrypted inferences that are homomorphic to the pruned, rearranged, and packed machine learning model.
10. Pruning a machine learning model based on the importance of neurons or weights via a processor; and Rearranging and packing the remaining neurons or weights of the pruned machine learning model via the processor to reduce the amount of ciphertext calculation under selected constraints A computer-implemented method comprising.
11. The computer-implemented method according to claim 10, further comprising the step of executing encrypted inferences that are homomorphic using the pruned, rearranged, and packed machine learning model.
12. The computer-implemented method according to claim 10, wherein the step of pruning the machine learning model has a step of pruning the weights of the machine learning model by setting weights having values not exceeding a threshold to zero.
13. The computer-implemented method according to claim 10, wherein the step of pruning the machine learning model has a step of pruning the neurons of the machine learning model.
14. The step of rearranging the machine learning model has a step of using balanced clustering, the computer-implemented method according to claim 10.
15. The step of rearranging the machine learning model has a step of alternately rearranging rows and columns of a weight matrix corresponding to weights between layers of the machine learning model until convergence is detected, the computer-implemented method according to claim 10.
16. The computer-implemented method according to claim 10, further comprising a step of expanding the pruned and rearranged machine learning model and un-pruning zero values within a pruned packing shape with partially zero values.
17. The computer-implemented method according to claim 10, further comprising a step of simulating the pruned and packed machine learning model and obtaining a latency score and a memory score associated with a plurality of packing shapes and pruning thresholds, wherein the pruned, rearranged, and packed machine learning model has a pruning threshold and a packing shape that minimize an objective function based on the selected constraints.
18. A computer program product for pruning and packing a machine learning model, comprising a computer-readable storage medium having program code embodied thereon, the program code causing a processor to: prune a machine learning model based on the importance of neurons or weights; and rearrange and pack the remaining neurons or weights of the pruned machine learning model to reduce the amount of ciphertext calculation under selected constraints and being executable by the processor therefor.
19. The computer program product according to claim 18, further comprising program code executable by the processor to set weights having values not exceeding a threshold to zero.
20. The computer program product according to claim 18, further comprising program code executable by the processor to rearrange the machine learning model using a heuristic.
21. The computer program product according to claim 18, further comprising program code executable by the processor to rearrange the machine learning model using balanced clustering.
22. The computer program product according to claim 18, further comprising program code executable by the processor to alternately rearrange rows and columns of a weight matrix corresponding to weights between layers of the machine learning model until convergence is detected.
23. The computer program product according to claim 18, further comprising program code executable by the processor to retrain the pruned machine learning model and run the pruned machine learning model to obtain an accuracy score for the pruned machine learning model associated with a particular pruning threshold.
24. The computer program product according to claim 18, further comprising program code executable by the processor to simulate the pruned and packed machine learning model and obtain a latency score and a memory score associated with a plurality of packing shapes and pruning thresholds, wherein the pruned, rearranged, and packed machine learning model has a pruning threshold and a packing shape that minimize an objective function based on the selected constraints.
25. The computer program product according to claim 18, further comprising program code executable by the processor to perform homomorphically encrypted inference using the pruned, rearranged, and packed machine learning model.