Method and System for Accelerating Inference of Fully Homomorphic Encryption Neural Network Based on Resource Reuse

By optimizing the resource of fully homomorphic encryption neural network on FPGA, using parallel and flow optimization and buffer multiplexing technology, the problems of large computing overhead and slow operation speed of FHE-CNN are solved, and significant acceleration effect and energy consumption optimization are achieved.

CN116048811BActive Publication Date: 2025-06-10SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310113879.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-06-10
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

The existing fully homomorphic encrypted neural network (FHE-CNN) has problems with huge computing overhead and slow operation in actual applications, especially in the optimization of specific application levels and storage performance bottlenecks.

Method used

A fully homomorphic encryption neural network inference acceleration method based on resource multiplexing is proposed. By optimizing all homomorphic encryption operations and hardware resource allocation on FPGA, parallel and flow optimization and homomorphic basic operation module multiplexing, and buffer multiplexing of different granularity in on-chip storage space is reduced to reduce the time for inference encrypted data.

Benefits of technology

It effectively improves the inference acceleration effect of fully homomorphic encrypted neural networks, reduces the problem of low utilization efficiency of computing resources and insufficient on-chip storage space, and significantly improves the inference speed and energy consumption ratio of encrypted neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116048811B_ABST
    Figure CN116048811B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and system for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse, including: obtaining information of the fully homomorphic encryption neural network to be accelerated and hardware resource information of an FPGA; and inputting into a pre-constructed hardware resource allocation model to obtain an optimal hardware resource configuration scheme for the fully homomorphic encryption neural network during arithmetic processing in the FPGA; wherein, the processing strategy of the hardware resource allocation model is: parallel and pipelining optimizations are performed for fully homomorphic encryption operations and each network layer of the fully homomorphic encryption neural network, and reuse of homomorphic basic operation modules is adopted between each network layer; meanwhile, for the on-chip storage space of the FPGA, on-chip buffer reuse at different granularities is performed based on the arithmetic division of the fully homomorphic encryption neural network; finally, with the goal of minimizing the time of the encrypted data for inference of the fully homomorphic encryption neural network, an optimal resource configuration scheme is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of fully homomorphic encryption applications, and in particular, relates to a method and system for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse. Background Art

[0002] The statements in this part merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] Using fully homomorphic encryption (FHE) technology to protect the data of a convolutional neural network (CNN) is a current popular direction. However, there are many obstacles in the combination of the two in practical applications, specifically including: on the one hand, fully homomorphic encryption only supports linear operations such as addition and multiplication, while a typical CNN contains non-linear activation functions such as ReLU; on the other hand, the computational overhead of encrypted data is huge. The operation of a ciphertext is 5 to 6 orders of magnitude more than the original operation, and the neural network itself has to perform multiple rounds of complex operations on multiple nodes, which leads to the fact that the fully homomorphic encryption convolutional neural network (FHE-CNN) needs to calculate a large amount of data and also requires a large amount of storage space to store and schedule the generated data, further resulting in a very slow running speed of the FHE-CNN. For example, the first work that combined FHE and CNN, CryptoNets, takes 205 seconds to infer an encrypted 5-layer neural network, and such a speed cannot meet the performance requirements in practical applications.

[0004] In the above context, there are many works on optimizing and accelerating FHE-CNN, specifically including:

[0005] (1) Works implemented based on the CPU, which only consider how to improve the algorithm to reduce ciphertext operations and thus reduce the inference time, without considering hardware optimization.

[0006] (2) Optimizations at a lower hardware level. Specifically, common hardware platforms include Graphics Processing Unit (GPU), Field-programmable Gate Array (FPGA), and Application Specific Integrated Circuit (ASIC). The inventors found that such works have the following problems:

[0007] 1) Some methods only optimize the fully homomorphic operations and fail to optimize at the level of combining fully homomorphic operations with specific applications;

[0008] 2) Lack of consideration for the performance bottleneck in the storage aspect of FHE-CNN and lack of optimization for the computing process. Summary of the Invention

[0009] To solve the above problems, the present disclosure provides a method and system for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse. The solution effectively improves the acceleration effect of the fully homomorphic encryption neural network inference based on the fully homomorphic encryption operation optimization method and the on-chip buffer reuse strategy.

[0010] According to the first aspect of the embodiments of the present disclosure, a method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse is provided, including:

[0011] Obtain the information of the fully homomorphic encryption neural network to be accelerated and the hardware resource information of the FPGA;

[0012] Input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model to obtain the optimal hardware resource configuration scheme for the fully homomorphic encryption neural network during operation processing in the FPGA;

[0013] Among them, the processing strategy of the hardware resource allocation model is: perform parallel and pipelining optimization for the fully homomorphic encryption operation and each network layer of the fully homomorphic encryption neural network, and reuse the homomorphic basic operation modules between each network layer; at the same time, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, perform on-chip buffer reuse at different granularities; finally, with the goal of minimizing the time of the encrypted data of the fully homomorphic encryption neural network inference, obtain the optimal resource configuration scheme, where the operations of the fully homomorphic encryption neural network are divided into three levels of operations: homomorphic basic operations, homomorphic operations, and network layer operations according to different granularities.

[0014] Further, the hardware resource allocation model is specifically expressed as follows:

[0015]

[0016]

[0017]

[0018] Among them, LAT lr is the running time of the lr layer, Layer represents all the layers included in the fully homomorphic encryption neural network, lr represents one of the layers, OP represents the set of all homomorphic operations, op represents a certain homomorphic operation among them, DSP max represents the number of DSPs owned by the FPGA development board, BRAM maxIndicates the total number of BRAMs owned by the FPGA development board, DSP op Is the number of DSPs occupied by this homomorphic operation, BRAM lr Is the number of BRAMs used by the lr layer.

[0019] Furthermore, each network layer operation includes several homomorphic operations, and each homomorphic operation includes several homomorphic basic operations. Among them, the operations of the fully homomorphic encryption neural network adopt a pipeline and parallel processing method, and are pipelined in units of the homomorphic basic operations.

[0020] Furthermore, parallel and pipeline optimizations are performed within each network layer of the fully homomorphic encryption operation and the fully homomorphic encryption neural network, and the reuse of homomorphic basic operation modules is adopted between network layers. Specifically: according to the order of homomorphic basic operations, homomorphic operations, and network layer operations, the parallelism within different granularity operations is set respectively, so that the parallel effects of different granularity operations are superimposed;

[0021] Or, for homomorphic basic operations, according to whether the coefficients of the RNS polynomial are traversed once or multiple times, the homomorphic basic operations are divided into two categories, and by setting different parallelisms for the two categories of operations, the running times of the two categories of operations are made similar;

[0022] Or, for the pipeline implementation of cross-homomorphic operations within a network layer, the same parallelism is set for different homomorphic operations, and based on the complexity of the data dependence relationship between RNS polynomials, the homomorphic operations are divided into two categories;

[0023] Or, for different network layer operations, according to whether they include KeySwitch operations, they are divided into two categories, and pipeline designs are carried out for different categories respectively.

[0024] Furthermore, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, on-chip buffer reuse at different granularities is carried out. Specifically: the on-chip buffer takes the space for storing one RNS polynomial as the storage unit, and the on-chip buffer is divided into two categories according to the classification of homomorphic basic operations; among them, the reuse of the buffer includes reuse within homomorphic operations, reuse within network layers, and reuse between network layers.

[0025] Furthermore, the homomorphic basic operations include; modular multiplication, modular addition, modular subtraction, modulo operation, fast number theoretic transform and its inverse transform;

[0026] Or, the homomorphic operations include plaintext-ciphertext addition, ciphertext addition, ciphertext multiplication, rescaling, relinearization, and rotation operations;

[0027] Or, the network layer includes a homomorphic convolutional layer, a homomorphic activation layer, and a homomorphic fully connected layer.

[0028] According to a second aspect of the embodiments of the present disclosure, there is provided a fully homomorphic encryption neural network inference acceleration system based on resource reuse, including:

[0029] A data acquisition unit, which is configured to acquire information of a fully homomorphic encryption neural network to be accelerated and hardware resource information of an FPGA;

[0030] A resource configuration unit, which is configured to input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model, and obtain an optimal hardware resource configuration scheme for the fully homomorphic encryption neural network to perform arithmetic processing in the FPGA;

[0031] Wherein, the processing strategy of the hardware resource allocation model is: parallel and pipelining optimizations are performed for fully homomorphic encryption operations and each network layer of the fully homomorphic encryption neural network, and reuse of homomorphic basic operation modules is adopted between each network layer; at the same time, for the on-chip storage space of the FPGA, on-chip buffer reuse at different granularities is performed based on the arithmetic division of the fully homomorphic encryption neural network; finally, with the goal of minimizing the time of the fully homomorphic encryption neural network inference encrypted data, an optimal resource configuration scheme is obtained, wherein the arithmetic of the fully homomorphic encryption neural network is divided into three levels of arithmetic: homomorphic basic operations, homomorphic operations, and network layer operations according to different granularities.

[0032] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program running on the memory, and when the processor executes the program, it implements the above-mentioned fully homomorphic encryption neural network inference acceleration method based on resource reuse.

[0033] According to a fourth aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned fully homomorphic encryption neural network inference acceleration method based on resource reuse.

[0034] Compared with the prior art, the beneficial effects of the present disclosure are:

[0035] (1) The present disclosure provides a fully homomorphic encryption neural network inference acceleration method based on resource reuse. The solution supports automated deployment from neural network applications to fully homomorphic encryption hardware optimization implementation. The fully homomorphic encryption technology adopted effectively and reliably protects the data security during the neural network inference process, greatly facilitating the process from neural networks to encrypted inference and hardware deployment, and achieving good deployment effects through hardware acceleration optimization with the goal of meeting actual application requirements;

[0036] (2) The proposed solution addresses the bottlenecks in FPGA hardware deployment, namely, the low utilization efficiency of computing resources and the insufficient on-chip storage space. It proposes the fully homomorphic operator optimization technology and the on-chip buffer reuse technology, which effectively alleviate the difficulties in the actual deployment process, give full play to the hardware advantages of FPGA, and achieve good optimization effects on the inference speed and energy consumption ratio of encrypted neural networks.

[0037] (3) The high-level synthesis technology adopted by the proposed solution has the advantages of flexible programming, easy implementation, short development cycle, strong portability, etc., which is convenient for exploring the design space for various application requirements. Specifically, it can evaluate the hardware deployment for different neural networks and different models of FPGA development boards, generate the optimal hardware resource configuration scheme for the running speed and generate the code for deployment. In addition, this method can also be adjusted for other optimization goals and has scalability.

[0038] Advantages of additional aspects of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The specification drawings forming a part of the present disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute an improper limitation to the present disclosure.

[0040] Figure 1 It is a schematic diagram of the pipeline of the NKS layer (i.e., the layer without the KeySwitch operation) described in the embodiments of the present disclosure;

[0041] Figure 2 It is a schematic diagram of the pipeline of the KS layer (i.e., the layer with the KeySwitch operation) described in the embodiments of the present disclosure;

[0042] Figure 3 It is a schematic diagram of the homomorphic matrix sub-multiplication described in the embodiments of the present disclosure, where "RO" represents rotation and "+" represents homomorphic addition;

[0043] Figure 4 (a) to Figure 4 (e) are schematic diagrams of buffer reuse within different homomorphic operations described in the embodiments of the present disclosure;

[0044] Figure 5 (a) to Figure 5 (b) are schematic diagrams of buffer reuse within the FHE-CNN network layer described in the embodiments of the present disclosure;

[0045] Figure 6 It is the overall design framework of the fully homomorphic encryption neural network inference acceleration method based on resource reuse described in the embodiments of the present disclosure;

[0046] Figure 7 (a) to Figure 7 (b) is a schematic diagram of the parallelism configuration result based on the Lola-MNIST network in the embodiments of the present disclosure;

[0047] Figure 8 is a schematic diagram of the optimization effect of on-chip storage in the embodiments of the present disclosure;

[0048] Figure 9 is a schematic diagram of the optimization effect of the on-chip computing unit DSP in the embodiments of the present disclosure;

[0049] Figure 10 is a schematic diagram of the effect of design space exploration based on the inference acceleration method in the embodiments of the present disclosure;

[0050] Figure 11 is a flowchart of a method for accelerating inference of a fully homomorphic encryption neural network based on resource reuse in the embodiments of the present disclosure; Detailed implementation manners

[0051] The following further describes the present disclosure in conjunction with the accompanying drawings and embodiments.

[0052] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0053] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0054] Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0055] Term explanation:

[0056] Fully Homomorphic Encryption is an algorithm in the field of cryptography, characterized by being able to perform arbitrary calculations on ciphertext without decryption.

[0057] High-level Synthesis (HLS) is a process of automatically converting a logical structure described in a high-level language into a circuit model described in a low-level abstraction language.

[0058] FPGA (Field Programmable Gate Array), a field-programmable logic gate array, is a semi-custom circuit that can be programmed on an existing integrated circuit to implement custom functions.

[0059] Neural Networks, a subset of machine learning, is used for classifying and predicting data.

[0060] Example 1:

[0061] The purpose of this example is to provide a method for accelerating the inference of fully homomorphic encryption neural networks based on resource reuse.

[0062] Currently, the mainstream technologies for computing privacy data include secure multi-party computation (MPC: Multi-party Computation) and homomorphic encryption (Homomorphic Encryption), etc. Homomorphic encryption technology is an encryption scheme based on cryptography and has more reliable security. Fully Homomorphic Encryption is a branch of homomorphic encryption that can support an infinite number of ciphertext additions and multiplications and has a broader application prospect. In addition to homomorphic addition and homomorphic multiplication, operations such as Rotate are collectively called homomorphic operations (Homomorphic Operation). The fully homomorphic encryption scheme has developed three generations of technologies. Among them, the second-generation fully homomorphic encryption schemes represented by the BGV scheme, BFV scheme, and CKKS scheme are currently the most efficient in operation and the most widely used, and are also more concerned by the academic and industrial communities. Currently, there are multiple mature open-source libraries, such as SEAL, Palisade, HElib, HEAAN, etc.

[0063] Due to various problems existing in the current optimization and acceleration of fully homomorphic encryption neural networks, this example proposes a method for accelerating the inference of fully homomorphic encryption neural networks based on resource reuse. The proposed scheme fully considers that FPGA has powerful raw data computing power and reconfigurability, is superior to GPU in terms of energy consumption, and is superior to ASIC in terms of flexibility. It is an ideal hardware solution for accelerating fully homomorphic applications. The proposed scheme combines the flexible programming characteristics of FPGA, uses high-level synthesis (HLS) tools, combines the application characteristics of FHE-CNN and the on-chip resource limitations of FPGA, and explores the design space of on-chip resource configuration to give the optimal resource configuration method to accelerate the computing speed of FHE-CNN.

[0064] As Figure 11 shown, the proposed scheme specifically includes:

[0065] Obtain the information of the fully homomorphic encryption neural network to be accelerated and the hardware resource information of the FPGA;

[0066] Input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model to obtain the optimal hardware resource configuration scheme for the fully homomorphic encryption neural network during arithmetic processing in the FPGA;

[0067] Among them, the processing strategy of the hardware resource allocation model is: perform parallel and pipelining optimizations within each network layer of the fully homomorphic encryption operation and the fully homomorphic encryption neural network, and reuse the homomorphic basic operation modules between each network layer; at the same time, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, perform on-chip buffer reuse at different granularities; finally, with the goal of minimizing the time for the fully homomorphic encryption neural network to infer encrypted data, obtain the optimal resource configuration scheme, where the operations of the fully homomorphic encryption neural network are divided into three levels of operations: homomorphic basic operations, homomorphic operations, and network layer operations according to different granularities.

[0068] Furthermore, the hardware resource allocation model is specifically expressed as follows:

[0069]

[0070]

[0071]

[0072] Among them, LAT lr is the running time of the lr layer, Layer represents all the layers included in the fully homomorphic encryption neural network, lr represents one of the layers, OP represents the set of all homomorphic operations, op represents a certain homomorphic operation among them, DSP max represents the number of DSPs owned by the FPGA development board, BRAM max represents the total number of BRAMs owned by the FPGA development board, DSP op is the number of DSPs occupied by this homomorphic operation, BRAM lr is the number of BRAMs used by the lr layer.

[0073] Furthermore, each network layer operation includes several homomorphic operations, and each homomorphic operation includes several homomorphic basic operations. Among them, the operations of the fully homomorphic encryption neural network adopt a pipelining and parallel processing method and are pipelined in units of the homomorphic basic operations.

[0074] Further, parallel and pipelining optimizations are performed within each network layer for fully homomorphic encryption operations and fully homomorphic encryption neural networks, and reuse of homomorphic basic operation modules is adopted between network layers. Specifically: According to the order of homomorphic basic operations, homomorphic operations, and network layer operations, the parallelism within operations of different granularities is set respectively, so that the parallel effects of operations of different granularities are superimposed;

[0075] Or, for homomorphic basic operations, according to whether the coefficients of the RNS polynomial are traversed once or multiple times, the homomorphic basic operations are divided into two categories, and by setting different parallelisms for the two categories of operations, the running times of the two categories of operations are made similar;

[0076] Or, for pipelining the cross-homomorphic operations within a network layer, the same parallelism is set for different homomorphic operations, and based on the complexity of the data dependency relationship between RNS polynomials, the homomorphic operations are divided into two categories;

[0077] Or, for different network layer operations, according to whether they include KeySwitch operations, they are divided into two categories, and pipeline designs are carried out for different categories respectively.

[0078] Further, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, on-chip buffer reuse under different granularities is carried out. Specifically: The on-chip buffer takes the space for storing one RNS polynomial as the storage unit, and the on-chip buffer is divided into two categories according to the classification of homomorphic basic operations; among them, the reuse of the buffer includes reuse within homomorphic operations, reuse within network layers, and reuse between network layers.

[0079] Further, the homomorphic basic operations include; modular multiplication, modular addition, modular subtraction, modulo operation, fast Fourier number theory transform and its inverse transform;

[0080] Or, the homomorphic operations include plaintext-ciphertext addition, ciphertext addition, ciphertext multiplication, rescaling, relinearization, and rotation operations;

[0081] Or, the network layers include homomorphic convolutional layers, homomorphic activation layers, and homomorphic fully connected layers.

[0082] Further, the information of the fully homomorphic encryption neural network mainly includes the number of terms of the homomorphic encryption parameter polynomial, and the hardware resource information of the FPGA mainly includes the rated DSP resource quantity and BRAM resource quantity.

[0083] Specifically, for the convenience of understanding, the following combines the accompanying drawings to elaborate on the solution described in this embodiment in detail:

[0084] The following elaborates on the solution described in this embodiment from three aspects in detail:

[0085] (1) Optimization of Fully Homomorphic Encryption Operations

[0086] Based on the combination of fully homomorphic encryption and neural networks, parallel and pipelining optimizations are designed for fully homomorphic encryption operations and within the FHE-CNN layer, and on this basis, the fully homomorphic module is reused. The specific process is as follows:

[0087] The operations after combining fully homomorphic encryption and neural networks are divided into three levels. The highest level contains multiple intermediate-level operations, and the intermediate level contains multiple lowest-level operations. Parallel strategies are designed for different levels respectively. The lowest-level operation is the homomorphic basic operation (hereinafter simply referred to as the basic operation), including modular multiplication (Modular mult), modular addition (Modularadd), modular subtraction (Modular sub), modulo operation, fast number-theoretic transform (NTT) and its inverse transform (INTT); the intermediate-level operations are homomorphic operations, including plaintext-ciphertext addition (PC-add), ciphertext-ciphertext addition (CC-add), plaintext-ciphertext multiplication (PC-mult), ciphertext-ciphertext multiplication (CC-mult), rescaling (Rescale), relinearization (Relinearize), rotation (Rotate). The relinearization and rotation operations are collectively referred to as the KeySwitch operation. The highest level is the operation of an FHE-CNN network layer, where the FHE-CNN layer refers to the homomorphic convolution layer, homomorphic activation layer, homomorphic fully connected layer, etc.

[0088] The parallelism within each of the three levels is set so that the parallel effects from the lowest level to the highest level are superimposed, combining fine-grained parallelism and coarse-grained parallelism. The lowest-level parallelism refers to the parallelism within the basic operation. The basic operation calculates on an RNS polynomial (in the FHE scheme, a ciphertext is split into several RNS polynomials for calculation). When there is only one computing unit for the basic operation, the computing unit traverses all the coefficients of the polynomial for operation; when there are multiple parallel computing units for the basic operation, each computing unit processes a part of the coefficients, which can reduce the calculation time of the basic operation. The intermediate-level parallelism is the parallelism within the homomorphic operation. The homomorphic operation calculates in units of ciphertexts. A ciphertext contains multiple RNS polynomials. The intermediate-level parallelism refers to calculating several RNS polynomials simultaneously, that is, paralleling several basic operation modules. Sometimes there are dependencies between RNS polynomials that prevent parallel calculation, and we will design differently for different situations later. The highest-level parallelism refers to the parallelism within an FHE-CNN layer. The FHE-CNN layer needs to process multiple ciphertexts. The highest-level parallelism refers to parallelly calculating several ciphertexts at a time, that is, paralleling several homomorphic operation modules.

[0089] Combined with parallel design, we design a pipeline within the FHE-CNN layer. The FHE-CNN layer contains multiple homomorphic operations, and each homomorphic operation contains multiple basic operations. We perform pipelining with basic operations as the unit. Since pipelining of basic operations is a fine-grained pipelining method, it is beneficial to improve the utilization rate of computing resources and reduce the use of storage space. To improve the pipelining efficiency, we need to configure the parallelism within basic operations to make the running time of each basic operation similar and reduce the bubbles in the pipeline (which means a certain module in the pipeline stops running). We divide basic operations into two categories according to the calculation process within the basic operation. One category is NTT and INTT, which need to traverse the coefficients of an RNS polynomial for multiple rounds of calculation, and the other category is the remaining basic operations, which only traverse the coefficients of the RNS polynomial for one round. Usually, different parallelisms need to be set for these two categories to achieve similar running times.

[0090] Since the pipelining within the FHE-CNN layer targets basic operations, we implement pipelining across homomorphic operations within the layer. We set the same intermediate-level parallelism for different homomorphic operations to enable data flow. Due to the data dependency relationship between the internal RNS polynomials of homomorphic operations, we divide homomorphic operations into two categories. One category is the KeySwitch operation, and the remaining is the other category of operations. Due to the complex data dependency relationship of the KeySwitch operation, which affects the pipelining of the KeySwitch operation across other operations, we further divide the FHE-CNN layer into two categories. One category is the layer containing the KeySwitch operation, called the "KS layer", and the other category is the layer without the KeySwitch operation, called the "NKS layer". We design the pipelining for these two types of layers separately.

[0091] As Figure 1 shown, it is the pipeline of the NKS layer. The horizontal axis represents time, the vertical axis represents homomorphic operations, and the arrows represent the data flow direction. Each square represents a basic operation. The pipeline interval is the time of one basic operation. The pipelining within Rescale extends to the surrounding homomorphic operations, reflecting the design of pipelining across operations.

[0092] As Figure 2 shown, it is the pipeline of the KS layer. The horizontal axis represents time, and the vertical axis represents homomorphic operations. Since the KeySwitch operation needs to perform calculations of multiple basic operations to calculate the next ciphertext, the pipeline interval of this scheme depends on the pipeline interval of the KeySwitch operation.

[0093] The KS layer is generally used for calculating homomorphic matrix multiplication. Among them, homomorphic vector multiplication requires multiple consecutive rotation and homomorphic addition operations, as Figure 3。Since each subsequent rotation can only be performed after the previous addition calculation is completed, these rotation operations cannot be pipelined. In this case, we can calculate multiple vector multiplications at once, and there is no data dependence between these vector multiplications. For example, the KeySwitch module calculates the first rotation of vector multiplication A, then the first rotation of vector multiplication B, and then the second rotation of vector multiplication A, and so on. This can utilize the pipelining effect of the KeySwitch module for acceleration.

[0094] For the FHE-CNN layers, we adopt the method of reusing basic operations. Since one layer is calculated before the next layer, there is no part that is calculated simultaneously between layers, so different layers can reuse the same basic operation module. Therefore, during the configuration process, we set the parallelism for the three levels of parallelism respectively, instantiate to generate the corresponding calculation circuits on the FPGA, and use the same circuit for different layers to achieve the final acceleration and improve the resource utilization rate.

[0095] (2) On-chip buffer reuse

[0096] Design on-chip buffer (Buffer) reuse for the on-chip storage space, and reduce the performance loss caused by off-chip data transmission through reuse, so as to achieve the acceleration effect; the specific process is as follows:

[0097] Take the space for storing one RNS polynomial as the storage unit for the on-chip buffer (Buffer). Since the two categories of basic operations usually have different parallelisms, the Buffer needs to be partitioned differently. According to different partitions, the Buffer is divided into two categories. One is the Buffer for NTT / INTT operations, denoted as "Bn", and the other is the Buffer for the remaining basic operations, denoted as "Bb". Because the number of partitions for NTT / INTT is more, Bn can be reused for Bb, while Bb cannot be reused for Bn.

[0098] We divide the on-chip buffer reuse optimization technology into three levels of reuse. The lowest level is the reuse within a homomorphic operation, the middle level is the buffer reuse within the FHE-CNN layer, and the highest level is the reuse between different FHE-CNN layers.

[0099] The lowest-level reuse means that within a homomorphic operation, the same Buffer is used as much as possible for calculation, reducing data transfer and data transmission with the off-chip cache space. Figure 4 (a) to Figure 4 (e) show the Buffer usage results of each homomorphic operation. Taking Figure 4 (a) as an example, the input and output of CC-add reuse the Bb1 Buffer; takingFigure 4 Taking (c) as an example, in addition to the input and output Bb1, and the data read from off-chip (DDR), KeySwitch also requires Bn1 to store the data during the calculation process.

[0100] Intermediate-level reuse means that adjacent operations within the FHE-CNN layer are calculated using the same Buffer. When it is impossible to calculate using the same Buffer according to the data dependency relationship, different Buffers can be used for calculation, and non-adjacent Buffers can be reused. Figure 5 (a) to Figure 5 (a) to (b) show the Buffer reuse of the NKS layer and the KS layer. For the NKS layer, the data to be calculated is read from off-chip, and after the PC-mult operation, the data is stored on Bn1. The Rescale operation is performed on Bn1, and the result is added to Bb1. Then, the same operation is performed on multiple ciphertexts, and Bb1 is responsible for storing the accumulated result. For the KS layer, the data is read from off-chip and Bb1 to PC-mult, and the result is stored in Bn1. The Rescale operation is performed on Bn1, and the result is used as the input of KeySwitch. Bn2 is used as the intermediate cache of KeySwitch, and the result is stored in Bb1. Bb1 is reused here, and Bn1 stores the accumulated result of CC-add, and Bn1 is also reused.

[0101] The highest-level reuse means that different FHE-CNN layers share the same set of Buffers. According to the characteristic that one layer is calculated and then the next layer is calculated, except that the output of this layer to the input of the next layer is the same data and a Buffer is needed to transfer this data, the calculation processes of the two layers have no intersection. Therefore, the Buffers of one layer can be reused on another layer, and the total amount of Buffers finally used is the maximum value of the Buffer usage of each layer.

[0102] (3) Construction of the hardware resource allocation model

[0103] Based on the above two steps, model the resource configuration of FHE-CNN on the FPGA. Through the characteristics of the given FHE-CNN and the given FPGA, perform design space exploration to obtain the best-performing configuration plan; the specific process is as follows:

[0104] Model the latency optimized for the fully homomorphic operation:

[0105] The number of cycles for a type of basic operation including NTT / INTT is as follows:

[0106]

[0107] Among them, N represents the number of terms of the ciphertext polynomial, and nc NTT represents the parallelism inside the NTT. The NTT operation needs to traverse N numbers log 2 N times, and each NTT core calculates two numbers each time.

[0108] The cycle numbers of the operations other than NTT / INTT in the basic operations are as follows:

[0109]

[0110] Among them, N represents the number of terms of the ciphertext polynomial, and p represents the parallelism inside the operation.

[0111] To calculate the cycle numbers within the FHE-CNN layer, first obtain the cycle of the pipeline interval. Since the pipeline interval is obtained from the maximum value of the basic operations, there is:

[0112] LAT b = max{LAT basic , AT NTT}

[0113]

[0114] Among them, LAT b represents the maximum cycle of a ciphertext polynomial for basic operations, and PI represents the cycle of the pipeline interval. P intra represents the parallelism inside the homomorphic operation, and L represents the level of the ciphertext.

[0115] Furthermore, the cycle calculations for the KS layer and the NKS layer are as follows:

[0116]

[0117]

[0118] Among them, and respectively represent the number of parallel modules of the homomorphic operation, N on represents the number of input ciphertexts, L represents the level of the ciphertext, and PI represents the cycle of the pipeline interval.

[0119] Model the DSP resources used for the homomorphic operation:

[0120]

[0121] Among them, P inter and P intra respectively represent the parallelism inside the homomorphic operation op and within the FHE-CNN layer, Indicates the minimum number of DSP resources required without parallelism.

[0122] Model the on-chip buffers used:

[0123] Model the number of BRAMs used by the KS layer and the NKS layer respectively. Since the buffer is divided into two categories, Bn and Bb, calculations are performed separately for these two categories. The number of BRAMs for the KS layer is:

[0124]

[0125]

[0126] Where, and represent the constants of the BRAMs used by the KS layer buffer, and Bn and Bb represent the number of BRAMs used by these two types of buffers. and represent the parallelism of the homomorphic operations inside the KS layer and the parallelism inside the homomorphic operations of this layer.

[0127] The number of BRAMs for the NKS layer is:

[0128]

[0129]

[0130] Where, and represent the constants of the BRAMs used by the NKS layer buffer, and represent the parallelism of the homomorphic operations inside the NKS layer and the parallelism inside the homomorphic operations.

[0131] Construct an equation for the resource configuration of FHE-CNN on the FPGA. By extracting the network parameters of FHE-CNN as input, as well as the available resources of the FPGA, through the resource and storage model, the parallelism of each layer and inside the homomorphic operations is obtained, so as to obtain the fastest inference time for encrypted data. The general formula is:

[0132]

[0133]

[0134]

[0135] Where, OP represents the set of all homomorphic operations, and op represents a certain homomorphic operation among them. Layer represents all the layers included in the neural network, and lr represents one of the layers. Where lr ∈ {KS, NKS}, LATlr is LAT KS and LAT NKS is one of them, while LAT KS and LAT NKS both are obtained by summing LAT b The difference lies in the specific calculation of the summation, which is due to the different calculation methods of different KS and NKS. For BRAM resources, BRAM lr is BRAM KS and BRAM NKS is one of them, while BRAM KS = Bn KS + Bb KS , BRAM NKS = Bn NKS + Bb NKS .

[0136] DSP max represents the number of DSPs owned by this FPGA development board, and BRAM max represents the total number of BRAMs owned by this FPGA development board. When the sum of DSPs used in each homomorphic operation is less than DSP; and the number of BRAMs in the layer that uses the most BRAMs is less than the total number of BRAMs, calculate the sum of the times of each layer, which is the inference time of the entire network, and minimize it.

[0137] Model the latency optimized for fully homomorphic operations:

[0138] The number of cycles for a category of basic operations that include NTT / INTT is as follows:

[0139]

[0140] Among them, N represents the number of terms of the ciphertext polynomial, and nc NTT represents the parallelism inside NTT. The NTT operation needs to traverse N numbers log 2 N times, and each NTT core calculates two numbers each time.

[0141] The number of cycles for the remaining operations in the basic operations except NTT / INTT is as follows:

[0142]

[0143] Among them, N represents the number of terms of the ciphertext polynomial, and p represents the parallelism inside the operation.

[0144] To calculate the number of cycles within the FHE-CNN layer, first obtain the number of cycles of the pipeline interval. Since the pipeline interval is obtained from the maximum value of the basic operations, there is:

[0145] LAT b = max{LAT basic , LAT NTT}

[0146]

[0147] where LAT b represents the maximum period of the basic operations of a ciphertext polynomial, and PI represents the period of the pipeline interval. P intra represents the degree of parallelism inside the homomorphic operation, and L represents the level of the ciphertext.

[0148] Furthermore, the periods of the KS layer and the NKS layer are calculated as follows:

[0149]

[0150]

[0151] where and respectively represent the number of parallel modules of the homomorphic operation, N in represents the number of input ciphertexts, L represents the level of the ciphertext, and PI represents the period of the pipeline interval.

[0152] Model the DSP resources used in the homomorphic operation:

[0153] DSP = P inter · P intra · Const DSP

[0154] where P inter and P intra respectively represent the degrees of parallelism inside the homomorphic operation and within the FHE-CNN layer, and Const DSP represents the minimum number of DSP resources required without parallelism.

[0155] Model the on-chip buffers used:

[0156] Model the number of BRAMs used in the KS layer and the NKS layer respectively. Since the Buffer is divided into two categories, Bn and Bb, calculations are performed for these two categories separately. The number of BRAMs in the KS layer is:

[0157]

[0158]

[0159] where and Constants representing the BRAMs used in the KS layer buffer, where Bn and Bb represent the number of BRAMs used for these two types of buffers. and represent the parallelism of the homomorphic operations within the KS layer and the parallelism within the homomorphic operations of this layer.

[0160] The number of BRAMs in the NKS layer is:

[0161]

[0162]

[0163] Among them, and are constants representing the BRAMs used in the NKS layer buffer, and represent the parallelism of the homomorphic operations within the NKS layer and the parallelism within the homomorphic operations.

[0164] Construct equations for the resource configuration of FHE-CNN on FPGA. By extracting the network parameters of FHE-CNN as inputs, along with the available resource quantity of the FPGA, through the resource and storage model, the parallelism within each layer and within the homomorphic operations is obtained, thereby obtaining the time for the fastest inference of encrypted data. The general formula is:

[0165]

[0166]

[0167]

[0168] Among them, OP represents the set of all homomorphic operations, and op represents a certain homomorphic operation among them. Layer represents all the layers included in this neural network, and lr represents one of the layers. DSP max represents the number of DSPs owned by this FPGA development board, and BRAM max represents the total number of BRAMs owned by this FPGA development board. Under the condition that the sum of the DSPs used by each homomorphic operation is less than the DSP; and the number of BRAMs in the layer that uses the most BRAMs is less than the total number of BRAMs, find the sum of the times of each layer, which is the inference time of the entire network, and minimize it.

[0169] Furthermore, to prove the effectiveness of the solution described in this embodiment, the following relevant experimental verifications were carried out:

[0170] To test the effect of the method described in this embodiment, this embodiment uses a specific neural network to test the performance and demonstrate the optimization effect of this method.

[0171] Specifically, the method in this embodiment adopts Figure 6 The design framework. First, extract the information of the homomorphic encryption application combined with the neural network and the hardware resources of the given FPGA. The homomorphic encryption application information includes the number of terms N of the homomorphic encryption parameter polynomial, the ciphertext modulus Q, and the small modulus q i , where Q = ∏ 0≤i<L q i , L represents the small modulus, and the hardware resources of the FPGA include the rated number of DSP resources, the number of BRAM resources, and the specifications of the board, which are used as the input of the dedicated accelerator generation framework. Based on the parameterized homomorphic encryption operator library and the two technologies of on-chip computing and storage resource management, an integer linear programming model is constructed to perform automated design space exploration to achieve the optimal design, and the optimal hardware configuration scheme is used as the output. After generating the corresponding bitstream file through the Vivado HLS tool, it is burned on the FPGA to obtain the accelerated implementation of this application.

[0172] Experiment 1:

[0173] The neural network selected in this example is Lola-MNIST. It is a five-layer network used to predict the MNIST dataset. The description of this network is shown in Table 1:

[0174] Table 1 Lola-MNIST Network Description

[0175]

[0176] After combining with fully homomorphic encryption, the number of homomorphic operations contained in each layer of the obtained FHE-CNN, as well as the final inference accuracy and the model data size, are shown in Table 2:

[0177] Table 2 Basic Information of Lola-MNIST Network

[0178]

[0179] Experimental tests are carried out on two low-power FPGA development boards to verify the feasibility of the deployment and implementation of FHE-CNN on embedded FPGAs. Specifically, one is the mid-range FPGA ALINX ACU9EG (with Xilinx ZynqUltraScale+MPSoC XCZU9EG device), which has 2,520 DSP units and 32.1 Mbit on-chip BRAM. The other is the high-end FPGA ALINX ACU15EG (with Xilinx Zynq UltraScale+MPSoC XCZU15EG), which has 3,528 DSP units, 26.2 Mbit on-chip BRAM, and 31.5 Mbit on-chip URAM.

[0180] After combining the homomorphic network, each layer is classified into KS layers and NKS layers. Among them, Cnv1 is an NKS layer, Act1 is an NKS layer, Fc1 is a KS layer, Act2 is an NKS layer, and Fc2 is a KS layer. The input amounts of the ciphertext for each layer are 25, 1, 275, 1, and 70 respectively. The homomorphic parameter is taken as N = 8192, and the ciphertext level of the Cnv1 layer of FHE-CNN is 6, and then it decreases by 1 after passing through each layer until Fc2 is 2.

[0181] The solution space of this FHE-CNN is as follows: The lowest-level parallel solution space is from 1 to 32, the middle-level parallel solution space is from 1 to 6, and the highest-level parallel solution space is from 1 to infinity. We traverse all solutions in the solution space and finally obtain that on the FPGA ALINX ACU9EG board, the parallelism of the basic operations NTT / INTT is 4, and the parallelism of other basic operations is 1; the parallelism inside the homomorphic operation, PC-mult is 3, CC-Mult is 1, Rescale is 3, and KeySwitch is 3; the parallelism within the FHE-CNN layer, except for 2 parallel KeySwitch modules, the others are all 1 module. The frequency of the board is set to 100 MHz, and the final time for on-board verification is 0.24 seconds. For the FPGA ALINX ACU15EG board, the parallelism of the basic operations NTT / INTT is 4, and the other basic operations are 1; the parallelism inside the homomorphic operation, PC-mult is 3, CC-Mult is 1, Rescale is 3, and KeySwitch is 3; the parallelism within the FHE-CNN layer, except for 3 parallel KeySwitch modules, the others are all 1 module. The frequency of the board is set to 100 MHz, and the final time for on-board verification is 0.19 seconds. Compared with the state-of-the-art CPU implementation (the implementation of the Lola scheme), the speed is increased by 11.58 times, and the energy consumption ratio is increased by 1019.04 times.

[0182] Table 3 Acceleration Results

[0183]

[0184] In addition, for the implementation on different FPGAs, we give the solutions of the integer linear programming to finally obtain the inter-parallelism and intra-parallelism, as shown in Figure 7 (a) and Figure 7(b). It can be seen from this that the parallelism setting of ACU15EG for KeySwitch operation is higher than that of ACU9EG. This is because the resources of ACU15EG are higher than those of ACU9EG, which can increase the parallelism of KeySwitch and thus reduce the latency. The parallelism of the CC-mult operation for both boards is 1. This is because the number of calls to the CC-mult module is small. Even if the parallelism is reduced, the impact on the total latency is not significant. Therefore, the resources saved by the low parallelism of CC-mult are used for other bottleneck operations.

[0185] To prove the optimization effects of the fully homomorphic operator optimization technology and the on-chip buffer reuse technology, we implemented a control group without using these two optimization means for comparative experiments. The experimental results of on-chip buffer reuse are as Figure 8 . It can be seen that the Fc1 layer occupies the most inference time. Through the BRAM reuse method, the BRAM used by the FC1 layer increased from 25.8% to 84.8%, and thus the inference speed of this layer increased by 6.63 times.

[0186] The results of the fully homomorphic operator optimization technology are shown in Figure 9 . After operator optimization and resource reuse, the usage of DSP resources used separately for each layer has increased, thus increasing the parallelism of homomorphic operations. Combined with the on-chip buffer reuse technology, the inference time of each layer has been reduced, verifying the optimization effect of this technology.

[0187] In addition, this acceleration optimization framework adopts high-level synthesis technology, which has the characteristics of being flexible and easy to program and short development time. Combining the flexibility of the acceleration optimization framework's own configurable programming, it is possible to explore the design space for FPGA development boards with different resources, so as to find the optimal design. We conducted experiments on this, Figure 10 showing the designs generated by this optimization framework when the number of BRAMs for BRAM-Latency ranges from 350 to 1500. Each point represents a configuration, and the configuration points on the red line reach the Pareto optimum. When the BRAM resource quantity is at a low level, there are fewer available optimal configuration points. This is because even in the case of the lowest parallelism, some on-chip BRAMs are still required to store temporary data.

[0188] Experiment 2:

[0189] This experiment uses the Lola-Cifar network to test the effect of the FPGA accelerator design method for fully homomorphic encryption neural network inference. Lola-Cifar is also a 5-layer network. The difference compared with Lola-MNIST is that the weight scale of this network is larger, the number of homomorphic operations required is more, and the input data is larger. The description of this network is shown in Table 4:

[0190] Table 4 Description of Lola-Cifar Network

[0191]

[0192] Combined with fully homomorphic encryption, the number of homomorphic operations contained in each layer of the resulting FHE-CNN, as well as the final inference accuracy and the model data size, are shown in Table 5:

[0193] Table 5 Basic information of the Lola-Cifar network

[0194]

[0195] The FPGA development boards used are ALINX ACU9EG (with Xilinx Zynq UltraScale+ MPSoC XCZU9EG device) and ALINX ACU15EG (with Xiliinx Zynq UltraScale+ MPSoC XCZU15EG) respectively. For the hardware optimization and acceleration framework of the neural network with fully homomorphic encryption, the optimal configuration solution is automatically generated, and the acceleration results are shown in Table 6:

[0196] Table 6 Acceleration results

[0197]

[0198] Among them, KS, λ, N, and log Q respectively represent the number of KeySwitch operations, the security level, the degree of the polynomial of the homomorphic encryption parameter, and the modulus bit width of the homomorphic encryption parameter. TDP represents the power of the used hardware platform for this implementation. Through actual on-board measurement, the final experimental optimization result has a 13.49-fold increase in inference speed and a 1187.12-fold increase in energy consumption ratio compared with the implementation of the Lola scheme.

[0199] Through Experiment 1 and Experiment 2, it is proved that our acceleration framework can achieve a speed increase of more than 10 times compared with CPU acceleration and an energy consumption ratio increase of more than 1000 times for the acceleration of multiple FHE-CNNs. In addition, the framework customizes and generates a configuration scheme for two different FPGA boards, generates the Pareto optimal solution of performance-resources, and obtains the optimal configuration result. The framework realizes the automatic generation of the optimal hardware resource configuration for a given neural network and a given FPGA development board, providing a solution for the hardware optimization and deployment of FHE-CNN.

[0200] Example 2:

[0201] The purpose of this embodiment is to provide a fully homomorphic encryption neural network inference acceleration system based on resource reuse.

[0202] A fully homomorphic encryption neural network inference acceleration system based on resource reuse includes:

[0203] A data acquisition unit, which is used to acquire information of a fully homomorphic encryption neural network to be accelerated and hardware resource information of an FPGA;

[0204] A resource configuration unit, which is used to input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model, and obtain an optimal hardware resource configuration scheme for the fully homomorphic encryption neural network to perform arithmetic processing in the FPGA;

[0205] Wherein, the processing strategy of the hardware resource allocation model is: parallel and pipeline optimization is carried out for fully homomorphic encryption operations and each network layer of the fully homomorphic encryption neural network, and reuse of homomorphic basic operation modules is adopted between each network layer; at the same time, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, on-chip buffer reuse at different granularities is carried out; finally, with the goal of minimizing the time for the fully homomorphic encryption neural network to infer encrypted data, an optimal resource configuration scheme is obtained, wherein the operations of the fully homomorphic encryption neural network are divided into three levels of operations: homomorphic basic operations, homomorphic operations, and network layer operations according to different granularities.

[0206] Specifically, the system in this embodiment corresponds to the method in Embodiment 1, and its technical details have been described in detail in Embodiment 1, so they will not be elaborated here.

[0207] In more embodiments, there is also provided:

[0208] An electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.

[0209] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0210] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.

[0211] A computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by the processor, the method in Embodiment 1 is completed.

[0212] The method in the first embodiment can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0213] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0214] The above-described method and system for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse provided by the above embodiment can be realized and have broad application prospects.

[0215] The above are only the preferred embodiments of this disclosure and are not used to limit this disclosure. For those skilled in the art, this disclosure can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse, characterized in that, it includes: Obtain the information of the fully homomorphic encryption neural network to be accelerated and the hardware resource information of the FPGA; Input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model to obtain the optimal hardware resource configuration scheme when the fully homomorphic encryption neural network performs arithmetic processing in the FPGA; Among them, the processing strategy of the hardware resource allocation model is: parallel and pipelining optimization is performed within each network layer of the fully homomorphic encryption operation and the fully homomorphic encryption neural network, and the reuse of homomorphic basic operation modules is adopted between each network layer; at the same time, for the on-chip storage space of the FPGA, based on the arithmetic division of the fully homomorphic encryption neural network, on-chip buffer reuse at different granularities is performed; finally, with the goal of minimizing the time of the encrypted data for the inference of the fully homomorphic encryption neural network, the optimal resource configuration scheme is obtained, where the arithmetic of the fully homomorphic encryption neural network is divided into three levels of arithmetic: homomorphic basic operation, homomorphic operation, and network layer operation according to different granularities; The parallel and pipelining optimization within each network layer of the fully homomorphic encryption operation and the fully homomorphic encryption neural network, and the reuse of homomorphic basic operation modules between each network layer are specifically: according to the order of homomorphic basic operation, homomorphic operation, and network layer operation, the parallelism within different granularity operations is set respectively, so that the parallel effects of different granularity operations are superimposed; Or, for the homomorphic basic operation, according to whether the coefficients of the RNS polynomial are traversed once or multiple times, the homomorphic basic operation is divided into two categories, and by setting different parallelisms for the two categories of operations, the running times of the two categories of operations are made similar; Or, for the pipelining of cross-homomorphic operations within the network layer, the same parallelism is set for different homomorphic operations, and based on the complexity of the data dependence relationship between the RNS polynomials, the homomorphic operations are divided into two categories; Or, for different network layer operations, according to whether they include KeySwitch operations, they are divided into two categories, and pipeline designs are performed for different categories respectively.

2. The method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse according to claim 1, characterized in that, the hardware resource allocation model is specifically expressed as follows: Among them, is the running time of the layer, represents all the layers included in the fully homomorphic encryption neural network, represents one of the layers, represents the set of all homomorphic operations, represents a certain homomorphic operation among them, represents the number of DSPs owned by the FPGA development board, represents the total number of BRAMs owned by the FPGA development board, is the number of DSPs occupied by this homomorphic operation, is the number of BRAMs used by the layer.

3. The method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse according to claim 1, characterized in that, Each network layer operation includes several homomorphic operations, and each homomorphic operation includes several homomorphic basic operations. Among them, the arithmetic of the fully homomorphic encryption neural network adopts a pipelining and parallel processing method, and the pipelining is performed in units of the homomorphic basic operation.

4. The method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse according to claim 1, characterized in that, For the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, on-chip buffer reuse at different granularities is performed. Specifically: The on-chip buffer is taken as the storage unit with the space for storing an RNS polynomial, and the on-chip buffer is divided into two categories according to the classification of homomorphic basic operations; among them, the reuse of the buffer includes reuse within homomorphic operations, reuse within network layers, and reuse between network layers.

5. A method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse as described in claim 1, characterized in that, the homomorphic basic operations include; modular multiplication, modular addition, modular subtraction, modulo operation, fast number theory transform and its inverse transform; or, the homomorphic operations include plaintext-ciphertext addition, ciphertext addition, ciphertext multiplication, rescaling, relinearization, and rotation operations; or, the network layers include homomorphic convolutional layers, homomorphic activation layers, and homomorphic fully connected layers.

6. A method for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse as described in claim 1, characterized in that, the information of the fully homomorphic encryption neural network mainly includes the number of polynomial terms of the homomorphic encryption parameters, and the hardware resource information of the FPGA mainly includes the rated DSP resource quantity and the BRAM resource quantity.

7. A system for accelerating the inference of a fully homomorphic encryption neural network based on resource reuse, characterized in that, comprises: a data acquisition unit, which is used to acquire the information of the fully homomorphic encryption neural network to be accelerated and the hardware resource information of the FPGA; a resource configuration unit, which is used to input the information of the fully homomorphic encryption neural network and the hardware resource information into a pre-constructed hardware resource allocation model to obtain the optimal hardware resource configuration scheme for the fully homomorphic encryption neural network to perform arithmetic processing in the FPGA; wherein, the processing strategy of the hardware resource allocation model is: parallel and pipelining optimizations are performed for the fully homomorphic encryption operations and within each network layer of the fully homomorphic encryption neural network, and reuse of homomorphic basic operation modules is adopted between each network layer; meanwhile, for the on-chip storage space of the FPGA, based on the operation division of the fully homomorphic encryption neural network, on-chip buffer reuse at different granularities is performed; finally, with the goal of minimizing the time of the fully homomorphic encryption neural network to infer encrypted data, the optimal resource configuration scheme is obtained, wherein the operations of the fully homomorphic encryption neural network are divided into three levels of operations: homomorphic basic operations, homomorphic operations, and network layer operations according to different granularities; The parallel and pipelining optimizations are performed for the fully homomorphic encryption operations and within each network layer of the fully homomorphic encryption neural network, and reuse of homomorphic basic operation modules is adopted between each network layer, specifically: according to the order of homomorphic basic operations, homomorphic operations, and network layer operations, the parallelism within different granularity operations is set respectively, so that the parallel effects of different granularity operations are superimposed; or, for homomorphic basic operations, according to whether the coefficients of the RNS polynomial are traversed in one round or multiple rounds, the homomorphic basic operations are divided into two categories, and by setting different parallelisms for the two categories of operations, the running times of the two categories of operations are made similar; Alternatively, for the pipeline that implements cross-homomorphic operations within the network layer, the same degree of parallelism is set for different homomorphic operations, and the homomorphic operations are divided into two categories based on the complexity of the data dependency relationship between RNS polynomials; Alternatively, for different network layer operations, they are divided into two categories according to whether they include KeySwitch operations, and pipeline designs are respectively performed for different categories.

8. An electronic device, comprising a memory, a processor, and a computer program running on the memory, wherein, when the processor executes the program, it implements a method for accelerating inference of a fully homomorphic encryption neural network based on resource reuse according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium, having a computer program stored thereon, wherein, when the program is executed by a processor, it implements a method for accelerating inference of a fully homomorphic encryption neural network based on resource reuse according to any one of claims 1-6.

Citation Information

Patent Citations

  • Fully homomorphic encryption deep learning reasoning method and system based on FPGA

    CN112699384A

  • Configuring reduced instruction set computer processor architecture to execute fully homomorphic encryption algorithms

    CN114631284A