Method for accelerating bootstrapping in cryptographic application
By processing input ciphertexts in batches and using a small on-chip cache for bootstrap key elements, the accelerator architecture addresses memory bandwidth limitations and computational inefficiencies in FHE bootstrapping, achieving improved performance and throughput.
Patent Information
- Application Number
- JP2023198464
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-06-03
AI Technical Summary
Current fully homomorphic encryption (FHE) schemes face significant performance bottlenecks due to the costly bootstrapping operation, particularly in polynomial multiplication, which is exacerbated by memory bandwidth limitations when loading large bootstrap key coefficients.
The proposed solution involves an accelerator architecture that processes input ciphertexts in batches, allowing for the amortization of loading bootstrap key elements across multiple iterations. This approach reduces memory requirements and computational load by using a small on-chip bootstrap cache memory to store key elements, enabling efficient loading and processing.
This method significantly reduces memory bandwidth bottlenecks and enhances computational efficiency during the bootstrapping operation, achieving higher throughput with moderate off-chip memory bandwidth requirements.
Smart Images

Figure 2025084508000001_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of encryption. More specifically, the present invention relates to a solution for accelerating the bootstrapping operation.
Background Art
[0002] Machine learning (ML) is driven by the availability of abundant data and has advanced rapidly in recent years, leading to new applications from autonomous driving to medical diagnosis. In many applications, the ML model is developed by one party and made available to users as a cloud service. For example, in such a context of cloud computing, users take a certain risk when uploading raw data to an untrusted cloud server. Therefore, it is necessary to provide sufficient security to protect the users' data. A promising new technology that has emerged in the field of data security is fully homomorphic encryption (FHE), which enables performing homomorphic calculations on encrypted data (ciphertext) without learning further information about that data. In other words, it is not necessary to decrypt the data first. Thus, the client encrypts the data with FHE before sending it to the cloud. Next, the cloud service calculates the FHE program on the encrypted data without obtaining information about the input and returns the (still encrypted) result to the client. Only the client can finally decrypt and obtain the result. The methods for implementing FHE have been improved over the years until they became practical.
[0003] FHE algorithms are often run on cloud computing servers. However, the computations are slow, sometimes even several orders of magnitude slower than non-encrypted computations. This remains true today, despite significant improvements in FHE schemes and algorithms in recent years. To avoid the speed limitations of FHE, designers are shifting focus from general-purpose CPUs to more specialized hardware implementations. For example, ASIC emulation on advanced technology nodes promises better FHE acceleration, but these ASICs can take years to be manufactured and made available. Furthermore, they are typically specialized for a limited range of parameter sets. FPGA-based implementations can be developed more quickly than ASIC implementations, are flexible in changing parameter sets, and can be easily deployed to FPGA-based cloud instances while facilitating large-scale speedups. As a result, they have become a popular target for FHE acceleration.
[0004] Typically, the execution of FHE schemes is bottlenecked by the so-called bootstrapping procedure. FHE ciphertexts contain noise for security purposes in the encryption. Furthermore, each FHE operation increases the noise present in the ciphertext. FHE is resilient to this noise as long as a certain threshold noise level is not exceeded. Therefore, FHE schemes need to periodically call the bootstrapping procedure, which reduces the amount of noise in the ciphertext and keeps it below the threshold noise level so that further computations can be performed. Bootstrapping is one of the costly operations in FHE computations.
[0005] The main cost of bootstrapping is thus generally in polynomial multiplication in FHE schemes. The ciphertext data on which computations are performed in FHE schemes are large polynomials (length N) from a specific scheme-dependent polynomial ring. Typical operations on these polynomials include addition and multiplication. Addition is linear in the length of the polynomial (O(N) operations), while multiplication has a quadratic cost (O(N2 ) operations).
[0006] Since one of the most costly operations in the FHE scheme is multiplication, speeding up the multiplication operation can substantially contribute to reducing the computational overhead. To speed up such multiplications, various well-studied algorithms are available. This can be achieved by using multiplication algorithms such as Toom-Cook multiplication, Karatsuba multiplication, or fast Fourier transform (FFT) for fast polynomial multiplication, taking advantage of specific properties of polynomials, and usually FFT is the fastest option.
[0007] Polynomial multiplication using FFT is not exact and may introduce some noise into the calculations. Quantization noise can occur because FFT operates on real numbers and is approximated using floating-point or fixed-point representations. As described above, all currently available FHE schemes have inherent noise that increases with each operation. The additional FFT quantization noise is added to the noise inherent in the FHE scheme, but is somewhat tolerable. However, as mentioned above, if there is too much noise, the bootstrapping will fail, so very careful handling is required when using FFT.
[0008] Two important FHE schemes are the FHEW scheme disclosed in the paper "FHEW: Bootstrapping homomorphic encryption in less than a second" (L. Ducas et al., Eurocrypt, pp. 617-640, 2015) and the TFHE scheme described in "TFHE: Fast Fully Homomorphic Encryption Over the Torus" (I. Chillotti et al., J. Cryptol. 33, 34-91, 2020). Both FHEW and TFHE enable the use of homomorphic Boolean algebras, such as NAND, XOR, and XNOR logic gates, in addition to homomorphic addition and multiplication.
[0009] Schemes like TFHE and FHEW have revisited the bootstrapping approach and made it cheaper, but are essentially linked to homomorphic computation. In these schemes, most of the homomorphic operations require bootstrapping of the ciphertexts immediately, i.e., after each encrypted gate. These are characterized by much faster bootstrapping algorithms compared to previous generations of FHE schemes. Furthermore, the bootstrapping in TFHE is a general-purpose tool that can be further "programmed" by any function applied to the ciphertext, e.g., non-linear activation functions in ML neural networks. This approach is called programmable bootstrapping (PBS) and constitutes the main cost of TFHE homomorphic computation. PBS, which accounts for up to 99% of the encrypted gate computations (encrypted NAND, XOR, XNOR, …), is the main target of TFHE's high-throughput hardware acceleration.
[0010] Some more detailed information is provided regarding the Torus Fully Homomorphic Encryption (TFHE) scheme and its operation. Torus Fully Homomorphic Encryption is a homomorphic encryption scheme based on the Learning With Errors (LWE) problem. This operates on elements defined over the real torus
Number
Number
[0011] TFHE further describes two variant ciphertexts. First, there is a generalized version (TGLWE), where e and μ are
Number
Number
Number
Number
Number
Number
Number
Number
[0012] Bootstrapping aims to reduce the noise of the ciphertext. For security reasons, bootstrapping homomorphically decrypts the ciphertext within the encrypted domain. This means that we hope to calculate b - a·s = e + μ homomorphically, and more specifically, since it is a "programmable" bootstrapping, it means that we hope to further calculate the function f(μ) on the data. To achieve this programmable bootstrapping, first, a "test" polynomial
Number
Number
Number
Number
Number
Number
[0013] Collectively, different TGGSW ciphertexts BK1,..., BKn that encrypt one secret coefficient s 1 , …, s n respectively are known as bootstrap key elements (also called bootstrap key coefficients), and together they form the bootstrap key. The result of the above operation is the TGLWE accumulator ACC that is "blindly" rotated at the position of the secret quantity b - a s, from which the output TLWE ciphertext c out can be easily extracted. A high-level overview of the calculations performed in PBS in the TFHE scheme is given in the following algorithm.
Table 1
[0014] Thus, the bootstrap operation requires two main inputs, namely, the input ciphertext coefficients a 1 , …, n and the bootstrap key coefficients BK 1 , …, BK n and the bootstrap key BK containing them. In each iteration i = 1, …, n of the bootstrap operation, one of the two elements is required. The ciphertext coefficient a i is relatively small in size and thus easy to accommodate. In contrast, the bootstrap key coefficients
Number
[0015] As shown above, the TFHE programmable bootstrap is mainly summarized in the iterative calculation of the outer product
Number
[0016] However, the FHE scheme requires polynomial multiplication modulo X N + 1, which performs a negative circular convolution rather than a conventional circular convolution. In a circular convolution (with N coefficients), coefficients that are outside the boundary (at position i > N) are cycled to the first coefficient (at position i - N). Conversely, in a negative circular convolution, these coefficients are not only cycled but also negated, i.e., negatively wrapped. This negative circular convolution has a period of 2N and thus a simple implementation would require an FFT of size 2N.
[0017] The cost of negative cyclic FFT for real input data can be reduced in various ways. The fact that the FFT is computed for complex numbers provides the first opportunity for optimization. Since the input polynomial is purely real and has an imaginary component equal to zero, an FFT optimized from real to complex (r2c) can be used, which achieves approximately a two-fold improvement in speed and memory usage. This is the approach adopted by the TFHE and FHEW software libraries that compute a size-2N r2c FFT. Another possible optimization is to use a "twiddle" preprocessing step to compute a negative cyclic FFT that can have a period and size of 2N instead of a normal FFT with period and size N. During the twiddle, the coefficients of the input polynomial a are multiplied by the so-called twiddle coefficients that are powers of the 2N-th root of unity ψ = ω 2N to the power of.
Number
Number
[0018] Two prior art accelerators that accelerate TFHE bootstrapping are MATCHA disclosed in the paper "MATCHA: A fast and energy-efficient accelerator for fully homomorphic encryption over the torus" (L. Jiang et al., ACM / IEEE DAC, pp. 235-240, 2022) and Ye et al. disclosed in the paper "FPGA acceleration of fully homomorphic encryption over the torus" (T. Ye et al., IEEE HPEC, pp. 1-7, 2022). MATCHA is constructed based on a classical CPU approach. This includes a set of TGGSW clusters with external product cores operating from a register file. As a result, MATCHA becomes a bottleneck due to data movement and cache memory access contention. Ye et al. includes a pipelined unit for computing CMUX. Each pipeline instance of Ye et al. includes SRAM that stores a single coefficient BK i After consuming all the coefficients, the next coefficient is loaded from off-chip memory. In practice, the off-chip memory bandwidth is limited, and loading the next coefficient is a major throughput bottleneck of the design.
[0019] Therefore, there is a need for an improved accelerator for bootstrapping that further reduces the computational and memory requirements compared to currently known accelerator solutions. Summary of the Invention Problems to be Solved by the Invention
[0020] It is an object of embodiments of the present invention to provide a device and method with reduced memory requirements and computational load for performing a bootstrapping operation in an encryption application. Means for Solving the Problems
[0021] The above object is achieved by the solution according to the present invention.
[0022] In a first aspect, the present invention relates to a method for performing a bootstrap operation in an encryption application. The method comprises - receiving, in an accelerator, one or more input ciphertexts used in the encryption application to be bootstrapped and repeatedly processing one or more accumulator variables as a function of a part of the input ciphertexts, whereby each accumulator variable is linked to one input ciphertext and each iteration step is performed in sequence for each of the one or more accumulator variables, receiving and processing; - multiplying, within each iteration, the processed accumulator variables by bootstrap key elements belonging to a bootstrap key comprising a plurality of bootstrap key elements, the bootstrap key elements being obtained from a bootstrap cache memory within the accelerator, multiplying; - while performing the multiplication in sequence for each of the one or more accumulator variables, loading, from an external memory to the bootstrap cache memory, the next bootstrap key element of a plurality of bootstrap key elements to be used in the next iteration of the bootstrap operation.
[0023] The proposed solution actually enables the bootstrap operation to be performed in an efficient way, thereby avoiding memory bottlenecks. By creating from a part of the input ciphertext batch, it becomes possible to load the various bootstrap key coefficients at a much slower pace than the prior art solutions. In the present invention, the accumulation iterations are completed first for the entire batch and only then is the next iteration with the next bootstrap key coefficient started. Thus, more time is available to load the (relatively large) bootstrap key coefficients. This is one of the main advantages of the present invention.
[0024] In a preferred embodiment, the bootstrap cache memory is SRAM memory.
[0025] In some embodiments, one or more input ciphertexts are supplied from further memory within the accelerator. Alternatively, in other embodiments, one or more input ciphertexts are supplied from external memory.
[0026] In a preferred embodiment, the vector elements and / or the bootstrap key elements and / or the accumulator variables are represented in fixed point.
[0027] In an advantageous embodiment, the proposed method for performing the bootstrap operation is for use in a torus fully homomorphic encryption scheme.
[0028] In another aspect, the present invention relates to a program executable on a programmable device that, when executed, includes instructions for performing the above method.
[0029] In yet another aspect, the present invention relates to an accelerator for performing a bootstrap operation in an encryption application. The accelerator is - a preprocessing block configured to receive one or more input ciphertexts to be bootstrapped and iteratively process one or more accumulator variables as a function of a portion of the input ciphertexts, whereby each accumulator variable is linked to one input ciphertext and each iteration step is performed in sequence for each of the one or more accumulator variables; - an arithmetic unit configured to multiply, within each iteration, a processed accumulator variable by a bootstrap key element belonging to a bootstrap key that includes a plurality of bootstrap key elements. The accelerator further comprises a bootstrap cache memory for storing the bootstrap key elements to be multiplied, and is adapted to load the next bootstrap key element of a plurality of bootstrap key elements to be used in the next iteration of the bootstrap operation into the bootstrap cache memory while sequentially performing multiplication operations on each of one or more accumulator variables.
[0030] In a preferred embodiment, the arithmetic unit comprises a plurality of cascaded calculation stages, and each calculation stage is operable simultaneously on different parts or different operations related to one or more accumulator variables.
[0031] Advantageously, the plurality of cascaded calculation stages form a pipeline.
[0032] In a preferred embodiment, fixed-point representation is used for vector elements and / or bootstrap key elements and / or accumulator variables.
[0033] In some embodiments, the bootstrap cache memory is arranged to store two or more next bootstrap key elements to be used.
[0034] In an advantageous embodiment, the accelerator is implemented on a field-programmable gate array or an application-specific integrated circuit (ASIC). In one embodiment, the accelerator is implemented on an FPGA and is optimized to efficiently use digital signal processing (DSP) units, look-up tables (LUTs) and block RAMs (BRAMs) within the FPGA fabric. In some embodiments, the accelerator is implemented entirely in software.
[0035] In some embodiments, the arithmetic unit has forward and inverse FFT or NTT calculation stages, and the forward FFT or NTT has a higher throughput than the inverse FFT or NTT.
[0036] In one aspect, the present invention relates to a computing system comprising the above-described accelerator and a memory external to the accelerator, the memory being arranged to store one or more bootstrap keys for use in a bootstrap operation.
[0037] In another aspect, the present invention relates to the use of the above-described accelerator in a cloud computing service.
[0038] For the purpose of summarizing the specific objects and advantages achieved by the present invention as compared to the prior art, the specific objects and advantages of the present invention are described hereinabove. It should be understood, of course, that not all such objects or advantages may be achieved in accordance with any particular embodiment of the present invention. Thus, for example, those skilled in the art will recognize that the present invention may be embodied or practiced in a manner that achieves or optimizes one advantage or group of advantages taught herein without necessarily achieving other objects or advantages that may be taught or suggested herein.
[0039] The above and other aspects of the present invention will be apparent from and will be elucidated with reference to the embodiments described hereinafter.
Brief Description of the Drawings
[0040] Next, the present invention will be further described by way of example with reference to the accompanying drawings, in which like reference numerals refer to like elements in the various drawings.
[0041]
Figure 1
Figure 2
Figure 3
Figure 4
Best Mode for Carrying Out the Invention
[0042] The present invention is described with respect to specific embodiments and with reference to specific drawings, but the present invention is not limited thereto and is limited only by the claims.
[0043] Furthermore, the terms first, second, and the like in this specification and the claims are used to distinguish similar elements and are not necessarily used to describe an order in time, space, ranking, or any other manner. Such terms are interchangeable under appropriate circumstances, and it should be understood that the embodiments of the present invention described herein are operable in an order other than those described or illustrated herein.
[0044] It should be noted that the term "comprising" used in the claims should not be construed as being limited to the means recited thereafter and does not exclude other elements or steps. Therefore, it is construed as specifying the presence of the recited features, integers, steps, or components referred to, but does not exclude the presence or addition of one or more other features, integers, steps, or components, or groups thereof. Therefore, the scope of the expression "a device comprising means A and B" should not be limited to a device consisting only of components A and B. It means, with respect to the present invention, that the only relevant components of the device are A and B.
[0045] References to "one embodiment" or "an embodiment" throughout this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, although they may. Further, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, as will be apparent to those skilled in the art from this disclosure.
[0046] Similarly, in the description of exemplary embodiments of the present invention, it should be understood that various features of the present invention may be grouped together in a single embodiment, figure, or description for the purpose of streamlining the disclosure and facilitating understanding of one or more of the various inventive aspects. However, this method of disclosure should not be construed as reflecting that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects are less than all of the features of the single disclosed embodiment described above. Accordingly, the claims that follow the detailed description are expressly incorporated into this detailed description, and each claim stands on its own as a separate embodiment of the present invention.
[0047] Furthermore, although some embodiments described herein may include some features included in other embodiments and may not include other features, as will be understood by those skilled in the art, combinations of features of different embodiments are within the scope of the present invention and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0048] Note that the use of particular terms when describing some features or aspects of the present invention should not be construed as implying that the terms are redefined herein to include any particular characteristics of the features or aspects of the invention with which the terms are associated.
[0049] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0050] From the introduction, it has already become clear that in encryption applications, methods and devices for accelerating the bootstrap operation are needed.
[0051] In a first aspect, the present invention discloses a method for performing a bootstrap operation based on a fully homomorphic encryption (FHE) scheme. In an advantageous embodiment, the scheme is a torus FHE scheme (TFHE). The present invention proposes an approach to amortize the loading of bootstrap key elements for each iteration of the bootstrap process. The necessary huge memory bandwidth bottleneck encountered when performing bootstrap with prior art schemes is thus overcome.
[0052] In a further aspect, the present invention proposes an accelerator device having an architecture adapted to execute a novel bootstrap approach. In a preferred embodiment, the accelerator is implemented on a field programmable gate array (FPGA). In one embodiment, the accelerator is implemented on an FPGA and optimized to efficiently use digital signal processing (DSP) units, look-up tables (LUTs), and block RAMs (BRAMs) within the FPGA fabric. In an alternative embodiment, an ASIC implementation of the accelerator is provided. In some embodiments, the accelerator implemented on an FPGA or ASIC is made available as a cloud computing instance.
[0053] The proposed approach is based on the inventors' observation that the memory bottleneck problem can be solved, or at least mitigated, by processing the various parts of the ciphertext processed in the encryption application in batches of a pre-defined size q, whereby the same iterative step is used and thus the same bootstrap key factor BK of the bootstrap key i is collected for the vector elements that require it. In a preferred embodiment, the value of q is between 8 and 32, although the invention is clearly not limited thereto. Thus, assuming q ciphertexts each having n parts, each batch is a vector
Number
Number
Number
Number
Number
[0054] Bootstrap key element (i.e., coefficient) BK i is stored in the on-chip bootstrap cache memory. The bootstrap key element BK i and each vector element of
Number
[0055] An accelerator microarchitecture adapted to implement the method of the present invention is presented here in more detail. Figure 2 shows a basic block scheme of an accelerator 1 according to the present invention. The main calculations are performed in a loop that is repeated a given number of times (typically n times, once for each of the n coefficients in the ciphertext). In the loop, the accumulator variable moves from the preprocessing block 5 to the arithmetic unit 4 and the postprocessing unit 6 and then back to the preprocessing block to start the subsequent iteration. The computing unit is preferably pipelined, and data related to q different parts of the ciphertext being batch processed is distributed across the computing unit.
[0056] Accelerator 1 comprises a preprocessing block 5 where some preprocessing is performed for each iteration, together with q public elements a (j) , j = 0,..., q - 1, various parts of the ciphertext that make up
Number
Number
Number
Number
[0057] The bootstrap key elements are larger in size than the ciphertext portion. In a preferred embodiment, the various bootstrap key elements BK i are stored in off-chip SDRAM memory. In some embodiments, the off-chip SDRAM memory is DDR (Double Data Rate) SDRAM. In other embodiments, the off-chip memory is HBM (High Bandwidth Memory) SDRAM.
[0058] As shown in FIG. 2, device 1 further includes an arithmetic unit 4 that takes care of creating the product (i.e., the outer product i ) of the outputs of the preprocessing block having each bootstrap key element BK [Number] ). The device further includes a small bootstrap cache memory 3 arranged to store at least two bootstrap key elements. The bootstrap key element BK for use in the outer product at iteration i iIt is retrieved from the on-chip bootstrap cache memory 3, which can be SRAM memory in some embodiments. In an implementation on a field programmable gate array (FPGA), the SRAM memory can be constructed using block RAM (BRAM) or ultra RAM (URAM). The bootstrap cache memory 3 has a size sufficient to load at least one further bootstrap key element while another bootstrap key element is being used in the multiplication operation. In such embodiments where the cache memory holds two bootstrap key elements, the load can be performed in a ping-pong manner. In some embodiments, the cache memory size is sufficient to store two or more further bootstrap key elements to be loaded. This can be advantageous, for example, to buffer the delay and throughput differences between on-chip cache memory and off-chip memory. In these embodiments, the loading of the bootstrap key elements can be performed in a first-in first-out manner.
[0059] The accelerator further comprises a post-processing block 6, where the result of the multiplication is received and post-processed to prepare it for the next iteration. The latter result is then applied again to the pre-processing block to start iteration i+1.
[0060] The accelerator operates as follows. During initialization, the inputs to the pre-processing block are set to carefully selected values (which may depend on external inputs in some cases). In some embodiments, this value is equal to F.X from equation (2). In other embodiments, the value can just be F, but the multiplication with X is only performed at the end. To perform the calculation, in each iteration i, different a and BK are used. While iteration i is being calculated, BK is already loaded in preparation for the next iteration. After a certain number (e.g., not necessarily, but n times) of iterations, the output of the post-processing block is returned, possibly after calculating the final output processing. -b is equal to. In other embodiments, the value can just be F, but the multiplication with X -b is only performed at the end. To perform the calculation, in each iteration i, different a i and BK i are used. While iteration i is being calculated, BK i+1 is already loaded in preparation for the next iteration. After a certain number (e.g., not necessarily, but n times) of iterations, the output of the post-processing block is returned, possibly after calculating the final output processing.
[0061] As described above, the process is implemented in a batch manner. The entire iteration (i.e., the preprocessing for collecting input coefficients, the multiplication in the arithmetic unit, and the postprocessing for accumulating the intermediate results of the multiplication operation) is pipelined, and the calculation is batch-processed for q ciphertext coefficients in one iteration.
Number
[0062] Figure 3 shows a preferred embodiment where the encryption application is a torus fully homomorphic encryption (TFHE) scheme. In this case, the accelerator is adapted to perform a controlled MUX (CMUX) operation based on the above formula (5), and a i takes on various values within the batch.
Number
Number
Number
[0063] The accelerator is constructed as a streaming processor with a wide data path and high throughput in a preferred embodiment. In some embodiments, the accelerator has a pipeline structure with a plurality of high - throughput calculation stages directly cascaded, along with a simplified control logic and routing network as compared to prior - art solutions. In this architecture, data flows directly from one calculation stage to the next without being read from / written to the central memory. Such a design enables very efficient utilization of arithmetic units in various calculation stages during the bootstrapping process. In an alternative embodiment, the accelerator is constructed according to a more CPU - like approach where the arithmetic unit fetches data and writes it to memory.
[0064] The accelerator computes a fixed sequence of preprocessing, arithmetic (i.e., outer product), and postprocessing. Instead of dividing the accelerator into sub-units for each of the operations sequenced to be executed by the register file, the accelerator according to embodiments of the present invention constructs a fixed sequence with directly cascaded computational stages. The stages probably balance throughput in a simple way. Each stage operates at the same throughput and processes several polynomial coefficients per clock cycle, called the streaming width. The stages are interconnected by a simple fixed pipeline with a static delay, avoiding complex control logic and simplifying the routing path.
[0065] The accelerator is constructed to achieve the highest possible bootstrap throughput. A preferred optimization metric is throughput / area (TP / A). In designs where the throughput is lower but the TP / A is higher, the highest throughput can be achieved by instantiating multiple copies. Generally, the throughput / area (TP / A) of the computational stages increases with the streaming width. This is, in one embodiment, the motivation to instantiate only a single accelerator with a high streaming width, as opposed to many accelerators with a smaller streaming width.
[0066] In some embodiments of the accelerator, the arithmetic unit comprises (k + 1)l forward FFT operations with (k, l encrypted parameters), but only (k + 1) inverse FFT operations. To obtain a maximum utilization design, it is necessary to balance the FFT throughput and IFFT throughput of the processing elements. In some embodiments, this balance is achieved by instantiating an FFT unit with a higher throughput than the IFFT unit. Two possible options for achieving higher throughput are to instantiate l times more FFT building blocks than IFFT building blocks (which can be considered a dot product unrolled architecture), or to instantiate an FFT block that is l times larger than the streaming width (which can be considered an FFT unrolled architecture). Figure 4 shows these two options for calculating the outer product in the arithmetic unit.
[0067] The dot product unroll architecture (left) represents a more obvious choice for parallel processing. In the FFT unroll architecture (right), throughput is balanced by instantiating an FFT that is l times larger than the streaming width of the IFFT. These two options utilize different types of "loop unrolling" within the outer product. In the former, first the dot product is loop unrolled before the FFT is unrolled, while in the latter, the FFT is loop unrolled to the maximum extent. The FFT unroll architecture is more complex than the dot product unroll architecture. First, since the polynomial coefficients that must be added are temporally separated here over different clock cycles, the multiply-accumulate operations must be replaced by MACs. Second, the inverse FFT can only start processing after the complete MAC has been completed, requires a parallel-in serial-out (PISO) block that double-buffers the MAC output and matches the throughput. Third, and most importantly, the FFT block can be difficult to unroll and implement for arbitrary throughput, and supporting two FFT blocks with different throughputs requires additional engineering effort. On the other hand, the FFT unroll architecture thus features fewer FFT units that can utilize a higher streaming width. This favors the common (and often overlooked) trend of pipelined FFTs, which typically feature a throughput / area ratio that increases significantly as the streaming width increases. At the extreme, a fully parallel FFT is a circuit with only constant multiplication and fixed routing paths that features throughput increases of up to 300% per DSP or per LUT on an FPGA.
[0068] The arithmetic unit calculates a number of polynomial multiplications and accumulations. To achieve high efficiency, these polynomial multiplications can be implemented using the FFT or NTT algorithm. In some embodiments of the accelerator, the forward FFT or NTT has a higher throughput than the inverse FFT or NTT. The algorithm can be implemented in an iterative or pipelined manner. In a preferred embodiment, a pipelined FFT including log(N) stages connected in series is instantiated. The main advantage of these architectures is to process a continuous data flow and is well suited for a fully streamed outer product design.
[0069] The vector elements and / or the bootstrap key can be represented in various ways, for example, in floating-point notation (single-precision or double-precision), block floating-point or fixed-point notation when using the FFT, or in integer notation when using the NTT. In an advantageous embodiment, the vector elements and / or the bootstrap key are represented in fixed-point notation and the accelerator uses the FFT. The fixed-point representation is determined by the number of bits in the representation and the scaling factor. In fixed-point representation, the fractional part has a fixed number of bits.
[0070] Although the present invention has been illustrated and described in detail in the drawings and the foregoing description, such illustration and description should be regarded as illustrative or exemplary and not restrictive. The foregoing description has described some embodiments of the present invention in detail. However, it will be understood that the present invention may be practiced in many ways, however detailed the foregoing may be set forth in text. The present invention is not limited to the disclosed embodiments.
[0071] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may perform the functions of several items recited in the claims. The fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid state medium supplied together with or as part of other hardware, but it may also be distributed in other forms, such as via the Internet or other wired or wireless communication systems. Any reference signs in the claims should not be construed as limiting the scope thereof.
Claims
1. A method for performing a bootstrap operation in an encryption application, comprising: - receiving, in an accelerator, one or more input ciphertexts used in the encryption application to be bootstrapped, and iteratively processing one or more accumulator variables with a function of a portion of the input ciphertexts, whereby each accumulator variable is linked to one input ciphertext and each iteration step is performed in sequence for each of the one or more accumulator variables; - multiplying, within each iteration, the processed accumulator variable by a bootstrap key element belonging to a bootstrap key comprising a plurality of bootstrap key elements, wherein the bootstrap key elements are obtained from a bootstrap cache memory within the accelerator; - loading, from an external memory into the bootstrap cache memory, a next bootstrap key element of the plurality of bootstrap key elements to be used in a next iteration of the bootstrap operation, while performing the multiplication in sequence for each of the one or more accumulator variables. A method for performing a bootstrap operation.
2. The method for performing a bootstrap operation according to claim 1, wherein the bootstrap cache memory is a SRAM memory.
3. The method for performing a bootstrap operation according to claim 1, wherein the one or more input ciphertexts are supplied from a further memory within the accelerator.
4. The method for performing a bootstrap operation according to claim 1, wherein the one or more input ciphertexts are supplied from an external memory.
5. The method for performing a bootstrap operation according to claim 1, wherein the vector elements and / or the bootstrap key elements and / or the accumulator variables are represented in fixed-point.
6. The method for performing a bootstrap operation according to claim 1, wherein the encryption application is a torus fully homomorphic encryption scheme.
7. A program executable on a programmable device, including instructions for performing the method according to claim 1 when executed.
8. An accelerator for performing a bootstrap operation in an encryption application, wherein the accelerator is - configured to receive one or more input ciphertexts to be bootstrapped and iteratively process one or more accumulator variables with a function of a part of the input ciphertexts, whereby each accumulator variable is linked to one input ciphertext and each iteration step is performed in sequence for each of the one or more accumulator variables, a preprocessing block (5); - an arithmetic unit (4) configured to multiply, within each iteration, the processed accumulator variable by a bootstrap key element belonging to a bootstrap key including a plurality of bootstrap key elements; - The accelerator further includes a bootstrap cache memory for storing the bootstrap key elements to be multiplied, and is adapted to load the next bootstrap key element of the plurality of bootstrap key elements used in the next iteration of the bootstrap operation into the bootstrap cache memory while sequentially performing a multiplication operation for each of the one or more accumulator variables. Accelerator.
9. The accelerator according to claim 8, wherein the arithmetic unit comprises a plurality of cascaded calculation stages, and each calculation stage operates on a different part of the one or more accumulator variables.
10. The accelerator according to claim 8, wherein the plurality of cascaded calculation stages form a pipeline.
11. The accelerator according to claim 8, further comprising a memory for storing the one or more input ciphertexts.
12. The accelerator according to claim 8, wherein a fixed-point representation is used for the one or more input ciphertexts and / or the bootstrap key elements and / or the accumulator variables.
13. The accelerator according to claim 8, wherein the bootstrap cache memory is arranged to store two or more next bootstrap key elements to be used.
14. The accelerator according to claim 8, implemented in a field programmable gate array or an application specific integrated circuit.
15. The accelerator according to claim 8, wherein the arithmetic unit has forward and reverse FFT or NTT calculation stages, and the forward FFT or NTT has a higher throughput than the reverse FFT or NTT.
16. A computing system comprising the accelerator according to claim 8 and a memory external to the accelerator, wherein the memory is arranged to store one or more bootstrap keys for use in a bootstrap operation.
17. Use of the accelerator according to claim 14 in a cloud computing service.