Method for designing device for encrypting application

By dividing polynomial multiplication into multiple data processing stages and optimizing data representation, the problem of high multiplication operation cost in FHE is solved and the calculation efficiency is improved.

CN119968610APending Publication Date: 2025-05-09KATHOLIEKE UNIV LEUVEN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070569.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-26
Filing Date
2023-08-23
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Fully homomorphic encryption (FHE) faces the problem of high computational overhead in practical applications, especially the high cost of multiplication operations, which affects the overall computing speed.

Method used

The optimized parameter values ​​are determined to implement polynomial multiplication by dividing the polynomial multiplication into at least two data processing stages of the interconnect and defining constraints for each stage of the data representation and output signal.

Benefits of technology

This method allows each data processing stage to have its optimized parameter set, improves the execution efficiency of polynomial multiplication and reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968610A_ABST
    Figure CN119968610A_ABST
Patent Text Reader

Abstract

The invention relates to a method of designing a device for performing multiplication of a polynomial in a cryptographic application. The method comprises:-dividing execution of a multiplication of a polynomial onto at least two data processing stages of the device, where at least one data processing stage is configured to receive an input operand to perform the multiplication and at least one data processing stage is configured to provide an output signal of the multiplication, defining, for each data processing stage, one or more parameters related to a representation of data to be processed by the data processing stage, defining, for the output signal, one or more constraints, determining, for each data processing stage, a value for the one or more parameters individually, wherein the one or more constraints are taken into account,-applying the values for the one or more parameters determined in each of the data processing phases to the device that performs a multiplication of a polynomial in a cryptographic application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to the field of encryption. More particularly, the present invention relates to a method of designing a device for use in encryption applications. Background Art

[0002] In the context of cloud computing, for example, users face certain risks when uploading raw data to untrusted cloud servers. Therefore, it is necessary to provide sufficient security to protect the user's data. A promising new technology has emerged in the field of data security, namely fully homomorphic encryption (FHE), which allows people to perform homomorphic computations on the encrypted data (ciphertext) without knowing additional information about the data. In other words, there is no need to decrypt the data first. Over the years, methods for performing FHE have been improved to the point where they can be used in practice.

[0003] FHE algorithms are usually executed on cloud computing servers. However, the computation is slow. Typically, the bottleneck for the execution of FHE schemes is multiplication operations. The ciphertext data on which computations are performed in FHE schemes are large polynomials (of length N) from some scheme-dependent polynomial ring. Typical operations on these polynomials include addition and multiplication. While addition is linear in the length of the polynomial (O(N) operations), multiplication has a quadratic cost (O(N) operations) when using the general direct technique (also known as textbook multiplication). 2 ) operation).

[0004] One of the main challenges facing FHE in practical applications is its computational overhead. Since one of the most expensive operations in FHE schemes is multiplication, making multiplication faster can be very helpful in reducing the computational overhead. This can be achieved by exploiting specific properties of polynomials. Various well-studied algorithms can be used to speed up this multiplication, including number theoretic transformation (NTT), Toom-Cook multiplication, or Karatsuba multiplication. Multiplication using NTT is generally the most efficient of these algorithms, but it also imposes the most stringent conditions on the polynomial ring used. Therefore, it cannot be used in every FHE scheme.

[0005] FHE schemes typically use NTT for fast polynomial multiplication when the underlying ring structure allows NTT. Two notable exceptions where NTT is not applied are the FHEW scheme disclosed in the paper "FHEW: Bootstrapping homomorphic encryption in less than a second" (L. Ducas et al., Eurocrypt, pp. 617-640, 2015) and the TFHE scheme described in "TFHE: Fast Fully Homomorphic Encryption Over the Torus" (I. Chillotti et al., J. Cryptol. 33, 34-91, 2020). The ring structures of both of them prohibit the use of NTT. Instead, these schemes can use Toom-Cook multiplication, Karatsuba multiplication, or fast Fourier transform (FFT) for fast polynomial multiplication, where the typical FFT is the fastest option. The FFT transform is similar to the NTT transform, but the polynomial ring condition on the FFT transform is less stringent. Both FHEW and TFHE support the use of homomorphic Boolean algebra, for example, in addition to homomorphic addition and multiplication, there are NAND logic gates, XOR logic gates, and XNOR logic gates.

[0006] For security purposes, FHE ciphertexts include noise in the encryption. Furthermore, each FHE operation increases the noise present in the ciphertext. The additional FFT quantization noise can be tolerated to some extent. The quantization noise may stem from the fact that the FFT works with real numbers and is approximated with floating or fixed point representations. FHE can tolerate this noise as long as a certain threshold noise level is not exceeded. Therefore, the FHE scheme must periodically call bootstrapping operations to reduce the amount of noise in the ciphertext so that the noise amount remains below the threshold noise level. However, too much noise can cause bootstrapping to fail, and therefore, the use of FFTs needs to be very careful.

[0007] Schemes such as TFHE and FHEW perform bootstrap after each cryptographic gate. Compared to previous generations of FHE schemes, they feature faster bootstrapping algorithms and additionally allow programming with arbitrary functions applied to the ciphertext. The cryptographic gate computation (encrypted NAND, XOR, XNOR, etc.) is dominated by the expensive (programmable) bootstrapping ((P)BS) computation. Therefore, hardware acceleration of the bootstrapping application is needed.

[0008] PBS involves continuously multiplying ciphertext data in the so-called "outer product". This outer product defines the multiplication between the General-LWE ciphertext and the General-GSW ciphertext. General Learning with Errors (LWE) and General-GSW (Gentry, Sahai, and Waters) are both fully homomorphic encryption schemes. Conceptually, a General-LWE ciphertext is a vector of k+1 polynomials; a General-GSW ciphertext is a matrix of (k+1)|×(k+1) polynomials. In TFHE, these polynomials are defined on the real torus, i.e., their coefficients lie in the interval (0, 1]. In FHEW, these polynomials have integer coefficients.

[0009] The first step of the “outer product” is to decompose each polynomial in the General-LWE ciphertext into I polynomials. Then, the outer product in the decomposed format can be expressed as a vector-matrix multiplication of polynomials. The output is also a General-LWE ciphertext.

[0010]

[0011] In this vector-matrix multiplication, the dimension of the vector (factored GLWE ciphertext) is 1×(k+1)|, the dimension of the matrix (General-GSW ciphertext) is (k+1)|×(k+1), and the dimension of the output is 1×(k+1). The variables k and I are encrypted FHE parameters.

[0012] As already mentioned, an important feature of FHE applications is that they have an inherent cryptographic mathematical noise present in the ciphertext data. When the ciphertext data is decrypted, the noise is rounded off and the correct result is retrieved. Similarly, the outer product c(X) = (c0(X(,c1(X),...,c k The GLWE output ciphertext of c(X) contains cryptographic mathematical noise. Polynomial multiplications can be efficiently computed using Fast Fourier Transforms (FFTs). In this case, due to this inherent mathematical noise, the additional FFT quantization noise in c(X) can be tolerated even in an encrypted setting. The amount of tolerable noise depends on the encryption parameter set. For example, in the above-mentioned paper by I. Chillotti, a noise formula was derived for the cryptographic noise present in c(X).

[0013] Real numbers can be represented with finite precision in a variety of ways. On the CPU, the typical approach is to use single-precision floating point numbers or double-precision floating point numbers. The precision is limited by the size of the mantissa, and the dynamic range is limited by the size of the exponent. Due to the integration of the floating point unit (FPU) in the CPU, this approach is effective, and therefore it is a typical selected representation for software designers. The implementation of the TFHE and FHEW schemes mentioned above has been limited to double-precision floating point FFTs because it was found that single-precision FFTs introduce too much noise. It has been found that double-precision floating point FFTs keep the amount of noise introduced at a sufficiently small level. The fixed-point representation is determined by the number of bits in the representation and the scaling factor. In a fixed-point representation, the mantissa has a fixed number of bits.

[0014] In hardware, double precision arithmetic is expensive to implement and is therefore preferably avoided if not absolutely required. A much cheaper alternative is fixed point arithmetic. In state-of-the-art solutions, a single representation is chosen and used throughout the architecture. Furthermore, the specific properties of the input data in the encryption setting and the specific requirements for noise at the output have not been exploited so far.

[0015] In the paper “MATCHA: A Fast and Energy-Efficient Accelerator for FullyHomomorphic Encryption over the Torus” (L. Jiang et al., 59th Annual Design Automation Conference, July 2022, pre-published on the Internet on February 17, 2022), a hardware accelerator for processing TFHE gates is presented that outperforms accelerators that frequently call expensive double-precision floating-point FFT and IFFT cores in terms of efficiency. To fully exploit the error tolerance of TFHE, polynomial multiplications are accelerated by using approximate multiplication-free integer FFTs and IFFTs that only require additions and binary shifts. Although the approximate FFTs and IFFTs introduce errors in each ciphertext, the ciphertext can still be decrypted correctly because the errors can be rounded off along with the noise during decryption. The integer representation can be viewed as a scaled version of the fixed-point representation to remove the decimal point.

[0016] Furthermore, in the "MATCHA" approach, a single representation is chosen for both FFT and IFFT. Verifying that the noise remains sufficiently small is done by performing multiple tests on decryption failures. However, in the sequence of steps that constitute the polynomial multiplication, using a single representation ignores the fact that each operation contributes differently to the output noise, depending on the data processed. Furthermore, matching the FFT noise to the mathematical outer product noise allows for finer-grained noise control and optimization than just testing on decryption failures.

[0017] Therefore, there is room for improvement in methods of designing devices for use in encryption applications. Summary of the invention

[0018] It is an object of embodiments of the present invention to provide a method for designing an apparatus for performing multiplication of polynomials in cryptographic applications, wherein the peculiarities of data in cryptographic applications are taken into account and wherein a degree of flexibility is built in.

[0019] The above objects are achieved by the solution according to the invention.

[0020] In a first aspect, the present invention relates to a method of designing an apparatus for performing multiplication of polynomials in cryptographic applications. The method first divides the execution of the multiplication of the polynomials into at least two interconnected data processing stages, wherein at least one data processing stage is configured to receive input operands to perform the multiplication, and at least one data processing stage is configured to provide an output signal of the multiplication. The method further comprises:

[0021] - defining for each data processing stage one or more parameters related to the representation of the data to be processed by this data processing stage,

[0022] - defining one or more constraints on an output signal of said multiplication,

[0023] - determining values ​​for said one or more parameters individually for each data processing stage, taking into account said one or more constraints,

[0024] - applying the values ​​for said one or more parameters determined in each of said data processing stages to said device for performing multiplication of polynomials in a cryptographic application.

[0025] The proposed solution indeed allows to design the various data processing stages of a device for performing multiplication of polynomials so that each stage has its own, preferably optimized, set of parameters. The set of parameters thus determined allows to represent the variables used when performing the multiplication in an optimal way given a representation type (for example fixed point or floating point) so as to satisfy one or more constraints on the output signal.

[0026] Preferably, the input operand is a vector of polynomials. The second input operand is then a matrix of polynomials, which can be pre-computed into the FFT domain.

[0027] In a preferred embodiment, the one or more parameters include one or more of {bit width, dynamic range, size of integer part, size of fractional part, position of decimal point}.

[0028] Advantageously, the device is reconfigurable.

[0029] In one embodiment, the method of designing a device includes the step of decomposing the polynomial multiplication into a plurality of subtasks, wherein each data processing stage performs at least one subtask.

[0030] In an embodiment of the method, the division into at least two data processing stages depends on the hardware implementation of the device. In an advantageous embodiment, the device has an architecture comprising a plurality of successive stages forming a pipeline. The device can then be applied in a field programmable gate array.

[0031] In a preferred embodiment, a maximum noise level is considered as one of the constraints.

[0032] Advantageously, the multiplication of the polynomials is part of a fully homomorphic encryption scheme.

[0033] In some embodiments, the step of determining parameter values ​​individually for each data processing stage is performed by means of simulation.

[0034] In an embodiment of the method, the step of determining the value individually for each data processing stage is performed by determining one or more attributes of the data to be processed, wherein the determination of the one or more attributes of the data to be processed is based on a model of a noise source affecting the one or more constraints on the one or more attributes and / or based on at least one attribute of an input of a subtask of the data processing stage.

[0035] In one embodiment, one source of noise may result from removing bits on the least significant bit side of the input to a subtask of the data processing stage.One source of noise may result from discarding bits on the most significant bit side of the input to a subtask of the data processing stage.

[0036] In some implementations, the multiplication of the polynomials is performed using negative circular convolution.

[0037] In a preferred embodiment, the representation of the data is a fixed point representation.

[0038] In order to summarize the present invention and its advantages over the prior art, certain objects and advantages of the present invention are described above. Of course, it should be understood that not all such objects or advantages can be achieved according to any particular embodiment of the present invention. Thus, for example, those skilled in the art will recognize that the present invention can be implemented or performed in a manner that achieves or optimizes one advantage or a group of advantages as taught herein without having to achieve other objects or advantages taught or suggested herein.

[0039] The above and other aspects of the invention will be apparent from and elucidated with reference to one or more embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention will now be further described, by way of example, with reference to the accompanying drawings, in which like reference numerals refer to like elements throughout the various figures.

[0041] Figure 1 A block diagram showing an accelerator structure on which the method of the present invention can be applied.

[0042] Figure 2 Shows Figure 1 An expanded version of the block diagram of .

[0043] Figure 3 The butterfly structure commonly used in FFT implementation is shown.

[0044] Figure 4 The decomposition into subtasks and the addition of building blocks for intermediate variables in the third example are shown.

[0045] Figure 5 A three step procedure for performing multiplication is shown.

[0046] Figure 6 One possible alternative architecture for the device is shown. DETAILED DESCRIPTION

[0047] The present invention will be described with respect to particular embodiments and with reference to certain drawings but the invention is not limited thereto but only by the claims.

[0048] In addition, the terms "first", "second", etc. in the specification and claims are used to distinguish similar elements, and are not necessarily used to describe a sequence in time, space, order, or any other manner. It should be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention described herein are capable of operating in sequences other than those described or illustrated herein.

[0049] It will be noted that the term "comprising" used in the claims should not be interpreted as being limited to the means listed thereafter; it does not exclude other elements or steps. It should therefore be interpreted as specifying the presence of the features, integers, steps, or components mentioned, but does not exclude the presence or addition of one or more other features, integers, steps or components or groups thereof. Thus, the scope of the expression "a device comprising means A and means B" should not be limited to devices comprising only components A and B. This means that, for the purposes of the present invention, the relevant components of the device are only A and B.

[0050] Throughout the specification, references to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least one embodiment of the invention. Thus, throughout the specification, the phrase "in one embodiment" or "in an embodiment" does not necessarily refer to the same embodiment at various locations, but may refer to the same embodiment. Furthermore, as will be apparent to one of ordinary skill in the art from this disclosure, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0051] Similarly, it should be understood that in the description of exemplary embodiments of the present invention, in order to streamline the present disclosure and help understand one or more different inventive aspects, various features of the present invention are sometimes combined together in a single embodiment, figure, or its description. However, this method of disclosure should not be interpreted as reflecting the intention that the claimed invention requires more features than those explicitly stated in each claim. On the contrary, as reflected in the attached claims, the inventive aspects are less than all the features of the aforementioned single embodiment disclosed. Therefore, the claims attached to the detailed description are hereby expressly incorporated into this detailed description, wherein each claim itself is a separate embodiment of the present invention.

[0052] In addition, although some embodiments described herein include other features not included in some other embodiments, the combination of features of different embodiments is intended to be within the scope of the present invention and form different embodiments, as understood by those skilled in the art. For example, in the appended claims, any claimed embodiment can be used in any combination.

[0053] It should be noted that the use of a particular term in describing certain features or aspects of the present invention should not be understood as implying that the term is redefined herein to include only any specific characteristics of the feature or aspect of the present invention associated with the term.

[0054] In the description provided herein, many specific details are set forth. However, it is to be understood that embodiments of the present invention can be practiced without these specific details. In other cases, well-known methods, structures, and techniques are not shown in detail in order not to obscure the understanding of this description.

[0055] The present invention discloses a novel method for designing a device for performing polynomial multiplication in cryptographic applications.

[0056] In order to achieve an efficient implementation of polynomial multiplication to perform FFT operations, the above-mentioned feature can be exploited, namely that in FHE schemes, the calculation introduces a certain amount of noise into the ciphertext that needs to be secure. Therefore, these schemes can also tolerate approximate calculations that include quantization noise. It will also be noted that encryption works with uniformly distributed random values, which limits the dynamic range of the coefficients on which the calculations are based. This makes it easier to predict the necessary bit representation range (i.e., word length) required to represent the coefficients. The noise tolerated by FFT in encryption applications depends on the encryption parameter set, not on the application constraints. The encryption parameter set can be selected to tolerate more or less noise. The inventors have realized that these insights provide the opportunity to perform multiplication in such a way that the intermediate variables applied when performing the operation have an optimized representation.

[0057] More specifically, the present invention presents a method for designing an apparatus for performing multiplication of polynomials in cryptographic applications. In the method, the multiplication is first divided into at least two data processing stages to be implemented using optimized data representations. Next, one or more parameters are established to determine how the variables used in the data processing stages are represented when performing multiplication in the cryptographic application. Although the mathematical description of the algorithm performing such operations typically assumes infinite precision of the variables, the hardware implementation (just like the software implementation) needs to determine the specific representation type of the variables, such as fixed point, floating point, block floating point, and the precision of the selected representation. This choice will have an impact on the cost and accuracy of the implementation.

[0058] Once the choice of representation type for each variable has been made, the exact parameters of the representation must still be decided. For example, for a fixed-point representation, the least significant bit (LSB) and the most significant bit (MSB) represented need to be decided. A floating-point representation is parameterized by the minimum and maximum possible values ​​of both the mantissa and the exponent (hence, the bit size (word length)). For example, the MSB value is characterized by whether overflow is likely to occur.

[0059] By way of example, the parameterization of MSB and LSB is as follows. Assume that the value a has 8 integer bits and 16 fractional bits, and the value b has 4 integer bits and 8 fractional bits. The full-precision result c=a×b then has 8+4=12 integer bits and 16+8=24 fractional bits. In this case, it can be said that the MSB of c is at position 12 and the LSB is at position -24. In an implementation, a parameterization is performed for the variable c: MSBc and LSBc. By truncating the bits at the MSB, i.e. performing a rescaling, (e.g., selecting MSBc=11), a certain overflow probability P is allowed. overflow This value depends on the distribution of the values ​​a and b, and in practice all MSBs are set so that, for example, P overflow <2 -64 By truncating the bits at the LSB side, some noise due to quantization is introduced, which increases the output noise. Note that any form of rounding can be considered instead of truncation.

[0060] Importantly, the representation of the variables has an impact on the cost of the implementation, as a larger range of possible values ​​makes the computation more expensive, but also results in a more accurate implementation. For a typical implementation, there will be design constraints on the accuracy of the output variables. These constraints are known in advance and can be, for example, the maximum noise variance introduced or the maximum probability of having a variable overflow (e.g., 2 -20 or 2 -40 ) in the form of, but not limited to, the above. The goal is to determine parameters that allow for an efficient implementation while satisfying the design constraints for the output signal.

[0061] A general overview of the proposed method is now provided, which is applied to an algorithm to be implemented in a device to perform polynomial multiplication on an input operand in a cryptographic application. The polynomial multiplication operation is decomposed into a plurality of subtasks. The device according to the present invention comprises at least two interconnected data processing stages. Each data processing stage performs at least one subtask. At least one data processing stage is configured to receive an input operand to perform a multiplication. At least one data processing stage is configured to provide an output signal representing the result of the multiplication of the polynomial. Any subtask that does not generate an output of the multiplication operation will produce a variable, which is referred to as an intermediate variable in this specification. Any intermediate variable is then used as an input to one or more subsequent subtasks. The input operand may contain one or two (all) variables of the variables to be multiplied. For example, when performing a multiplication of two variables, the two variables may form an input operand. In an advantageous embodiment, the input operand is a vector of polynomials. One or more constraints are imposed on certain properties of the resulting output of the operation (e.g., the maximum value of the variance of the introduced noise). The goal is to determine valid parameters for the representation of the output variable and the intermediate variable in the implementation.

[0062] At a high level, the design method of the present invention includes the following main steps. For each data processing stage of the device, one or more parameters related to the way of representing the data to be processed in this stage (i.e., intermediate variables) are defined. One or more constraints that the output of the polynomial multiplication operation must satisfy are also defined. In the next step, for each data processing stage, specific values ​​for each parameter in the representation are determined separately so that all constraints are satisfied. Preferably, the values ​​for these parameters are determined in an optimal manner, that is, the values ​​for these parameters are determined according to a suitable optimization function. These parameter values ​​are then used in each data processing stage of the device to implement the multiplication of the polynomials in the device.

[0063] The preferred method for determining the parameter values ​​for each data processing stage is as follows. One or more attributes of the variable output by one of the subtasks of the data processing stage under consideration are determined. These attributes of the variable can be a model of a noise source that affects one or more constraints on the attributes of the signal output by this subtask and / or at least one attribute based on the input of the subtask. The model is constructed according to the parameters of the representation to be determined. This operation is repeated for all intermediate variables of all data processing stages and for the output variables. For each data processing stage, determine the specific value for each parameter in the representation so that the one or more attributes meet all constraints. Then the parameter values ​​as determined in the corresponding data processing stage are used in the device for performing the multiplication of polynomials.

[0064] Modeling the noise sources can be performed as follows. The algorithm for performing the multiplication is decomposed into various subtasks, and intermediate variables are identified. For each subtask, the input-output behavior of the attributes relative to the constraints is determined (e.g., the introduced noise, the scaling factor of the input to the output, ...). For each of these variables, the type of representation is selected (i.e., fixed point, floating point, ...). The parameters of this representation (e.g., such as the highest representable value, the lowest representable value, ...) are still undetermined and are initially retained as symbolic variables, i.e., variables that do not yet have a specific value.

[0065] To analyze the execution of the algorithm, one makes use of an additional building block that is introduced at the location of each intermediate variable in the scheme of multiplications. This building block represents the effect of the finite precision of the representation of the variables, but has no effect on the algorithm itself. To achieve this, the additional building block has an input-output behavior that relates the symbolic parameters to the imposed constraints.

[0066] The algorithm is performed step by step from input to output, and a model is built for each constraint. To this end, one or more relevant properties (e.g., noise variance, input variance, maximum possible value, ...) are determined at the input of the subtask or at the input of the input operand, and then these properties are propagated to the output. For each subtask, the input properties are converted into corresponding output properties using input-output behavior. As a result, for each constraint, a model of the properties of the variable is obtained, on which the constraints are imposed according to the symbolic parameters.

[0067] In the next step of the method, the value of each symbolic parameter in the symbolic parameters is determined so that the constraints are satisfied. Preferably, these values ​​are selected so that the implementation cost is as low as possible. One way to achieve this is to select a cost function for each symbolic parameter that models the implementation cost for a given value of the parameter. A solution can then be found by solving an optimization problem, wherein a value is determined that reduces the total cost function while complying with all constraints and preferably minimizes the total cost function.

[0068] Finally, the selected values ​​for the parameters as determined are applied to instantiate the design of the various data processing stages of the device for the encryption application.

[0069] In other embodiments, the adjustment of the data representation can be performed with the aid of simulation. In these embodiments, it is assumed that the contribution of each computational stage to the output noise is independent. First, each computational stage is implemented with very high precision, resulting in negligible output noise. Next, a single stage is adjusted (e.g., the bit width of the fixed-point representation is reduced), and the resulting output noise is measured and compared to the high-precision case. At a certain threshold bit width, the output noise due to the adjusted stage is close to the tolerable output noise (e.g., within X bits). This threshold bit width is then considered the "optimal" bit width for that particular stage. This operation is repeated until a threshold bit width is determined for each particular stage. Finally, each stage is implemented with its threshold bit width, and the total output noise of all stages is collectively measured. If the output noise exceeds the tolerable output noise, the entire process is repeated for a larger threshold bit width (i.e., the tolerance of X bits is increased).

[0070] In another embodiment, the adjustment of the data representation is performed in a similar manner, except that a cost function is assigned to each of the adjusted stages. Stages with higher computational costs are first adjusted to their minimum bit width. Subsequently, subsequent stages are adjusted while the costly stages are kept at the adjusted bit width. In this way, the costly stages are implemented with the absolute minimum precision, while other costly stages are implemented with higher precision.

[0071] An example is now presented in which the proposed method is demonstrated for designing a hardware architecture to accelerate the vector-matrix multiplication given in equation (1) as mentioned above. The architecture utilizes FFT blocks to accelerate polynomial multiplication.

[0072] Figure 1 A block diagram of an advantageous accelerator structure is presented in . The architecture has clear similarities to streaming DSP processor architectures, with large fully pipelined, directly cascaded computation stages, and highly simplified control logic. The accelerator is envisioned to achieve maximum throughput / area, with maximally expanded arithmetic units, eliminating control logic and hard-coded routing paths.

[0073] The acceleration of polynomial multiplication can be achieved through the convolution theorem:

[0074] c=a×b=IFFT(FFT(a)·FFT(b)) (2)

[0075] Repeat formula (1) here:

[0076]

[0077] Where c=[c0 c1...c k ] and a=[a0 a1...a(k+1)l+1 ] represent vectors of polynomials of dimensions 1×(k+1) and 1×(k+1)I, respectively, and B represents a matrix of polynomials of dimensions (k+1)l×(k+1). The FFT-based multiplication is performed according to formula (2) by using FFT to convert the input polynomial [a0 a1 ... a (k+1)l-1 ] is converted into another representation. In this domain, the point-by-point comparison with the polynomial B can be performed i,j The multiplication operations (pre-calculated in the FFT domain) are performed (N operations are performed in pairs in formula (2)). The accumulation step of the vector-matrix product is preferably performed in the FFT domain. Afterwards, an inverse FFT (IFFT) is required to convert the result back to the original representation. The FFT and IFFT conversion operations are typically the most expensive operations in FFT-based multiplication, requiring O(N.log(N)) operations, where N is the number of coefficients in the polynomial. The number of coefficients determines the depth and width of the FFT.

[0078] FFT-based multiplication operates on complex numbers, where both the real and imaginary parts are real numbers, while other multiplication algorithms use integers. When finite precision is used to represent real numbers, the calculation of the multiplication is not always exact and may be noisy, i.e., a small error δ may be introduced:

[0079] FFT -1 (FFT(a).FFT(b))=c+δ

[0080] As already mentioned, in addition to the (mathematical) noise already present in the formula for security reasons, a certain degree of noise introduced by the FFT can be tolerated in FHE. This means that the size of the noise δ needs to be considered very carefully. FHE implementations impose strict limits on the introduced noise. If too much noise is introduced by the FFT in the polynomial multiplication, the calculation will fail and return incorrect results.

[0081] Figure 2 Shows Figure 1 1 is an expanded version of the exemplary accelerator hardware architecture with pipeline structure depicted in FIG. This architecture is advantageously applied in a field programmable gate array (FPGA). The matrix input B in Formula 1 i,j can be precomputed in the FFT domain. Figure 1 and Figure 2 It can also be seen that input B i,jis available in the FFT domain. The FFT component (12) computes the FFT of the vector a applied to the input of the structure. The multiply-accumulate (MAC) component (13) accumulates the components of the point-by-point operation in the FFT domain. An inverse FFT component (14) is provided to put the output vector back from the FFT domain. In formula 1, there are (k+1).I forward FFTs and (k+1) inverse FFTs, so these blocks require different throughputs. All blocks are connected as a pipeline, as already mentioned.

[0082] from Figure 1 and Figure 2 As can be seen, the architecture can be divided into multiple stages (functional blocks). There are FFT blocks, MAC blocks, and inverse FFT blocks. There may also be multiple stages within each of these blocks. The MAC has a multiplication stage and an addition stage. The multiplication itself is a complex multiplier (out = (a + bj)(c + dj)), and this operation also has multiple stages. Each stage has specific data representation parameters, such as (fixed point, floating point, number of bits, rounding mode), which will affect the noise present in the output c (c = c0, c1). The FFT block has log2 (N) stages of "butterfly" units. These butterfly units perform butterfly operations, which are an important part of the FFT transform. Figure 3 A schematic diagram of a butterfly operation is depicted. The butterfly operation with its two inputs and two outputs is well known in the implementation of FFT algorithms and recursively decomposes a discrete Fourier transform of composite size N=rm into r smaller transforms of size m, where r is the cardinality of the transform. These smaller DFTs are then combined via a butterfly of size r, which itself is a DFT of size r (performed m times on the corresponding outputs of the sub-transforms), pre-multiplied by a root of unity called a twiddle factor.

[0083] Since FHE applications have inherent cryptographic mathematical noise present in the polynomial c, and considering that additional FFT-based rounding noise can be tolerated, the Figure 2 Each stage in the architecture can use a different data representation. The available space for some extra noise allows for different bit widths, rounding modes, fixed point formats, etc. This applies to each stage in the architecture.

[0084] There are various ways to determine the data representation for each stage. Each of these approaches starts with a property of the output signal (the vector c(X) = (c0(X), c1(X)), e.g., a maximum tolerable noise level derived from a mathematical formula for the noise present in the FHE scheme under consideration. In addition to the maximum noise level, other constraints may be imposed, e.g., constraints related to overflow. In practice, the maximum tolerable noise level is always one of the constraints.

[0085] The variables in the Fast Fourier Transform are complex numbers, and therefore the variables in the butterfly operation are also complex numbers. However, typically it can be assumed that the distribution and properties of the real and imaginary parts are the same. In this case, one can only focus on the properties of the real part in the analysis.

[0086] In the example considered here, only the maximum noise variance constraint is imposed on each butterfly structure. For example, the goal is to select a value for the position of the least significant bit (LSB) of each intermediate variable. The LSB and variance of the noise and signal are considered as related properties. After dividing the algorithm into subtasks, at 、v c 、v d Additional building blocks are added to the solution, such as Figure 4 These blocks model the inaccuracies due to the limited range of the representation. For example, the noise introduced by truncating the least significant bits is modeled by assuming that the truncated LSB bits are independently uniformly distributed.

[0087] The multiplication subtask in the butterfly operation can be simplified by exploiting the knowledge of the rotation input properties. real and the imaginary part t imag An interesting property of the rotation factor t is that t real 2 +t imag 2 = 1. Given a number x whose real and imaginary parts have the same variance, multiplying x by the rotation factor t does not change the variance, i.e. var(xt) = var(x). This can be easily derived as follows:

[0088]

[0089] For multiplication, the input-output behavior can thus be described as LSB xt =LSB x +LSB t and σ 2 noise,out =σ 2 noise,x +σ 2 x σ 2 noise,t (Assuming the inputs have zero mean). The input-output behavior of the adder block for adding variables in1 and in2 is given by the following formula:

[0090] LSB out =min(LSB in1 ,LSB in2 ) and σ 2 noise,out=σ 2 noise,in1+ o 2 noise,in2

[0091] The model can then be computed starting with the rotated input, with reduced accuracy due to the finite representation (except for rotations of 1 and –1 and i and –i). Symbolic parameters are shown in bold to distinguish them from other parameters or properties.

[0092] LSB=LSB t

[0093] σ 2 =1

[0094] σ 2 noise =2 2LSBt / 12

[0095] For the intermediate variable v at After the additional building blocks, we get the real part of the product:

[0096] LSB=LLSB at

[0097] σ 2 =σ 2 a

[0098] σ 2 noise =σ 2 nois,a +2 2LSBt / 12*σ 2 a +ramp(2 2LSBt -2 2(LSBa) / 12

[0099] For the imaginary part, the expression results are similar. The ramp() function returns 0 for negative input values ​​and returns the input value for positive input values. This function is used because additional noise only occurs when the relevant bits are discarded. If the LSB v <LSB in This is what happens.

[0100] For the intermediate variable v c After the additional building blocks, we can write:

[0101] LSB out =LSB c

[0102] σ 2 c =σ 2a +σ 2 b

[0103] σ 2 noise,out =σ 2 noise,c =σ 2 noise,a +2 2LSBt / 12*σ 2 a +σ 2 noise,b

[0104] +ramp(2 2LSBt -2 2(LSBa) / 12+ramp(2 2LSBc -2 2min(LSBat,LSBb) ) / 12

[0105] Next, we determine the values ​​of the symbolic parameters. Each parameter should now be fixed to a value that satisfies the constraint σ 2 noise,out ≤σ 2 maxnoise One way to achieve this is to construct a cost function, e.g., a function in which all parameters are equally costed according to their bit width. An optimizer that optimizes the cost function under given constraints can then be used to find efficient parameter values. In some embodiments, the cost function can of course be changed to a function that more closely represents the cost of implementation.

[0106] Note that for simple computations, such as the additions performed in the MAC phase, the approach using models as outlined above can also be applied. Again, the input-output behavior is described first. Once this is done, the model can be examined from start to finish to determine the properties associated with the constraints in terms of the symbolic parameters. The intermediate computations at the various nodes can then be written down. Next, the values ​​of the symbolic parameters are determined. The constraint function σ is derived 2 noise,out ≤σ 2 maxnoise Each parameter is fixed to a value such that this constraint is satisfied.

[0107] Polynomial multiplication using FFT as discussed in the example above can also be viewed as the computation of the inner product between the input (a vector of polynomials) and the pilot key (also a vector of polynomials). In order to process this multiplication efficiently, a three-step procedure can be used: FFT, coefficient product and accumulation, and inverse FFT. Figure 5Such a procedure is depicted in . Note that, in contrast to typical FFT-based multiplications, the second multiplication term in the figure (i.e., the boot key) does not undergo an explicit FFT operation. This is because the input is known in advance (as already noted above in equation (1) with respect to matrix b), and thus the FFT can be precomputed with very high accuracy, implying that this particular FFT does not need to be considered.

[0108] First consider the product-accumulation operation. This operation is performed in a coefficient manner, and therefore similar multiplication and addition operations discussed in the above example can be used to model. As already mentioned, FFT operation and inverse FFT operation mainly include multilayer butterfly operations. Therefore, the analysis in the above example is applied to the model that produces FFT operation and IFFT operation on these butterfly operations. In some embodiments of polynomial multiplication, different types of butterfly operations (radix 2, radix 4, ...) can be used for realization, but the analysis of these butterfly operations can be carried out in a similar manner to the analysis in the above example. By combining the building blocks discussed previously, a noise model of polynomial multiplication based on full FFT can be constructed.

[0109] One challenge is that a large number of operations need to be determined, and therefore a large number of parameters need to be determined. In order to reduce the number of parameters in the model, similar parameters can be grouped together. In this example, there is a high degree of parallelism and structure that can be used to achieve this. For example, similar parameters of variables in the same "layer" of the FFT (i.e., variables that have undergone the same number of butterfly operations) or variables after multiplication operations in product-accumulation can be grouped together. This will reduce the number of parameters from approximately O((V+1)N / 2log2(N / 2)) (where V is the vector length and N is the number of coefficients in the polynomial) to approximately O(2log2(N / 2)), because this is approximately the number of layers in the proposed algorithm.

[0110] Furthermore, it will be noted that typical cryptographic applications like FHE require performing negative circular convolutions, rather than traditional circular convolutions. In circular convolutions (with N coefficients), the coefficients that are out of bounds (at positions i>N) are cycled around to the first coefficient (at positions iN). In contrast, in negative circular convolutions, the coefficients are not only cycled back, but also negated. To achieve this, many implementations of cryptographic algorithms perform so-called twist and fold steps at the beginning and end of the algorithm, which explain the negative circular behavior. These steps can be found in Figure 1 and Figure 2They are represented by multiplications before the FFT and after the IFFT, respectively. This twisting and folding step involves additional packing of the input and multiplication with a complex number (twist factor). Packing takes two integers a, b and combines them into a complex number a+bi. This operation typically does not generate any noise. The additional multiplication operation can be modeled using the method described previously. Post-processing after the IFFT is the inverse operation. Negative circular convolution can be performed in a separate data processing stage.

[0111] As already mentioned, the type of representation for the variables must be chosen. In FFT schemes that perform polynomial multiplications as contemplated by the present invention, a fixed-point representation is advantageously chosen. The method described above can then be applied to determine the parameters of the fixed-point representation of the variables that appear when the device performs the operation, while satisfying the imposed constraints.

[0112] Given the optimal parameters obtained as described, a device, e.g., a hardware circuit, can be built using fixed-point arithmetic for these parameters. In practice, it may be advantageous to have a library of parameterized hardware circuit implementations where the fixed-point bit width is a common parameter. A circuit can be selected to match the input type, and these parameters set at "circuit synthesis time" to match the desired output noise delta.

[0113] As already mentioned above, an alternative way to determine the parameter values ​​is by performing simulations. Specifically, this may be as follows. As in the example above, the calculation and determination of the parameters of the data representation are first divided into a set of processing stages. For example, the data processing stages can be selected as higher level operations, such as forward FFT, MAC, and inverse FFT. The parameters depend on the choice of stage, for example, a = forward FFT bit width, b = MAC bit width, c = inverse FFT bit width. First, the architecture is simulated using very large bit widths for a, b, and c, resulting in an output noise variance σ 2 noise,out <<σ 2 maxnoise can be ignored. Next, assign the cost function to the parameters 'a', 'b', and 'c'. A simple cost function is applicable here: Since the forward FFT is more than the inverse FFT ( Figure 1 ), so the cost of assigning to 'a' is higher than 'b'. Also, since the inverse FFT operation is more expensive than the MAC operation, the cost of assigning to 'b' is higher than 'c'. Next, the output noise is simulated by first sweeping the most expensive parameter 'a'. Decrease 'a' until σ 2 noise,out ≈σ 2 maxnoiseThis is considered as the minimum value of 'a'. Taking into account a small margin, for example, 'a' is chosen to be 1 bit greater than the minimum value and the next parameter 'b' is processed in a similar manner. This process is repeated until values ​​are found for all parameters.

[0114] Given the optimal parameters, a device (e.g., a hardware circuit) can be simulated using the obtained parameter set. The output noise δ is measured and the output noise can be compared to a floating-point reference implementation. It is verified that the output noise meets the noise bounds determined above (e.g., standard deviation 2). An FPGA bitstream can be created for the circuit with the optimal fixed-point parameter set determined in the method proposed above. The FPGA bitstream allows for the acceleration of the FHE bootstrap procedure, which involves many (thousands) of iterations of polynomial vector multiplications.

[0115] In a first preferred example, the architecture is as follows Figure 1 The architecture is implemented in the manner depicted in . The architecture is similar to a streaming DSP processor with fully pipelined units directly cascaded. Data flows from one processing stage to the next. The purpose of the invention is to implement each data processing stage with a different optimized data representation. Advantageously, in a streaming architecture, the conversion between data representations can be hardwired in dedicated connections between streaming processing stages. In addition, the architecture has simplified control logic. It is envisioned to have maximum throughput / area with maximally expanded algorithmic units, thereby eliminating control logic and hard-coded routing paths.

[0116] Figure 6 Another example of an architecture that may be considered (based on Matcha) is depicted in . In this embodiment, the architecture resembles a more classical CPU, with a central memory file and connected arithmetic processing units. Figure 6 In the structure of FIG. 1 , the computation components (FFT, MAC, inverse FFT) can be connected to the memory file by means of a crossbar switch. As in the previous example, a suitable data representation will be determined for each stage of the structure (i.e., FFT block, various MAC blocks, and inverse FFT). In addition, the central memory can be regarded as a separate write-back stage, and its implementation and data representation can be optimized similarly.

[0117] Although the present invention has been illustrated and described in detail in the accompanying drawings and the foregoing description, such illustration and description should be regarded as illustrative or exemplary, rather than restrictive. The foregoing description details certain embodiments of the present invention. However, it will be understood that no matter how detailed the foregoing text shows, the present invention can be practiced in many ways. The present invention is not limited to the disclosed embodiments.

[0118] Those skilled in the art can understand and realize other variations of the disclosed embodiments when practicing the claimed invention according to the study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "one" or "an" does not exclude multiple. A single processor or other unit can implement the functions of multiple items recorded in the claims. The fact that specific measures are recorded in mutually different dependent claims does not itself indicate that the combination of these measures cannot be advantageously utilized. The computer program can be stored / distributed on a suitable medium, such as an optical storage medium or solid-state medium provided with or as part of other hardware, but can also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems. Any figure mark in the claims should not be interpreted as limiting the scope.

Claims

1. A method for designing a device for performing multiplication of polynomials in cryptographic applications, the method comprising: - dividing the performance of the multiplication of the polynomials over at least two data processing stages of the device, wherein at least one data processing stage is configured to receive an input operand to perform the multiplication and at least one data processing stage is configured to provide an output signal of the multiplication, - defining for each data processing stage one or more parameters related to the representation of the data to be processed in said data processing stage, - defining one or more constraints for said output signal, - determining values ​​for said one or more parameters individually for each data processing stage, taking into account said one or more constraints, - applying the values ​​for the one or more parameters determined in each of the data processing stages in the device performing the multiplication of the polynomials in the cryptographic application.

2. The method for designing a device according to claim 1, wherein: The input operand is a vector of polynomials.

3. The method for designing a device according to claim 1 or 2, wherein: The one or more parameters include one or more of {bit width, dynamic range, size of integer part, size of fractional part, position of decimal point}.

4. A method for designing a device according to any one of the preceding claims, wherein: The device is reconfigurable.

5. A method for designing a device according to any one of the preceding claims, comprising the step of decomposing the multiplication into a plurality of subtasks, wherein: Each data processing stage executes at least one subtask.

6. A method for designing a device according to any one of the preceding claims, wherein: The apparatus has an architecture comprising a plurality of sequential stages forming a pipeline.

7. The method for designing a device according to claim 6, wherein: The device is implemented in a field programmable gate array.

8. A method for designing a device according to any one of the preceding claims, wherein: The maximum noise level is considered as one of the constraints.

9. A method for designing a device according to any one of the preceding claims, wherein: The multiplication of the polynomials is part of a fully homomorphic encryption scheme.

10. A method of designing a device according to any one of the preceding claims, wherein: The step of determining the values ​​for the one or more parameters individually for each data processing stage is performed by means of simulation.

11. The method for designing a device according to any one of claims 1 to 9, wherein: The step of determining the value individually for each data processing stage is performed by determining one or more attributes of the data to be processed, wherein the determination of the one or more attributes of the data to be processed is based on a model of noise sources affecting the one or more constraints on the one or more attributes and / or based on at least one attribute of an input to a subtask of the data processing stage.

12. The method for designing a device according to claim 11, wherein: One source of noise arises from removing bits on the least significant bit side of the inputs to the subtasks of the data processing stage.

13. The method for designing a device according to claim 11 or 12, wherein: One source of noise arises from discarding bits on the most significant bit side of the input to a subtask of the data processing stage.

14. A method of designing a device according to any one of the preceding claims, wherein: The multiplication of the polynomials is performed using negative circular convolution.

15. A method of designing a device according to any one of the preceding claims, wherein: The representation of the data is a fixed point representation.