TFHE-oriented programmable bootstrap hardware implementation system

By designing a programmable bootstrap hardware implementation system for TFHE, and using the number theory transform unit to calculate polynomials in parallel, the existing TFHE solution has solved the problem of large computing resource occupancy in the blind rotation step, and achieved efficient and low-area hardware implementation, improving the success rate and performance efficiency of decryption.

CN120162083APending Publication Date: 2025-06-17TSINGHUA UNIVERSITY
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510146098.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When implementing the blind rotation step, the existing TFHE scheme consumes a large amount of computing resources, takes a long time, and it is difficult to fully utilize parallelism, resulting in high hardware unit usage, low decryption success rate, and poor performance and area efficiency.

Method used

A programmable bootstrap hardware implementation system for TFHE is designed. Through the combination of rotary registers, functional calculation modules and vector calculation modules, the number theory transformation unit is used to decompose the polynomial into subpolynomials in parallel to calculate, reducing the hardware unit usage, and reducing the SRAM usage through a streaming processing solution.

Benefits of technology

It realizes efficient blind rotation operation, reduces the use of hardware units, improves the success rate of decryption, improves performance and area efficiency, and is significantly better than existing solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162083A_ABST
    Figure CN120162083A_ABST
Patent Text Reader

Abstract

The invention discloses a TFHE-oriented programmable bootstrap hardware implementation system, and the system comprises a vector calculation module which is used for executing an analog-to-digital switching step in a TFHE-oriented programmable bootstrap algorithm on an LWE ciphertext; the rotary register is used for caching and rearranging the LWE ciphertext; the function calculation module is used for executing a blind rotation step in a TFHE-oriented programmable bootstrap algorithm on the rearranged LWE ciphertext and the analog-to-digital switching result to obtain a blind rotation result; the vector calculation module is also used for executing a sample extraction step and a key switching step in a TFHE-oriented programmable bootstrap algorithm on the blind rotation result to obtain a decrypted ciphertext; when the function calculation module executes the blind rotation step, the polynomial in the rearranged LWE ciphertext and the analog-digital switching result is decomposed into a plurality of sub-polynomials through a number theory conversion unit, and the sub-polynomials are subjected to parallel calculation. The method can reduce the use amount of hardware units, and is high in decryption success rate, high in performance and small in area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of hardware design and computer data processing, and particularly to a hardware implementation system for programmable bootstrapping for TFHE. Background Art

[0002] This section aims to provide background or context for the embodiments of the present invention described in the claims. The description herein is not admitted to be prior art merely by virtue of being included in this section.

[0003] Fully Homomorphic Encryption (FHE), as a special advanced privacy protection encryption scheme, can directly perform linear operations on encrypted data (ciphertext) without decryption. In the current situation where private data is continuously shared, processed, and stored online, privacy protection technology has become an increasingly important aspect of current information technology. FHE can perform online calculations on sensitive data in this scenario without revealing its actual data, thus ensuring the privacy and security of its information. Its application scenarios include private machine learning, secure genomic analysis, secure databases, etc.

[0004] Although FHE has powerful functions, all existing FHE schemes usually bring some major computational challenges when used in practical applications. First, in order to perform operations on ciphertext, the computational complexity and memory requirements are quite large. For example, when using existing FHE schemes to perform operations on a simple neural network with 10 layers, the number of calculations can exceed 1.1×10 10The operand requires a computation that is 100,000 times more computationally intensive than that on unencrypted data (plaintext). This significant computational difference is due to the algorithmic characteristics of FHE, which is based on the Learning with Errors (LWE) problem. A single plaintext typically needs to be encrypted into a vector of length N or a polynomial of degree N (N is usually taken as 1024 or larger). Therefore, computations need to be performed on vectors or polynomials rather than scalar data, which not only leads to an increase in the number of operations but also an increase in memory occupancy. The second computational challenge is due to the noise characteristics in LWE-based FHE schemes. To securely hide data, noise needs to be added during encryption. Whenever a computation is performed on the ciphertext, the noise increases. At a certain point, the noise level in the ciphertext exceeds the threshold, at which point the correct result cannot be decrypted. In this case, the noise level needs to be reduced by performing a process called Bootstrapping. However, Bootstrapping is a high-cost operation that typically involves many polynomial operations such as polynomial multiplication, addition, rotation, and Number Theoretic Transform (NTT), which makes Bootstrapping the computational bottleneck of FHE. Therefore, designing a dedicated FHE accelerator to accelerate the FHE algorithm, especially the Bootstrapping part, is crucial for the practical application of FHE.

[0005] TFHE (The Fully Homomorphic Encryption over the Torus) is an efficient FHE scheme based on Programmable Bootstrapping (PBS), and it performs the Bootstrapping operation more efficiently than other schemes. The PBS operation of TFHE can be split into the following steps: Modulus Switching (MS), Blind Rotation (BR), Sample Extraction (SE), and Key Switching (KS). Among these steps, the Blind Rotation step is the most computationally resource-intensive and time-consuming step. This step involves modular multiplication and modular accumulation operations between matrices in a looped computation. The unit of operation is a vector of polynomial elements, and it also includes a large number of Number Theoretic Transform operations, making its parallel implementation difficult. Most existing technical solutions improve the parallelism between different ciphertext operations by measures such as increasing functional unit modules, which leads to additional hardware overhead, and the underutilized pipelines also result in low hardware utilization.

[0006] There is currently a need for a hardware implementation scheme for programmable Bootstrapping for TFHE that fully utilizes the internal parallelism of the PBS algorithm to achieve greater compactness and efficiency.

[0007] There are currently two main types of solutions for the implementation of programmable bootstrapping for TFHE. One is to address the loop problem that cannot be unfolded during the blind rotation process by increasing additional parallelism. This problem is alleviated by introducing the concepts of device-level and core-level parallel processing, thereby enhancing the average parallelism size of a single blind rotation. To achieve core-level parallel processing, this solution requires adding dedicated hardware units to utilize the available parallelism and amortize the cost of sequentially processing ciphertexts in each blind rotation iteration. The other solution is an accelerator relative to traditional CPU architectures, which uses an innovative microarchitecture, such as instantiating directly cascaded high-throughput computing stages and adopting a simplified control logic and routing network, that is, a design idea similar to a stream processor. This solution can utilize the arithmetic units 100% during the PBS process, but the drawback of this solution is that it uses fixed-point FFT to calculate polynomial multiplication. Although it avoids complex double-precision multiplication calculations, the precision is too low, resulting in excessive noise and a high probability of decryption failure.

[0008] In summary, there is currently a need for a hardware implementation solution for programmable bootstrapping for TFHE that can reduce the usage of hardware units, has a high decryption success rate, high performance, and low area. Summary of the Invention

[0009] An embodiment of the present invention provides a hardware implementation system for programmable bootstrapping for TFHE, which can reduce the usage of hardware units, has a high decryption success rate, high performance, and low area. The system includes: a rotation register, at least one functional calculation module, and a vector calculation module, where,

[0010] The vector calculation module is used to perform the modulus switching step in the programmable bootstrapping algorithm for TFHE on the LWE ciphertext to obtain a modulus switching result;

[0011] The rotation register is used to cache and rearrange the LWE ciphertext to obtain a rearranged LWE ciphertext and send it to a functional calculation module;

[0012] The functional calculation module is used to perform the blind rotation step in the programmable bootstrapping algorithm for TFHE on the rearranged LWE ciphertext and the modulus switching result to obtain a blind rotation result;

[0013] The vector calculation module is further used to: perform the sample extraction step and the key switching step in the programmable bootstrapping algorithm for TFHE on the blind rotation result to obtain a decrypted ciphertext;

[0014] Wherein, when performing the blind rotation step, the functional calculation module decomposes the polynomials in the rearranged LWE ciphertext and the modulus switching result into multiple sub-polynomials through a number theory transformation unit, calculates the multiple sub-polynomials in parallel, and calculates the blind rotation result based on the multiple sub-polynomials.

[0015] In the embodiment of the present invention, when the functional computing module executes the blind rotation step, the rearranged LWE ciphertext and the polynomials in the modulus switching result are decomposed into multiple sub-polynomials, and the multiple sub-polynomials are calculated in parallel, thereby reducing the usage of hardware units for calculating the multiple sub-polynomials. In addition, the number theory transform unit is used instead of the FFT. By utilizing the characteristics of the number theory transform without truncation and without error, the decryption success rate of the system is high, the performance is high, and the area is low. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:

[0017] Figure 1 It is a schematic structural diagram of a hardware implementation system for programmable bootstrapping for TFHE in the embodiment of the present invention;

[0018] Figure 2 It is another schematic structural diagram of a hardware implementation system for programmable bootstrapping for TFHE in the embodiment of the present invention;

[0019] Figure 3 It is a schematic structural diagram of the functional computing module in the embodiment of the present invention;

[0020] Figure 4 It is a schematic structural diagram of the NTT unit in the embodiment of the present invention;

[0021] Figure 5 It is a schematic position diagram of the timing scheduling of the vector computing module relative to the functional computing module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.

[0023] The embodiment of the present invention deeply disassembles the PBS algorithm and couples it with the hardware design, constructs a scheme with extremely high hardware utilization rate, selects appropriate modulus parameters at the same time, and designs a TFHE programmable bootstrapping hardware module with high performance, low area, and low decryption failure probability by utilizing the characteristics of the number theory transform without truncation and without error. At the same time, a streaming processing scheme is used in the overall architecture to reduce the usage of SRAM. This scheme is superior to the existing schemes in terms of performance and area performance.

[0024] First, the concepts, terms, and variables involved in the embodiments of the present invention are explained.

[0025] In the embodiments of the present invention, represents the ring of integers modulo q, represents the polynomial ring modulo X N +1, represents the integer polynomial ring modulo X with coefficients in N +1. Conventional lowercase letters represent elements in R or R q , such as c; bold lowercase letters represent column vectors with elements in R or R q , such as a, and a[i] represents the i-th element in the vector; bold uppercase letters represent matrices, such as A. In the programmable bootstrapping algorithm for TFHE, the message can be various types of data, including boolean values, integers, or fixed-point (real) numerical values. Before encryption, the programmable bootstrapping algorithm for TFHE performs encoding of the message, mapping it to the plaintext space. Then, the programmable bootstrapping algorithm for TFHE encrypts the plaintext into a ciphertext, which can be presented in the form of a vector, polynomial, or set of polynomials. The ciphertext in TFHE consists of elements called tori, and the tori is the set of real numbers in the interval [0, 1). In implementation, these torus elements are actually represented as the discretized torus T q , that is, the fixed-point representation of the torus where q is the modulus of the ciphertext. Several types of ciphertexts are used in the programmable bootstrapping algorithm for TFHE, including LWE ciphertext, GLWE ciphertext, and GGSW ciphertext.

[0026] LWE ciphertext is mainly used to encrypt plaintexts or messages of scalar types. If n is defined as the dimension of the LWE ciphertext and p is the modulus of the plaintext, and the key then the LWE ciphertext of the plaintext m ∈ T p is defined as In this case, a1, a2,..., a n are LWE masks, which are n random integers sampled from the set of integers modulo q , representing the discretized torus, and where e is a small noise term. In implementation, the LWE ciphertext can be represented as an (n + 1)-tuple consisting of scalar elements.

[0027] GLWE ciphertext is used to encrypt plaintexts or messages of polynomial types. The ring polynomial is defined as a polynomial with modulus q and modulo polynomial x N +1. If k is defined as the dimension of the GLWE ciphertext and the key is the polynomial then the GLWE ciphertext of the plaintext is defined as:

[0028]

[0029] In this case, A1(x),..., A k (x) are GLWE masks, and where E(x) is a small noise term. In implementation, a GLWE ciphertext can be represented as a (k + 1)-dimensional vector composed of polynomials.

[0030] GGSW ciphertext: Let l be the level, β be the decomposition basis, and the key be GGSW encryption is an extension of GLWE ciphertext. A GGSW ciphertext is defined as In implementation, a GGSW ciphertext can be represented as a matrix of (k + 1)l × (k + 1) elements composed of polynomials.

[0031] Several types of ciphertexts are used in TFHE calculations. First, LWE ciphertexts are used to encrypt messages. The bootstrapping key (BSK) is derived from the key encrypted by GGSW and is used to reset the noise in the LWE ciphertext during the bootstrapping operation. Each key (s1, s2,..., s n ) used for message encryption is encrypted into a GGSW ciphertext. Therefore, the BSK consists of n GGSW ciphertexts: where l b is the level of the BSK. The key-switching key (KSK) is used to homomorphically convert one key to another in the TFHE scheme. It consists of kN × l k LWE ciphertexts, where and l k is the level of the KSK. In addition, the test polynomial (TP) is a GLWE ciphertext that stores all function values of an arbitrary function f(m). Finally, the accumulative ciphertext (ACC) is another GLWE ciphertext specifically used to store the accumulative result of polynomials in the blind rotation step. We refer to ACC i-1 as the ACC input and ACC i as the ACC output, representing the input and output during the iterative process.

[0032] The programmable bootstrapping (PBS) step of TFHE is a process of homomorphically performing a decryption operation on the operated LWE ciphertext c (such as gate-level operations like AND, OR, NOT, etc.) to obtain a ciphertext with refreshed noise.

[0033] Table 1 shows the pseudocode of the programmable bootstrapping algorithm for TFHE in the embodiments of the present invention.

[0034] Table 1

[0035]

[0036]

[0037] Among them, in the above algorithm, in the first step, the modulus switching scales the ciphertext, and calculates The second step is a rotation operation, which is a homomorphic operation in the ciphertext domain where the plaintext vector rotates. Assume the ciphertext c1, and the plaintext vector after decrypting and decoding with the corresponding key is (m0, m1,..., m N / 2-1 ), then after performing the automorphism operation of rotating k steps, the ciphertext c2 is obtained, and the plaintext vector after decrypting and decoding with the corresponding key is (m k , m k+1 ,..., m N / 2-1 , m0, m1,..., m k-1 ). The third to twelfth steps are blind rotation steps, which convert the LWE ciphertext into a noise-free GLWE ciphertext through n iterations of the outer product. Among them, the decomposition step is based on B g as the basis, and decomposes the input into l numbers. The twelfth to fourteenth steps are sample extraction steps, which are operations to extract the sample LWE ciphertext after refreshing (homomorphic decryption) from the blind rotation result. The sixteenth to eighteenth steps are key switching steps, which switch the key of the extracted LWE ciphertext to the same as the key of the original LWE ciphertext.

[0038] Figure 1 FIG.

[0039] The vector calculation module is used to perform the modulus switching step in the programmable bootstrapping algorithm for TFHE on the LWE ciphertext to obtain the modulus switching result;

[0040] The rotation register is used to cache and rearrange the LWE ciphertext to obtain the rearranged LWE ciphertext, and send it to a functional calculation module;

[0041] The functional calculation module is used to perform the blind rotation step in the programmable bootstrapping algorithm for TFHE on the rearranged LWE ciphertext and the modulus switching result to obtain the blind rotation result;

[0042] The vector calculation module is also used to: perform the sample extraction step and the key switching step in the programmable bootstrapping algorithm for TFHE on the blind rotation result to obtain the decrypted ciphertext;

[0043] Among them, when the functional calculation module performs the blind rotation step, through the number theory transformation unit, the polynomials in the rearranged LWE ciphertext and the modulus switching result are decomposed into multiple sub-polynomials, multiple sub-polynomials are calculated in parallel, and the blind rotation result is calculated according to the multiple sub-polynomials.

[0044] Figure 1 In it, through the top-level abstract representation of each functional module, the structure and interaction between components within the system are clearly depicted. The overall architecture is constructed around the computing requirements of TFHE, with the goal of achieving efficient hardware acceleration.

[0045] The functional computing module is one of the core components of the entire architecture and is responsible for executing the Blind Rotation step (corresponding to steps 3 - 12 in Table 1) in the programmable bootstrapping algorithm for TFHE. The Blind Rotation step is a key step in the programmable bootstrapping algorithm for TFHE and is mainly used to achieve phase adjustment during ciphertext operations. Through hardware acceleration design in the system embodiments of the present invention, the functional computing module can significantly improve the execution speed of Blind Rotation, thereby optimizing the overall computing performance of homomorphic encryption.

[0046] The vector computing module undertakes the Modulus Switching step (corresponding to step 1 in Table 1), the Sample Extraction step (corresponding to steps 13 - 14 in Table 1), and the Key Switching step (corresponding to steps 16 - 18 in Table 1) in the programmable bootstrapping algorithm for TFHE. These steps are crucial in the programmable bootstrapping algorithm for TFHE and are directly related to the format conversion of ciphertexts and the compatibility of encrypted data. Through modular hardware design, the vector computing module can effectively reduce latency and improve operation efficiency on the basis of the parallelization and pipelining processing of the aforementioned functional computing module.

[0047] The system integrates a rotation register, which is responsible for caching and rearranging the input LWE ciphertext (corresponding to step 2 in Table 1). The rotation register improves the overall efficiency of computing by reducing data transfer latency, especially in high-dimensional NTT operations.

[0048] Figure 2 This is another schematic structural diagram of the hardware implementation system for the programmable bootstrapping for TFHE in the embodiments of the present invention. Refer to Figure 2 and the system further includes:

[0049] A first register bank for storing the LWE ciphertext used in instruction calculation;

[0050] A second register bank for storing intermediate calculation results and decrypted ciphertexts, where the intermediate calculation results include modulus switching results, blind rotation results, and decrypted ciphertexts.

[0051] The above register file plays a role in data caching and exchange in the entire system. As an intermediate storage unit, it temporarily stores the operation results from the functional computing module and the vector computing module. The efficient design of the register file not only improves the speed of data exchange but also optimizes the utilization rate of hardware resources.

[0052] Specifically, the vector computing module is also used to: send the modulus switching result to the second register file;

[0053] The functional computing module is also used to: read the modulus switching result from the second register file and send the obtained blind rotation result to the second register file;

[0054] The vector computing module is also used to: read the blind rotation result from the second register file and send the obtained decrypted ciphertext to the second register file.

[0055] See Figure 2 , the system further includes:

[0056] A scheduler, which is used to analyze the multiple received instructions and send the instructions to the controller according to the execution order of the instructions;

[0057] A controller, which is used to, after receiving an instruction, fetch the LWE ciphertext used for instruction calculation from the first register file and send it to the rotation register and the vector computing module.

[0058] Specifically, the scheduler may receive multiple instructions, but there are order or other conditional restrictions among the multiple instructions. The scheduler needs to determine the next instruction to be executed and then send it to the controller. In this way, after receiving the instruction, the controller fetches the LWE ciphertext used for instruction calculation from the first register file and sends it to the rotation register and the vector computing module, thus starting the calculation. The controller is also responsible for monitoring and regulating the cooperation status among the modules. The coordinated operation of the two ensures the stability and computing efficiency of the entire system.

[0059] See Figure 2 , the system further includes an on-chip network, which is used for data transmission between the rotation register, the functional computing module, the vector computing module, the first register file, and the second register file.

[0060] The on-chip network is responsible for coordinating data communication among the modules. It can achieve high-bandwidth and low-latency communication capabilities to ensure that different functional modules can share calculation results and intermediate data at the lowest cost. The efficiency of the on-chip network is directly related to the throughput and scalability of the system.

[0061] Through the modular design of the hardware architecture of the system in the embodiments of the present invention and the introduction of the on-chip interconnection network, the contradiction between computationally intensive tasks and communication overhead can be balanced. The functional computing module and the vector computing module have clear division of labor, and at the same time, efficient data interaction is carried out through two register files, significantly improving the operating efficiency of the programmable bootstrapping algorithm for TFHE on hardware. In addition, the intelligent design of the scheduler and the controller further enhances the adaptive ability of the system in complex homomorphic encryption tasks.

[0062] In the outer product calculation of the programmable bootstrapping algorithm for TFHE, although the fast Fourier transform (FFT) is often used to accelerate the calculation, by selecting a prime modulus p that satisfies p ≡ 1 mod 2N and making it as close as possible to the ciphertext modulus q, the number theoretic transform (NTT) can be used to replace the FFT. This method not only eliminates the dependence of the FFT on double-precision floating-point multiplication, but also significantly improves the calculation accuracy, thereby enhancing the efficiency and accuracy of encryption calculations, especially in the high-performance homomorphic encryption hardware implementation.

[0063] In the programmable bootstrapping algorithm for TFHE, the blind rotation step is the most time-consuming, and its core includes a large number of basis conversions and finite field operations. These calculations rely on bit-level operations in homomorphic encryption, specifically involving the vectorized multiplication of matrices with polynomials as unit elements to achieve phase adjustment and functional rotation in the encrypted domain. The key to optimizing the blind rotation step lies in reducing the complexity of ciphertext operations, and at the same time, by streamlining the size of the blind rotation key (BSK), reducing the storage and calculation overhead, thereby significantly improving the overall operation efficiency and reducing resource consumption. This optimization is particularly important for efficient hardware implementation.

[0064] Therefore, the system in the embodiments of the present invention designs an efficient hardware implementation. Figure 3 For the structural schematic diagram of the functional computing module in the embodiments of the present invention, see Figure 3 , the functional computing module includes a basis decomposition unit, at least one number theoretic transform unit (NTT), a multiply-accumulate unit, and at least one inverse number theoretic transform unit (INTT);

[0065] The basis decomposition unit is used to decompose the polynomials in the rearranged LWE ciphertext and the modulus switching result into multiple sub-polynomials, and send the multiple sub-polynomials to a number theoretic transform unit;

[0066] The number theoretic transform unit is used to calculate multiple sub-polynomials in parallel to obtain the polynomials in the transform domain;

[0067] The multiply-accumulate unit is used to perform accumulation operations and modular operations on multiple polynomials in the transform domain to obtain the polynomials for performing multiply-accumulate operations in the transform domain;

[0068] The inverse number theory transform unit is used to restore the polynomial obtained by performing multiplication and accumulation operations in the transform domain to the standard domain to obtain the blind rotation result. Among them, the polynomial in the standard domain after multiple loops is used as the blind rotation result.

[0069] Figure 4 This is the structural schematic diagram of the NTT unit in the embodiment of the present invention. Refer to Figure 4 , each number theory transform unit includes two number theory transform modules, and each number theory transform module includes multiple number theory transform cores;

[0070] One sub-polynomial corresponds to multiple number theory transform cores in one number theory transform unit.

[0071] In one embodiment, the controller is further configured to fetch the constants used for instruction calculation from the first register bank and send them to the number theory transform unit;

[0072] Each number theory transform core includes:

[0073] ROM, which is used to store the constants required for number theory transform calculation;

[0074] RAM, which is used to store the intermediate calculation data during the calculation of multiple sub-polynomials;

[0075] The butterfly unit is used to perform the calculation of multiple sub-polynomials.

[0076] In the above embodiment, the constants required for number theory transform calculation include the twiddle factor, etc. Storing the constants in the ROM can provide fast access to reduce the memory access overhead. The RAM supports efficient data exchange. The butterfly unit, as an operation engine, focuses on basic operations such as modular multiplication and modular addition to construct a single NTT operation. In the embodiment of the present invention, one sub-polynomial corresponds to multiple number theory transform cores in one number theory transform unit. Specifically, each sub-polynomial corresponds to four NTT cores. By decomposing the calculation tasks into small-scale parallel operations, the processing performance of high-dimensional tasks is significantly improved. It should be noted that in the embodiment of the present invention, the calculation of the LWE ciphertext of one instruction corresponds to one NTT unit. And in the embodiment of the present invention, the system may include multiple functional calculation units, so the system can process the decryption requirements of multiple LWE ciphertexts in parallel.

[0077] The sequential execution process of the basis decomposition module, NTT unit, multiply-accumulate unit, and INTT unit is a pipeline process. The multiply-accumulate unit located in the middle section of the pipeline is used to perform accumulation operations on multiple polynomials in the transform domain and handle modular operations to ensure that the result is within the transform domain, obtaining the polynomial for multiplication and accumulation operations in the transform domain. Its design is adapted to pipeline operations and can achieve high-throughput real-time processing.

[0078] The INTT unit is responsible for restoring the polynomial that performs multiplication and accumulation operations in the transform domain to the standard domain to obtain the blind rotation result. Each INTT unit operates independently and supports parallel computing.

[0079] The computing function module realizes the scheduling and storage management of the data stream through two register files to ensure the high-speed transmission and sharing of data among computing units.

[0080] Due to the high hardware cost of implementing a fully pipelined NTT in traditional designs, the in-place NTT method is usually adopted. However, this design makes it difficult to fully pipeline each row of the pipeline, resulting in the running time of the NTT being longer than the computing time of the multiplication and accumulation module, causing the multiplication and accumulation module to be idle, thereby reducing the utilization rate of hardware resources. To solve this problem, the embodiments of the present invention adopt a design that multiplexes a single multiplication and accumulation unit among multiple NTT groups.

[0081] For example, in this system, the coefficients of a polynomial (with a degree of 1024) are distributed in 8 registers (here referring to the registers in the RAM). Since the NTT calculation requires log2 1024 = 10 layers of butterfly operations, and each layer of operation requires 1024 / 8 = 128 clock cycles, a total of 128×10 clock cycles are required to complete the entire NTT calculation. The multiplication and accumulation module only needs 1024 / 8 = 128 clock cycles to complete one calculation. Therefore, by multiplexing a single multiplication and accumulation unit among multiple NTT groups, the hardware area is effectively reduced, the overall hardware utilization efficiency is improved, and the balance of operation performance is ensured at the same time.

[0082] In one embodiment, the controller is further configured to:

[0083] Control the time for the vector calculation module to execute the sample extraction step.

[0084] In one embodiment, the controller is specifically configured to: after receiving an instruction, after a preset time interval after the previous instruction, take out the LWE ciphertext used for the instruction calculation from the first register file, and the preset time interval is equal to the clock cycle for the multiplication and accumulation module to complete one calculation.

[0085] Specifically, the controller receives instructions sequentially, but it is possible that the current instruction has not been completed yet, and two more instructions are received successively. In order for multiple instructions to be executed and to maximize the utilization rate and execution efficiency of the system, the controller makes the instructions be spaced apart by a preset time interval, and the preset time interval is equal to the clock cycle for the multiplication and accumulation module to complete one calculation.

[0086] For example, the data stream and task decomposition of the present invention are as follows: Every eight registers with a depth of 128 (registers in the RAM) can store 1,024 coefficients of a polynomial, which is called a limb. Two registers correspond to one BFU. The input polynomial is decomposed into 4 limbs and stored in 32 registers. This structure is defined as one lane. For each of the aforementioned number-theoretic transform units, it includes two number-theoretic transform modules. An LWE ciphertext consists of k + 1 polynomials (k = 1 in this design). Therefore, an RLWE ciphertext requires two lanes, a total of 64 registers (registers in the RAM), which is defined as one group. One group corresponds to one NTT unit. Due to the requirements of matrix multiplication in the systolic array, one group corresponds to one row of the array. The calculations between different NTT units are independent, but there are periodic differences in the data stream. As Figure 5 is a schematic diagram of the position of the vector calculation module relative to the functional calculation module during the timing scheduling in the embodiment of the present invention. It shows the data throughput and latency matching of the functional calculation module. The parallelism between different NTT units (corresponding to Figure 5 different groups in) lies in the different ciphertexts being processed. The parallelism within the same group is divided according to the ciphertext polynomials to be processed. The purpose is to ensure that each round of NTT takes 128 cycles to match the subsequent operations. Different LWE ciphertexts (polynomials) are required to be input with an interval of 128 cycles. In this way, each polynomial can be processed in a pipelined manner with 128 cycles as a unit in the multiply-accumulate unit. In order to achieve the requirement of an interval of 128 cycles for input between LWE ciphertexts, the controller needs to take out the LWE ciphertext used for instruction calculation from the first register bank 128 cycles after the previous instruction. Through reasonable flow control and buffer design, continuous seamless connection between input and output is achieved.

[0087] Compared with the traditional ASIC implementation method of the programmable bootstrapping algorithm for TFHE, the system proposed in the embodiment of the present invention has a significant reduction in hardware resource consumption. When using the NTT unit compared with the double-precision FFT, the number of bits of the multiplier is reduced from 64 bits to 32 bits. The processing architecture of this system reduces the SRAM area of the calculation module required by the existing work by more than 67.2%, and reduces the number of multipliers by more than 42.9%, achieving a higher throughput rate and better area and time performance.

[0088] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A hardware implementation system for programmable bootstrapping of TFHE, characterized in that: It includes a rotation register, at least one function calculation module and a vector calculation module, wherein: A vector calculation module, used for executing a modulus switching step in a programmable bootstrap algorithm for TFHE on the LWE ciphertext to obtain a modulus switching result; A rotation register is used to cache and rearrange the LWE ciphertext, obtain the rearranged LWE ciphertext, and send it to a functional computing module; A functional calculation module, used for performing a blind rotation step in a programmable bootstrap algorithm for TFHE on the rearranged LWE ciphertext and the modulus switching result, to obtain a blind rotation result; The vector calculation module is also used to: perform the sample extraction step and the key switching step in the programmable bootstrap algorithm for TFHE on the blind rotation result to obtain the decrypted ciphertext; Among them, when executing the blind rotation step, the functional calculation module decomposes the polynomials in the rearranged LWE ciphertext and the modulus switching result into multiple sub-polynomials through the number theory transformation unit, calculates the multiple sub-polynomials in parallel, and calculates the blind rotation result according to the multiple sub-polynomials.

2. The system according to claim 1, characterized in that Also includes: A first register file is used to store LWE ciphertext used for instruction calculation; The second register stack is used to store intermediate calculation results and decrypted ciphertext, wherein the intermediate calculation results include modulus switching results, blind rotation results and decrypted ciphertext.

3. The system according to claim 2, characterized in that The vector calculation module is also used to: send the modulus-to-digital switching result to the second register file; The function calculation module is further used to: read the modulus switching result from the second register file, and send the obtained blind rotation result to the second register file; The vector calculation module is also used to: read the blind rotation result from the second register file, and send the obtained decrypted ciphertext to the second register file.

4. The system according to claim 2, characterized in that Also includes: The scheduler is used to analyze the received multiple instructions and send the instructions to the controller according to the execution order of the instructions; The controller is used for taking out the LWE ciphertext used for instruction calculation from the first register stack after receiving the instruction, and sending it to the rotation register and vector calculation module.

5. The system according to claim 2, characterized in that The controller is also used to: Controls when the vector calculation module performs the sample extraction step.

6. The system according to claim 3, characterized in that It also includes an on-chip interconnect network for data transmission between the rotation register, the function calculation module, the vector calculation module, the first register file and the second register file.

7. The system according to claim 4, characterized in that The functional computing module includes a basis decomposition unit, at least one number theory transformation unit, a multiplication and accumulation unit and at least one number theory inverse transformation unit; a basis decomposition unit, for decomposing a polynomial in the rearranged LWE ciphertext and the modulus switching result into a plurality of sub-polynomials, and sending the plurality of sub-polynomials to a number theory transformation unit; A number theory transformation unit, used for calculating multiple sub-polynomials in parallel to obtain a polynomial in a transformation domain; A multiplication-accumulation unit, used for performing accumulation operation and modular operation on a plurality of polynomials in the transform domain to obtain a polynomial for performing multiplication-accumulation operation in the transform domain; The number theory inverse transform unit is used to restore the polynomials subjected to multiplication and accumulation operations in the transform domain to the standard domain to obtain the polynomials in the standard domain, wherein the polynomials in the standard domain after multiple cycles are used as the blind rotation results.

8. The system according to claim 7, characterized in that Each number theory transformation unit includes two number theory transformation modules, and each number theory transformation module includes multiple number theory transformation cores; One subpolynomial corresponds to multiple number theory transformation cores in one number theory transformation unit.

9. The system according to claim 7, characterized in that The controller is also used to fetch constants used for instruction calculation from the first register file and send them to the number theory transformation unit; Each number theory transformation core consists of: ROM, used to store constants required for number theory transformation calculations; RAM, used to store intermediate calculation data when calculating multiple sub-polynomials; Butterfly unit, used to perform multiple sub-polynomial calculations.

10. The system according to claim 7, characterized in that The controller is specifically used to: after receiving an instruction, fetch the LWE ciphertext used for instruction calculation from the first register stack after a preset time interval after the previous instruction, wherein the preset time interval is equal to the clock cycle of the multiplication and accumulation module to complete one calculation.

Citation Information

Cited By

  • TFHE fully homomorphic bootstrap operation acceleration method and device

    CN120729504A

  • A method and apparatus for accelerating TFHE fully homomorphic bootstrap computation

    CN120729504B

  • Shift register array-based TFHE fully homomorphic operation bootstrap acceleration method

    CN122293305A