Paillier homomorphic encryption hardware accelerator based on residual number system and number theory transformation and acceleration method of Paillier homomorphic encryption hardware accelerator
By combining RNS and NTT's hardware accelerator architecture, the problems of high computing latency, low energy efficiency, and poor parallel scalability of Paillier homomorphic encryption accelerators are solved, and high-throughput and high-energy-efficiency Paillier homomorphic acceleration are achieved, which is suitable for cloud computing and edge AI acceleration.
Patent Information
- Application Number
- CN202511061311.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-10
AI Technical Summary
Existing Paillier homomorphic encryption accelerators have shortcomings in terms of high computational latency, low energy efficiency, poor parallel scalability, and lack of structured support for residual number systems (RNS), making it difficult to meet the high throughput requirements of cloud computing and edge devices.
It adopts a hardware accelerator architecture based on the residual number system (RNS) and number theoretic transform (NTT), including the RNS conversion unit, the NTT fast multiplication unit, CRT reconstruction and modular exponential operation, supports parallel modular reduction, fast convolution multiplication and dynamic modular basis control, and realizes pipeline operation in combination with a global controller.
It significantly improves the throughput of large integer operations, increases energy efficiency by 20 to 50 times, supports multiple homomorphic operation types, is suitable for cloud-based privacy computing and edge AI acceleration, and has flexible scalability and system-level deployment friendliness.
Smart Images

Figure CN120768529A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of cryptographic accelerator design and information security hardware implementation, and particularly relates to a large integer operation optimization method for Paillier homomorphic encryption, and in particular to a high-throughput homomorphic operation accelerator based on a residue number system (RNS) and number theory transform (NTT) architecture and an implementation method thereof. BACKGROUND
[0002] With the rapid expansion of cloud computing, artificial intelligence and big data scenarios, the demand for data privacy protection is increasingly urgent. Homomorphic encryption technology, as an important way to realize "ciphertext domain computing", has important application prospects in medical, financial, government and multi-party secure computing scenarios.
[0003] Among them, the Paillier homomorphic encryption algorithm has the structural advantage of supporting additive homomorphism and constant multiplication homomorphism, and has become a widely used basic cryptographic module in cloud computing. The basic operations in its encryption domain include: large integer modular multiplication, modular power (exponentiation), modular inverse, modular subtraction, and other extended arithmetic.
[0004] However, as the key size grows to 1024 bits, 2048 bits or even 4096 bits, the complexity of Paillier's homomorphic operation increases rapidly, posing a serious challenge to existing software and hardware:
[0005] High operation delay and limited throughput: large integer multiplication, modular power and other operations require multiple rounds of carry multiplication and conditional judgment, resulting in pipeline stalls, cache misses and other efficiency bottlenecks in general-purpose processors;
[0006] Low energy efficiency and high power consumption: general-purpose CPUs or GPUs are not optimized for modular operations, and their bit width utilization is low, which cannot exploit the underlying modular parallelism and reusability, resulting in linear or even exponential increase in energy consumption with key size;
[0007] Poor scalability and resource waste: existing specialized circuits mostly use serial Montgomery multiplication, lack scalable data paths and multi-core support, and are difficult to meet the dual requirements of low latency and high concurrency on edge and cloud.
[0008] To alleviate the above bottlenecks, some research has introduced the residue number system (RNS) and number theory transform (NTT) to parallelize large integer operations. This type of method maps integer operations to parallel small integer calculations in multiple modular domains, and combines NTT to achieve fast convolution and multiplication acceleration. However, current related implementations mostly remain at the software level, and have the following problems:
[0009] Only suitable for static data structures, unable to adaptively optimize the dynamic characteristics of ciphertext at runtime;
[0010] The lack of a complete hardware-level architecture integration design makes efficient deployment in FPGAs or ASICs difficult.
[0011] The lack of system scheduling strategies and data flow management makes it difficult to meet end-to-end high-throughput acceleration requirements.
[0012] Therefore, there is an urgent need for a configurable hardware acceleration solution for Paillier homomorphic encryption that integrates RNS and NTT, and supports data flow optimization, dynamic modular basis control, and multi-stage pipeline scheduling mechanisms to achieve efficient processing of large integer ciphertexts in on-chip systems. Summary of the Invention
[0013] This paper addresses the following deficiencies of existing Paillier homomorphic encryption accelerators and proposes a novel hardware acceleration method and architecture:
[0014] High computational latency: Existing acceleration solutions primarily rely on a serial Montgomery modular multiplication architecture, which results in numerous pipeline bubbles and idle resources, making it difficult to meet the demands of batch large integer processing.
[0015] Low energy efficiency: Traditional modular multiplication designs require frequent high-order carry and intermediate reduction operations, resulting in significant power consumption, especially in large-scale homomorphic multiplication.
[0016] Poor parallel scalability: Large integer operations on general-purpose CPU / GPU platforms are limited by the RNS and FFT / NTT acceleration capabilities, resulting in insufficient batch processing efficiency.
[0017] Lack of structured support for residual number systems (RNS): Traditional schemes are generally unable to efficiently convert and restore Paillier ciphertext in the RNS domain, limiting overall throughput.
[0018] To address the above issues, the present invention constructs a Paillier homomorphic encryption processor architecture based on the Residual Number System (RNS) + Number Theoretic Transform (NTT) acceleration, which includes the following key technical modules:
[0019] Residual Number System (RNS) conversion unit:
[0020] The large integer Paillier ciphertext Convert to RNS representation under multiple small models;
[0021] Selected module base They are mutually prime, with a bit width of 16 to 32 bits, and the number of modular bases m is usually 4 to 16;
[0022] Supports parallel modular reduction paths to meet high throughput requirements.
[0023] NTT fast multiplication unit:
[0024] Each NTT operation with length L (512, 1024 or 2048) is performed in each modulus field, realizing fast convolution multiplication;
[0025] The core is composed of a butterfly unit, and the operation complexity is O(L log L), which greatly reduces the delay;
[0026] The same NTT core supports forward and reverse NTT switching, improving hardware resource reuse rate;
[0027] Supporting in-place transformation and inter-layer buffer scheduling two modes, the recommended width of the butterfly array is L / 2.
[0028] CRT reconstruction and modulus exponent operation:
[0029] After performing inverse NTT (INTT) on the NTT multiplication result, the Chinese remainder theorem (CRT) is used to restore the large integer result;
[0030] The modulus exponent operation adopts a windowed lookup table mechanism, and the window size k is 4-8 bits, and 2^k power value items are pre-stored in the ROM;
[0031] At the same time, it supports modulus subtraction, modulus inverse and other operations, all based on parallel paths in RNS domain and standard arithmetic restoration strategy.
[0032] Accelerator hardware structure:
[0033] It includes: global controller, modular mul-acc unit, RNS conversion cluster, NTT core cluster, CRT reconstruction cluster, coefficient SRAM, window lookup table ROM, NTT cache SRAM, data SRAM, DDR controller and PHY interface, DMA&PCIe endpoint, central NOC;
[0034] The control module schedules each calculation process through microinstructions, supporting pipelining operation;
[0035] Each module is interconnected through a 256-bit wide x 8 channel NoC bus to realize high-bandwidth and low-delay data exchange;
[0036] Peripheral interface supports PCIe Gen3 standard and above, seamlessly integrated with host system;
[0037] The system supports dynamic expansion of module base, configurable NTT length and window function precision.
[0038] Supported homomorphic operation types:
[0039] Homomorphic addition: modular addition in RNS domain;
[0040] Homomorphic multiplication: NTT domain multiplication convolution + INTT + CRT;
[0041] Modular exponentiation: NTT domain fast power + window lookup table acceleration;
[0042] Modular subtraction / inverse: Euclidean algorithm or inverse element method in small modulus domain + CRT reorganization.
[0043] The RNS + NTT acceleration architecture proposed in the application has the following significant advantages compared with the prior art:
[0044] Significant improvement in throughput: the large integer operation is divided into multiple small modulus channels and executed in parallel, combined with NTT convolution acceleration, the single-chip homomorphic multiplication throughput is improved by 3-8 times;
[0045] Very high energy efficiency: the NTT core supports pipelining execution, reduces the redundant operation of traditional modular multiplication, and the energy efficiency can be improved by 20-50 times under a 12nm process;
[0046] Flexible scalability: supports adjusting the number of module bases m, NTT length L, and window size k, facilitating dynamic performance-area-power trade-off in different scenarios;
[0047] Comprehensive functions: supports the complete operation set (addition, multiplication, exponent, modular inverse, modular subtraction) required for Paillier encryption, suitable for cloud privacy computing, federated learning, security outsourcing and other applications;
[0048] System-level deployment friendly: can be deployed as an independent chip, PCIe acceleration card, edge module, etc., and supports domestic CPU platforms (such as Feiteng and Kunpeng) system adaptation. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 It is a schematic diagram of the overall structure and data flow of the application.
[0050] Figure 2 It is a structural schematic diagram of the RNS conversion cluster (RNS Conversion Cluster) described in the embodiment of the application.
[0051] Figure 3The structural schematic diagram of the NTT core cluster in the embodiment of the present application.
[0052] Figure 4 The structural schematic diagram of the modular mul-acc unit in the embodiment of the present application.
[0053] Figure 5 The structural schematic diagram of the CRT reconstruction cluster in the embodiment of the present application. DETAILED DESCRIPTION
[0054] The present application provides a Paillier homomorphic encryption acceleration system combining a residual number system (RNS) and a number theory transform (NTT), and a system architecture thereof is shown in Figure 1 The system includes multiple parallel modules, adopts an on-chip high-speed interconnection structure, and supports a complete data path from input ciphertext to output result.
[0055] As shown in Figure 1 The system includes the following core modules:
[0056] DMA&PCIe endpoint: PCIe high-speed communication with a host system, supporting DMA data transfer.
[0057] Data SRAM: Central shared data buffer, buffering input / output of each stage.
[0058] RNS conversion cluster: Parallelly performing modulo m_i operation, realizing large integer to RNS component conversion.
[0059] NTT core cluster: Performing NTT forward transform and inverse transform, constituting a frequency domain transform core.
[0060] Coefficient SRAM: Storing ω, ω⁻¹, Mᵢ, tᵢ and other coefficients required for NTT and CRT calculation.
[0061] NTT cache SRAM: Storing intermediate level data in the NTT or INTT process.
[0062] Modular mul-acc unit: Realizing multiplication and multiplication-addition operation in the NTT domain.
[0063] Window lookup table ROM (Window ROM) (optional): Pre-stores a table of windowed modular exponents for table lookup multiplication.
[0064] CRT Reconstruction Cluster: performs multi-mode result merging and reconstructs the large integer corresponding to the Paillier ciphertext.
[0065] DDR controller and PHY interface (DDR Controller & PHY) (optional): External large-capacity data cache for data exchange in high-throughput scenarios.
[0066] The steps that the system goes through to perform a complete Paillier homomorphic operation include: data exchange between the host and the accelerator chip, analog conversion within the chip, number theory transformation, multiplication and addition calculation, CRT reconstruction and optional post-processing steps. Figure 1 The structure diagram and flow chart of the system detail the data flow at each stage.
[0067] Host data transmission and RNS conversion phase:
[0068] Step 1: The host transfers the Paillier ciphertext data to be processed to the DMA&PCIe endpoint through the PCIe interface;
[0069] Step 2: The received ciphertext is written to Data SRAM via DMA;
[0070] Step 3: The RNS Conversion Cluster reads the ciphertext from the Data SRAM and performs modular reduction on the modules m1 to m_k.
[0071] Step 4: The RNS converted mode domain components (C1, C2, ..., C_k) are written back to the Data SRAM for subsequent processing.
[0072] NTT forward transformation stage:
[0073] Step 5: The RNS component is read from the Data SRAM and sent to the NTT Core Cluster.
[0074] Step 6: The coefficient SRAM provides the root value of ω or ω⁻¹ required for NTT operation.
[0075] Step 7: The intermediate transformation results are temporarily stored in the NTT Cache SRAM;
[0076] This stage transforms the time domain data in the modulo domain into a frequency domain representation in order to perform element-wise convolution.
[0077] Homomorphic multiplication / multiply-add calculation stage:
[0078] Step 8: The Modular Mul-Acc Unit reads the frequency domain data from the NTT Cache SRAM.
[0079] Step 9: (Optional) If this is a modular exponentiation operation, obtain the precomputed power value table from the Window ROM.
[0080] Step 10: After performing the multiplication or multiply-accumulation operation, the result is written back to the Data SRAM.
[0081] NTT inverse transformation and RNS reconstruction stage:
[0082] Step 11: Send the result from Data SRAM back to the NTT Core Cluster to perform inverse NTT;
[0083] Step 12: Write the RNS component output by the inverse transform into the Data SRAM;
[0084] Step 13: The CRT Reconstruction Cluster reads all components and prepares to perform CRT synthesis.
[0085] Step 14: The coefficient SRAM provides the Mᵢ and tᵢ constants.
[0086] Step 15: The CRT result is written into the Data SRAM to obtain the final output in the Paillier ciphertext space.
[0087] Output and post-processing stage:
[0088] Step 16: The processed data is written back to the DMA&PCIe endpoint by the Data SRAM;
[0089] Step 17: Transmit the data back to the host via PCIe for reading by upper-layer software.
[0090] Optional paths and mechanisms:
[0091] Window ROM path: Enabled when performing modular exponential operations;
[0092] Dynamic module base adjustment: The global controller triggers adjustments based on bit density monitoring, increasing or decreasing the number of RNS paths.
[0093] DDR path: In the large data batch processing scenario, the Data SRAM data can be transferred to the external DDR.
[0094] The system significantly improves the operation throughput while ensuring high-precision security through the above-mentioned pipelined data path and parallel module design, and is suitable for Paillier homomorphic acceleration requirements in privacy computing, financial security, AI encryption reasoning and other scenarios.
[0095] The timing flow of the homomorphic acceleration operation processing of the application is as follows:
[0096] Ciphertext input (Input Ciphertext): The host sends Paillier ciphertext to the chip end through the PCIe interface, and the data is first received through DMA&PCIe Endpoint (DMA&PCIe Endpoint) and written into the on-chip cache Data SRAM. The task of this stage is to complete the loading of the ciphertext on the chip, preparing for the subsequent module conversion.
[0097] RNS conversion (RNS Conversion): The large integer ciphertext loaded from the SRAM enters the RNS conversion cluster (RNS Conversion Cluster) and is split into several small modulus components. Each parallel path realizes C mod mᵢ according to the modulus m1, m2, …m_k through multiplication, shift and modulus subtraction, forming the RNS represented ciphertext component. After conversion, the result is written into the Data SRAM again.
[0098] Forward NTT: Each RNS component is sent from the Data SRAM to the NTT core cluster (NTT Core Cluster) and through the butterfly structure and the precomputed ω coefficients in the Twiddle ROM (Twiddle ROM), the fast number theory positive transformation (NTT) of the component in the modulus domain is completed, and the frequency domain represented data (Ĉ) is obtained. The transformation process is L-level pipeline calculation, supporting parallel multiple points. The result is temporarily stored in the NTT cache SRAM (NTT Cache SRAM).
[0099] Homomorphic Mul: The NTT domain represented data (Ĉ) is sent to the modular multiplication-accumulation unit (Modular Mul-Acc Unit). According to the operation type, the ciphertext×ciphertext multiplication or ciphertext×constant multiplication-addition path is selected. If the modulus exponent is involved, the precomputed power value will also be read from the window lookup table ROM (Window ROM). All operations are performed in the small modulus domain in the form of modular multiplication+modular addition, and the output result is written back to the Data SRAM.
[0100] Inverse NTT Transform: The frequency domain results are loaded from SRAM back into the NTT Core Cluster. The system switches to INTT mode (using ω⁻¹) to perform the inverse number-theoretic transform, restoring the data from the frequency domain to the time domain. This stage also uses a butterfly pipeline structure, and the output is written back to the data SRAM.
[0101] CRT Reconstruction: The RNS components in each mode domain are read out again and fed into the CRT Reconstruction Cluster. Here, a parallel multiplier array is used to calculate yᵢ × Mᵢ × tᵢ, which is then reduced to a large integer using a tree-based addition network. The final result forms the complete Paillier homomorphic output.
[0102] Output Result: The reconstructed large integer result is written to Data SRAM and returned to the host memory through DMA & PCIe Endpoint.
[0103] This patented system consists of multiple independent functional modules, each with specialized arithmetic processing capabilities, that collaborate to perform the core operations of Paillier homomorphic encryption. This section details the structure and implementation of the RNS Conversion Cluster, NTT Core Cluster, Modular Multiply-Accumulate Unit, and CRT Reconstruction Cluster.
[0104] like Figure 2 As shown in the figure, the RNS Conversion Cluster is responsible for converting the Paillier large integer ciphertext received from the host into the Residual Number System (RNS) representation. Its module structure includes:
[0105] Input Interface: Reads the complete large integer ciphertext from the Data SRAM and broadcasts it to each modular conversion path, acting as a data distributor and the starting point of the entire modular reduction process.
[0106] Constant ROM Bank: Provides modulus for each modulo path Related precomputed constants, such as multiplication inverses, multiplication constants (multiply-subtract optimization), and segmented lookup table values for fast division.
[0107] Mod Path #1 ~ Mod Path #k: Each path processes a large integer modulo The conversion includes the following core components:
[0108] 1) Multiplier: Execution Mod The multiplication part involved in fast modular operations supports pipeline concurrent execution;
[0109] 2) The Modular Subtractor performs the final modular reduction, i.e. Mod , ensuring that the result falls on ;
[0110] 3) Shifter: Used when using the "multiplication + right shift" method to approximate modular reduction.
[0111] Dynamic Basis Controller: Dynamically enables / disables some modular paths based on ciphertext density or accuracy requirements, enabling runtime dynamic adjustment of the number of modular bases k, and working in conjunction with the Global Controller.
[0112] Output Aggregator: collects the output results of each module path (RNS component ), packaged and sent back to Data SRAM.
[0113] This module achieves high-speed conversion of large integers into multiple module domain representations through parallel design, laying the foundation for subsequent NTT processing.
[0114] like Figure 3 As shown in Figure 1, the NTT Core Cluster is the core array for number theory transformations (NTT / INTT) in the system. The main structure includes:
[0115] Butterfly Stage (Butterfly Array): Consists of L / 2 butterfly units, each of which performs a modular addition and modular multiplication.
[0116] Twiddle ROM: stores precomputed ω or ω⁻¹ root coefficients.
[0117] NTT Cache SRAM: used to cache transformation results in the middle layer.
[0118] NTT Controller: Supports NTT / INTT dual-mode switching and can dynamically switch directions according to the operation type.
[0119] Supported points: Supports two NTT lengths of 1024 and 2048.
[0120] Pipeline depth and performance: Each transformation stage consists of four pipeline stages, with an overall throughput of approximately 0.8 MTransforms / s.
[0121] This module improves the efficiency of polynomial multiplication by converting time domain to frequency domain (NTT) to convolution to time domain (INTT), and is the key to accelerating homomorphic multiplication.
[0122] like Figure 4 As shown in Figure 1, the Modular Mul-Acc Unit is used to perform the core operations of the Paillier ciphertext, including homomorphic multiplication, multiplication-addition, and modular exponential operations. Its structural components include:
[0123] Input path: Receives input data from NTT Cache SRAM or Data SRAM.
[0124] Dual multiplication engine: supports different operation paths such as ciphertext × ciphertext, ciphertext × constant, and exponential multiplication.
[0125] Window ROM interface (optional): When the modular exponential function is enabled, connect the Window ROM as the multiplication constant lookup table source.
[0126] Modular addition and accumulation link: provides element-by-element modular addition operations and supports convolutional accumulation.
[0127] Output path: The calculation results are written back to Data SRAM.
[0128] Control logic: The global controller performs unified scheduling to achieve switching between different computing paths.
[0129] This module directly determines the latency and throughput of Paillier homomorphic multiplication, and optimizing its structure is one of the focuses of this invention.
[0130] like Figure 5 As shown in Figure 1, the CRT Reconstruction Cluster is responsible for combining components distributed across multiple modular domains into a large integer, i.e., performing Chinese Remainder Theorem (CRT) reconstruction. The structure is as follows.
[0131] Input interface: receives RNS components from Data SRAM, and Mᵢ and tᵢ from Coefficient SRAM.
[0132] Multiplication array: performs yᵢ × Mᵢ × tᵢ in parallel.
[0133] Tree addition network: reduces each multiplication result in stages to the final value.
[0134] Output interface: writes the reconstruction result back to Data SRAM.
[0135] Parallelization optimization: equips the pipeline multiplier and addition array to make full use of the on-chip DSP resources.
[0136] This module ensures that the parallel computation in the RNS domain is finally correctly aggregated into the Paillier integer result, and is a key component in the output stage.
[0137] The Paillier homomorphic encryption hardware accelerator described in the application is implemented based on a 12nm FinFET process in design, has excellent energy efficiency indicators and process realizability. Under the typical configuration of a modulus number k=8, an NTT length L=2048, and a window size k=6, the resource consumption of the system logic unit is about 1.2 million LUTs, and the register resource is about 2.5 million flip-flops (FFs). In terms of the calculation module, the system integrates 2048 18×18-bit DSP multiplier units, which are mainly used for high-concurrency operation support of the NTT butterfly unit array and the modulus domain multiplication and addition engine.
[0138] In terms of on-chip storage resources, the system contains about 8Mb SRAM capacity, which is used to store the NTT primitive root and inverse root coefficient table, the window function lookup table, and the RNS modulus conversion and NTT intermediate value cache. Among them, the NTT coefficient and window function table can support configuration switching of multiple L and k parameters to meet different Paillier key lengths and precision requirements. The NoC bus adopts a 256-bit×8-channel bidirectional high-bandwidth structure to realize high-speed data transmission and multicast support between modules, and can support data reduction and broadcast operation modes.
[0139] The performance test results show that under the condition of a 800MHz main frequency, the average delay of a single Paillier homomorphic multiplication operation is about 2.5 microseconds (µs), and the peak throughput can reach about 400,000 times per second (400kOps / s). Under this configuration, the energy efficiency performance is about 0.05 microjoules per operation (0.05µJ / Op), which can achieve an order of magnitude performance and energy efficiency improvement compared with the traditional serial implementation based on Montgomery modular multiplication.
[0140] In terms of deployment, the accelerator supports deployment in a PCIe slot on a cloud server motherboard as a PCIe expansion card. It can also be packaged as an edge AI computing module and embedded in a lightweight device platform. Furthermore, the invention supports multi-chip horizontal expansion architectures, enabling the construction of larger acceleration clusters through a NoC mesh or multi-card interconnection to serve larger volumes of privacy-focused computing tasks, demonstrating strong engineering feasibility and application scalability.
[0141] In order to achieve efficient deployment and programmable call of the Paillier homomorphic encryption hardware accelerator described in the present invention in an actual system, the entire acceleration system is equipped with a complete interface protocol and software support mechanism. The underlying communication uses PCIe Gen3 and above as a high-speed data channel between the host and the acceleration chip, and the DMA engine realizes high-speed transmission and reception of ciphertext and task triggering, supports the continuous transmission of batch data blocks, and effectively reduces software interrupt overhead and bus waiting delay. Configuration and status management are carried out through the AXI-Lite or APB bus. The host can dynamically read the register status of each functional module of the accelerator and set operating parameters such as the number of modular bases, NTT length, root coefficient table address, task control bits, etc.
[0142] On the software side, the present invention provides an OpenCL driver, C / C++ encapsulation library, and a callable API interface, enabling host programs to complete operations such as ciphertext loading, computation scheduling, and encryption / decryption process configuration through a unified function call method. The accelerator provides a task queue management mechanism, allowing asynchronous scheduling of multiple tasks and automatic result transmission. Furthermore, the accompanying software development toolchain includes an automatic generator for root coefficient tables and window function lookup tables, supporting parameterized configuration of different key length and modulus-based combinations, further enhancing the system's portability and deployment flexibility.
[0143] In addition, to facilitate cloud integration and edge deployment, the present invention also supports embedding the accelerator into the GPU / FPGA slot in the form of a PCIe expansion card, or packaging it into the edge AI platform as a SoC module. The system integrates performance counters and error detection registers to support refined monitoring of indicators such as the execution efficiency, cache hit rate, and bus bandwidth utilization of each stage of the accelerator during operation, providing feedback data support for subsequent software optimization and hardware upgrades. The overall solution has good openness and compatibility, and can be smoothly connected with existing cloud security platforms, edge gateways, privacy computing clusters and other systems to meet the needs of actual engineering deployment.
Claims
1. A hardware acceleration method for accelerating large integer computations using Paillier homomorphic encryption, characterized in that: The steps include: 1.1 The input large integer ciphertext is converted to the residual number system (RNS) according to the preset modular basis set B={m1,m2,...,m_k}, where each m_i is a coprime number, k is an integer not less than 4, and the bit width of each m_i is between 32 and 64 bits; 1.2 Perform a number theoretic transform (NTT) on each RNS component x_i in parallel, where the length n of the NTT is a power of 2 and n∈{1024, 2048}; 1.3 Perform element-wise convolution multiplication on the transformed components in parallel in the NTT domain to achieve homomorphic multiplication between ciphertexts or homomorphic multiplication and addition of ciphertext and plaintext constants; 1.4 Perform inverse number theoretic transformation (INTT) on the convolution products in parallel to restore them to the multiplication or multiplication-addition results in each module domain; 1.5 Reconstruct all INTT output results based on the Chinese Remainder Theorem (CRT) to restore complete large integers, and support homomorphic addition, multiplication, modular exponentiation, modular inverse and modular subtraction operations; 1.6 The reconstructed result is pruned into an output that conforms to the Paillier ciphertext space format through the most significant bit (MSB) truncation or dynamic modulus adjustment mechanism.
2. The method according to claim 1, characterized in that The dynamic modular basis adjustment mechanism is used to dynamically increase or decrease the number of prime numbers in the modular basis set B according to the bit density of the data to be processed during operation, so as to achieve a dynamic trade-off between throughput and hardware resource overhead.
3. The method according to claim 1 or 2, characterized in that The NTT and INTT operations are both implemented through a unified butterfly computing unit, and the forward or reverse mode can be switched through a configuration register to reduce hardware resource redundancy.
4. The method according to any one of claims 1 to 3, characterized in that Before executing steps 1.2 and 1.4, the primitive root coefficients and the inverse root coefficients are loaded from the pre-computed root coefficient table, respectively, to accelerate the modular multiplication operation in the butterfly transform.
5. The method according to any one of claims 1 to 4, characterized in that The CRT reconstruction step includes: calculating y_i·M_i·t_i for the remainder y_i under each module domain and performing parallel accumulation, where Mi_i is the product of all modules except m_i, that is, , t_i is the multiplicative inverse of Mi_i modulo mi_i.
6. A hardware accelerator device for implementing the method according to claim 1, characterized in that: include: 6.1RNS conversion unit, including k configurable modulo m_i conversion channels, used to convert input ciphertext into RNS components; 6.2NTT / INTT butterfly computation unit array, each unit supports forward or inverse number theory transformations through register configuration; 6.3 Modular domain multiplication and addition unit, used to perform convolution multiplication or multiplication and addition operations in NTT domain in parallel; 6.4CRT reconstruction unit, integrating parallel M_i·t_i multipliers and adder-accumulator links; 6.5 root coefficient memory, used to store pre-calculated NTT primitive roots and INTT inverse root coefficients; 6.6 Control unit, used to schedule the operation process of the above modules and control the implementation of module-based adjustment strategy.
7. The device according to claim 6, characterized in that The butterfly computation unit array and the modular domain multiplication and addition unit are interconnected via a configurable cross-connection network to support variable-length number-theoretic transformation operations.
8. The device according to claim 6 or 7, characterized in that The system further includes a read-only memory (ROM) for storing a pre-computed windowed modular exponentiation table to support fast modular exponentiation by k-bit windows.
9. The device according to any one of claims 6 to 8, characterized in that It further includes a dynamic module base adjustment module for dynamically increasing or decreasing the active module according to the instructions issued by the control unit, and synchronously updating the root coefficient table.
10. The device according to any one of claims 6 to 9, characterized in that The accelerator is implemented on a 28nm to 7nm process node and exchanges high-bandwidth data with a host system through a PCIe Gen3 or higher rate interface.
Citation Information
Cited By
Paillier hardware accelerator based on distributed multi-core pipeline
CN121864279A