High-performance and configurable Diithium hardware coprocessor
By designing a mod-multiplied high parallel NTT structure, reconfigurable SHA-3 unit, multi-channel sampler and data storage unit in the Dilithium hardware coprocessor, the problems of large hardware overhead and high resource consumption in the prior art are solved, and a high performance and high area efficiency Dilithium hardware coprocessor is realized.
Patent Information
- Application Number
- CN202510204491.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
AI Technical Summary
While improving performance, the existing Dilithium hardware coprocessors have high hardware overhead, resulting in high resource consumption and low area efficiency.
A high-performance, configurable Dilithium hardware coprocessor is designed. The optimized design of the computing unit includes a mod-multiplied high-parallel NTT structure, a reconfigurable SHA-3 unit, a multi-channel sampler and a data storage unit. By flexibly configuring NTT and INTT transformations, polynomial operations and other functions, it can achieve low hardware overhead and high performance.
It achieves lower hardware overhead, improves performance while maintaining resource consumption and improves area efficiency.
Smart Images

Figure CN120144529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Dilithium coprocessors, and specifically refers to a high-performance and configurable Dilithium hardware coprocessor. Background Art
[0002] Threats such as security attacks, privacy leaks, and trust deficiencies that are increasingly prominent in network information systems have also driven the continuous innovation and improvement of information security technologies. The core of information security technologies lies in cryptographic algorithms, and the security of these algorithms depends on the complexity of a series of mathematical problems, such as integer factorization and discrete logarithm problems. However, quantum computers can process these problems at an exponential speed using quantum mechanical properties such as entanglement and superposition, and this property is also considered the key to the next major revolution in the computing field. In 2015, technology companies such as Google and IBM accelerated their exploration of quantum computing technologies, and domestic scientists have also made many important breakthroughs in the field of quantum computing.
[0003] This cutting-edge computing power will pose a serious threat to current encryption technologies that rely on traditional mathematical problems. Therefore, the cryptography field needs to gradually migrate existing encryption infrastructures to post-quantum secure alternatives that can resist quantum computing attacks. To address this challenge, the National Institute of Standards and Technology (NIST) in the United States launched a public competition in 2016 to solicit post-quantum cryptographic algorithms that can withstand potential attacks from quantum computers. As a platform for digital communication, hardware facilities need to integrate reliable encryption protocols to ensure the privacy and integrity of their own data.
[0004] Due to its high efficiency in security and flexibility, Dilithium is applicable to many fields that require data protection:
[0005] (1) Network security and identity authentication: In digital identity authentication and online transactions, Dilithium can be used to generate and verify digital signatures of users, ensuring data integrity and non-repudiation;
[0006] (2) Blockchain technology: In blockchain networks, the quantum-resistant characteristics of Dilithium provide higher security for distributed ledgers and smart contracts;
[0007] (3) Internet of Things (IoT) devices: IoT devices usually face problems of resource constraints and high security requirements, and the high efficiency of Dilithium makes it an ideal choice.
[0008] Meanwhile, as the NIST-recommended lattice digital signature standardization solution, after making progress in security assessment, Dilithium needs to focus on the efficiency of its deployment and implementation to adapt to application scenarios such as high-frequency signature verification. The Dilithium coprocessor is a hardware accelerator designed to improve the execution efficiency and security of the Dilithium digital signature algorithm. It is specifically optimized for the computationally intensive operations in the Dilithium algorithm and provides a series of hardware support functions.
[0009] In the early evaluation of Dilithium, many scholars proposed novel methods for pure hardware implementation. SONI et al. used a hardware design method based on High-Level Synthesis (HLS) to flexibly map the C specifications of Dilithium and qTESLA to FPGA implementation and analyze the algorithm. Subsequently, RICCI et al. presented the first complete VHDL-based hardware design of Dilithium, hardwareizing all functional components by consuming a large amount of resources and comparing it with C-based and HLS-based implementations. LAND et al. optimized the NTT / INTT unit by efficiently using a dedicated Digital Signal Processor (DSP) and incorporated the scaling factor n of INTT into the rotation factor to reduce the consumption of Look Up Table (LUT) and Flip Flop (FF) resources. Further, BECKWITH et al. re-planned the polynomial coefficients and rotation factors, used 4 butterfly units to process two layers of NTT / INTT each time to reduce the memory access frequency, and adopted a multi-stage pipeline to efficiently process data at each stage. AIKATA et al. aimed at low area, shared the Keccak core and polynomial multiplier among algorithms to implement a unified architecture for Dilithium and Saber / Kyber. ZHAO et al. used a Radix-2 multi-path delay converter to reduce the memory access of NTT, performed fine-grained parallel operations on independent data, reduced the overall clock cycles, and achieved a compact and high-performance hardware structure. GUPTA et al. used a single compact Keccak core and area-optimized NTT units for devices with strict area or power constraints to implement a lightweight coprocessor and tested it on the Xilinx Zynq 7000 platform. Summary of the Invention
[0010] The technical problem to be solved by the present invention is to provide a high-performance and configurable Dilithium hardware coprocessor with lower hardware overhead, improved performance while maintaining resource consumption, and high area efficiency, aiming at the deficiencies mentioned in the above background technology.
[0011] To solve the above technical problems, the technical solution provided by the present invention is: a high-performance and configurable Dilithium hardware coprocessor, and the optimized design of the arithmetic unit includes the following structures: a modular multiplication high-parallel NTT structure, a reconfigurable SHA-3 unit supporting SHAKE-128 and SHAKE-256, a multi-channel sampler, and a data storage unit.
[0012] Further, the modular multiplication high-parallel NTT structure is composed of an address generation unit, a control unit, a rotation factor preprocessing unit, a coefficient transposition unit, a reconfigurable low-latency radix-4 CT / GS unit, and an address synchronization unit;
[0013] Relying on the adjustment of the mode control signal Mode, this hardware structure can flexibly configure the data paths of NTT transformation, INTT transformation, polynomial addition operation, polynomial subtraction operation, and polynomial multiplication operation, so as to execute different polynomial operations. In addition, through the RAM selection signal, chip selection control of two RAM units can be realized;
[0014] In particular, three RAM units are introduced, including a rotation factor RAM, a coefficient storage RAM, etc.
[0015] Further, the SHA-3 unit adopts two extendable output functions XOFS listed in the SHA-3 standard: SHAKE-128 and SHAKE-256. This implementation uses a Keccak core to perform the Keccak permutation within 24 clock cycles. This core is shared by the two hash functions SHAKE128 and SHAKE256 through time-division multiplexing, and is configured as the SHAKE128 or SHAKE256 mode through the mode selection signal;
[0016] The Keccak module includes two parts: a data path and a controller. The data path is composed of units such as logical operations, input / output caches, and selectors, and can complete message padding, Keccak-f, and data absorption and extrusion.
[0017] Further, the multi-channel sampler is combined with the Keccak unit to complete the coefficient generation function. The hash provides bit-stream data of the central binomial distribution, while the sampler is responsible for processing these bit-streams into a data set that obeys a specific distribution.
[0018] After the hash configuration is enabled, the seed is loaded and data is filled, and then iterative operations are performed. After all 24 rounds of iterations are completed, truncation is configured according to the output data and cached. When the generated bit-stream does not meet the preset requirements, the iterative function will be continuously requested for re-operation. The sampler starts the sampling operation when a certain amount of bit-stream is cached to obtain data streams supporting different distributions.
[0019] Furthermore, the data storage unit adopts the method of static storage, and uses RAM to store the pre-computed TW;
[0020] Select the split storage method to implement TW storage, and split the required 5 rotation factors into 3 rotation factors and 2 rotation factors into two parts, and store them in two RAMs;
[0021] Furthermore, the data storage unit adopts the following method to maximize the use efficiency of RAM resources:
[0022] 1. On the premise of ensuring data integrity and security, use the original but data-expired RAM for new data storage strategy;
[0023] 2. Adopt the in-place NTT / INTT structure, which can avoid additional RAM requirements;
[0024] 3. Evaluate in detail the KeyGen, Sign and Verify three stages under the two parameter sets of PARAMS II and PARAMS III, and finally determine to select 5, 10 and 5 RAM groups with different depths to store polynomial data.
[0025] After adopting the above structure, the present invention has the following advantages: The implementation scheme of this design optimizes the operation unit, has lower hardware overhead, improves performance while maintaining resource consumption, and has high area efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 is a schematic structural diagram of a high-performance and configurable Dilithium hardware coprocessor;
[0027] Figure 2 is a schematic structural diagram of a modulo multiplication high-parallel NTT structure of a high-performance and configurable Dilithium hardware coprocessor;
[0028] Figure 3 is a schematic structural diagram of a low-latency radix-4 GS butterfly pipeline structure of a high-performance and configurable Dilithium hardware coprocessor;
[0029] Figure 4 is a schematic structural diagram of a SHA-3 unit of a high-performance and configurable Dilithium hardware coprocessor;
[0030] Figure 5 is a schematic structural diagram of a multi-channel sampler of a high-performance and configurable Dilithium hardware coprocessor;
[0031] Figure 6 It is a schematic diagram of the data storage unit structure of a high-performance and configurable Dilithium hardware coprocessor. Specific implementation mode
[0032] Through in-depth analysis and research on the Dilithium algorithm theory, this design proposes a high-performance, configurable Dilithium hardware coprocessor that supports all security levels. The highlights of the design solution we submitted this time are as follows:
[0033] Optimization of the arithmetic unit. We propose a novel modulo multiplication high-parallel NTT structure in the polynomial arithmetic unit, avoiding the memory write-back conflict problem caused by the deep pipeline design;
[0034] A reconfigurable SHA-3 unit design with optimized round constants is proposed, which supports two functions of SHAKE-128 and SHAKE-256, reduces the area required for separate implementation, and has configurable output bit width;
[0035] A sampler coupled to the hash unit and supporting the central binomial distribution is proposed, avoiding the cache hardware overhead in the sampler;
[0036] By optimizing data storage and introducing an efficient memory allocation mechanism, the BRAM resource consumption is effectively reduced;
[0037] In addition, units such as Makehint and Decompse are optimized to meet the requirements of low resource consumption and high performance scenarios.
[0038] The present invention will be further described in detail below with reference to the accompanying drawings.
[0039] Combined with the attached Figures 1-6 , 1. Modulo multiplication high-parallel NTT
[0040] This structure consists of an address generation unit, a control unit, a rotation factor preprocessing unit, a coefficient transposition unit, a reconfigurable low-latency radix-4 CT / GS unit, and an address synchronization unit (Add_regs). Depending on the adjustment of the mode control signal Mode, this hardware structure can flexibly configure the data paths of NTT transformation, INTT transformation, polynomial addition operation, polynomial subtraction operation, and polynomial multiplication operation, so as to execute different polynomial operations. In addition, through the RAM selection signal, chip selection control of two RAM units can be realized. In particular, for the convenience of explaining the overall polynomial operation process, this design specially introduces three RAM units, including rotation factor RAM, coefficient storage RAM, etc. These are all part of the storage array in the actual hardware implementation. For the simplicity of the illustration of this structure, it is not shown in Figure 2The data path that clearly represents the polynomial addition and subtraction operations is shown. The following sections will detail the functions of each circuit unit and the specific working mechanism of the data path.
[0041] The NTT / INTT operation is the core unit of polynomial multiplication. Among them, the radix-4 CT / GS butterfly operation is the key link of the NTT / INTT operation. Therefore, the operation efficiency of the radix-4 CT / GS butterfly operation unit will directly affect the operation performance of the entire polynomial multiplication. In addition, the formulas of the radix-4 CT butterfly operation and the radix-4 GS butterfly operation are extremely similar. When designing the hardware of the radix-4 CT / GS butterfly operation unit, the characteristics of these two operations should be fully considered to achieve partial reuse of the underlying hardware units, thereby improving the resource utilization efficiency. In summary, this design aims at high performance and low resource consumption, and proposes a compact and multifunctional circuit structure by combining pipeline technology and reconfigurable technology, which supports low-latency radix-4 CT butterfly structure, low-latency radix-4 GS butterfly structure, matrix multiplication structure, polynomial addition structure and polynomial subtraction structure, and realizes function switching through the chip selection signal. First, 5 modular multiplication units, 4 modular addition units, and 4 modular subtraction units are instantiated as the underlying units. On the basis of the above underlying units, reconfigurable technology is used to realize the reuse of the underlying basic units, as Figures 3-6 shown. At the same time, in order to shorten the critical data path and improve the calculation speed, the intermediate data of the calculation steps is isolated hierarchically by inserting registers, and the delays of different operation units are balanced, thereby realizing pipelining. In Figure 3 the low-latency radix-4 CT butterfly pipeline structure shown, the pipeline depth is 9, where the multiplier is implemented with two-stage pipelining, the modular reduction unit is implemented with four-stage pipelining, and both modular addition and modular subtraction are implemented with one-stage pipelining. Therefore Figure 3 the low-latency radix-4 GS butterfly pipeline structure shown has a pipeline depth of 10.
[0042] 2. Reconfigurable SHA-3 Unit Supporting SHAKE-128 and SHAKE-256
[0043] This design adopts two extendable output functions XOFS (SHAKE-128 and SHAKE-256) listed in the SHA-3 standard. These two functions have different bit rates r and capacities c, but their sum is always a fixed value of 1600. Among them, r of SHAKE-128 = 1344, and r of SHAKE-256 = 1088. In order to reduce the overall hardware resource consumption, this implementation uses a Keccak core to perform the Keccak permutation within 24 clock cycles. This core is shared by the two hash functions SHAKE128 and SHAKE256 through time-division multiplexing and can be configured as SHAKE128 or SHAKE256 mode through the mode selection signal.
[0044] The Keccak module consists of two parts: a data path and a controller. The data path is composed of units such as logical operations, input / output caches, and selectors, and can complete message padding, Keccak-f
[1600] , as well as data absorption and extrusion. The controller is designed strictly according to the algorithm flow, and generates timing control signals according to different working modes and external requirements to flexibly adjust the working process of the data path.
[0045] 3. Multi-channel sampler
[0046] The sampler can be combined with the Keccak unit to complete the coefficient generation function. The hash provides bitstream data of the central binomial distribution, while the sampler is responsible for processing these bitstreams into a dataset that follows a specific distribution. After the hash configuration is enabled, the seed is loaded and data is filled, and then iterative operations are performed. After all 24 rounds of iteration are completed, truncation is performed according to the output data configuration, and caching is carried out. When the generated bitstream does not meet the preset requirements, a request for re-operation will be continuously sent to the iterative function. The sampler starts running the sampling operation when a certain amount of the bitstream is cached, and obtains data streams that support different distributions. At the same time, the sampling amount is also controlled by the output threshold, and when the cache amount is insufficient, it cannot be started.
[0047] 4. Data storage unit
[0048] The design consideration is the split storage of the twiddle factor (TW). In different hardware design scenarios for performance requirements, there are two methods for generating TW: dynamic calculation and static storage. This design adopts the static storage method, using RAM to store the pre-calculated TW. Compared with dynamic calculation, it can avoid introducing additional modular multiplication operation resources. Storing 1 TW at each address of the RAM is the simplest and most direct storage method. One low-latency radix-4 CT butterfly operation requires the use of 5 different TWs, which requires 3 36-bit True Dual-Port RAMs (synthesized into 6 18K BARM resources) to increase the system bandwidth and keep the response speed consistent with the system. Adopting the method of storing multiple data at a single address may be a better approach. Storing 5 TWs at a single address, with a total of 5×22 = 110 bits, only requires 1 144-bit Simple Dual-Port RAM (synthesized into 2 36K BARM resources) for storage. However, since the 4 stages of NTT will cycle through 1, 4, 16, and 64 groups of different TWs respectively, storing the above 85 groups of TWs in sequence in the RAM results in a low resource utilization rate of the RAM. The row utilization rate can reach 110 / 144≈76.4%, while the depth utilization rate can only reach 85 / 512≈16.6%. To further reduce the use of RAM resources, this design selects the split storage method to implement TW storage, splitting the required 5 twiddle factors into 3 twiddle factors and 2 twiddle factors into two parts and storing them in two RAMs. The advantage of the above split storage is that the maximum bit width of the twiddle factor to be stored at a single address is only 66. Therefore, the idle addresses of the RAM resources used in the Sign, KeyGen, and Verify stages can be reused for storage. To solve this problem, this design selects to share a RAM for TW2 and w1, and a RAM for TW1 and A. Since w1 and A do not involve NTT / INTT transformations, there will be no port conflicts. The above optimization makes full use of the idle RAM for TW storage, effectively avoiding the introduction of new RAM resources and further improving the resource utilization rate.
[0049] To maximize the utilization efficiency of RAM resources, it is not advisable to solely rely on adding new RAM for storing intermediate polynomials. Instead, this design emphasizes the strategy of using the original but data-expired RAM for storing new data while ensuring data integrity and security. Through this reuse mechanism, the demand for RAM resources can be significantly reduced, further enhancing its utilization efficiency. In addition, this design also deliberately adopts the in-place NTT / INTT structure, which can avoid additional RAM requirements. To meet this requirement, this design has carefully evaluated the KeyGen, Sign, and Verify phases under two parameter sets, PARAMS II and PARAMS III, and finally determined to select 5, 10, and 5 RAM groups with different depths to store polynomial data. Among them, since s1 and s2 need to sample 8 coefficients per clock cycle and need to use two ports simultaneously for writing coefficients, RAM2 and RAM4 are designed as True Dual-Port RAM, which will consume 3 BRAM resources after synthesis, and the other RAMs are all Simple Dual-Port RAM. Table 3-4 shows the RAM allocation mechanism in the KeyGen, Sign, and Verify phases.
[0050] Finally, we implemented the proposed Dilithium optimization scheme on the Xilinx Artix-7 FPGA platform and compared it with the state-of-the-art research. Experiments show that the hardware implementation of the designed Dilithium signature scheme has lower hardware overhead.
[0051] The above describes the present invention and its implementation manners, and this description is not restrictive. The actual structure is not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
Claims
1. A high-performance, configurable Dilithium hardware coprocessor, characterized by: The optimized design of the operation unit includes the following structures: a highly parallel NTT structure for modular multiplication, a reconfigurable SHA-3 unit supporting SHAKE-128 and SHAKE-256, a multi-channel sampler, and a data storage unit.
2. A high performance, configurable Dilithium hardware coprocessor according to claim 1, characterized in that: The modular multiplication high parallel NTT structure is composed of an address generation unit, a control unit, a rotation factor preprocessing unit, a coefficient transposition unit, a reconfigurable low-latency base-4CT / GS unit, and an address synchronization unit; By adjusting the mode control signal Mode, the hardware structure can flexibly configure the data paths of NTT transformation, INTT transformation, polynomial addition, polynomial subtraction and polynomial multiplication, so as to perform different polynomial operations. In addition, the chip selection control of the two RAM units can be realized through the RAM selection signal; In particular, three RAM units are introduced, including a rotation factor RAM, a coefficient storage RAM, and the like.
3. A high performance, configurable Dilithium hardware coprocessor according to claim 1, characterized in that: The SHA-3 unit uses two scalable output functions XOFS listed in the SHA-3 standard: SHAKE-128 and SHAKE-256. The implementation uses a Keccak core to perform Keccak permutations within 24 clock cycles. The core is shared by the two hash functions SHAKE128 and SHAKE256 in a time-division multiplexing manner and is configured to SHAKE128 or SHAKE256 mode through a mode selection signal. The Keccak module consists of two parts: datapath and controller. The datapath consists of logic operations, input and output buffers, selectors and other units, which can complete message filling, Keccak-f, and data absorption and extrusion.
4. A high performance, configurable Dilithium hardware coprocessor according to claim 3, characterized in that: The multi-channel sampler and the Keccak unit are combined to complete the coefficient generation function. The hash provides the bit stream data of the central binomial distribution, and the sampler is responsible for processing these bit streams into data sets that obey a specific distribution; After the hash configuration is enabled, the seed is loaded and the data is filled, and then iterative operations are performed. After all 24 rounds of iterations are completed, the output data is truncated according to the configuration and cached. When the generated bit stream does not meet the preset requirements, it will continue to request the iterative function to calculate again; the sampler starts running the sampling operation when a certain amount of bit stream is cached to obtain a data stream that supports different distributions.
5. A high performance, configurable Dilithium hardware coprocessor according to claim 1, characterized in that: The data storage unit adopts a static storage method and uses RAM to store TW calculated in advance; Select the split storage method to implement TW storage, splitting the required 5 rotation factors into 3 rotation factors and 2 rotation factors Two parts and stored in two RAMs.
6. A high performance, configurable Dilithium hardware coprocessor according to claim 5, characterized in that: The data storage unit adopts the following method to maximize the efficiency of RAM resource usage:
1. A strategy to use the original RAM with invalid data to store new data while ensuring data integrity and security; 2. Use in-place NTT / INTT structure, which can avoid additional RAM requirements; 3. Detailed evaluation of the three stages of KeyGen, Sign and Verify under the two parameter sets of PARAMS II and PARAMS III, and finally the selection of 5, 10 and 5 RAM groups with different depths to store polynomial data.