A lattice cipher co-processor circuit compatible with Kyber and Dilithium algorithms

By designing a grid-code coprocessor circuit that is compatible with Kyber and Dilithium algorithms, the connection of the modular computing unit is optimized, solving the problem of low operation efficiency of Kyber and Dilithium algorithms on the same processor, and achieving efficient and flexible hardware resource utilization.

CN118963700BActive Publication Date: 2025-07-22HUAZHONG UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410964710.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2025-07-22
Estimated Expiration
2044-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently realize the Kyber and Dilithium algorithms running efficiently on the same processor, and the implementation of the modular computing hardware is complex, resulting in high resource overhead and increased design costs.

Method used

A grid cryptographic coprocessor circuit compatible with Kyber and Dilithium algorithms is designed, including finger fetching units, hashing units, polynomial units and storage units. It adopts a polynomial operation array and a multiplexable mode multiplier to optimize the connection of the operation unit and supports algorithmic operations of multiple security levels.

Benefits of technology

It realizes efficient operation of Kyber and Dilithium algorithms on the same processor, improves hardware efficiency and flexibility, supports algorithm processes of multiple security levels, and reduces hardware resource overhead and design costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118963700B_ABST
    Figure CN118963700B_ABST
Patent Text Reader

Abstract

The present invention discloses a lattice cipher co-processor circuit compatible with Kyber and Dilithium algorithms, including: an instruction fetch unit, a hash unit, a polynomial unit, and a storage unit; the instruction fetch unit is used to obtain instructions transmitted from the outside and cause the hash unit and the polynomial unit to operate respectively according to the instructions; the hash unit is used to generate polynomial data of Kyber algorithm and Dilithium algorithm; the polynomial unit is used to accelerate the operation of the polynomial data; the storage unit is used to store various data output by the hash unit and the polynomial unit. By designing the hash unit and the polynomial unit in the embodiments of the present invention, the hash unit and the polynomial unit can operate Kyber algorithm and Dilithium algorithm, so that the embodiments of the present invention can not only support three different security-level Kyber key encapsulation processes, but also support Dilithium digital signature processes. While providing high hardware efficiency, it maintains flexible programmability and has good application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of post-quantum information security algorithms, digital signal processing, and circuit implementation, and particularly relates to a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms. Background Art

[0002] The explosive growth of the computing performance of quantum computers poses a huge threat to the existing security infrastructure based on traditional public-key cryptosystems. In order to resist the brute-force attacks of gradually mature quantum computers, the National Institute of Standards and Technology (NIST) in the United States has been soliciting proposals for post-quantum cryptography (PQC) globally since 2016. Among them, Kyber and Dilithium are respectively a key encapsulation mechanism (KEM) based on the M-LWE (Module Learning-With-Errors) lattice problem and a digital signature (DS) algorithm based on the M-SIS (Module Short-Integer-Solution) lattice problem.

[0003] The operations of lattice-based post-quantum cryptography in the ring domain cover operations such as modular addition, subtraction, and multiplication between polynomials, matrices, and vectors, and are mainly realized by performing multiple basic modular operations between coefficients. The hardware implementation results of modular operations directly affect the resource overhead and operation performance of the entire post-quantum cryptosystem. If the modular domain operation functions are implemented separately for each algorithm, it will consume a large amount of hardware structures and bring higher design costs.

[0004] At the same time, the third-generation hash operation widely used in Kyber and Dilithium algorithms involves input and output formats with different bit widths, which easily leads to complex hardware connections and reduces the operating efficiency of the system. Therefore, it is very important to refine common core operators, disassemble different operation steps, and design a reconfigurable hash unit and polynomial operation unit with multi-mode fusion through schemes such as dividing key timing paths, sharing operation data paths, and optimizing the connection relationship of operation units to achieve breakthroughs in computing performance and higher hardware efficiency. Summary of the Invention

[0005] The technical problem to be solved by the present invention is that, in order to enable the Kyber and Dilithium algorithms to run efficiently and harmoniously on the same processor, the present invention provides a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms.

[0006] To solve the above technical problems, an embodiment of the present invention provides a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms, including: an instruction fetch unit, a hash unit, a polynomial unit, and a storage unit;

[0007] The instruction fetch unit is used to obtain instructions transmitted from the outside and cause the hash unit and the polynomial unit to operate respectively according to the instructions;

[0008] The hash unit is used to generate polynomial data of Kyber algorithm and Dilithium algorithm;

[0009] The polynomial unit is used to accelerate the operation of the polynomial data;

[0010] The storage unit is used to store various data output by the hash unit and the polynomial unit;

[0011] Wherein, the polynomial unit includes a plurality of polynomial operation arrays; the polynomial operation arrays include: 9 polynomial data acceleration operation methods of c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b.

[0012] Preferably, the connection mode of the plurality of polynomial operation arrays in the polynomial unit is parallel connection;

[0013] Wherein the plurality of polynomial operation arrays are used to divide the polynomial data into a plurality of polynomial series data, and each polynomial operation array accelerates each polynomial series data.

[0014] Preferably, the polynomial operation array includes: a first memory, a second memory, a third memory, a first selector, a second selector, a third selector, a fourth selector, a fifth selector, a sixth selector, a seventh selector, an eighth selector, a modular adder, a modular subtractor, a first modular divider, a second modular divider, and a reusable modular multiplier;

[0015] Among them, the first memory input end is connected to the output end of the first selector; the first memory output end is connected to the second selector, the third selector, the fourth selector, and the input end of the first modulo divider; the output end of the second selector is connected to the third selector and the input end of the modulo adder; the output end of the third selector is connected to the input end of the modulo subtractor; the output end of the first modulo divider is connected to the input end of the fourth selector; the output end of the modulo adder is connected to the input ends of the first selector and the fourth selector; the input end of the reusable modular multiplier is connected to the output ends of the second memory and the third memory; the input end of the second memory is connected to the output end of the fifth selector; the output end of the reusable modular multiplier is connected to the input ends of the sixth selector, the seventh selector, the eighth selector, and the second modulo divider; the output end of the sixth selector is connected to the input end of the modulo subtractor; the output end of the seventh selector is connected to the input end of the modulo adder; the output end of the modulo subtractor is connected to the input ends of the fifth selector and the eighth selector; the output end of the second modulo divider is connected to the input end of the eighth selector.

[0016] Preferably, the polynomial operation array includes four input ports, which are respectively used for inputting data a, b, c, and d; the selection ends of the first selector, the second selector, the third selector, the fourth selector, the fifth selector, the sixth selector, the seventh selector, and the eighth selector are connected to the instruction fetch unit, and are used to determine the data at the input ends of the first selector, the second selector, the third selector, the fourth selector, the fifth selector, the sixth selector, the seventh selector, and the eighth selector to be output.

[0017] Among them, the data a is connected to the input end of the third memory; the data b is connected to the input end of the fifth selector; the data c is connected to the input ends of the first selector and the second selector; the data d is connected to the input ends of the sixth selector and the seventh selector; through the instruction of the instruction fetch unit, the polynomial operation array outputs the nine kinds of polynomial data acceleration operation results of c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b through the fourth selector and the eighth selector.

[0018] Preferably, the reusable modular multiplier is used to perform modular multiplication operations with moduli of 3329 and 8380417; the reusable modular multiplier includes a reusable modular reducer and a Kyber modular reducer; the reusable modular reducer completes 1 modular multiplication operation with a modulus of 3329 or 1 modular multiplication operation with a modulus of 8380417 per cycle; the Kyber modular reducer completes 1 modular multiplication operation with a modulus of 3329 per cycle; the reusable modular reducer and the Kyber modular reducer are connected in parallel.

[0019] Preferably, the multiplexing modularizer includes: a common part, a Kyber part, and a Dilithium part;

[0020] The common part includes a first adder, a second adder, a first subtractor, and a fourth memory; the output ends of the first adder and the first subtractor are connected to the input end of the second adder; the output end of the second adder is connected to the input end of the fourth memory;

[0021] The Dilithium part includes a first accumulator, a first subtracter, a second subtractor, a third subtractor, a first shifter, a fifth memory, and a sixth memory; the output ends of the first accumulator and the first subtracter are connected to the input end of the second subtractor; the output end of the second subtractor is connected to the input end of the first shifter; the output ends of the first shifter and the fourth memory are connected to the input end of the fifth memory; the output end of the fifth memory is connected to the sixth memory; wherein the shifter is used to multiply the input data by a multiple of the number 2;

[0022] The Kyber part includes a second accumulator, a second subtracter, a fourth subtractor, a fifth subtractor, a second shifter, a seventh memory, and an eighth memory; the output end of the second accumulator is connected to the input end of the second shifter; the output ends of the second shifter and the second subtracter are connected to the input end of the fourth subtractor; the output end of the fourth subtractor is connected to the input end of the seventh memory; the output ends of the seventh memory and the fourth memory are connected to the input end of the fifth subtractor; the output end of the fifth subtractor is connected to the input end of the eighth memory.

[0023] Preferably, the Kyber modularizer includes a plurality of registers; the registers are used to store all the modular values after splitting the output value of the multiplication operation of the reusable modular multiplier, and select the modular value stored in the register through the input value; the Kyber modularizer obtains the modular multiplication operation result with a modulus of 3329 by performing an addition operation on the modular values output by the registers.

[0024] Preferably, the hash unit includes a rejection sampler, a central sampler, an extended mask sampler, and an in-sphere sampler; the rejection sampler, the central sampler, the extended mask sampler, and the in-sphere sampler are connected in parallel; the rejection sampler has a parallelism of 4; the central sampler has a parallelism of 16; the extended mask sampler has a parallelism of 8;

[0025] Among them, the rejection sampler and the central sampler are used to generate polynomial data of the Kyber algorithm; the rejection sampler, the extended mask sampler, and the in-sphere sampler are used to generate polynomial data of the Dilithium algorithm.

[0026] Preferably, the instruction fetching unit includes a hash unit control area and a polynomial unit control area; the hash unit control area and the polynomial unit control area are connected in parallel;

[0027] The hash unit control area is used to control the operation of the hash unit according to an externally input instruction, and after the hash unit completes one operation, it feeds back to the outside world to obtain the next instruction;

[0028] The polynomial unit control area is used to control the operation of the polynomial unit according to an externally input instruction, and after the polynomial unit completes one operation, it feeds back to the outside world to obtain the next instruction.

[0029] Preferably, the storage unit includes two independent memories and three true dual-port memories;

[0030] The true dual-port memory is used to store the intermediate coefficients generated by the hash unit and the polynomial unit;

[0031] The independent memory is used to store the Kyber NTT rotation factors and Dilithium NTT rotation factors generated by the polynomial unit.

[0032] Implementing the embodiments of the present invention has the following beneficial effects:

[0033] (1) In the embodiments of the present invention, by adopting a polynomial unit capable of performing 9 kinds of polynomial data acceleration operation methods, namely c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b, all the polynomial data acceleration operation methods of the Kyber algorithm and the Dilithium algorithm are summarized. And by adopting a hash unit capable of generating the polynomial data of the Kyber algorithm and the Dilithium algorithm, the embodiments of the present invention can not only support the Kyber key encapsulation process with three different security levels, but also support the Dilithium digital signature process. While providing high hardware efficiency, it maintains flexible programmability and has good application prospects.

[0034] (2) In the embodiments of the present invention, by designing a reusable modular multiplier, modular multiplication operations with moduli of 3329 and 8380417 can be performed. The modular multiplication operations of the Kyber algorithm and the Dilithium algorithm are satisfied in one modular multiplier, greatly improving the operation efficiency and flexibility of the polynomial data in the polynomial unit.

[0035] (3) By designing the instruction fetch unit in the embodiments of the present invention, the embodiments of the present invention can perform two independent decode-execute-writeback pipelines, namely, the instruction fetch unit - hash unit - storage unit and the instruction fetch unit - polynomial unit - storage unit. When executing different instructions or algorithms, the corresponding data paths and coefficient storage modes will adaptively change to make the most efficient use of the computing bandwidth of the coprocessor. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0037] Figure 1 It is a circuit architecture diagram of a lattice cipher coprocessor compatible with Kyber and Dilithium algorithms provided by the embodiments of the present invention;

[0038] Figure 2 It is an architecture diagram of a polynomial operation array in a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms provided by the embodiments of the present invention;

[0039] Figure 3 It is an architecture diagram of a reusable modular multiplier in a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms provided by the embodiments of the present invention;

[0040] Figure 4 It is a schematic diagram of the operation of three true dual-port memories in a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms provided by the embodiments of the present invention;

[0041] Figure 5 It is a schematic diagram of the operation of two independent memories in a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0043] Such as Figure 1As shown in the figure, a lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms disclosed in this embodiment includes: an instruction fetch unit 10, a hash unit 20, a polynomial unit 30, and a storage unit 40. The instruction fetch unit 10 is used to obtain instructions transmitted by the external ICB bus and make the hash unit 20 and the polynomial unit 30 operate respectively according to the instructions. So that the lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms can perform fine-grained operations of hash sampling or polynomial operations. The hash unit 20 is used to generate polynomial data of Kyber algorithm and Dilithium algorithm. The polynomial unit 30 is used to accelerate the operation of the polynomial data. The storage unit 40 is used to store various types of data output by the hash unit 20 and the polynomial unit 30.

[0044] The input end of the instruction fetch unit 10 is connected to the external ICB bus and is tightly coupled with an external RISC-V processor through the ICB bus, and is used to receive R-Type instructions in the RV32I instruction set and support the ICB bus protocol. The output end of the instruction fetch unit 10 is connected to the hash unit 20 and the polynomial unit 30. Among them, the hash unit 20 and the polynomial unit 30 are connected in parallel, that is, the hash unit 20 and the polynomial unit 30 do not interfere with each other and are directly controlled by the same previous device. The output ends of the hash unit 20 and the polynomial unit 30 are connected to the storage unit 40, and various types of data output by the hash unit 20 and the polynomial unit 30 are stored through the storage unit 40.

[0045] Among them, the instruction fetch unit 10 includes a pre-decoding area (IF), a hash unit control area 110, and a polynomial unit control area 120. The hash unit control area 110 and the polynomial unit control area 120 are connected in parallel. The pre-decoding area is used to pre-decode the R-Type instruction to determine whether the R-Type instruction enters the hash unit control area 110 or the polynomial unit control area 120.

[0046] The hash unit control area 110 is used to control the operation of the hash unit 20 according to an externally input instruction, and after the hash unit 20 completes one operation, it feeds back to the outside world to obtain the next instruction. Specifically, the hash unit control area 110 includes: a hash decoder (Hash Decoder) and a hash address controller (HashAddr Control). The pre-decoding area pre-decodes the R-Type instruction to obtain the pre-decoded instructions SHAInst. and Poly Inst. The pre-decoded instruction SHAInst. will be transmitted into the hash decoder for decoding, and the 15-bit address numbers rs / rd are obtained and transmitted to the hash address controller. The hash address controller will control the operation of the hash unit 20 through the 15-bit address numbers rs / rd, and after the hash unit 20 completes one operation, it feeds back to the outside world to obtain the next instruction.

[0047] The polynomial unit control area 120 is used to control the operation of the polynomial unit according to an externally input instruction, and after the polynomial unit completes one operation, it feeds back to the outside world to obtain the next instruction. The structure of the polynomial unit control area 120 is similar to that of the hash unit control area 110, and includes: a polynomial decoder (Poly Decoder) and a polynomial address controller (PolyAddr Control). The functions of the polynomial decoder (Poly Decoder) and the polynomial address controller (PolyAddr Control) are the same as those of the hash decoder (Hash Decoder) and the hash address controller (HashAddr Control), and will not be introduced one by one here. Finally, the polynomial address controller will control the operation of the polynomial unit 30 through the 15-bit address numbers rs / rd, and after the polynomial unit 30 completes one operation, it feeds back to the outside world to obtain the next instruction. The lattice cipher coprocessor circuit compatible with the Kyber and Dilithium algorithms can perform two independent decode-execute-write-back pipelines of fetch unit-hash unit-storage unit and fetch unit-polynomial unit-storage unit through the design of the fetch unit 10. When executing different instructions or algorithms, the corresponding data paths and coefficient storage modes will adaptively change to make the most efficient use of the computing bandwidth of the coprocessor.

[0048] The hash unit 20 includes a Keccak core 210, a temporary memory (FIFO), a rejection sampler (UniformSampler) 220, a central sampler (CBD Sampler) 230, an expand mask sampler (Expand Mask) 240, and a sample in ball sampler (Sample In Ball) 250. The input end of the Keccak core 210 is connected to the hash address controller and operates by transmitting data through the hash address controller. The output end of the Keccak core 210 is connected to the temporary memory, and the temporary memory temporarily stores the output data of the Keccak core 210. The rejection sampler 220, the central sampler 230, the expand mask sampler 240, and the sample in ball sampler 250 are connected in parallel. The rejection sampler has a parallelism of 4. The central sampler has a parallelism of 16. The expand mask sampler has a parallelism of 8. The sample in ball sampler 250 is independent. The output end of the temporary memory is connected to the input ends of the rejection sampler 220, the central sampler 230, the expand mask sampler 240, and the sample in ball sampler 250, and the rejection sampler 220, the central sampler 230, the expand mask sampler 240, and the sample in ball sampler 250 output the output data of the Keccak core 210. After the rejection sampler 220, the central sampler 230, the expand mask sampler 240, and the sample in ball sampler 250 complete sampling, under the action of the selector of the MenArbitrator, the polynomial data of the Kyber algorithm and the Dilithium algorithm to be generated are determined and output to the storage unit 40. Thus, one operation of the hash unit 20 is completed.

[0049] Among them, the rejection sampler 220 and the central sampler 230 are used to generate the polynomial data of the Kyber algorithm. The rejection sampler 220, the expand mask sampler 240, and the sample in ball sampler 250 are used to generate the polynomial data of the Dilithium algorithm. The operation of the rejection sampler 220, the central sampler 230, the expand mask sampler 240, and the sample in ball sampler 250 is regulated by the hash address controller.

[0050] The polynomial unit (PLU) 30 includes a plurality of polynomial operation arrays 310. In this embodiment, there are 8 polynomial operation arrays 310, but this does not limit the number of the polynomial operation arrays 310 in the present invention. The number of the polynomial operation arrays 310 depends on specific circumstances. The polynomial operation array 310 includes 9 polynomial data acceleration operation methods: c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b. The 9 polynomial data acceleration operation methods summarize all the polynomial data acceleration operation methods of the Kyber algorithm and the Dilithium algorithm, enabling the polynomial unit 30 to be compatible with the polynomial data acceleration operations of the Kyber algorithm and the Dilithium algorithm. The connection mode of the multiple polynomial operation arrays 310 in the polynomial unit 30 is parallel connection. The multiple polynomial operation arrays 310 are used to divide the polynomial data into multiple polynomial series data, and each polynomial operation array 310 accelerates each polynomial series data. The polynomial series data is one column or multiple columns of the polynomial data, which is formed by evenly dividing the polynomial data into columns. In this embodiment, there are 8 polynomial operation arrays 310. The polynomial data will be evenly divided into 8 columns, and each column of polynomial data is a polynomial series data. Every two of the polynomial operation arrays 310 form a group, and the two polynomial operation arrays 310 in a group of polynomial operation arrays 310 respectively process the polynomial series data of odd and even sequences.

[0051] The storage unit 40 includes three true dual-port memories (RAM1-3) 410 and two independent memories (K / DROM) 420. The true dual-port memory 410 is used to store the intermediate coefficients generated by the hash unit 20 and the polynomial unit 30. The independent memory 420 is used to store the Kyber NTT rotation factors and Dilithium NTT rotation factors generated by the polynomial unit. The Sample RAM1 in the true dual-port memory 410 is mainly used to store the coefficients sampled by the hash unit 20, and the Poly RAM2 / 3 is mainly used to store the intermediate values in the polynomial operations. The DROM is used to store the Dilithium NTT rotation factors. The KROM is used to store the Kyber NTT rotation factors.

[0052] Please refer to Figure 2, the polynomial operation array 310 includes: a first memory 301, a second memory 302, a third memory 303, a first selector 304, a second selector 305, a third selector 306, a fourth selector 307, a fifth selector 308, a sixth selector 309, a seventh selector 311, an eighth selector 312, a modulo adder 313, a modulo subtractor 314, a first modulo divider 315, a second modulo divider 316, and a reusable modulo multiplier 317.

[0053] Among them, the modulo adder is used for modulo addition operations. The modulo subtractor is used for modulo subtraction operations. The modulo divider is used for modulo division operations. The reusable modulo multiplier is used for modulo multiplication operations.

[0054] The input end of the first memory 301 is connected to the output end of the first selector 304. The output end of the first memory 301 is connected to the input ends of the second selector 305, the third selector 306, the fourth selector 307, and the first modulo divider 315. The output end of the second selector 305 is connected to the input ends of the third selector 306 and the modulo adder 313. The output end of the third selector 306 is connected to the input end of the modulo subtractor 314. The output end of the first modulo divider 315 is connected to the input end of the fourth selector 307. The output end of the modulo adder 313 is connected to the input ends of the first selector 304 and the fourth selector 307. The input end of the reusable modulo multiplier 317 is connected to the output ends of the second memory 302 and the third memory 303. The input end of the second memory 302 is connected to the output end of the fifth selector 308. The output end of the reusable modulo multiplier 317 is connected to the input ends of the sixth selector 309, the seventh selector 311, the eighth selector 312, and the second modulo divider 316. The output end of the sixth selector 309 is connected to the input end of the modulo subtractor 314. The output end of the seventh selector 311 is connected to the input end of the modulo adder 313. The output end of the modulo subtractor 314 is connected to the input ends of the fifth selector 308 and the eighth selector 312. The output end of the second modulo divider 316 is connected to the input end of the eighth selector 312. The selection ends of the first selector 304, the second selector 305, the third selector 306, the fourth selector 307, the fifth selector 308, the sixth selector 309, the seventh selector 311, and the eighth selector 312 are connected to the fetch unit 10. The sel[9:0] selection signal output through the fine-grained operation of the fetch unit 10 is used to determine the data at the input ends of the first selector 304, the second selector 305, the third selector 306, the fourth selector 307, the fifth selector 308, the sixth selector 309, the seventh selector 311, and the eighth selector 312. That is, the 7-bit Poly_func instruction output by the polynomial address controller controls the polynomial operation array 310 to output the 9-bit sel[9:0] selection signal. The sel[9:0] selection signal is distributed at the selection ends of 8 selectors. When there are only two input data at the input end of the selector, the sel[9:0] selection signal splits out 1 bit signal to control the selection. One of the two input data is selected through the only 0 / 1 in the 1 bit signal. When there are only three input data at the input end of the selector, the sel[9:0] selection signal splits out 2 bit signals to control the selection. One of the three input data is selected through 00 / 01 / 11 in the 2 bit signals. Among them, the first memory 301 includes five memories, and the five memories are connected in series to form the first memory 301.

[0055] The polynomial operation array 310 includes four input ports, which are respectively used for inputting data a, b, c, and d. Among them, the data a is connected to the input end of the third memory 303. The data b is connected to the input end of the fifth selector 308. The data c is connected to the input ends of the first selector 304 and the second selector 305. The data d is connected to the input ends of the sixth selector 309 and the seventh selector 311. The polynomial operation array 310 outputs nine kinds of polynomial data acceleration operation results, namely c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b, through the instruction of the instruction fetch unit, that is, the sel[9:0] selection signal, by the fourth selector 307 and the eighth selector 312.

[0056] See Figure 3 , the reusable modular multiplier 317 is used to perform modular multiplication operations with moduli of 3329 and 8380417. The modular multiplication operations of the Kyber algorithm and the Dilithium algorithm are satisfied in one modular multiplier. The reusable modular multiplier 317 includes a multiplication operation part, a reusable modular reducer 320, and a Kyber modular reducer 330. The multiplication operation part is used to perform multiplication operations on the input data. The reusable modular reducer 320 and the Kyber modular reducer 330 are used to perform modular operations on the multiplication operation results. The reusable modular reducer 320 completes 1 modular multiplication operation with a modulus of 3329 or 1 modular multiplication operation with a modulus of 8380417 per cycle. The Kyber modular reducer 330 completes 1 modular multiplication operation with a modulus of 3329 per cycle. The reusable modular reducer and the Kyber modular reducer are connected in parallel.

[0057] The reusable modular reducer 320 includes: a common part, a Kyber part, and a Dilithium part. The common part includes a first adder 318, a second adder 321, a first subtractor 319, and a fourth memory 322. The output ends of the first adder 318 and the first subtractor 319 are connected to the input end of the second adder 321. The output end of the second adder 321 is connected to the input end of the fourth memory 322. The adder is used to perform addition operations. The subtractor is used to perform subtraction operations.

[0058] The Dilithium part includes a first accumulator 323, a first subtractor 324, a second subtractor 325, a third subtractor 326, a first shifter 327, a fifth memory 328, and a sixth memory 329. The output ends of the first accumulator 323 and the first subtractor 324 are connected to the input end of the second subtractor 325. The output end of the second subtractor 325 is connected to the input end of the first shifter 327. The output ends of the first shifter 327 and the fourth memory 322 are connected to the input end of the fifth memory 328. The output end of the fifth memory 328 is connected to the sixth memory 329. Among them, the shifter is used to multiply the input data by a multiple of the number 2, so that the input data is shifted to the left by the number of digit positions of the multiple in the multiple of the number 2.

[0059] Among them, the accumulator is used to perform an accumulation operation on the input data. The subtractor is used to perform a subtraction operation on the input data.

[0060] The Kyber part includes a second accumulator 331, a second subtractor 332, a fourth subtractor 333, a fifth subtractor 334, a second shifter 335, a seventh memory 336, and an eighth memory 337. The output end of the second accumulator 331 is connected to the input end of the second shifter 335. The output ends of the second shifter 335 and the second subtractor 332 are connected to the input end of the fourth subtractor 333. The output end of the fourth subtractor 333 is connected to the input end of the seventh memory 336. The output ends of the seventh memory 336 and the fourth memory 322 are connected to the input end of the fifth subtractor 334. The output end of the fifth subtractor 334 is connected to the input end of the eighth memory 337.

[0061] The multiplexing modulo reducer 320 utilizes the modular property 2 of Dilithium 23 ≡2 13 -1 to split the modulo reduction process:

[0062] a = 2 23 a[45:23]+a[22:0]

[0063] = 2 13 a[45:23]-a[45:23]+a[22:0]

[0064] = 2 23 a[45:33]+2 13 a[32:23]-a[45:23]+a[22:0]

[0065] = 2 13 (a[32:23]+a[42:33]+a[45:43])

[0066] -a[45:23]-a[45:33]-a[45:43]+a[22:0]

[0067] Among them, a in the above splitting has no relation to the input data a in the four input ports of the polynomial operation array 310. The final range of a is within [-q, 2q]. The q is the modulus, q = 8380417. Therefore, only one additional -q judgment operation is required. Similarly, the modulus of Kyber also has the property of 2 12 ≡2 9 +2 8 -1. After adopting the same modular reduction splitting method, a modular reduction operation expression based on shift addition is obtained, which will not be listed one by one here. Among them, a part of the operation structures in the modular reduction operation expressions of the Kyber algorithm and the Dilithium algorithm are the same. Therefore, this part can be reused to form the shared part (Reused MR). The remaining parts are the Kyber part (KyberMR) and the Dilithium part (Dilithium MR) respectively.

[0068] The Kyber modular reducer 330 includes a plurality of registers 338. The registers 338 are used to store all the modular reduction values after splitting the output value of the multiplication operation of the reusable multiplier, and select and output the modular reduction values stored in the registers through the input values. The Kyber modular reducer 330 obtains the modular multiplication operation result with the modulus of 3329 by performing an addition operation on the modular reduction values output by the registers 338. Specifically, the Kyber modular reducer 330 divides the high 12-bit of the multiplication operation result output by the multiplication operation part into three groups by using the divide-and-conquer modular reduction algorithm. Since each group is 4-bit data, each group has 16 possible modular reduction values. These possible modular reduction values are pre-stored in a plurality of registers 338, that is, stored in ROM_H, ROM_M, and ROM_L. The input multiplication operation result serves as the input value of the registers 338. By adding four groups of modular reduction values to obtain a 14-bit result sum, and then repeatedly using the divide-and-conquer method to obtain the high 2-bit modular reduction value from the registers 338, that is, the high 2-bit modular reduction value in ROM_C. Finally, by performing a judgment of subtracting the modulus of 3329 and modular addition on the low 12-bit data of the result sum, the modular multiplication operation of a group of Kyber coefficients can be completed.

[0069] The reusable multiplier 317 can complete the modular multiplication operation of 2 pairs of Kyber coefficients or 1 pair of Dilithium coefficients per cycle. The modular multiplication operations of the Kyber algorithm and the Dilithium algorithm are satisfied in one multiplier, greatly improving the operation efficiency and flexibility of the polynomial data in the polynomial unit.

[0070] Specifically, the Kyber modular reducer 330 splits the modular reduction process using a divide-and-conquer modular reduction algorithm:

[0071] Input: a, b ∈ [0, q], q = 3329.

[0072] Output: M = a × b mod q.

[0073] 1: c[23:0] = a[11:0] × b[11:0]

[0074] 2: ROM H = c[23:20] × 2 20 mod q

[0075] ROM M = c[19:16] × 2 16 mod q

[0076] ROM L = c[15:12] × 2 12 mod q

[0077] 3: sum[13:0] = (ROM H + ROM M ) + (ROM L + c[11:0])

[0078] 4: ROM c = sum[13:12] × 2 12 mod q

[0079] 5: M = ROM C + (sum[11:0] mod q) mod q

[0080] See Figure 4 、 Figure 5 There are multiple memory access routes between the storage unit 40 and the hash unit 20 and the polynomial unit 30. Among them, the three true dual-port memories 410 are Sample RAM1 (96 × 1024), Poly RAM2 (96 × 2048), and Poly RAM3 (96 × 1024) respectively. Each row of the three true dual-port memories 410 can store 8 Kyber coefficients or 4 Dilithium coefficients. The lattice cryptography coprocessor circuit compatible with the Kyber and Dilithium algorithms adopts an odd-even coefficient order storage strategy to store the coefficients of Kyber or Dilithium in the dual-port memory 410.

[0081] Since the polynomial data dimensions of both the Kyber algorithm and the Dilithium algorithm are 256, every 256 / x (where x is the number of coefficients stored in a row) memory addresses can be divided into a vector address (Bank), and the instruction operands point to these vector addresses. At the same time, the hash unit 20 and the polynomial unit 30 are matched with the bandwidth of the true dual-port memory 410. Considering the number of polynomial operation arrays 310, the Dilithium NTT rotation factors are stored through the DROM and the Kyber NTT rotation factors are stored through the KROM. For example, when the bit width of the K / DROM is 96-bit, a polynomial operation array 310 with a parallelism of 8 is used, and polynomial operation arrays 310 with other parallelisms cannot be used.

[0082] The above-disclosed is only a preferred embodiment of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms, characterized in that, Including: An instruction fetch unit, a hash unit, a polynomial unit, and a storage unit; The instruction fetch unit is used to obtain instructions transmitted from the outside and cause the hash unit and the polynomial unit to operate respectively according to the instructions; The hash unit is used to generate polynomial data of the Kyber algorithm and the Dilithium algorithm; The polynomial unit is used to accelerate the operation of the polynomial data; The storage unit is used to store various types of data output by the hash unit and the polynomial unit; Among them, the polynomial unit includes a plurality of polynomial operation arrays; the polynomial operation arrays include: 9 polynomial data acceleration operation methods of c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b; The polynomial operation array includes: a first memory, a second memory, a third memory, a first selector, a second selector, a third selector, a fourth selector, a fifth selector, a sixth selector, a seventh selector, an eighth selector, a modular adder, a modular subtractor, a first modular divider, a second modular divider, and a reusable modular multiplier; Among them, the input end of the first memory is connected to the output end of the first selector; the output end of the first memory is connected to the input ends of the second selector, the third selector, the fourth selector, and the first modular divider; the output end of the second selector is connected to the input ends of the third selector and the modular adder; the output end of the third selector is connected to the input end of the modular subtractor; the output end of the first modular divider is connected to the input end of the fourth selector; the output end of the modular adder is connected to the input ends of the first selector and the fourth selector; the input end of the reusable modular multiplier is connected to the output ends of the second memory and the third memory; the output end of the reusable modular multiplier is connected to the input ends of the sixth selector, the seventh selector, the eighth selector, and the second modular divider; the output end of the sixth selector is connected to the input end of the modular subtractor; the output end of the seventh selector is connected to the input end of the modular adder; the output end of the modular subtractor is connected to the input ends of the fifth selector and the eighth selector; the output end of the second modular divider is connected to the input end of the eighth selector.

2. The lattice cipher co-processor circuit compatible with Kyber and Dilithium algorithms according to claim 1, characterized in that, The connection method of the multiple polynomial operation arrays in the polynomial unit is parallel connection; Among them, the multiple polynomial operation arrays are used to divide the polynomial data into multiple polynomial series data, and each polynomial operation array accelerates each polynomial series data.

3. The lattice cryptography coprocessor circuit compatible with Kyber and Dilithium algorithms according to claim 1, characterized in that, The polynomial operation array includes four input ports, which are respectively used to input data a, b, c, and d; the selection ends of the first selector, the second selector, the third selector, the fourth selector, the fifth selector, the sixth selector, the seventh selector, and the eighth selector are connected to the instruction fetch unit, and are used to determine the data input to the input ends of the first selector, the second selector, the third selector, the fourth selector, the fifth selector, the sixth selector, the seventh selector, and the eighth selector. Among them, the data a is connected to the third memory input terminal; the data b is connected to the fifth selector input terminal; the data c is connected to the input terminals of the first selector and the second selector; the data d is connected to the input terminals of the sixth selector and the seventh selector; through the instruction of the instruction fetch unit, the polynomial operation array outputs, through the fourth selector and the eighth selector, the nine kinds of polynomial data acceleration operation results of c + a·b, c - a·b, a·b, 1 / 2(c + d), c + d, 1 / 2(c - d)·a, (c - d)·a, c - d, and c + d - a·b.

4. The lattice cipher co-processor circuit compatible with Kyber and Dilithium algorithms according to claim 1, wherein The reusable modular multiplier is used to perform modular multiplication operations with moduli of 3329 and 8380417; the reusable modular multiplier includes a reusable modular reducer and a Kyber modular reducer; the reusable modular reducer completes 1 modular multiplication operation with a modulus of 3329 or 1 modular multiplication operation with a modulus of 8380417 per cycle; the Kyber modular reducer completes 1 modular multiplication operation with a modulus of 3329 per cycle; the reusable modular reducer and the Kyber modular reducer are connected in parallel.

5. The lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms according to claim 4, characterized in that The reusable modular reducer includes: a common part, a Kyber part, and a Dilithium part; The common part includes a first adder, a second adder, a first subtractor, and a fourth memory; the output terminals of the first adder and the first subtractor are connected to the input terminal of the second adder; the output terminal of the second adder is connected to the input terminal of the fourth memory; The Dilithium part includes a first accumulator, a first accumulative subtractor, a second subtractor, a third subtractor, a first shifter, a fifth memory, and a sixth memory; the output terminals of the first accumulator and the first accumulative subtractor are connected to the input terminal of the second subtractor; the output terminal of the second subtractor is connected to the input terminal of the first shifter; the output terminals of the first shifter and the fourth memory are connected to the input terminal of the fifth memory; the output terminal of the fifth memory is connected to the sixth memory; wherein the shifter is used to multiply the input data by a multiple of the number 2. The Kyber part includes a second accumulator, a second accumulative subtractor, a fourth subtractor, a fifth subtractor, a second shifter, a seventh memory, and an eighth memory; the output terminal of the second accumulator is connected to the input terminal of the second shifter; the output terminals of the second shifter and the second accumulative subtractor are connected to the input terminal of the fourth subtractor; the output terminal of the fourth subtractor is connected to the input terminal of the seventh memory; the output terminals of the seventh memory and the fourth memory are connected to the input terminal of the fifth subtractor; the output terminal of the fifth subtractor is connected to the input terminal of the eighth memory.

6. The lattice cryptography coprocessor circuit compatible with Kyber and Dilithium algorithms according to claim 4, characterized in that, The Kyber modular reducer includes a plurality of registers; the registers are used to store all the modular reduction values after splitting the output value of the multiplication operation of the reusable modular multiplier, and select and output the modular reduction values stored in the registers through the input value; the Kyber modular reducer obtains the modular multiplication operation result with a modulus of 3329 by performing an addition operation on the modular reduction values output by the registers.

7. The lattice cipher co-processor circuit compatible with Kyber and Dilithium algorithms according to claim 1, characterized in that The hash unit includes a rejection sampler, a central sampler, an extended mask sampler, and an in-sphere sampler; the rejection sampler, the central sampler, the extended mask sampler, and the in-sphere sampler are connected in parallel; the rejection sampler has a parallelism of 4; the central sampler has a parallelism of 16; the extended mask sampler has a parallelism of 8; Among them, the rejection sampler and the central sampler are used to generate polynomial data of the Kyber algorithm; the rejection sampler, the extended mask sampler, and the in-sphere sampler are used to generate polynomial data of the Dilithium algorithm.

8. The lattice cryptographic co-processor circuit compatible with Kyber and Dilithium algorithms according to claim 1, characterized in that, The instruction fetch unit includes a hash unit control area and a polynomial unit control area; the hash unit control area and the polynomial unit control area are connected in parallel; The hash unit control area is used to control the operation of the hash unit according to an external input instruction, and after the hash unit completes one operation, it feeds back to the outside to obtain the next instruction; The polynomial unit control area is used to control the operation of the polynomial unit according to an external input instruction, and after the polynomial unit completes one operation, it feeds back to the outside to obtain the next instruction.

9. The lattice cipher coprocessor circuit compatible with Kyber and Dilithium algorithms according to claim 1, characterized in that, The storage unit includes two independent memories and three true dual-port memories; The true dual-port memory is used to store the intermediate coefficients generated by the hash unit and the polynomial unit; The independent memory is used to store the Kyber NTT rotation factors and Dilithium NTT rotation factors generated by the polynomial unit.

Citation Information

Patent Citations

  • Configurable lattice cryptography processor for the quantum-secure internet of things and related techniques

    US20200265167A1