Post-quantum cryptography vector processor based on RISC-V architecture

By designing a post-quantum encryption vector processor based on the RISC-V architecture, employing a three-stage pipeline structure and specific modules, the throughput and parallelism issues of the post-quantum encryption algorithm are solved, achieving efficient polynomial field computation and nonlinear decoding, and adapting to encryption and decryption operations with different security requirements.

CN119483931BActive Publication Date: 2025-10-28BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411533066.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-30
Publication Date
2025-10-28
Estimated Expiration
2044-10-30

AI Technical Summary

Technical Problem

Post-quantum cryptography algorithms, when implemented on traditional computers, have insufficient throughput, are not parallelizable enough, have low configurability, cannot efficiently handle complex computations in the polynomial field and nonlinear decoding problems, and lack customization in cryptographic systems.

Method used

A post-quantum encryption vector processor based on a RISC-V architecture is designed. It adopts a three-stage pipelined structure, including a decoding module, a CSR register module, a vector-registers module, and an execution module. Combined with SHA-3, Sample, and NTT modules, it realizes hardware-software co-operational vector encryption and supports polynomial field operations and nonlinear decoding.

Benefits of technology

It improves the throughput and parallelism of the post-quantum encryption algorithm, realizes high-speed, configurable encryption and decryption operations, adapts to different security requirements and application scenarios, and solves the problems of complex computation in the polynomial field and nonlinear decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119483931B_ABST
    Figure CN119483931B_ABST
Patent Text Reader

Abstract

This application proposes a RISC-V architecture post-quantum encryption vector processor, which relates to the field of information security technology. The vector processor is divided into three stages of pipelines. The first stage pipeline includes a decoding module, the second stage pipeline includes a CSR register module, a Vector-registers module, and an execution module. The third stage pipeline includes a load / store module. The decoding module receives custom vector instructions distributed by the main processor, decodes them, and outputs the decoding results to the CSR register module and the execution module; the CSR register module configures the parameters of the execution module according to the decoding results; the Vector-registers module stores vector data; the execution module implements different functions according to the decoding results; and the load / store module implements data exchange between the processor and the first-level cache, completing data loading and storage operations. The present invention adopts the above scheme to realize a software and hardware collaborative, vectorized, high-speed post-quantum encryption algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information security technology, and in particular to a RISC-V architecture-based post-quantum cryptographic vector processor. Background Art

[0002] In the 5G era, massive numbers of devices need to be securely connected to the edge of communication networks, while emerging quantum computers can easily crack traditional public-key cryptography (RSA, ECC, etc.). The threat of quantum computers lies in their ability to utilize quantum parallelism and quantum entanglement to solve problems that are difficult to handle on traditional computers, one of which is factoring large integers into their prime factors. Therefore, traditional encryption algorithms based on integer factorization problems may fail under the attack of quantum computers. To counter the threat of quantum computing technology, post-quantum cryptography (PQC) has emerged.

[0003] Post-quantum cryptography is also secure on classical computers. These algorithms are primarily based on other mathematical problems, such as the lattice problem. Post-quantum cryptography does not utilize any quantum properties; all post-quantum cryptography algorithms can be used on conventional computing platforms based on standard semiconductor technology, without requiring specialized quantum hardware to encrypt data.

[0004] RISC-V is a novel, simple, and open-source reduced instruction set architecture. It avoids the historical baggage of backward compatibility required by x86 and ARM instruction set architectures, and boasts advantages such as concise and well-organized instruction coding, strong scalability, and comprehensive open-source software toolchain support. Since its release, it has received close attention from academia and industry. Currently, due to evolving circumstances, processor localization has become an increasingly important issue. Designing domestically produced processors with an independent, secure, and controllable RISC-V instruction set architecture is beneficial for enhancing processor information security.

[0005] Compared to traditional encryption algorithms, lattice-based cryptography typically involves polynomial rings and polynomial equations. The complexity of the polynomial field makes lattice-based cryptography highly secure against attacks from both traditional and potential quantum computers. However, it also has higher computational complexity, consuming more resources than traditional encryption algorithms. Traditional scalar processors operate primarily on scalars (single data elements). This means that for parallel computing algorithms, each computational step can only process one data element, unable to process multiple data elements simultaneously. Furthermore, general-purpose vector processors are typically designed to flexibly adapt to various computational tasks, which may prevent high optimization for specific post-quantum algorithms. When designing proprietary encoding / decoding schemes, the implementation of post-quantum encryption algorithms should be flexible enough to adapt to different application scenarios and security requirements, and possess a degree of customizability, allowing the system configuration and algorithm to be adjusted according to actual needs.

[0006] In summary, modern post-quantum cryptography is a supercomputing problem that combines complex computational problems in multinomial fields, nonlinear decoding problems requiring fault-tolerant learning, and custom encoding / decoding problems requiring privatization. Therefore, its encryption and decryption implementation cannot be accomplished through a deterministic platform. This processor, based on the coupling of RISC-V and core computing, achieves software-based post-quantum encryption and decryption. Summary of the Invention

[0007] This application aims to at least partially address one of the technical problems in the related art.

[0008] Therefore, the purpose of this application is to propose a RISC-V architecture post-quantum encryption vector processor, which solves the problems of insufficient throughput, lack of parallelization, and low configurability in the implementation of post-quantum encryption algorithms. It can realize a hardware-software co-operated, vectorized, and high-speed post-quantum encryption algorithm, and solves the problems of complex polynomial field computation, nonlinear decoding, and pattern customization in post-quantum encryption.

[0009] To achieve the above objectives, this application proposes a RISC-V architecture-based post-quantum cryptography vector processor. This vector processor consists of a three-stage pipeline: the first stage includes a decoding module; the second stage includes a CSR register module, a vector-registers module, and an execution module; and the third stage includes a load / store module.

[0010] The decoding module is used to receive custom vector instructions distributed by the main processor, decode them, and output the decoding results to the CSR register module and the execution module.

[0011] The CSR register module is used to configure the parameters of the execution module based on the decoding result;

[0012] The Vector-registers module is used to store vector data;

[0013] The execution module is used to implement different functions based on the decoding results;

[0014] The load / store module is used to implement data exchange between the processor and the L1 cache, and to complete data loading and storage operations.

[0015] Optionally, in one embodiment of this application, the parameters of the execution module include: the modulus q in the finite field, the pre-calculated value μ in Barrett reduction, the bit width of the modulus q, the boundary value in rejection sampling, and the parameter k in binomial sampling.

[0016] Optionally, in one embodiment of this application, the execution module includes an SHA-3 module, a Sample module, and an NTT module.

[0017] Optionally, in one embodiment of this application, the SHA-3 module includes a padding module, an iterative operation module, a control module, and a truncation module.

[0018] Optionally, in one embodiment of this application, the Sample module includes a rejection sampler and a binomial sampler, wherein,

[0019] Reject sampler, used to achieve uniform sampling;

[0020] A binary sampler is used to sample the error distribution.

[0021] Optionally, in one embodiment of this application, the rejection sampler includes a comparator and a Barrett modulo reduction module. The rejection sampler is specifically used to: discard sampled data that exceeds the limit, and when the boundary is a modulus q, discard sampled data that is not uniformly distributed on Rq, set the boundary to a multiple of the modulus q, and perform Barrett reduction again.

[0022] Optionally, in one embodiment of this application, the binomial sampler includes a Hamming weight calculation module and a modular reduction module. The Hamming weight calculation module is used to calculate the Hamming weights of two input data through a configurable addition tree.

[0023] Optionally, in one embodiment of this application, the NTT module is used to accelerate polynomial multiplication. The NTT module includes 32 parallel butterfly operation units, an input permutation network, and an output permutation network. The NTT module processes 64 16-bit data per clock cycle.

[0024] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0025] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0026] Figure 1 This is a schematic diagram of the structure of a RISC-V architecture-based quantum encryption vector processor provided in Embodiment 1 of this application;

[0027] Figure 2 This is a schematic diagram of the SHA-3 module according to an embodiment of this application;

[0028] Figure 3 This is a schematic diagram of the structure of the vector SAMPLE module in an embodiment of this application;

[0029] Figure 4 This is a schematic diagram of the structure of the vector NTT module according to an embodiment of this application;

[0030] Figure 5 This is a schematic diagram of the NTT permutation network according to an embodiment of this application;

[0031] Figure 6 This is a schematic diagram of the butterfly-shaped arithmetic module according to an embodiment of this application;

[0032] Figure 7 This is a schematic diagram of the Barrett modular multiplier according to an embodiment of this application; Detailed Implementation

[0033] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0034] The following describes an embodiment of the RISC-V architecture post-quantum cryptography vector processor of this application with reference to the accompanying drawings.

[0035] Figure 1 This is a schematic diagram of a RISC-V architecture-based quantum encryption vector processor provided in Embodiment 1 of this application.

[0036] like Figure 1As shown, this RISC-V architecture-based quantum cryptography vector processor is divided into three pipelines. The first pipeline includes a decoding module, the second pipeline includes a CSR register module, a vector-registers module, and an execution module, and the third pipeline includes a load / store module.

[0037] The decoding module is used to receive custom vector instructions distributed by the main processor, decode them, and output the decoding results to the CSR register module and the execution module.

[0038] The CSR register module is used to configure the parameters of the execution module based on the decoding result;

[0039] The Vector-registers module is used to store vector data (polynomial coefficients and temporary calculation results);

[0040] The execution module is used to implement different functions based on the decoding results;

[0041] The load / store module is used to implement data exchange between the processor and the L1 cache, and to complete data loading and storage operations.

[0042] The RISC-V architecture post-quantum encryption vector processor of this application solves the problems of insufficient throughput, lack of parallelization, and low configurability in the implementation of post-quantum encryption algorithms. It can realize a hardware-software co-operated, vectorized, and high-speed post-quantum encryption algorithm, and solves the problems of complex polynomial field calculations, nonlinear decoding, and pattern customization in post-quantum encryption.

[0043] Optionally, in one embodiment of this application, the parameters of the execution module include:

[0044] csrModulusq: Modulus q in a finite field;

[0045] csrBarrettu: The pre-calculated value μ in Barrett reduction;

[0046] csrModulusLen: Bit width of the modulus q;

[0047] csrBound: Reject boundary values ​​in sampling;

[0048] csrBinomialk: The parameter k in binomial sampling.

[0049] Optionally, in one embodiment of this application, the execution module includes an SHA-3 module, a Sample module, and an NTT module.

[0050] Optionally, in one embodiment of this application, a hash function, also known as a hashing function, is a one-way cryptographic system that can convert an input message of arbitrary length into a fixed-length digest output and is irreversible. SHA-3 is the latest generation of hash functions, which can construct hash functions and generate random number bit streams. Configurable SHA-3 modules include a padding module, an iterative operation module (Transformation Round), a control module (control), and a truncating module.

[0051] Specifically, if Figure 2 As shown, after SHA-3 starts, the random number seed first enters the padding module through a register. The padding module pads the random number seed according to the rule "pad10*1". After padding, zeros are added to the end of the data to make the data length 1600 bits. Next, the data enters the iterative operation module through a multiplexer (mux). The iterative operation module performs iterative operations on the data, which include five steps: θ, ρ, π, χ, and τ. After the five steps are completed, the data enters the dmux through a register. The data output from the dmux will then enter the iterative operation module through the mux again for iterative operation. This process needs to be repeated 24 times. After 24 rounds of iterative operation, the data enters the truncation module through the mux, which truncates the output to a specific length.

[0052] Optionally, in one embodiment of this application, the vector sample module includes a vector rejection sampler and a vector binomial sampler, wherein,

[0053] Reject sampler, used to achieve uniform sampling;

[0054] A binary sampler is used to sample the error distribution.

[0055] Optionally, in one embodiment of this application, the rejection sampler includes a comparator and a Barrett modulo reduction module. The rejection sampler enables uniform sampling, and sampled data exceeding a threshold must be discarded. When the threshold is a modulus q, the sampled data is either uniformly distributed across Rq or discarded. To reduce the rejection rate to an acceptable level, the threshold is set to a multiple of the modulus q, and then Barrett reduction is performed again.

[0056] Optionally, in one embodiment of this application, in addition to rejection sampling, another sampling process in lattice ciphers is error distribution sampling, which can be implemented using a discrete Gaussian sampler. The binomial sampler includes a Hamming weight calculation module and a modular reduction module. The Hamming weight calculation module calculates the Hamming weights of the two input data using a configurable addition tree. The parameter k can be configured to 1, 2, 4, 8, and 16 to meet different application requirements.

[0057] Specifically, after the random number is generated by the SHA-3 module, it enters the vector sample module for sampling. For example... Figure 3 As shown, the rejection sampling process is as follows: (1) 32 groups of 32-bit random data enter the vector rejection sampling module; (2) The input data enters the comparator, which compares the input data with the preset boundary value to determine whether the data is valid; (3) The data is input into the Barrett reduction module to reduce the computational complexity of the data modulo operation and output a 16-bit processing result; (4) Valid data is marked as valid by outputting through mux.

[0058] The process of binary sampling is as follows: (1) Input 32 sets of 16-bit random data vectors Data a and Data b; (2) Data a and Data b are processed by the Hamming weight calculation module to calculate the Hamming weight of vectors a and b; (3) Perform the modulo operation of the Hamming weight difference to obtain the sampled data; (4) Output 32 sets of 16-bit sampled data.

[0059] Optionally, in one embodiment of this application, the vector NTT module, used to accelerate polynomial multiplication, includes 32 parallel butterfly operation units (modular multiplication, modular subtraction, and modular addition), an input permutation network, an output permutation network, and processes 64 16-bit data per clock cycle.

[0060] Specifically, the vector NTT module, such as Figure 4 As shown, it includes a permutation network and 32 parallel butterfly units. The permutation network is used to organize and sort the inputs and outputs of the butterfly units, and the 32 parallel butterfly units perform NTT operations, thus enabling the parallel processing of two 512-bit vectors.

[0061] (1) Permutation networks, which can solve permutation problems through software-based position swapping, require hundreds of cycles per permutation. Therefore, a parallel hardware-based permutation network is adopted. For example... Figure 5 As shown, to achieve 100% hardware resource utilization, a configurable permutation network was designed within the NTT core. The permutation network proposed in this embodiment mainly consists of five independent permutation networks, with the PermConfig signal used to select a specific output and initial input of these five networks. The aforementioned permutation network is placed in front of the butterfly unit, and an inverse permutation network is also needed after the butterfly unit to restore the natural order of the data.

[0062] (2) Butterfly-shaped unit, such as Figure 7 As shown, the Barrett method is implemented, supporting configurability of the modulus q. The bit width of the pre-computed values ​​μ and q required by the Barrett algorithm can be set. Figure 6These are represented as csrBarrettU and csrModulusLen, respectively. Because it uses the universal Barrett method, it fully supports arbitrary moduli within 16 bits. The butterfly element workflow is as follows:

[0063] Bit reversal is too resource-intensive for hardware implementation, especially when the dimension n is large. This design uses a hybrid DIT / DIF butterfly structure. Since the INTT calculation always follows the NTT calculation in polynomial multiplication, bit reversal can be avoided by applying the following procedure.

[0064] (a) Setting ctrl=1 will put the butterfly unit in DIF mode.

[0065] (b) Perform NTT calculation, and output the data in reverse bit order.

[0066] (c) Perform vector arithmetic calculations.

[0067] (d) Setting ctrl=0 will put the butterfly unit in DIT mode.

[0068] (e) Perform INTT calculation directly, because the input data is already in reverse order after NTT calculation, and the output data is in natural order.

[0069] For NTT calculations, 64 16-bit data are processed per clock cycle.

[0070] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0071] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0072] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0073] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0074] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0075] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it includes one or a combination of the steps of the method embodiments.

[0076] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0077] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A RISC-V architecture-based post-quantum cryptography vector processor, characterized in that, The vector processor is divided into three pipelines. The first pipeline includes a decoding module; the second pipeline includes a CSR register module, a Vector-registers module, and an execution module; and the third pipeline includes a load / store module. The decoding module is used to receive custom vector instructions distributed by the main processor, decode them, and output the decoding results to the CSR register module and the execution module. The CSR register module is used to configure the parameters of the execution module according to the decoding result; The Vector-registers module is used to store vector data; The execution module is used to perform different functions based on the decoding result; The load / store module is used to implement data exchange between the processor and the L1 cache, and to complete data loading and storage operations.

2. The post-quantum encryption vector processor as described in claim 1, characterized in that, The parameters of the execution module include: the modulus q in the finite field, the pre-calculated value μ in Barrett reduction, the bit width of the modulus q, the boundary value in rejection sampling, and the parameter k in binomial sampling.

3. The post-quantum cryptography vector processor as described in claim 1, characterized in that, The execution module includes an SHA-3 module, a Sample module, and an NTT module.

4. The post-quantum cryptography vector processor as described in claim 3, characterized in that, The SHA-3 module includes a padding module, an iterative operation module, a control module, and a truncation module.

5. The post-quantum encryption vector processor as described in claim 3, characterized in that, The Sample module includes a rejection sampler and a binomial sampler, wherein, The rejection sampler is used to achieve uniform sampling; The binary sampler is used to sample the error distribution.

6. The post-quantum cryptography vector processor as described in claim 5, characterized in that, The rejection sampler includes a comparator and a Barrett modulo reduction module. Specifically, the rejection sampler is used to: discard sampled data that exceeds the limit, and when the boundary is a modulus q, discard sampled data that is not uniformly distributed on Rq, set the boundary to a multiple of the modulus q, and perform Barrett reduction again.

7. The post-quantum cryptography vector processor as described in claim 5, characterized in that, The binomial sampler includes a Hamming weight calculation module and a modular reduction module. The Hamming weight calculation module is used to calculate the Hamming weights of two input data through a configurable addition tree.

8. The post-quantum cryptography vector processor as described in claim 3, characterized in that, The NTT module is used to accelerate polynomial multiplication. The NTT module includes 32 parallel butterfly operation units, an input permutation network, and an output permutation network. The NTT module processes 64 16-bit data per clock cycle.

Citation Information

Patent Citations

  • RISC-V-based lattice password processing system, method and equipment and storage medium

    CN112748929A

  • RISC-V-based processor special for post-quantum cryptography algorithm

    CN116432765A