Post-quantum cryptographic algorithm processor and system-on-chip comprising same
By using component recombination technology and operator fusion technology to design the post-quantum cryptographic processor on low-cost terminal devices, the deployment problem of post-quantum cryptographic algorithms on low-cost terminal devices is solved, and the effect of efficient acceleration and low power consumption is achieved.
Patent Information
- Application Number
- PCT/CN2023/138004
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-05
AI Technical Summary
The prior art is difficult to efficiently deploy post-quantum cryptography algorithms on low-cost terminal devices, especially in the case of large computing volume, long key length and complex data flow, resulting in slow operation speed, high power consumption and large resource consumption.
Using component recombination technology and operator fusion technology, a post-quantum cryptographic processor is designed, including finger fetching, decoding, execution and writeback units. By recombining the same components, the core operations of post-quantum cryptographic algorithms with different computing types and bit widths are supported, and the post-quantum cryptographic processor, bus protocol and peripherals are integrated in the on-chip system, providing a variety of connection methods and flexible configuration options.
By reducing hardware resource overhead and reducing the number of cycles required by the algorithm, the operation speed of the post-quantum cryptography algorithm is improved, and the acceleration effect of low power consumption and low resource overhead is achieved, and more application scenarios are supported.
Smart Images

Figure CN2023138004_05062025_PF_FP_ABST
Abstract
Description
A post-quantum cryptographic algorithm processor and its system-on-chip Technical Field
[0001] The present invention belongs to the technical field of post-quantum cryptographic algorithm hardware acceleration and on-chip systems, and specifically relates to a post-quantum cryptographic algorithm processor and a system on a chip thereof. Background Art
[0002] Post-quantum cryptography is a new generation of public-key cryptography algorithms designed to mitigate quantum computer attacks. Traditional public-key cryptography algorithms (such as RSA, Diffie-Hellman, and elliptic curves) are currently based on difficult mathematical problems such as large integer factorization and discrete logarithms (and elliptic curve variants). However, with the rapid development of quantum computer technology and the emergence of efficient quantum algorithms (such as Shor's algorithm), sufficiently large and stable quantum computers are expected to be able to crack these problems in polynomial time, posing a threat to the security of traditional public-key cryptography algorithms. To ensure information security, the cryptography community has, after years of extensive research and discussion, developed new post-quantum cryptography standards to gradually replace traditional cryptographic algorithms, which are becoming less secure.
[0003] However, there are some difficulties in deploying post-quantum cryptographic algorithms on existing hardware devices, especially low-cost terminal devices. Compared with traditional public key cryptographic algorithms, post-quantum cryptographic algorithms have a larger amount of computation and more complex calculation forms. The longer key length and computational scale also lead to complex planning of data flow and storage space, making applications with high real-time requirements difficult to implement. In order to improve the running speed of post-quantum cryptographic algorithms, the academic community has conducted extensive research and optimized and accelerated software and hardware. However, these designs are mainly aimed at high-performance and high-cost server equipment, and have problems such as high power consumption, low circuit reuse and high resource consumption. As for hardware acceleration solutions for post-quantum cryptographic algorithms, which are mainly used in low-performance, low-cost terminal devices such as the Internet of Things, there are still some gaps.
[0004] To address this issue, researchers have recently focused on the development of system-on-chips (SoCs) to provide hardware acceleration solutions for post-quantum cryptographic algorithms. SoCs integrate multiple functional modules and components on a single chip, enabling more efficient and compact processing capabilities. SoCs can integrate components such as post-quantum cryptographic processors, bus protocols, and other peripherals to accelerate computation and implement specific applications of post-quantum cryptographic algorithms. By utilizing component reorganization and operator fusion technologies, post-quantum cryptographic processors can use the same components to support core operations of post-quantum cryptographic algorithms of varying computation types and bit widths. SoCs also offer multiple connectivity options and flexible configuration options to meet diverse application needs.
[0005] Summary of the Invention
[0006] To solve the problems in the prior art, the present invention provides a post-quantum cryptographic processor, comprising: an instruction fetch unit, a decoding unit, an execution unit, and a write-back unit;
[0007] The instruction fetch unit includes an instruction interface module that serves as an interface for exchanging data with an instruction storage unit outside the processor, a prefetch module for pre-reading multiple instructions in the instruction interface module and caching them, and an instruction dispatch module for dispatching instructions cached in the prefetch module; the instructions dispatched by the instruction dispatch module include read-write instructions and non-read-write instructions;
[0008] The decoding unit receives the instruction sent by the instruction distribution module and translates the received instruction into a specific control signal; then the control signal and the data inside the decoding unit are sent to the execution unit;
[0009] The execution unit performs specific calculations or memory access operations based on the control signals and data sent by the decoding unit; then sends the calculation results and the control signals generated by the execution unit to the write-back unit;
[0010] The write-back unit writes the calculation result sent by the execution unit back to the decoding unit according to the control signal.
[0011] Furthermore, the decoding unit includes a register module and a decoder module; the register module is used to store data, the decoder module translates the received instructions, reads the required data from the register module, and then sends the control signal and the data to the execution unit.
[0012] Furthermore, the pre-processing module in the execution unit is used to perform data transformation of polynomial operations and hash algorithms, and to perform preparatory work before calculation;
[0013] The processing module is composed of multiple identical general computing modules and multiple hash computing modules, and is used to implement various computing operations. The processing module in the execution unit selects the connection mode of each module in the processing module according to the control signal sent by the decoding unit to support computing operations of different bit widths and different operation types required by the post-quantum cryptography algorithm and obtain the corresponding operation results.
[0014] The post-processing module is used to perform subsequent data transformation processing on the calculation results of the polynomial operation and the hash algorithm; after the data transformation processing, the calculation results and the control signal generated by the execution unit are sent to the write-back unit.
[0015] Furthermore, the general computing module includes a multiplier module for implementing multiplication-related operations, a carry-save adder module for implementing modular operations and logical operations, an adder module for implementing addition and subtraction operations, and a shifter for shift operations; the hash calculation module consists of a carry-save adder module, an adder module, and a shifter module.
[0016] Furthermore, the processing module is composed of eight identical general computing modules and two hash computing modules, which are used to implement various computing operations;
[0017] Each general computing module contains a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-preserve adder module for implementing modular operations and logical operations, two 32-bit adder modules for implementing addition and subtraction operations, and a 32-bit shifter for shift operations;
[0018] Each hash calculation module consists of a 32-bit carry-save adder module, two 32-bit adder modules and a 32-bit shifter module.
[0019] Furthermore, the execution unit also includes a read-write control unit, which is connected to the decoding unit and the data storage unit outside the post-quantum cryptographic processor; the read-write control unit receives the read-write related control signal sent by the decoding unit, and performs specific memory access operations on the data storage unit outside the post-quantum cryptographic processor through the control signal sent by the decoding unit, and writes the result obtained by accessing the cache of the data storage unit back to the decoding unit according to the control signal generated by the read-write control unit.
[0020] Furthermore, the write-back unit includes a non-read-write pipeline write-back unit, which is connected to the decoding unit. The non-read-write pipeline write-back unit writes the calculation result sent by the execution unit back to the corresponding register in the decoding unit according to the control signal generated by the execution unit.
[0021] The present invention also provides a system on chip for running the post-quantum cryptographic processor, comprising: a system bus, the post-quantum cryptographic processor connected to the system bus, a storage unit connected to the post-quantum cryptographic processor, and an external device interface connected to the system bus;
[0022] The post-quantum cryptographic processor communicates with other devices on the chip system through the system bus, and also communicates with external devices connected to the external device interface through the system bus, and reads and writes data and instructions to the storage unit connected to the post-quantum cryptographic processor through the instructions received from the external device.
[0023] Furthermore, the system on chip further comprises: a cache active pre-fetch / pre-store data function module combined with DMA connected to the post-quantum cryptographic processor, the cache active pre-fetch / pre-store data function module combined with DMA comprises
[0024] A cache portion includes a cache controller, a cache SRAM connected to the cache controller, a cache usage status record table connected to the cache controller, a communication port of the post-quantum cryptographic processor core connected to the cache controller, and a communication port of a DMA controller connected to the cache controller; the cache controller reads and writes the cache SRAM according to the cache usage status record table; the cache usage status record table records the cache units and operation types in the cache SRAM being read / written by the post-quantum cryptographic processing core and DMA, and can automatically detect and avoid data read / write conflicts before the post-quantum cryptographic processor core reads and writes;
[0025] The DMA part includes a DMA controller, a transfer instruction list connected to the DMA controller, a communication port of the post-quantum cryptographic processor core connected to the DMA controller, a communication port of the cache controller connected to the DMA controller, an external memory read / write port connected to the DMA controller, and a cache SRAM read / write port connected to the DMA controller; the DMA controller performs read and write operations on the external memory and cache SRAM by reading the transfer instruction list; the DMA controller also modifies the transfer instruction list according to the control signals of the post-quantum cryptographic processor and the cache controller; the DMA part is used to accept move operations requested by the post-quantum cryptographic processor core and the cache controller, write cached data back to the external memory or read external memory data into the cache SRAM, and can automatically detect and avoid data read and write conflicts before the DMA controller reads and writes the cache SRAM. The transfer instruction list can store multiple transfer instructions arranged by the post-quantum cryptographic processor core and execute them in sequence. The transfer instruction list also has multiple emergency task bars with the highest priority, which are used to handle the failure of the post-quantum cryptographic processor to access the cache SRAM and urgently read data from the external memory through DMA.
[0026] Furthermore, the system bus supports the AXI 32-bit bus protocol, which further supports the APB general peripheral protocol. The APB general peripheral protocol is then used to connect to the SPI, UART, GPIO, and I2C peripheral protocols. The AXI bus protocol accesses the storage units connected to the post-quantum cryptographic processor in 32 bits.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] (1) The present invention adopts component reorganization technology and operator fusion technology, using the same basic resources to support different types of post-quantum cryptographic algorithms through reorganization, which greatly reduces the overhead of hardware resources. It also reduces the number of cycles required for the algorithm through operator fusion, thereby improving the running speed of the post-quantum cryptographic algorithm.
[0029] (2) The present invention uses 8 or 10 configurable parallel computing units to accelerate the key operations in the post-quantum cryptographic algorithm - hash algorithm and polynomial calculation; when in non-hash operation, only 8 computing units are turned on, thereby saving power consumption during operation; the 10 computing units process a total of 320 bits of data, which matches the amount of data in the hash algorithm, thereby reducing the number of additional data accesses, solving the data dependency problem of the entire algorithm operation process, and achieving an acceleration effect of more than 2 times with a 20% overhead.
[0030] (3) The present invention implements an entire system-on-chip. Compared to directly using a 128-bit bus width, a 32-bit bus width saves resources, power consumption, and other overheads for the entire system. When a peripheral device needs to access or operate a 128-bit data structure, it can be completed through the designed data access module in at most four time cycles; and the 32-bit bus does not affect the execution efficiency of the processor. Through this design, the present invention reduces the power consumption and resource overhead of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] FIG1 is a diagram of the architecture of a post-quantum cryptographic algorithm processor according to the present invention;
[0032] FIG2 is an architectural diagram of a parallel computing unit module in the present invention;
[0033] FIG3 is an architecture diagram of the system on chip in the present invention;
[0034] FIG4 is a diagram of a cache structure in the present invention;
[0035] FIG5 is a structural diagram of a DMA transfer instruction list in the present invention;
[0036] Figure 6 is a flowchart of the cache controller;
[0037] FIG7 is a flowchart of the DMA controller operation process; FIG.
[0038] FIG8 shows the replacement rule and update rule of data blocks based on the least recently used (PLRU) strategy of the binary tree. DETAILED DESCRIPTION
[0039] The present invention will be further described and illustrated below in conjunction with specific embodiments. The embodiments are merely illustrative of the present disclosure and do not limit its scope. The technical features of the various embodiments of the present invention may be combined accordingly, provided that there is no conflict between them.
[0040] This invention proposes a post-quantum cryptographic algorithm processor and its system-on-chip (SoC). This invention aims to accelerate post-quantum cryptographic algorithm computations with reduced power consumption and resource overhead by leveraging the unique characteristics of single-instruction, multiple-data (SIMD) technology and combining it with the parallelism, high reusability, and low power consumption of technologies such as computational component reorganization and operator fusion. Furthermore, the SoC design enables this invention to be applied to a wider range of scenarios, supporting a wider range of applications while accelerating post-quantum cryptographic algorithms.
[0041] (1) Processor part
[0042] As shown in Figure 1, an embodiment of the present invention provides a RISC-V-based dedicated processor for post-quantum cryptographic algorithms. The dedicated processor adopts a superscalar pipeline design with separation of reading, writing and computing. It has two pipelines that can be executed in parallel: reading and writing and computing (non-reading and writing). Each pipeline has four parts: instruction fetch, decoding, execution, and write back that are executed in sequence.
[0043] In a specific embodiment of the present invention, the processor specifically includes:
[0044] The instruction fetch unit, the read-write and calculation pipelines (non-read-write pipelines) reuse the same instruction fetch unit. This part is used to obtain instructions from the instruction cache outside the processor and distribute the instructions to the decoding parts of the two pipelines of the processor after pre-analysis. The instruction fetch unit includes an instruction interface module that serves as an interface to exchange data with the instruction storage unit outside the processor, a pre-fetch module for pre-reading multiple instructions and caching them, and an instruction dispatch module. The instruction dispatch module adopts a sequential multi-issue scheme, that is, the two instructions read sequentially are sent to the first decoder and the second decoder of the decoding unit respectively. This part transmits non-read-write instructions to the calculation pipeline and transmits read-write instructions to the read-write pipeline;
[0045] The decoding unit mainly includes four register modules and two decoder-type modules, which are used to process instructions sent by the instruction fetch unit, translate these instructions into specific control signals, read the data in the register modules, and then send these control signals and data to the execution parts of the two pipelines; the register modules include a general register module for storing 32-bit data, a parallel register group module for storing 64-bit data, and a parameter register module for storing 32-bit parameters. These modules are shared by the two pipelines; the decoder-type modules include a first decoder module belonging to the calculation pipeline and a second decoder module belonging to the read-write pipeline. These two decoder modules respectively translate the instructions received by the corresponding pipeline, and each reads the required data from the required register module, and then each sends the control signals and data of the corresponding pipeline to the execution unit of its own pipeline;
[0046] The execution unit is used to perform specific calculations or memory access operations based on the control signals and data sent from the decoding unit, and then send the calculation results and control signals to the write-back parts of the two pipelines. The execution unit of the computing pipeline includes a configurable parallel computing unit module. The configurable parallel computing unit module adopts a mixed bit-width design and supports parallel computing of data with multiple bit widths up to 320 bits. The configurable parallel computing unit module uses operator fusion technology to combine and reuse the hardware required for different basic operations to implement various basic operations and combined operations. The execution part of the read-write pipeline includes the first stage of the read-write control unit module, which is used to access the data cache outside the processor based on the control signals and data sent from the decoding unit, while the rest of the module operates in the write-back unit of the read-write pipeline.
[0047] The write-back unit is used to write the calculation results sent by the calculation pipeline and the read-write pipeline execution unit back to the corresponding register module contained in the decoding unit according to the control signal, wherein the write-back unit of the read-write pipeline includes the second stage of the read-write control unit module, which is used to return the results obtained by accessing the external data cache in the first stage to the processor and store them in the register module.
[0048] Furthermore, as shown in FIG1 , the decoding unit of the RISC-V-based post-quantum cryptography algorithm dedicated processor also includes three parts, wherein the first decoder is used to temporarily store the signals and data to be sent by the access finger part to the decoding part of the computing pipeline, the second decoder is used to temporarily store the signals and data to be sent by the access finger part to the decoding part of the read-write pipeline, and the read-write control unit (first stage) is used to temporarily store the signals and data to be sent by the execution part of the read-write pipeline to the write-back part. The read-write control unit (second stage) is used to access the external data cache in the read-write control unit (first stage) and return the results to the decoding unit of the post-quantum cryptography algorithm processor and store them in the register module.
[0049] As shown in FIG2 , the configurable parallel computing unit of the execution unit of the post-quantum cryptography algorithm processor in an embodiment of the present invention adopts component reorganization technology and operator fusion technology to implement the core operations of the post-quantum cryptography algorithm of different calculation types and bit widths using the same components. This module includes:
[0050] Preprocessing module, used to perform data transformation before calculation for NTT (fast number theoretic transformation) and Keccak (hash) algorithms;
[0051] The processing module, including the main computing resources, is composed of 8 identical sub-unit modules for implementing arithmetic operations and an additional 2 identical sub-unit modules for accelerating the Keccak algorithm; (most configurable parallel computing units use 8 sub-unit modules for acceleration, while hash algorithms need to be configured as 10 sub-unit modules to solve data dependency and other problems) Each sub-unit module contains a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-preserve adder module for implementing modular operations and logical operations, two 32-bit adder modules for implementing addition and subtraction, and a 32-bit shifter for shift operations; these modules have selectable connection methods, and several modules in the processing module have multiple connection methods, which are used to support the calculation operations of different operations with different bit widths required in the post-quantum cryptography algorithm. According to the different instructions obtained by the instruction fetch part, the decoding unit translates the specific control signal, and the processing module in the execution unit will select the corresponding connection method according to the control signal, adopting different connection methods to perform different calculation operations with different bit widths to obtain the calculation results. By configuring the use of 8 sub-unit modules or 10 sub-unit modules, this processor can be used to accelerate the NTT and Keccak algorithms;
[0052] The post-processing module is used to transform the data after the NTT and Keccak algorithms are calculated.
[0053] (2) System on Chip
[0054] As shown in Figure 3, an embodiment of the present invention provides a system-on-chip (SoC) that runs a dedicated processor for post-quantum cryptography algorithms. The example shown uses a two-level bus architecture, AXI4 and APB, with a data bit width of 32 bits. Internally, the AXI4 bus is used for interconnection, and communication with low-speed peripheral devices is achieved via the APB bus. The APB bus is connected to the AXI4 bus via an AXI2APB bridge.
[0055] Devices mounted on the APB bus can be divided into two categories: low-speed peripheral interfaces and control modules. Low-speed peripheral interfaces include: GPIO (General Purpose Input Output), UART (Universal Asynchronous Receiver / Transmitter), I2C (Inter Integrated Circuit), and SPI (Serial Peripheral Interface) host interfaces; control modules include SoC Control, FLL Control, Timer, Event Unit, etc. These devices all act as APB slaves and are interconnected with the APB bus through the bus model's IP multiplexing module. The AXI4 bus connects the processor core, instruction RAM, data RAM, Flash, SPI slave interface, and advanced debug module (Adv. Debug Unit). The processor core, SPI slave interface, and advanced debug module are connected to the bus as AXI masters; the two RAMs are connected to the bus as AXI slaves.
[0056] In the example shown, the processor can communicate with external devices via various protocols, such as UART, I2C, and SPI, to read and write data. When the SoC boots, instructions are written from external devices (such as Flash memory) via the UART / SPI protocol, sequentially over the APB and AXI buses, to the instruction memory cells. Data is also written to the data memory cells in the same manner. This example distinguishes instructions and data by address partitioning, writing to different memory cells based on the target address.
[0057] Because the AXI bus only implements 32 bits, when the processor core or external device accesses the instruction storage unit or data storage unit, the data access module will process and cache it. When writing data continuously, the data access module will cache the first arriving data until the cached data reaches the required bit width and then write it all at once to the memory. When reading data, the data access module will adjust and filter the read data and ultimately send it to the AXI bus in 32-bit format. In the example shown, the instruction storage unit bit width is 64 bits to support the processor core to read two 32-bit instructions at a time; the data storage unit bit width is 128 bits to support the processor core's custom read and write instructions to read and write 128 bits of data at a time. When the read and write bit width is less than 128 bits, it is cached according to the above rules.
[0058] (3) Cache active prefetch / prestore data module combined with DMA (direct memory access)
[0059] The module includes a cache part and a DMA part.
[0060] The cache part includes a cache controller, cache SRAM, a cache usage status record table, and communication ports of the processor core and DMA controller. The cache part implements a four-channel group associative scheme based on the binary tree idea of least recently used (PLRU) replacement, write return and write allocation strategy, supporting all basic functions of the cache. As shown in Figure 4, the cache part of the example of the present invention is a four-way group associative structure with a cache space of 8kB and 128 groups. Each group has four channels, and each channel has two flag bits. The Valid bit is used to indicate whether the channel has data, and the Dirty bit is used to indicate whether the data of the channel has been written. The cache usage status record table records the cache units and operation types that the processing core and DMA are reading / writing. The cache workflow is shown in Figure 6. After receiving an instruction from the processor core, it first determines whether the address meets the requirements and processes the address that meets the requirements: if it is a data read / write instruction, first match the channel and data to see if they are valid, then check the cache usage status record table to find out whether DMA is reading or writing to the same location. If so, let DMA process it first, and finally write and read data. If data needs to be transferred in or out of the external memory outside the on-chip system, the required address, data, and data transfer direction are written to the emergency status line of the DMA transfer instruction list; if it is a custom data pre-store / pre-fetch instruction, the instruction is written to the emergency status normal line of the DMA transfer instruction list and waits for DMA processing.
[0061] The DMA component includes a DMA controller, a transfer instruction list, communication ports with the processor core and cache controller, external memory read / write ports, and cache SRAM read / write ports. It accepts data transfer requests from the processor core and cache controller, writing cached data back to external memory or reading data from external memory into cache SRAM. As shown in Figure 5, the DMA transfer instruction list can store four pre-store / pre-fetch instructions assigned by the processor core, which are executed sequentially. There is also a top-priority urgent task column to handle situations where the post-quantum cryptography processor fails to access cache SRAM and urgently needs to read data from external memory via DMA. Each row in the transfer instruction list stores the address, data length, and transfer direction. The DMA workflow is shown in Figure 7. The DMA loop checks the transfer instruction list, prioritizing urgent rows over regular rows. If there is a need to transfer data, the cache usage status table is first checked to determine whether a processor core is currently reading or writing to the same location. If so, the processor core is allowed to read or write first before the data transfer begins. After the data transfer, the address and data length of the corresponding row in the transfer instruction list are updated. To maximize efficiency, DMA will automatically select other channels in the same group when the processor core reads or writes a channel. In various situations, various conflicts may occur. The specific scheduling options are:
[0062] When DMA is about to write data from a channel in the cache back to the external memory, the following unexpected situations may occur:
[0063] 1. The processor core is reading the channel (Dirty=0), and the read is hit: the CPU reads the data; because the data of this channel has not been changed, the DMA does not need to move it.
[0064] 2. The processor core is reading the channel (Dirty=1), and the read is hit: the CPU and DMA read data simultaneously without interfering with each other.
[0065] 3. The processor core is ready to move data into the channel (Dirty=0), and the read fails: the CPU instructs the DMA to move the data it needs; the DMA does not need to move.
[0066] 4. The processor core is ready to move data into the channel (Dirty = 1, the channel has data that has not been written back), and the read fails: the CPU instructs DMA to move the data in the channel to Flash, and then instructs DMA to move the data it needs.
[0067] 5. The processor core is writing to this channel (Dirty = 0), and the write is successful: the CPU completes the write, and the DMA moves the data again. After writing, the data becomes dirty and needs to be moved.
[0068] 6. The processor core is writing to this channel (Dirty=1), and the write is successful: the CPU completes writing, and the DMA moves the data again.
[0069] 7. The processor core prepares to write to the channel (Dirty = 0), and the write fails: If the bit width of the write data is not 128 bits, the CPU needs to first instruct the DMA to move the data it needs, and then write the data; because the data in this channel has not been changed, the DMA does not need to move the data.
[0070] 8. The processor core prepares to write to the channel (Dirty = 1, the channel has data that has not been written back), and the write fails: the CPU instructs DMA to move the data in the channel to Flash. If the bit width of the write data is not 128 bits, the CPU needs to instruct DMA to move the data it needs and finally write the data.
[0071] When DMA is about to write data to a channel, the following unexpected situations may occur:
[0072] 1. The processor core is reading the channel (Dirty=0), and if a read hit occurs, DMA moves the data to another channel.
[0073] 2. The processor core is reading the channel (Dirty=1), and if a read hit occurs, DMA moves the data to another channel.
[0074] 3. The processor core prepares to move data into the channel (Dirty=0), and the read fails: the CPU instructs the DMA to move the data it needs; if the data to be moved by the DMA and the CPU are different, the DMA moves the data to another channel.
[0075] 4. The processor core prepares to move data into the channel (Dirty = 1, the channel has data that has not been written back), and the read fails: the CPU instructs the DMA to move the original data in the channel and then move the data it needs; if the data that the DMA and the CPU need to move are different, the DMA moves the data to another channel.
[0076] 5. The processor core is writing to this channel (Dirty=0), write hit: CPU writes; DMA moves data to another channel.
[0077] 6. The processor core is writing to this channel (Dirty=1), write hit: CPU writes; DMA moves data to another channel.
[0078] 7. The processor core prepares to write to the channel (Dirty = 0), and the write fails: If the bit width of the write data is not 128 bits, the CPU needs to first instruct the DMA to move the data it needs, and then write the data; if the data that the DMA and the CPU need to move are different, the DMA moves the data to another channel.
[0079] 8. The processor core prepares to write to the channel (Dirty = 1, the channel has data that has not been written back), and the write fails: the CPU instructs the DMA to move the original data in the channel. If the bit width of the write data is not 128 bits, the CPU needs to instruct the DMA to move the data it needs and finally write the data; if the data that the DMA and the CPU need to move are different, the DMA moves the data to another channel.
[0080] As shown in Figure 8, the replacement and update rules of the data block based on the binary tree least recently used (PLRU) strategy are implemented as follows: a binary tree structure is used to store the historical access order information of the data block. For a 4-way set associative cache, a three-bit register is set for each group, denoted as B[2:0].
[0081] The above-described embodiments only express several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be understood as limiting the scope of the patent of the present invention. The present invention can support a variety of post-quantum cryptographic algorithms and the various core operations contained therein, and correspondingly has dozens of dedicated extension instructions that have not yet been listed and corresponding combinations of various components in the processing module in the parallel computing unit module. For those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can be made, which all fall within the scope of protection of the present invention.
Claims
1. A post - quantum cryptographic processor, characterized in that, it includes: an instruction fetch unit, a decoding unit, an execution unit, and a write - back unit; The instruction fetch unit includes an instruction interface module that exchanges data with an instruction storage unit outside the processor as an interface, a pre - fetch module that pre - reads multiple instructions in the instruction interface module and caches them, and an instruction distribution module that distributes the instructions cached in the pre - fetch module; the instructions distributed by the instruction distribution module include read - write instructions and non - read - write instructions; The decoding unit receives the instructions emitted by the instruction distribution module, and translates the received instructions into specific control signals; then it sends the control signals and the data inside the decoding unit to the execution unit; The execution unit performs specific calculation or memory access operations according to the control signals and data sent by the decoding unit; then it sends the calculation results and the control signals generated by the execution unit to the write - back unit; The write - back unit writes back the calculation results sent by the execution unit to the decoding unit according to the control signals; The execution unit includes a configurable parallel computing unit module, and the configurable parallel computing unit module includes: a pre - processing module for performing polynomial operations and data transformation of hash algorithms to prepare for calculations; a processing module composed of multiple identical general - purpose computing modules and multiple hash computing modules for implementing various calculation operations; the processing module in the execution unit selects the connection method of each module in the processing module according to the control signals sent by the decoding unit to support calculation operations of different bit widths and different operation types required in the post - quantum cryptographic algorithm, and obtains corresponding operation results; and a post - processing module for performing subsequent data transformation processing on the calculation results of polynomial operations and hash algorithms; after the data transformation processing, it sends the calculation results and the control signals generated by the execution unit to the write - back unit.
2. The post - quantum cryptographic processor according to claim 1, characterized in that, the decoding unit includes a register module and a decoder module; the register module is used to store data, the decoder module translates the received instructions, reads the required data from the register module, and then sends the control signals and the data to the execution unit.
3. The post - quantum cryptographic processor according to claim 1, characterized in that, the configurable parallel computing unit module adopts component recombination technology and operator fusion technology, and can use the same components to implement the core operations of post - quantum cryptographic algorithms of different calculation types and bit widths.
4. The post - quantum cryptographic processor according to claim 1, characterized in that, the general - purpose computing module includes a multiplier module for implementing multiplication - related operations, a carry - save adder module for implementing modulo operations and logical operations, an adder module for implementing addition and subtraction operations, and a shifter for implementing shift operations; the hash computing module is composed of a carry - save adder module, an adder module, and a shifter module.
5. The post - quantum cryptographic processor according to claim 1, characterized in that, the processing module is composed of eight identical general - purpose computing modules and two hash computing modules for implementing various calculation operations; Each general computing module includes a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-save adder module for implementing modulo operations and logical operations, two 32-bit adder modules for implementing addition and subtraction operations, and a 32-bit shifter for implementing shift operations; Each hash computing module consists of a 32-bit carry-save adder module, two 32-bit adder modules, and a 32-bit shifter module.
6. The post-quantum cryptographic processor according to claim 1, characterized in that, the execution unit further includes a read-write control unit, and the read-write control unit is connected to the decoding unit and a data storage unit outside the post-quantum cryptographic processor; the read-write control unit receives control signals related to reading and writing sent by the decoding unit, and performs specific memory access operations on the data storage unit outside the post-quantum cryptographic processor through the control signals sent by the decoding unit, and writes back the result cached by accessing the data storage unit to the decoding unit according to the control signals generated by the read-write control unit.
7. The post-quantum cryptographic processor according to claim 2, characterized in that, the write-back unit includes a non-read-write pipeline write-back unit, and the non-read-write pipeline write-back unit is connected to the decoding unit. The non-read-write pipeline write-back unit writes back the calculation result sent by the execution unit to the corresponding register in the decoding unit according to the control signal generated by the execution unit.
8. A system-on-chip for running the post-quantum cryptographic processor according to any one of claims 1-7, characterized in that, it includes: a system bus, the post-quantum cryptographic processor connected to the system bus, a storage unit connected to the post-quantum cryptographic processor, and an external device interface connected to the system bus; The post-quantum cryptographic processor communicates with other devices of the system-on-chip through the system bus, and also communicates with external devices connected to the external device interface through the system bus, and reads and writes data and instructions to the storage unit connected to the post-quantum cryptographic processor through instructions received from the external devices.
9. The system-on-chip according to claim 8, characterized in that, it further includes: a cache active prefetch / prestore data function module combined with DMA connected to the post-quantum cryptographic processor, and the cache active prefetch / prestore data function module combined with DMA includes a cache part, including a cache controller, a cache SRAM connected to the cache controller, a cache usage status record table connected to the cache controller, a communication port of the post-quantum cryptographic processor core connected to the cache controller, and a communication port of the DMA controller connected to the cache controller; The cache controller reads and writes the cache SRAM according to the cache usage status record table; the cache usage status record table records the cache units and operation types in the cache SRAM being read / written by the post-quantum cryptographic processor core and the DMA, and can automatically detect and avoid data read / write conflicts before the post-quantum cryptographic processor core reads and writes. The DMA part includes a DMA controller, a transfer instruction list connected to the DMA controller, a communication port of the post-quantum cryptographic processor core connected to the DMA controller, a communication port of the cache controller connected to the DMA controller, an external memory read / write port connected to the DMA controller, and a cache SRAM read / write port connected to the DMA controller; the DMA controller reads and writes the external memory and the cache SRAM by reading the transfer instruction list; according to the control signals of the post-quantum cryptographic processor and the cache controller, the DMA controller also modifies the transfer instruction list; the DMA part is used to accept the transfer operations requested by the post-quantum cryptographic processor core and the cache controller, write the cache data back to the external memory or read the external memory data into the cache SRAM, can automatically detect and avoid data read / write conflicts before the DMA controller reads and writes the cache SRAM, the transfer instruction list can store multiple transfer instructions arranged by the post-quantum cryptographic processor core and execute them sequentially, and the transfer instruction list also has multiple highest-priority emergency task bars for handling the situation where the post-quantum cryptographic processor accesses the cache SRAM fails, and reads data from the external memory urgently through the DMA.
10. The system-on-chip according to claim 8, wherein, the system bus supports the AXI 32-bit bus protocol, and the AXI 32-bit bus protocol further supports the APB general peripheral protocol; then connects to the SPI, UART, GPIO, and I2C peripheral protocols through the APB general peripheral protocol; the access to the storage unit connected to the post-quantum cryptographic processor by the AXI bus protocol is all in 32 bits.
Citation Information
Patent Citations
Encryption-decryption coprocessor for SOC, implementing method and programming model thereof
CN101201811A
32-Bit triple-emission digital signal processor supporting SIMD
CN102750133A
Unified accelerator for classical and post-quantum digital signature schemes in computing environments
CN112152809A
RISC-V-based processor special for post-quantum cryptography algorithm
CN116432765A
Systems and methods for providing user authentication for quantum-entangled communications in a cloud environment
US20230353348A1
Cited By
Accelerated processing method and device for polynomial operation in post quantum cryptography
CN121441501A