A post-quantum cryptographic algorithm processor and its system-on-chip
By using component recombination technology and operator fusion technology to design a post-quantum cryptographic processor on low-cost terminal devices, the problems of large hardware resource consumption and high power consumption in the existing technology are solved, and efficient post-quantum cryptographic algorithm calculation and acceleration effect are achieved.
Patent Information
- Application Number
- CN202311594500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-11-27
AI Technical Summary
It is difficult for the prior art to efficiently deploy post-quantum cryptography algorithms on low-cost terminal devices, especially in application scenarios with large computing volume, long key length and high real-time requirements, there are problems such as large hardware resource consumption, high power consumption and low circuit multiplexing.
Using component recombination technology and operator fusion technology, a post-quantum cryptographic processor is designed to support different types of post-quantum cryptographic algorithm calculations by recombining the same components, reducing hardware resource overhead, and accelerating hashing algorithms and polynomial calculations through configurable parallel computing units.
It realizes efficient operation of post-quantum cryptography algorithm on low-cost terminal devices, reduces the overhead of hardware resources and power consumption, and improves the computing speed, achieving an acceleration effect of more than 2 times.
Smart Images

Figure CN117435251B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hardware acceleration of post - quantum cryptographic algorithms and system - on - chip technology, and particularly relates to a post - quantum cryptographic algorithm processor and its system - on - chip. Background Art
[0002] Post - quantum cryptographic algorithms are a new generation of public - key cryptographic algorithms designed to counter quantum computer attacks. Traditional public - key cryptographic algorithms (such as RSA, Diffie - Hellman, and elliptic curves) currently rely on mathematical problems such as integer factorization and discrete logarithms (and variants of elliptic curves). However, with the rapid development of quantum computer technology and the emergence of efficient quantum algorithms (such as Shor's algorithm), a sufficiently large and stable quantum computer is expected to break these problems in polynomial time, thus threatening the security of traditional public - key cryptographic algorithms. To ensure information security, the cryptographic community has, through years of extensive research and discussion, developed new post - quantum cryptographic algorithm standards to gradually replace the traditional cryptographic algorithms that will soon become insecure.
[0003] However, there are some difficulties in deploying post - quantum cryptographic algorithms on existing hardware devices, especially low - cost terminal devices. Compared with traditional public - key cryptographic algorithms, post - quantum cryptographic algorithms have a larger computational amount and more complex computational forms. The longer key length and computational scale also lead to complex planning of data streams and storage spaces, making it difficult to implement applications with high real - time requirements. To improve the running speed of post - quantum cryptographic algorithms, the academic community has conducted extensive research and optimized and accelerated both software and hardware. However, these designs mainly target high - performance and high - cost server devices, and have problems such as high power consumption, low circuit reuse rate, and large resource consumption. As for the hardware acceleration solutions for post - quantum cryptographic algorithms mainly applied to low - performance and low - cost terminal devices such as the Internet of Things, there are still some blank areas currently.
[0004] To solve this problem, in recent years, researchers have begun to focus on the development of system - on - chip to provide hardware acceleration solutions for post - quantum cryptographic algorithms. A system - on - chip is a collection of multiple functional modules and components integrated on a single chip, which can provide more efficient and more compact processing capabilities. In a system - on - chip, components such as a post - quantum cryptographic processor, bus protocol, and other peripherals can be integrated to achieve accelerated computing of post - quantum cryptographic algorithms and specific applications. By adopting component recombination technology and operator fusion technology, the post - quantum cryptographic processor can use the same components to support the core operations of post - quantum cryptographic algorithms with different computational types and bit widths. In addition, the system - on - chip can also provide various connection methods and flexible configuration options according to specific application requirements to adapt to a variety of other needs. Summary of the Invention
[0005] To solve the problems in the prior art, the present invention provides a post - quantum cryptographic processor, including: an instruction fetch unit, a decoder unit, an execution unit, and a write - back unit;
[0006] The instruction fetch unit includes an instruction interface module that exchanges data with an instruction storage unit outside the processor as an interface, a pre - fetch module for pre - reading multiple instructions in the instruction interface module and caching them, and an instruction distribution module for distributing the instructions cached in the pre - fetch module; the instructions distributed by the instruction distribution module include read - write instructions and non - read - write instructions;
[0007] The decoder unit receives the instructions emitted by the instruction distribution module, and translates the received instructions into specific control signals; then it sends the control signals and the data inside the decoder unit to the execution unit;
[0008] The execution unit performs specific calculation or memory access operations according to the control signals and data sent by the decoder unit; then it sends the calculation results and the control signals generated by the execution unit to the write - back unit;
[0009] The write - back unit writes the calculation results sent by the execution unit back to the decoder unit according to the control signals.
[0010] Further, the decoder unit includes a register module and a decoder module; the register module is used to store data, the decoder module translates the received instructions, reads the required data from the register module, and then sends the control signals and the data to the execution unit.
[0011] Further, the pre - processing module in the execution unit is used to perform polynomial operations and data transformation of the hash algorithm to prepare for the calculations;
[0012] The processing module is composed of multiple identical general - purpose computing modules and multiple hash computing modules, and is used to implement various computing operations; the processing module in the execution unit selects the connection method of each module in the processing module according to the control signals sent by the decoder unit to support the computing operations with different bit widths and different operation types required in the post - quantum cryptographic algorithm, and obtains the corresponding operation results;
[0013] The post - processing module is used to perform subsequent data transformation processing on the calculation results of the polynomial operation and the hash algorithm; after the data transformation processing, it sends the calculation results and the control signals generated by the execution unit to the write - back unit.
[0014] Further, the general - purpose computing module includes a multiplier module for implementing multiplication - related operations, a carry - save adder module for implementing modulo operations and logical operations, an adder module for implementing addition and subtraction operations, and a shifter for implementing shift operations; the hash computing module is composed of a carry - save adder module, an adder module, and a shifter module.
[0015] Further, the processing module is composed of eight identical general computing modules and two hash computing modules, and is used to implement various computing operations;
[0016] Each general computing module includes a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-save adder module for implementing modulo operations and logical operations, two 32-bit adder modules for implementing addition and subtraction operations, and a 32-bit shifter for implementing shift operations;
[0017] Each hash computing module is composed of a 32-bit carry-save adder module, two 32-bit adder modules, and a 32-bit shifter module.
[0018] Further, the execution unit further includes a read / write control unit, and the read / write control unit is connected to the decoding unit and the data storage unit outside the post-quantum cryptographic processor; the read / write control unit receives the control signals related to reading and writing sent by the decoding unit, and performs specific memory access operations on the data storage unit outside the post-quantum cryptographic processor through the control signals sent by the decoding unit, and writes back the results cached by accessing the data storage unit to the decoding unit according to the control signals generated by the read / write control unit.
[0019] Further, the write-back unit includes a non-read / write pipeline write-back unit, and the non-read / write pipeline write-back unit is connected to the decoding unit. The non-read / write pipeline write-back unit writes back the calculation results sent by the execution unit to the corresponding registers in the decoding unit according to the control signals generated by the execution unit.
[0020] The present invention also provides a system-on-chip for running the post-quantum cryptographic processor, including: a system bus, the post-quantum cryptographic processor connected to the system bus, a storage unit connected to the post-quantum cryptographic processor, and an external device interface connected to the system bus;
[0021] The post-quantum cryptographic processor communicates with other devices of the system-on-chip through the system bus, and also communicates with external devices connected to the external device interface through the system bus, and reads and writes data and instructions to the storage unit connected to the post-quantum cryptographic processor through instructions received from the external devices.
[0022] Further, the system-on-chip further includes: a cache active prefetch / presave data function module combined with DMA connected to the post-quantum cryptographic processor, and the cache active prefetch / presave data function module combined with DMA includes
[0023] The cache part includes a cache controller, a cache SRAM connected to the cache controller, a cache usage status record table connected to the cache controller, a communication port of the post-quantum cryptographic processor core connected to the cache controller, and a communication port of the DMA controller connected to the cache controller. The cache controller reads and writes the cache SRAM according to the cache usage status record table. The cache usage status record table records the cache units and operation types in the cache SRAM that the post-quantum cryptographic processing core and the DMA are reading / writing, and can automatically detect and avoid data read / write conflicts before the post-quantum cryptographic processor core reads and writes.
[0024] The DMA part includes a DMA controller, a transfer instruction list connected to the DMA controller, a communication port of the post-quantum cryptographic processor core connected to the DMA controller, a communication port of the cache controller connected to the DMA controller, an external memory read / write port connected to the DMA controller, and a cache SRAM read / write port connected to the DMA controller. The DMA controller reads and writes the external memory and the cache SRAM by reading the transfer instruction list. According to the control signals of the post-quantum cryptographic processor and the cache controller, the DMA controller also modifies the transfer instruction list. The DMA part is used to accept the transfer operations requested by the post-quantum cryptographic processor core and the cache controller, write the cache data back to the external memory or read the external memory data into the cache SRAM. It can automatically detect and avoid data read / write conflicts before the DMA controller reads and writes the cache SRAM. The transfer instruction list can store multiple transfer instructions arranged by the post-quantum cryptographic processor core and execute them sequentially. The transfer instruction list also has multiple highest-priority emergency task bars for handling the situation where the post-quantum cryptographic processor accesses the cache SRAM and fails, and reads data from the external memory urgently through the DMA.
[0025] Further, the system bus supports the AXI 32-bit bus protocol, and the AXI 32-bit bus protocol further supports the APB general peripheral protocol. Then, the SPI, UART, GPIO, and I2C peripheral protocols are connected through the APB general peripheral protocol. The access to the storage unit connected to the post-quantum cryptographic processor by the AXI bus protocol is all in 32 bits.
[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0027] (1) The present invention adopts the component recombination technology and the operator fusion technology, uses the same basic resources to support different types of post-quantum cryptographic algorithms through recombination, greatly reduces the hardware resource overhead, and reduces the number of cycles required by the algorithm through operator fusion, improving the running speed of the post-quantum cryptographic algorithm.
[0028] (2) The present invention employs 8 or 10 configurable parallel computing units to accelerate the key operations in post-quantum cryptographic algorithms, namely, the hash algorithm and polynomial calculation. When not in hash operation, only 8 computing units are enabled, thus saving power consumption during operation. The 10 computing units altogether process 320-bit data, which matches the data volume in the hash algorithm, thereby reducing the number of additional data accesses and solving the data dependence problem in the entire algorithm operation. It achieves an acceleration effect of more than 2 times with an overhead of 20%.
[0029] (3) The present invention implements the entire system-on-chip. Compared with directly using a 128-bit bus width, the 32-bit bus width saves the overhead of the entire system's resources, power consumption, etc. When the peripheral needs to access or operate on 128 bits, it can be completed through the designed data access module in at most 4 time cycles; and the 32-bit bus does not affect the execution efficiency of the processor. Through this design, the present invention reduces the power consumption and resource overhead of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is the architecture diagram of the post-quantum cryptographic algorithm processor of the present invention;
[0031] Figure 2 is the architecture diagram of the parallel computing unit module in the present invention;
[0032] Figure 3 is the architecture diagram of the system-on-chip in the present invention;
[0033] Figure 4 is the cache structure diagram of the present invention;
[0034] Figure 5 is the structure diagram of the DMA transfer instruction list in the present invention;
[0035] Figure 6 is the working flow chart of the cache controller;
[0036] Figure 7 is the working flow chart of the DMA controller;
[0037] Figure 8 is the replacement rule and update rule of the data block based on the binary tree-based pseudo least recently used (PLRU) strategy. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be further described and explained below in conjunction with the specific embodiments. The embodiments are only illustrative of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment in the present invention can be combined accordingly without conflict.
[0039] The present invention proposes a post - quantum cryptographic algorithm processor and its system - on - chip. The present invention aims to utilize the characteristics of single - instruction multiple - data, combined with the parallelism, high reusability, low power consumption and other characteristics of technologies such as computing component recombination technology and operator fusion technology, to accelerate the computing speed of post - quantum cryptographic algorithms with lower power consumption and resource overhead. In addition, the design of the system - on - chip enables the present invention to be applied to more application scenarios, supporting more applications while accelerating the post - quantum cryptographic algorithms.
[0040] (1) Processor part
[0041] As Figure 1 shown, an embodiment of the present invention provides a dedicated processor for post - quantum cryptographic algorithms based on RISC - V. The dedicated processor adopts a superscalar pipeline design with separate read - write and computing operations, having two pipelines that can execute in parallel, namely a read - write pipeline and a computing (non - read - write) pipeline. Each pipeline has four sequential parts: instruction fetch, decode, execute, and write - back.
[0042] In a specific example of the present invention, the processor specifically includes:
[0043] An instruction - fetch unit. The read - write and computing pipelines (non - read - write pipeline) share the same instruction - fetch unit. This part is used to obtain instructions from the instruction cache outside the processor and distribute the instructions to the decode parts of the two pipelines of the processor after pre - analysis. The instruction - fetch unit includes an instruction interface module that exchanges data with the instruction storage unit outside the processor as an interface, a pre - fetch module for pre - reading multiple instructions and caching them, and an instruction distribution module. The instruction distribution module adopts a sequential multi - issue scheme, that is, the two sequentially read instructions are respectively sent to the first decoder and the second decoder of the decode unit. This part sends non - read - write instructions to the computing pipeline and read - write instructions to the read - write pipeline;
[0044] A decode unit. This part mainly includes four register modules and two decoder - type modules, which are used to process the instructions sent by the instruction - fetch unit, translate these instructions into specific control signals, read the data in the register modules, and then send these control signals and data to the execute parts of the two pipelines; The register modules include a general - purpose register module for storing 32 - bit data, a parallel register bank module for storing 64 - bit data, and a parameter register module for storing 32 - bit parameters. These modules are shared by the two pipelines; The decoder - type modules include a first decoder module belonging to the computing pipeline and a second decoder module belonging to the read - write pipeline. These two decoder modules respectively translate the instructions received by the corresponding pipelines, read the required data from the required register modules, and then send the control signals and data of the corresponding pipelines to the execute units of their respective pipelines;
[0045] Execution unit, which is used to perform specific calculation or memory access operations according to the control signals and data sent by the decoding unit, and then send the calculation results and control signals to the write-back parts of the two pipelines. The execution unit of the calculation pipeline includes a configurable parallel computing unit module, which adopts a hybrid bit-width design and supports parallel computing of various bit-width data up to 320 bits; this configurable parallel computing unit module adopts the operator fusion technology to combine and reuse the hardware required for different basic operations to implement various basic operations and combined operations; the execution part of the read-write pipeline includes the first stage of the read-write control unit module, which is used to access the data cache outside the processor according to the control signals and data sent by the decoding unit, and the other parts of this module work in the write-back unit of the read-write pipeline;
[0046] Write-back unit, which is used to write back the calculation results sent by the execution units of the calculation pipeline and the read-write pipeline to the corresponding register modules included in the decoding unit according to the control signals. The write-back unit of the read-write pipeline includes the second stage of the read-write control unit module, which is used to return the results obtained by accessing the external data cache in the first stage to the processor and store them in the register modules.
[0047] Further, as Figure 1 shown, the decoding unit of the post-quantum cryptography algorithm specific processor based on RISC-V further includes 3 parts. The first decoder is used to temporarily store the signals and data that the instruction fetch part needs to send to the decoding part of the calculation pipeline. The second decoder is used to temporarily store the signals and data that the instruction fetch part needs to send to the decoding part of the read-write pipeline. The read-write control unit (the first stage) is used to temporarily store the signals and data that the execution part of the read-write pipeline needs to send to the write-back part. The read-write control unit (the second stage) is used to return the results obtained by accessing the external data cache in the read-write control unit (the first stage) to the decoding unit of the post-quantum cryptography algorithm processor and store them in the register modules.
[0048] As Figure 2 shown, in the embodiment of the present invention, the configurable parallel computing unit of the execution unit of the post-quantum cryptography algorithm processor is implemented by adopting the component recombination technology and the operator fusion technology, and the same components are used to implement the core operations of the post-quantum cryptography algorithm with different calculation types and bit-widths. This module includes:
[0049] Preprocessing module, which is used to perform pre-calculation data transformation on NTT (Fast Number-Theoretic Transform) and Keccak (Hash) algorithms;
[0050] The processing module, including the main computing resources, is composed of 8 identical sub-unit modules for implementing arithmetic operations and an additional 2 identical sub-unit modules for accelerating the Keccak algorithm; (Most configurable parallel computing units are accelerated by 8 sub-unit modules, while the hash algorithm requires 10 sub-unit modules to be configured to solve data dependency and other problems) Each sub-unit module contains a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-save adder module for implementing modulo and logical operations, two 32-bit adder modules for implementing addition and subtraction, and a 32-bit shifter for implementing shift operations; There are selectable connection methods between these modules. Several modules in the processing module have multiple connection methods, which are respectively used to support the computing operations in different operations with different bit widths required in the post-quantum cryptography algorithm. According to the different instructions obtained from the instruction fetch part and the different specific control signals translated by the decoding unit, the processing module in the execution unit will select the corresponding connection method according to the control signals and adopt different connection methods to perform different computing operations with different bit widths to obtain the operation results. By configuring the use of 8 sub-unit modules or 10 sub-unit modules, this processor can be used to accelerate the NTT and Keccak algorithms;
[0051] The post-processing module is used for data transformation of the calculated data of the NTT and Keccak algorithms.
[0052] (2) System on Chip
[0053] As Figure 3 shown, the embodiment of the present invention provides a system on chip (SoC) where the post-quantum cryptography algorithm dedicated processor runs. The shown example adopts a two-level bus structure of AXI4 and APB, and the data bit width is 32 bits for both. Internally, the AXI4 bus is used for interconnection, and communication with low-speed peripheral devices is achieved through the APB bus. The APB bus is connected to the AXI4 bus through an AXI2APB bridge.
[0054] Devices mounted on the APB bus can be divided into two categories: low-speed peripheral interfaces and control modules. Low-speed peripheral interfaces include: GPIO (General-Purpose Input / Output), UART (Universal Asynchronous Receiver / Transmitter), I2C (Inter-Integrated Circuit), SPI (Serial Peripheral Interface) host interfaces; control modules include SoC Control, FLL Control, Timer, Event Unit, etc. These devices all act as APB slaves and are interconnected with the APB bus through the IP reuse module of the bus model. The AXI4 bus connects the processor core, instruction RAM, data RAM, Flash, SPI slave interface, and advanced debug module (Adv.DebugUnit), etc. Among them, the processor core, SPI slave interface, and advanced debug module are connected to the bus as AXI masters; the two RAMs are connected to the bus as AXI slaves.
[0055] In the shown example, the processor can communicate with external devices through multiple communication protocols such as UART, I2C, SPI, etc. to read and write data. When the SoC starts, instructions will be written from external devices (such as storage devices like Flash) to the instruction storage unit through the UART / SPI protocol, passing through the APB bus and AXI bus in sequence. Data is also written to the data storage unit in the same way. In this example, instructions and data are distinguished by address division and written to different storage units according to different target addresses.
[0056] Since the AXI bus only implements 32 bits, when the processor core or external device accesses the instruction storage unit or data storage unit, the data access module will handle and cache this. When continuously writing data, the data access module will cache the data that arrives first until the cached data reaches the required bit width and then write it to the memory at once; when reading data, the data access module will adjust and filter the read data and finally send it to the AXI bus in the form of 32 bits. In the shown example, the bit width of the instruction storage unit is 64 bits to support the processor core to read two 32-bit instructions at once; the bit width of the data storage unit is 128 bits to support the custom read and write instructions of the processor core to read and write 128-bit data at once. When the read and write bit width is less than 128 bits, it is cached and processed according to the above rules.
[0057] (3) Cache active prefetch / presave data module combined with DMA (Direct Memory Access)
[0058] The module includes a cache part and a DMA part.
[0059] The cache part includes a cache controller, cache SRAM, a cache usage status record table, and communication ports for the processor core and the DMA controller. The cache part implements a four-way set-associative scheme with the least recently used (PLRU) replacement, write-back, and write-allocate policies based on the binary tree idea, supporting all basic functions of the cache. As Figure 4 shown, the cache part of the embodiment of the present invention is a four-way set-associative structure with an 8 kB cache space, 128 sets, four channels in each set, and two flag bits in each channel. The Valid bit is used to indicate whether there is data stored in the channel, and the Dirty bit is used to indicate whether the data in the channel has been written. The cache usage status record table records the cache units being read / written by the processing core and the DMA and the operation types. The working process of the cache is as Figure 6 shown. After receiving an instruction from the processor core, first, it is judged whether the address meets the requirements. For the address that meets the requirements, the following processing is performed: If it is a data read / write instruction, first, the channel and the data are matched for validity, and then the cache usage status record table is checked to know whether the DMA is performing read / write on the same location. If so, the DMA is allowed to process first, and finally, the data is written and read. If data needs to be transferred in or out from the external memory outside the on-chip system, the required address, data, and data transfer direction are written to the emergency status row of the transfer instruction list of the DMA; If it is a custom data pre-storage / pre-fetch instruction, the instruction is written to the normal row of the emergency status of the transfer instruction list of the DMA and waits for the DMA to process.
[0060] The DMA part includes a DMA controller, a transfer instruction list, communication ports with the processor core and the cache controller, external memory read / write ports, cache SRAM read / write ports, etc., and is used to accept the data transfer operations requested by the processor core and the cache controller, and write the cache data back to the external memory or read the external memory data into the cache SRAM. As Figure 5 shown, the transfer instruction list of the DMA can store four pre-storage / pre-fetch instructions arranged by the processor core and execute them in sequence. There is also a highest-priority emergency task bar, which is used to handle the situation where when the post-quantum cryptographic processor accesses the cache SRAM and fails, data needs to be urgently read from the external memory through the DMA. Each row of the transfer instruction list stores the address, data length, and transfer direction. The working process of the DMA is as Figure 7As shown, the DMA cyclically checks the transfer instruction list, preferentially checking the emergency rows and then the regular rows. If there is a need to transfer data, it first queries the cache usage status record table to find out whether the processor core is reading and writing to the same location. If so, it allows the processor core to perform the read and write first, and then starts to transfer the data. After transferring the data, it updates the address and data length of the corresponding row in the transfer instruction list. To improve efficiency as much as possible, the DMA will automatically select other channels in the same group when the processor core reads and writes a certain channel. Various conflicts will be encountered in various situations. The specific scheduling options are as follows:
[0061] The DMA is about to write data from a certain channel in the cache back to the external memory. In this case, the following several unexpected situations may occur:
[0062] 1. The processor core is reading this channel (Dirty = 0), and the read hits: the CPU reads the data; since the data in this channel has not been modified, the DMA does not need to move it.
[0063] 2. The processor core is reading this channel (Dirty = 1), and the read hits: the CPU and the DMA read the data simultaneously without interfering with each other.
[0064] 3. The processor core is about to move data into this channel (Dirty = 0), and the read misses: the CPU commands the DMA to move the data it needs; the DMA does not need to move it.
[0065] 4. The processor core is about to move data into this channel (Dirty = 1, and there is data in this channel that has not been written back), and the read misses: the CPU commands the DMA to move the data in the channel to the Flash, and then commands the DMA to move the data it needs.
[0066] 5. The processor core is writing to this channel (Dirty = 0), and the write hits: after the CPU finishes writing, the DMA moves the data. The data becomes dirty after writing and needs to be moved.
[0067] 6. The processor core is writing to this channel (Dirty = 1), and the write hits: after the CPU finishes writing, the DMA moves the data.
[0068] 7. The processor core is about to write to this channel (Dirty = 0), and the write misses: if the bit width of the data to be written is not 128 bits, the CPU needs to first command the DMA to move the data it needs, and then write the data; since the data in this channel has not been modified, the DMA does not need to move the data.
[0069] 8. The processor core is ready to write to the channel (Dirty = 1, there is data in the channel that has not been written back), write failure: The CPU commands the DMA to move the data in the channel to the Flash. If the bit width of the written data is not 128 bits, the CPU needs to command the DMA to move the data it needs, and finally write the data.
[0070] When the DMA is ready to write data to a certain channel, the following unexpected situations may occur:
[0071] 1. The processor core is reading the channel (Dirty = 0), read hit: The DMA moves the data to another channel.
[0072] 2. The processor core is reading the channel (Dirty = 1), read hit: The DMA moves the data to another channel.
[0073] 3. The processor core is ready to move data into the channel (Dirty = 0), read failure: The CPU commands the DMA to move the data it needs; if the data to be moved by the DMA and the CPU is different, the DMA moves the data to another channel.
[0074] 4. The processor core is ready to move data into the channel (Dirty = 1, there is data in the channel that has not been written back), read failure: The CPU commands the DMA to move the original data in the channel, and then move the data it needs; if the data to be moved by the DMA and the CPU is different, the DMA moves the data to another channel.
[0075] 5. The processor core is writing to the channel (Dirty = 0), write hit: The CPU writes; The DMA moves the data to another channel.
[0076] 6. The processor core is writing to the channel (Dirty = 1), write hit: The CPU writes; The DMA moves the data to another channel.
[0077] 7. The processor core is ready to write to the channel (Dirty = 0), write failure: If the bit width of the written data is not 128 bits, the CPU needs to first command the DMA to move the data it needs, and then write the data; if the data to be moved by the DMA and the CPU is different, the DMA moves the data to another channel.
[0078] 8. The processor core is ready to write to the channel (Dirty = 1, there is data in the channel that has not been written back), write failure: The CPU commands the DMA to move the original data in the channel. If the bit width of the written data is not 128 bits, the CPU needs to command the DMA to move the data it needs, and finally write the data; if the data to be moved by the DMA and the CPU is different, the DMA moves the data to another channel.
[0079] Such asFigure 8 As shown in Figure 8 , the replacement rule and update rule of data blocks based on the pseudo least recently used (PLRU) strategy of a binary tree are implemented as follows: Use a binary tree structure to save the historical access order information of data blocks. For a 4-way set-associative cache, set a three-bit register for each group, denoted as B[2:0].
[0080] The above embodiments only represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. The present invention can support multiple post-quantum cryptographic algorithms and various core operations included therein, and there are correspondingly dozens of dedicated extension instructions not listed yet and the combination manners of each component in the processing module of the corresponding parallel computing unit module. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. An on-chip system for running a post-quantum cryptographic processor, characterized in that, The post-quantum cryptographic processor includes: an instruction fetch unit, a decoding unit, an execution unit, and a write-back unit; The instruction fetch unit includes an instruction interface module that exchanges data with an instruction storage unit outside the processor as an interface, a prefetch module that pre-reads multiple instructions in the instruction interface module and caches them, and an instruction distribution module that distributes the instructions cached in the prefetch module; The instructions distributed by the instruction distribution module include read / write instructions and non-read / write instructions; The decoding unit receives the instructions emitted by the instruction distribution module, and translates the received instructions into specific control signals; Then it sends the control signals and the data inside the decoding unit to the execution unit; The execution unit performs specific calculation or memory access operations according to the control signals and data sent by the decoding unit; Then it sends the calculation results and the control signals generated by the execution unit to the write-back unit; The write-back unit writes back the calculation results sent by the execution unit to the decoding unit according to the control signals; The execution unit includes a configurable parallel computing unit module. The configurable parallel computing unit module adopts component recombination technology and operator fusion technology, and can use the same components to implement the core operations of post-quantum cryptographic algorithms with different calculation types and bit widths; The configurable parallel computing unit module consists of the following modules: A preprocessing module, which is used to perform polynomial operations and data transformation of hash algorithms, and prepare for calculations; A processing module, which consists of eight identical general computing modules and two hash computing modules, and is used to implement various calculation operations; The processing module in the execution unit selects the connection methods of each module in the processing module according to the control signals sent by the decoding unit to support the calculation operations with different bit widths and different operation types required in the post-quantum cryptographic algorithm, and obtains the corresponding operation results; Each general computing module contains a 32-bit multiplier module for implementing multiplication-related operations, a 32-bit carry-save adder module for implementing modulo operations and logical operations, two 32-bit adder modules for implementing addition and subtraction operations, and a 32-bit shifter for implementing shift operations; Each hash computing module consists of a 32-bit carry-save adder module, two 32-bit adder modules, and a 32-bit shifter module; A post-processing module, which is used to perform subsequent data transformation processing on the calculation results of polynomial operations and hash algorithms; After the data transformation processing, it sends the calculation results and the control signals generated by the execution unit to the write-back unit; The system-on-chip includes: a system bus, the post-quantum cryptographic processor connected to the system bus, a storage unit connected to the post-quantum cryptographic processor, an external device interface connected to the system bus, and a cache active prefetch / prefetch data function module combined with DMA connected to the post-quantum cryptographic processor; The post-quantum cryptographic processor communicates with other devices of the system-on-chip through the system bus, and also communicates with external devices connected to the external device interface through the system bus, and reads and writes data and instructions to the storage unit connected to the post-quantum cryptographic processor through the instructions received from the external devices; The cache active prefetch / presave data function module combined with DMA includes: The cache part includes a cache controller, a cache SRAM connected to the cache controller, a cache usage status record table connected to the cache controller, a communication port of the post-quantum cryptographic processor core connected to the cache controller, and a communication port of the DMA controller connected to the cache controller; the cache controller reads and writes the cache SRAM according to the cache usage status record table; the cache usage status record table records the cache units and operation types in the cache SRAM being read / written by the post-quantum cryptographic processing core and DMA, and can automatically detect and avoid data read / write conflicts before the post-quantum cryptographic processor core reads and writes; The DMA part includes a DMA controller, a transfer instruction list connected to the DMA controller, a communication port of the post-quantum cryptographic processor core connected to the DMA controller, a communication port of the cache controller connected to the DMA controller, an external memory read / write port connected to the DMA controller, and a cache SRAM read / write port connected to the DMA controller; the DMA controller reads and writes the external memory and the cache SRAM by reading the transfer instruction list; according to the control signals of the post-quantum cryptographic processor and the cache controller, the DMA controller also modifies the transfer instruction list; the DMA part is used to accept the transfer operations requested by the post-quantum cryptographic processor core and the cache controller, write the cache data back to the external memory or read the external memory data into the cache SRAM, can automatically detect and avoid data read / write conflicts before the DMA controller reads and writes the cache SRAM, the transfer instruction list can store multiple transfer instructions arranged by the post-quantum cryptographic processor core and execute them sequentially, and the transfer instruction list also has multiple highest-priority emergency task bars for handling the situation where the post-quantum cryptographic processor accesses the cache SRAM unsuccessfully, and reads data from the external memory urgently through DMA; The decoding unit includes a register module and a decoder module; the register module is used to store data, and the decoder module translates the received instructions, reads the required data from the register module, and then sends the control signals and the data to the execution unit; The execution unit further includes a read / write control unit, and the read / write control unit is connected to the decoding unit and the data storage unit outside the post-quantum cryptographic processor; the read / write control unit receives the read / write-related control signals sent by the decoding unit, and performs specific memory access operations on the data storage unit outside the post-quantum cryptographic processor according to the control signals sent by the decoding unit, and writes back the result cached by accessing the data storage unit to the decoding unit according to the control signals generated by the read / write control unit; The write-back unit includes a non-read / write pipeline write-back unit, and the non-read / write pipeline write-back unit is connected to the decoding unit. The non-read / write pipeline write-back unit writes back the calculation results sent by the execution unit to the corresponding registers in the decoding unit according to the control signals generated by the execution unit; The system bus supports the AXI 32-bit bus protocol, and the AXI 32-bit bus protocol further supports the APB general peripheral protocol; then, the APB general peripheral protocol is used to connect to the SPI, UART, GPIO, and I2C peripheral protocols; all accesses of the AXI bus protocol to the storage unit connected to the post-quantum cryptographic processor are performed in 32 bits.
Citation Information
Patent Citations
Caching policy in a multicore system on a chip (SOC)
CN108984428A
RISC-V-based processor special for post-quantum cryptography algorithm
CN116432765A
Data transceiving method, device and equipment and storage medium
CN117119074A
Hardware accelerator for cryptographic hash operations
TW201717573A
Computer system and data processing method
WO2023004762A1