Hardware acceleration device of SHA-3 encryption algorithm

By designing a hardware acceleration device based on the RISC-V extended instruction set and optimizing the communication between the coprocessor and the main processor, the problem of low resource consumption and high-speed processing of the SHA-3 encryption algorithm in edge artificial intelligence was solved, achieving efficient SHA-3 encryption processing and enhancing security and performance.

CN120893082APending Publication Date: 2025-11-04XI AN JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510742273.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to simultaneously achieve low resource consumption and high processing speed with the SHA-3 encryption algorithm in edge AI scenarios, and the external IP core approach results in high communication overhead, impacting overall efficiency.

Method used

Design a hardware acceleration device based on the RISC-V extended instruction set, including an instruction counter, instruction memory, integer register file, immediate extension unit, integer execution unit, integer memory access unit, SHA-3 module, and write-back unit. Optimize the communication mechanism between the coprocessor and the main processor, provide extended instructions dedicated to the SHA-3 algorithm, reduce resource usage, and improve processing speed.

Benefits of technology

It achieves efficient SHA-3 encryption processing in edge AI scenarios, which reduces resource consumption, improves processing speed, and enhances security and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120893082A_ABST
    Figure CN120893082A_ABST
Patent Text Reader

Abstract

A hardware acceleration device of an SHA-3 encryption algorithm based on an RISC-V extended instruction set comprises an instruction counter, an instruction memory, an integer register file, an immediate operand extension unit, an integer execution unit, an integer memory access unit, an SHA-3 module, a data memory and a write-back unit, the instruction counter controls the address of the currently executed instruction; the integer register file is a register file of an RISC-V integer instruction set, and the immediate operand extension unit is used for implementing symbol extension on immediate operand in an instruction; the integer execution unit and the integer access unit are used for an RISC-V shaping instruction, and the SHA-3 module is used for executing an expansion instruction; according to the device, the calculation step of the SHA-3 encryption algorithm is condensed into an expansion instruction. The device provides a special and efficient instruction set, and the architecture makes full use of hardware resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of artificial intelligence technology, and specifically relates to a hardware acceleration device for the SHA-3 encryption algorithm. Background Technology

[0002] In edge AI acceleration scenarios, deploying neural network models on application-specific integrated circuits (ASICs) can effectively reduce power consumption. However, these highly optimized models often have a more streamlined structure, potentially increasing the risk of reverse engineering attacks. To mitigate this risk, Physically Unclonable Functions (PUFs) can be used as hardware-level security credentials, acting as chip keys to protect them from unauthorized theft. The key generation process from the PUF involves using cryptographic primitives such as SHA-3 (Secure Hash Algorithm 3) through a Key Derivation Function (KDF) to ensure security. This combines the efficiency of ASICs with the security of PUFs, providing robust protection against key leakage while enjoying the advantages of low power consumption.

[0003] As an open-source instruction set architecture standard, RISC-V provides a flexible and highly customizable platform, making it possible to design specialized extension instructions for specific tasks. These dedicated extensions can significantly improve the processing performance of the SHA-3 algorithm. However, effectively balancing resource usage and performance metrics while pursuing high performance remains a pressing issue.

[0004] Current research generally faces the challenge of simultaneously achieving low resource consumption and high processing speed. To address the resource utilization issue, some methods attempt to accelerate data processing through V-expansion (vector expansion) and K-expansion (cryptographic expansion). While these expansions offer some performance gains, they are not specifically designed for cryptographic algorithms or SHA-3, leaving considerable room for optimization. Furthermore, to improve execution speed, some solutions employ external IP cores, utilizing coprocessors to accelerate the process. While this approach can speed up processing to some extent, its drawbacks include consuming bus resources and incurring high communication overhead between the processor and coprocessor, thus impacting overall efficiency. Summary of the Invention

[0005] To address the aforementioned issues, this disclosure provides a hardware acceleration device for the SHA-3 encryption algorithm based on the RISC-V extended instruction set. The device includes an instruction counter, an instruction memory, an integer register file, an immediate extension unit, an integer execution unit, an integer memory access unit, an SHA-3 module, a data memory, and a write-back unit. The instruction fetch stage includes an instruction counter and an instruction memory. The instruction memory contains compiled binary instructions, and the instruction counter controls the address of the currently executing instruction. The decoding stage includes an integer register file and an immediate extension unit. The integer register file is the register file of the RISC-V integer instruction set, and the immediate extension unit is used to sign-extend the immediate values ​​in the instructions. The decoded results are sent to the execution stage according to different encodings. The execution stage includes an integer execution unit, an integer memory access unit, and an SHA-3 module. The integer execution unit and the integer memory access unit are used for RISC-V integer instructions, and the SHA-3 module is used to execute extended instructions. The memory access stage includes a data memory. The write-back stage includes a write-back unit. This device condenses the calculation steps of the SHA-3 encryption algorithm into extended instructions.

[0006] This device not only aims to enhance the processing performance of SHA-3 through a carefully designed RISC-V extended instruction set, but also focuses on exploring a strategy that can ensure high processing speed while reducing resource usage. Specific measures include, but are not limited to: designing optimized extended instructions specifically for the SHA-3 algorithm and improving the communication mechanism between the coprocessor and the main processor to reduce overhead. In this way, it can provide more optimized options for both applications requiring efficient encryption operations and resource-constrained environments, thereby achieving better security and performance. Attached Figure Description

[0007] Figure 1 This is an overall circuit architecture diagram with an extended processor core provided in one embodiment of the present disclosure;

[0008] Figure 2 This is a flowchart of the Keccak-f[b] function provided in one embodiment of this disclosure;

[0009] Figure 3 This is a schematic diagram of a three-dimensional matrix provided in one embodiment of this disclosure;

[0010] Figure 4 This is a diagram of the initial storage format of the three-dimensional state matrix in SR provided in one embodiment of this disclosure;

[0011] Figure 5 This is a schematic diagram of the top-level interface of the SHA-3 module provided in one embodiment of this disclosure;

[0012] Figure 6This is a schematic diagram of the internal structure of the SHA-3 module provided in one embodiment of this disclosure. Detailed Implementation

[0013] In one embodiment, such as Figure 1 As shown, this disclosure provides a hardware acceleration device for the SHA-3 encryption algorithm based on the RISC-V extended instruction set. It includes an instruction counter, an instruction memory, an integer register file, an immediate extension unit, an integer execution unit, an integer memory access unit, an SHA-3 module, a data memory, and a write-back unit. The instruction fetch stage includes an instruction counter and an instruction memory. The instruction memory contains compiled binary instructions, and the instruction counter controls the address of the currently executing instruction. The decoding stage includes an integer register file and an immediate extension unit. The integer register file is the register file of the RISC-V integer instruction set, and the immediate extension unit is used to perform sign extension on the immediate values ​​in the instructions. The decoded results are sent to the execution stage according to different encodings. The execution stage includes an integer execution unit, an integer memory access unit, and an SHA-3 module. The integer execution unit and the integer memory access unit are used for RISC-V integer instructions, and the SHA-3 module is used to execute extended instructions. The memory access stage includes a data memory. The write-back stage includes a write-back unit. This device condenses the calculation steps of the SHA-3 encryption algorithm into extended instructions.

[0014] In this embodiment, Figure 1 This paper demonstrates a minimal overall architecture of a RISC-V processor core with extended instruction sets. The processor core supports the basic RISC-V integer instruction set and is extended to the SHA-3 instruction set. In addition to the modules for fetching, decoding, executing, accessing memory, and writing back for integer instructions, a dedicated SHA-3 module is added to execute extended instructions.

[0015] This device condenses the computation steps of the SHA-3 algorithm into extended instructions, enabling efficient execution of the algorithm at the hardware level. The SHA-3 algorithm is a permutation-based round algorithm, completing 24 identical rounds of operations within a single iteration. The transformation in each round is implemented by the Keccak-f[b] permutation function, which includes five algorithmic steps. .

[0016] The flow of the Keccak-f[b] function is as follows: Figure 2 As shown, the steps are executed sequentially. To accelerate hardware execution, the process specifically includes... The two steps are combined.

[0017] The input and output data for each step of the Keccak-f[b] function are as follows: Figure 3 As shown, all data are 1600 bits in length. This data forms a dimension... A three-dimensional state matrix, where each small square represents 1 bit of data. Given , Afterwards, along The blocks of data for the axes are called lanes. The three-dimensional state matrix contains 25 lanes, and each lane has 64 bits of data, as shown in the figure. It is 64. for , for The lane is called Or lane Specifically, if the three-dimensional state matrix is ​​denoted as... The corresponding lane is denoted as .

[0018] At the beginning of each round of the Keccak-f[b] function, the storage format of the three-dimensional state matrix in SR is as follows: Figure 4 As shown. The three-dimensional state matrix is ​​stored in SRs in units of lanes, with each SR storing data for exactly one lane. A total of 25 SRs are used, numbered SR0-SR24. Specifically, Figure 3 middle direction and The direction numbers are respectively , The lane, it is stored in the first Of the SRs, among them .

[0019] This device is used in the implementation of computational circuits for large-scale integrated circuits, with applications including hardware acceleration of encryption and decryption algorithms for FPGAs and encryption of neural network models for ASICs. Through RISC-V software programming, it can support a wider variety of encryption and decryption algorithms.

[0020] In another embodiment, the top-level interface of the SHA-3 module mainly includes input signals, output signals, and handshake signals with the front-end module.

[0021] In this embodiment, the top-level interface of the SHA-3 module is as follows: Figure 5 As shown in the figure, this diagram illustrates the interaction logic between the SHA-3 module and other external modules. It mainly contains three types of signals: the left part represents the input signals of this module, the right part represents the output signals of this module, and the top part represents the handshake signals with the preceding module.

[0022] In another embodiment, the input signal comes from a front-end decoding module and includes an input control signal, an input address, and input data. The input control signal contains a decoding result, which is used to represent the current operation type, operand, and immediate value. The input address and input data are both integer register read results, which are used to control the reading and writing of the status register and the passing of integer register parameters.

[0023] In this embodiment, the input signal comes from the front-end decoding module. The decoding module sends the decoded instruction, integer register read result, extended immediate value, and other information to the SHA-3 module based on the encoded field. Specifically, the input control signal contains the decoding result, used to indicate the current operation type, operands, immediate value, etc.; the input address and input data are both integer register read results, used to control the reading and writing of the status register and the passing of integer register parameters.

[0024] In another embodiment, the output signal is the result of a CSR instruction reading the status register.

[0025] In this embodiment, the CSR (Control and Status Register) related instructions in the RISC-V basic instruction set are used to read and write the status register. The SHA-3 output signal is the result of the CSR instruction reading the status register.

[0026] In another embodiment, the handshake signal is used to handle backpressure and pauses between the decoding module and the SHA-3 module.

[0027] In this embodiment, the handshake signal is used to handle backpressure and pauses between the decoding module and the SHA-3 module. When the instruction execution time exceeds one clock cycle, this signal can pause the execution of the upstream pipeline to ensure the correctness of dependent data. The pipeline continues execution after the extended instructions have been completed.

[0028] In another embodiment, the SHA-3 module includes a control module, a status register file, a temporary status register (TSR), an operand selector, and a functional arithmetic unit. The control module receives decoded information and generates control flow signals for the execution phase. The status register file contains 32 64-bit status registers, with indexes indicating the selected register addresses. The TSR provides operands, and the operand selector multiplexes operands from different sources and outputs them to the subsequent functional arithmetic unit. The functional arithmetic unit consists of multiple sub-modules, including an XOR unit, an AND unit, a shift unit, and a write-back selector.

[0029] In this embodiment, the internal structure of the SHA-3 module is as follows: Figure 6As shown. This module is managed by the control module and drives the various functional units through multiplexing and operation signals. It can be divided into the following main parts, including the control module (top), input / register resources (left side), and functional operation units (right side box).

[0030] The control module receives the decoded information and simultaneously generates control flow signals for the execution phase. When an instruction requires multiple iterations, the control module internally counts the iterations and generates different selection and operation signals for each submodule based on the iteration count. These signals determine which registers or memory locations to fetch operands from, what logical / shift operations to perform, and where to write the result back.

[0031] The status register file contains 32 status registers, with indexes indicating the address of the selected register. The TSR (Time Register) is a separately listed temporary status register that also provides operands. The operand selector multiplexes operands from different sources and outputs them to subsequent arithmetic units.

[0032] The arithmetic unit consists of multiple sub-modules, including an XOR unit (performing XOR operations), an AND unit (performing AND operations), and a shift unit (performing circular shifts). When one operand in the XOR unit is all 1s, the other operand can be inverted. Finally, the write-back unit writes the result back to the specified register or storage location.

[0033] The functional operation unit can be expanded to achieve better performance with more resources. For example, adding an extra inverting unit can reduce the load on the XOR unit. Adding a pre-circular right shift function of 1 bit to the XOR unit can speed up the process. The speed at which instructions are executed in each step.

[0034] In another embodiment, the extended instruction set includes five extended instructions specifically designed for the SHA-3 encryption algorithm: sha3_reduce_xor, sha3_shift_xor, sha3_ring_set, sha3_ring_mv, and sha3_inv_and_xor. These instructions can efficiently execute key steps in the SHA-3 algorithm. .

[0035] In this embodiment, the extended instruction set is shown in Table 1, containing a total of 5 SHA-3 extended instructions. These extended instructions utilize the reserved encoding space of RISC-V and operate on specially defined registers called SRs (State Registers). There are a total of 32 SRs (specifically SR0-SR31), each 64 bits long, and all are readable and writable registers. In addition, there is a TSR (Temporary State Register), whose length is the same as the SRs.

[0036]

[0037] Table 1

[0038] In another embodiment, the sha3_reduce_xor instruction performs a reduce XOR operation on the status register; the sha3_shift_xor instruction performs a shift XOR operation on the status register; the sha3_ring_set instruction configures the initial loop status register; the sha3_ring_mv instruction moves the status register cyclically and allows shift operations on the moved register data; and the sha3_inv_and_xor instruction performs inversion, AND, and XOR operations on the status register.

[0039] For this embodiment, the instructions are described in detail below:

[0040] sha3_reduce_xor (SHA-3 reduce XOR instruction): The sha3_reduce_xor instruction performs a reduce XOR operation on the status register. The complete instruction mnemonic is "sha3_reduce_xor srd, srs, rs(stride),imm_num", where srd refers to the destination status register, srs refers to the source status register, rs is an integer register storing the iteration step size, and imm_num is an immediate number representing the number of iterations for this instruction.

[0041] When this instruction is executed, it first performs an XOR operation on the two source status registers, with register indices srs and srs+stride, respectively. Next, if imm_num is not 0, it performs another XOR operation on the result of the previous XOR operation and SR(j+stride). Finally, it decrements imm_num by 1, increments j by stride, and repeats the reduction operation until imm_num is 0. The result of the reduction XOR operation is ultimately stored in the destination status register.

[0042] sha3_shift_xor (SHA-3 shift XOR instruction): The sha3_shift_xor instruction performs a shift XOR operation on the status register. The complete instruction mnemonic is "sha3_shift_xor srd, srs1, srs2, rs(stride), imm_num, bit_flag", where srd refers to the destination status register, srs1 refers to the source status register 1, srs2 is the source status register 2, rs is an integer register storing the iteration step size, imm_num is an immediate value representing the number of iterations for this instruction, and bit_flag is a 1-bit flag.

[0043] This instruction has two modes, with the specific behavior determined by the bit_flag value. In the first case, when bit_flag is 0, it loops through imm_num+1 basic iterations. In the i-th iteration ( First, the data in the srs2 status register is cyclically shifted right by 1 bit. Then, the shifted result is XORed with the srs1+stride*i status register. The resulting destination status register has a register index of srd+stride*i. In the second case, when bit_flag is 1, the basic iteration is executed 1+imm_num.1 times. In the i-th iteration... The XOR operation is performed on the two source status registers, with register indices srs1+stride*i and srs2 respectively. The result is placed in the destination status register with register index srd+stride*i.

[0044] sha3_ring_set (SHA-3 Ring Move Configuration Instruction): The sha3_ring_set instruction configures the initial ring status register. The complete instruction mnemonic is "sha3_ring_set srs". It sets the initial status register (SRS) to be used by subsequent sha3_ring_mv instructions. This instruction puts SRS into the TSR to assist subsequent ring move instructions.

[0045] sha3_ring_mv (SHA-3 circular shift execution instruction): The sha3_ring_mv instruction can circularly shift the status register and simultaneously allow shift operations on the shifted register data. The complete instruction mnemonic is "sha3_ring_mv srd, rs(shift_offset)", where srd refers to the destination status register, the source status register srs has been set by the previously executed sha3_ring_set instruction, and shift_offset is the number of bits to store the circular right shift.

[0046] The instruction first circularly shifts the data in the TSR to the right by `shift_offset` bits, while simultaneously putting the data in the srd into the TSR. Then, it stores the shift result back into the srd. This method avoids overwriting the data in the srd and ensures that the source data for the next `sha3_ring_mv` instruction is written into the TSR.

[0047] sha3_inv_and_xor (SHA-3 ternary instruction): The sha3_inv_and_xor instruction can perform inversion, AND, and XOR operations on the status registers. The complete instruction mnemonic is "sha3_inv_and_xor srd, srs, rs(stride), imm_num", where srd refers to the destination status register, srs refers to the source status register 1, source status register 2 is obtained by adding stride[15:0] to srs, source status register 3 is obtained by adding 2*stride[15:0] to srs, and imm_num is an immediate value representing the number of iterations of this XOR operation.

[0048] The instruction inverts source status register 2, performs a bitwise AND operation with source status register 3, XORs the result with source status register 1, and stores the result in the destination status register. The instruction iterates through imm_num+1 times, and in the i-th iteration... The srd and srs of the instruction are both incremented by stride[31:16].

[0049] In another embodiment, In the step, the input state matrix Output state matrix This process mainly consists of three steps:

[0050] First: Output 5 lanes, called... ,in For all that satisfy of Perform an XOR operation to obtain the corresponding... ;

[0051] Second: Output 5 lanes, called... ,in .Will The result of circularly shifting right by 1 bit, and Perform an XOR operation to obtain the corresponding... ;

[0052] Third: Output , for and The result of the XOR operation, where .

[0053] In this embodiment, The algorithm steps are explained as follows: In this step, the input state matrix... Output state matrix The format is as follows Figure 3 As shown, it mainly consists of three steps. The instruction mapping is as follows:

[0054] The SHA-3 extension instructions corresponding to Step 1 are:

[0055] sha3_reduce_xor SR25, SR0, rs(0x5), 0x4

[0056] sha3_reduce_xor SR26, SR1, rs(0x5), 0x4

[0057] sha3_reduce_xor SR27, SR2, rs(0x5), 0x4

[0058] sha3_reduce_xor SR28, SR3, rs(0x5), 0x4

[0059] sha3_reduce_xor SR29, SR4, rs(0x5), 0x4

[0060] Among them, SR25-SR29 respectively store to .

[0061] The SHA-3 extension instructions corresponding to Step 2-3 are:

[0062] sha3_shift_xor SR30, SR29, SR26, rs(0x0), 0x0, 1'b0

[0063] sha3_shift_xor SR0, SR0, SR30, rs(0x1), 0x4, 1'b1

[0064] sha3_shift_xor SR30, SR25, SR27, rs(0x0), 0x0, 1'b0

[0065] sha3_shift_xor SR5, SR5, SR30, rs(0x1), 0x4, 1'b1

[0066] sha3_shift_xor SR30, SR26, SR28, rs(0x0), 0x0, 1'b0

[0067] sha3_shift_xor SR10, SR10, SR30, rs(0x1), 0x4, 1'b1

[0068] sha3_shift_xor SR30, SR27, SR29, rs(0x0), 0x0, 1'b0

[0069] sha3_shift_xor SR15, SR15, SR30, rs(0x1), 0x4, 1'b1

[0070] sha3_shift_xor SR30, SR28, SR25, rs(0x0), 0x0, 1'b0

[0071] sha3_shift_xor SR20, SR20, SR30, rs(0x1), 0x4, 1'b1

[0072] Each pair of consecutive instructions forms a group, for a total of 5 groups. Group( )calculate ,in SR30 stores the corresponding .

[0073] In another embodiment, In the step, the input state matrix Output state matrix This step is for Perform a circular right shift operation on each lane to obtain ; In the step, the input state matrix Output state matrix This step moves The position of each lane, except; In the step, the input state matrix Output state matrix Output Calculation method: First, for Invert the result and Do and operate, then sum the results Perform an XOR operation to obtain the output lane, where In step ι, the input state matrix Output state matrix This step only applies to Perform an XOR operation with the 64-bit round constant.

[0074] In this embodiment, The algorithm steps are explained below: In the step, the input state matrix Output state matrix The format is as follows Figure 3 As shown. This step is for Perform a circular right shift operation on each lane to obtain Specifically, the number of shift bits is shown in Table 2.

[0075]

[0076] Table 2

[0077] In the step, the input state matrix Output state matrix The format is as follows Figure 3 As shown. This step moves... The position of each lane ( (Except for). The specific shift loop is shown in the linked list below, position The lane is not shifted, the start and end registers of the linked list are in the same position, and each lane appears exactly once.

[0078] (1, 0) --> (0, 2) --> (2, 1) --> (1, 2) --> (2, 3) --> (3, 3) --> (3,0) --> (0, 1) --> (1, 3) --> (3, 1) --> (1, 4) --> (4, 4) --> (4, 0) --> (0,3) --> (3, 4) --> (4, 3) --> (3, 2) --> (2, 2) --> (2, 0) --> (0, 4) --> (4,2) --> (2, 4) --> (4, 1) --> (1, 1) --> (1, 0)

[0079] The instruction mapping is as follows:

[0080] To accelerate computation, the two steps are combined during instruction mapping; each lane performs a shift operation immediately after a cyclic shift. The corresponding SHA-3 extended instructions are:

[0081] sha3_ring_set SR5

[0082] sha3_ring_mv SR2, 0x1

[0083] sha3_ring_mv SR11, 0x3

[0084] sha3_ring_mv SR7, 0x6

[0085] sha3_ring_mv SR13, 0xa

[0086] sha3_ring_mv SR18, 0xf

[0087] sha3_ring_mv SR15, 0x15

[0088] sha3_ring_mv SR1, 0x1c

[0089] sha3_ring_mv SR8, 0x24

[0090] sha3_ring_mv SR16, 0x2d

[0091] sha3_ring_mv SR9, 0x37

[0092] sha3_ring_mv SR24, 0x2

[0093] sha3_ring_mv SR20, 0xe

[0094] sha3_ring_mv SR3, 0x1b

[0095] sha3_ring_mv SR19, 0x29

[0096] sha3_ring_mv SR23, 0x38

[0097] sha3_ring_mv SR17, 0x8

[0098] sha3_ring_mv SR12, 0x19

[0099] sha3_ring_mv SR10, 0x2b

[0100] sha3_ring_mv SR4, 0x3e

[0101] sha3_ring_mv SR22, 0x12

[0102] sha3_ring_mv SR14, 0x27

[0103] sha3_ring_mv SR21, 0x3d

[0104] sha3_ring_mv SR6, 0x14

[0105] sha3_ring_mv SR5, 0x2c

[0106] The first instruction configures the initial loop status register, and the following 24 instructions perform shift operations cyclically, in the same order as the aforementioned linked list. The number of shift bits in each instruction is the result of taking the remainder after dividing by 64.

[0107] The algorithm steps are explained as follows: In this step, the input state matrix... Output state matrix The format is as follows Figure 3 As shown. Output The calculation method is shown below, where First of all, Invert the result and Do and operate, then sum the results Perform an XOR operation to obtain the corresponding lane in the output.

[0108] The instruction mapping is as follows:

[0109] This step has a relatively consistent operation and can be mapped to the following instruction:

[0110] sha3_inv_and_xor SR0, SR0, rs(0x10005), 0x18

[0111] This instruction takes a relatively long time to execute, but it can complete the calculations for all lanes.

[0112] The algorithm for step ι is explained as follows: In this step, the input state matrix... Output state matrix The format is as follows Figure 3 As shown. This step only applies to... Perform an XOR operation with the 64-bit round constant.

[0113] The instruction mapping is as follows:

[0114] sha3_shift_xor SR0, SR0, SR31, rs(0x0), 0x0, 1'b1

[0115] This instruction requires the wheel constant of the current round to be stored in SR31. Therefore, the wheel constant of the current round must be written into SR31 before the instruction is issued.

[0116] Although embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments and application fields described above. The specific embodiments described above are merely illustrative and instructive, and not restrictive. Those skilled in the art can make many other forms based on the guidance of this specification and without departing from the scope of protection of the claims of the present invention, and all of these are within the scope of protection of the present invention.

Claims

1. A hardware acceleration device for the SHA-3 encryption algorithm based on the RISC-V extended instruction set, comprising an instruction counter, an instruction memory, an integer register file, an immediate extension unit, an integer execution unit, an integer memory access unit, an SHA-3 module, a data memory, and a write-back unit, wherein, The instruction fetch stage includes an instruction counter and an instruction memory. The instruction memory contains compiled binary instructions, and the instruction counter controls the address of the currently executing instruction. The decoding stage includes an integer register file and an immediate extension unit. The integer register file is the register file of the RISC-V integer instruction set, and the immediate extension unit is used to perform sign extension on the immediate values ​​in the instructions. The decoded results of the instructions are sent to the execution stage according to different encodings. The execution stage includes an integer execution unit, an integer memory access unit, and an SHA-3 module. The integer execution unit and the integer memory access unit are used for RISC-V integer instructions, and the SHA-3 module is used to execute extended instructions. The memory access stage includes a data memory. The write-back stage includes a write-back unit. This device condenses the calculation steps of the SHA-3 encryption algorithm into extended instructions.

2. The apparatus according to claim 1, preferably, the top-level interface of the SHA-3 module includes an input signal, an output signal, and a handshake signal.

3. The apparatus according to claim 2, wherein the input signal comes from a front-end decoding module and includes an input control signal, an input address, and input data, wherein the input control signal contains a decoding result, used to represent the current operation type, operand, and immediate value, and the input address and input data are both integer register read results, used to control the reading and writing of the status register and the passing of integer register parameters.

4. The apparatus according to claim 2, wherein the output signal is the result of the CSR instruction reading the status register.

5. The apparatus according to claim 2, wherein the handshake signal is used to handle backpressure and pauses between the decoding module and the SHA-3 module.

6. The apparatus according to claim 1, wherein the SHA-3 module includes a control module, a status register file, a temporary status register (TSR), an operand selector, and a function arithmetic unit, wherein, The control module receives decoded information and generates control flow signals. The status register file contains 32 64-bit status registers, with the selected register address indicated by an index. The temporary status register (TSR) provides operands, and the operand selector multiplexes operands from different sources and outputs them to the subsequent functional arithmetic unit. The functional arithmetic unit consists of multiple sub-modules, including an XOR unit, an AND unit, a shift unit, and a write-back selector.

7. The apparatus according to claim 1, wherein the extended instruction set includes five extended instructions specifically designed for the SHA-3 encryption algorithm, namely sha3_reduce_xor, sha3_shift_xor, sha3_ring_set, sha3_ring_mv, and sha3_inv_and_xor, which can efficiently execute key steps in the SHA-3 algorithm. .

8. The apparatus according to claim 1, wherein the sha3_reduce_xor instruction can perform a reduce XOR operation on the status register; the sha3_shift_xor instruction can perform a shift XOR operation on the status register; the sha3_ring_set instruction is used to configure the initial loop status register; the sha3_ring_mv instruction can cyclically move the status register and simultaneously allow shift operations to be performed on the moved register data; and the sha3_inv_and_xor instruction can perform invert, AND, and XOR operations on the status register.

9. The apparatus according to claim 7, In the step, the input state matrix Output state matrix This process mainly consists of three steps: First: Output 5 lanes, called... ,in For all that satisfy of Perform an XOR operation to obtain the corresponding... ; Second: Output 5 lanes, called... ,in ,Will The result of circularly shifting right by 1 bit, and Perform an XOR operation to obtain the corresponding... ; Third: Output , for and The result of the XOR operation, where .

10. The apparatus according to claim 7, In the step, the input state matrix Output state matrix This step is for Perform a circular right shift operation on each lane to obtain ; In the step, the input state matrix Output state matrix This step moves The position of each lane, except; In the step, the input state matrix Output state matrix Output Calculation method: First, for Invert the result and Do and operate, then sum the results Perform an XOR operation to obtain the output lane, where In step ι, the input state matrix Output state matrix This step only applies to Perform an XOR operation with the 64-bit round constant.