Chip implementation device and method for ultra-lightweight Ascon hash algorithm

By designing a chip implementation device of the ultra-lightweight Ascon hashing algorithm, optimizing the permutation network and pre-computing optimization solution, the problem of difficult to balance throughput and resources in the existing technology of Ascon hashing algorithm chips is solved, and efficient and secure hashing calculations are realized, suitable for resource-constrained devices.

CN120180467APending Publication Date: 2025-06-20SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244798.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing Ascon hashing algorithm chip implementation has problems that throughput and chip resources are difficult to balance, which makes it difficult for traditional serialized state update strategies to meet the efficient hashing computing needs of resource-constrained devices.

Method used

A chip implementation device of an ultra-lightweight Ascon hashing algorithm is designed, including a control unit, a replacement network, an initialization unit, an information absorption unit and a hash generation unit. By optimizing the design of the replacement network, a pre-computation optimization and throughput optimization scheme, efficient hashing calculation is achieved.

Benefits of technology

The Ascon hashing algorithm chip with smaller area and faster area occupancy is realized, which is suitable for scenarios such as the Internet of Things and embedded systems with limited resources, improving the efficiency and performance of chip implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180467A_ABST
    Figure CN120180467A_ABST
Patent Text Reader

Abstract

The invention discloses a chip implementation device and method for an ultra-lightweight Ascon hash algorithm. The device comprises a control unit, a replacement network, an initialization unit, an information absorption unit and a hash generation unit. The control unit controls the Hash algorithm to sequentially go through an initialization stage, an information absorption stage and a Hash generation stage; the replacement network is used for updating the internal state S; the initialization unit is used for connecting and fixing an initial vector IV and 0 of 256 bits in an initialization state, so that an initial value is assigned to an internal state S; the information absorption unit is used for absorbing the updated internal state S into the input information M, and updating the internal state S through a permutation network in the absorption process; the Hash generation unit generates and outputs a Hash value by using the internal state S of the absorbed input information M, and updates the internal state S through the permutation network in the process of generating the Hash value. The device and the method disclosed by the invention are smaller in occupied area and fastest in speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an ultra-lightweight Ascon hash algorithm, and particularly to a chip implementation device and method for an ultra-lightweight Ascon hash algorithm. Background Art

[0002] With the continuous development of the information society, the problem of data security has become increasingly prominent. Especially in resource-constrained devices such as the Internet of Things and embedded systems, how to implement an efficient and secure hash algorithm has become the key to data protection technology. Although traditional hash functions provide high security, their computational complexity is relatively high, and they require a large amount of processor and storage resources, which are not suitable for embedded devices and Internet of Things terminals. Therefore, lightweight hash algorithms have become an important direction to solve this problem.

[0003] The Ascon algorithm is the lightweight encryption algorithm standard selected by the US National Institute of Standards and Technology. It is designed specifically for resource-constrained environments and has the characteristics of high efficiency, security, and flexibility. Ascon-Hash, Ascon-HashA, Ascon-Xof, and Ascon-XofA in the Ascon hash part inherit the security and efficiency of Ascon authenticated encryption and are suitable for scenarios where data integrity needs to be protected and forgery needs to be prevented.

[0004] The core design of the Ascon hash algorithm is based on the sponge structure. Its internal operation adopts an update operation of a 320-bit state and realizes efficient data processing through simple bit operations such as exclusive OR (XOR) and bitwise AND (AND). However, there are still significant bottlenecks in the existing chip implementation: the traditional serialization state update strategy makes it difficult to balance throughput and chip resources. Therefore, there is an urgent need for a configurable chip architecture with a small area and high throughput to break through the performance boundaries of existing solutions. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a chip implementation device and method for an ultra-lightweight Ascon hash algorithm to achieve the purpose of smaller area occupation and the fastest speed.

[0006] To achieve the above purpose, the technical solution of the present invention is as follows:

[0007] A chip implementation device for an ultra-lightweight Ascon hash algorithm includes a control unit, a permutation network, an initialization unit, an information absorption unit, and a hash generation unit;

[0008] The control unit controls the hash algorithm to go through an initialization stage, an information absorption stage, and a hash generation stage in sequence;

[0009] The permutation network is used to update the internal state S;

[0010] The initialization unit is used to fix the connection between the initial vector Ⅳ and 256 bits of 0 in the initialization state, so as to initialize the internal state S, and send it to the permutation network for 12 rounds of update to update the internal state S;

[0011] The information absorption unit is used to absorb the input information M into the updated internal state S, and update the internal state S through the permutation network during the absorption process. After all the input information M is absorbed, it is sent to the permutation network for update;

[0012] The hash generation unit generates and outputs a hash value by using the internal state S that absorbs the input information M, and updates the internal state S through the permutation network during the process of generating the hash value.

[0013] In the above solution, the control unit includes a counter and a state machine. The counter has 4 bits, and the counting range is from 0 to 15. It is cleared after reset or update completion, and increments by 1 each time a start signal for update is received. The counter records the number of update rounds to control the state transition; the state machine has a total of 4 states: idle state, initialization state, information absorption state, and hash generation state, which are used to indicate different stages.

[0014] In the above solution, the permutation network is denoted as p, which is composed of round constant addition, substitution layer, and linear diffusion layer, and is used to update the 320-bit internal state S; in the permutation network, the internal state S is divided into 5 64-bit data blocks; in the round constant addition stage, the third 64-bit data block is XORed with the round constant; in the substitution layer, NOT, AND, and XOR operations are performed on the data blocks; in the linear diffusion layer, shift and XOR operations are performed on each data block.

[0015] In a further technical solution, in the permutation network, the 320-bit internal state is evenly divided into 5 64-bit long data blocks (x0, x1, x2, x3, x4), and the internal state S = x0||x1||x2||x3||x4;

[0016] The operation of round constant addition is expressed as:

[0017] x2←x2⊕c r

[0018] If a total of 12 rounds of update are required, c r = 8'hf0 - (count value - 1) * 15;

[0019] If a total of 8 rounds of update are required, c r = 8'hb4 - (count value - 1) * 15;

[0020] The operation of the substitution layer is expressed as:

[0021]

[0022] The operation of the linear diffusion layer is expressed as:

[0023] x0←x0⊕(x0>>19)⊕(x0>>28)

[0024] x1←x1⊕(x1>>61)⊕(x1>>39)

[0025] x2←x2⊕(x2>>1)⊕(x2>>6)

[0026] x3←x3⊕(x3>>10)⊕(x3>>17)

[0027] x4←x4⊕(x4>>7)⊕(x4>>41)

[0028] Wherein, ‖ represents concatenating two binary numbers, ⊕ represents exclusive OR, represents taking the bitwise complement of x, >> represents circular right shift, c r is a round constant.

[0029] In the above solution, the information absorption unit divides the input information into several information blocks of r bits. If the last information block is less than r bits, it is padded with a bit string starting with 1 followed by several 0s to make its length equal to r; after taking the first r bits of the internal state S and performing exclusive OR with the first information block, it is sent to the permutation network for b rounds of update. After the update, the first r bits of the internal state S are taken and exclusive OR with the next information block, and so on, until after performing exclusive OR with the last information block, the internal state S is sent to the permutation network for 12 rounds of update.

[0030] In the above solution, the hash generation unit first obtains a 320-bit internal state S from the permutation network, and copies the first r bits of the 320-bit internal state S as the first hash block; then the 320-bit internal state S is sent to the permutation network for b rounds of update. After the update is completed, the first r bits of the 320-bit internal state S are copied as the second hash block, and the generation process of the second hash block is repeated until t hash blocks of length r bits are generated; t is the smallest integer greater than l / r. The t-th hash block is intercepted by l mod r bits, and is concatenated with the previous t - 1 hash blocks in sequence to obtain a hash value H of length l bits and output.

[0031] A chip implementation method of an ultra-lightweight Ascon hash algorithm, adopting the chip implementation device of an ultra-lightweight Ascon hash algorithm as described above, includes the following steps:

[0032] Step 1, initialization phase:

[0033] Under the control of the control unit, after the device receives the hash value generation start signal, it enters the initialization state. The initialization unit concatenates the initial vector IV with 256 bits of 0 as the initial value of the internal state S, and then sends the internal state S into the permutation network for 12 rounds of update. Each round of update sequentially passes through round constant addition, substitution layer, and linear diffusion layer. The counter records the number of update rounds. When the count value reaches 11, the initialization state ends, the counter is cleared, and the internal state S is sent to the message absorption unit;

[0034] Step 2, message absorption phase:

[0035] In the message absorption unit, the input message M is divided into several message blocks of r bits. If the last message block is less than r bits, it is padded with a bit string starting with 1 followed by several 0s to make its length equal to r; after taking the first r bits of the internal state S and performing exclusive OR with the first message block, after completion, the internal state S is sent into the permutation network for b rounds of update. Each round of update sequentially passes through round constant addition, substitution layer, and linear diffusion layer. The counter records the number of update rounds. When the count value is b - 1, it indicates that the message block is processed, and the counter is cleared; then take the first r bits of the internal state S and perform exclusive OR with the next message block. After completion, send the internal state S into the permutation network for b rounds of update. Repeat this process until after performing exclusive OR with the last message block, send the internal state S into the permutation network for 12 rounds of update. After the update is completed, the message absorption phase is completed;

[0036] Step 3, hash generation phase:

[0037] The hash generation unit first obtains a 320-bit internal state from the permutation network, copies the first r bits of the 320-bit internal state as the first hash block, and then sends the 320-bit internal state into the permutation network for b rounds of update. After the update is completed, copy the first r bits of the 320-bit internal state as the second hash block. Repeat the process of generating the second hash block until t hash blocks of length r bits are generated; t is the smallest integer greater than l / r. The t-th hash block intercepts l mod r bits. After concatenating all the hash blocks, a hash value H of length l bits is obtained. After outputting the hash value H, the hash generation phase is completed.

[0038] In the above solution, the ultra-lightweight Ascon hash algorithm includes algorithm Ascon-Hash, algorithm Ascon-HashA, algorithm Ascon-Xof, and algorithm Ascon-XofA, where r = 64 in the message absorption unit and the hash generation unit;

[0039] In the chip implementation device of algorithm Ascon-Hash, the initial vector value Ⅳ = 00400c0000000100 in the initialization unit, the required number of update rounds b = 12, and the generated hash value H is 256 bits long;

[0040] In the chip implementation device of the Ascon-HashA algorithm, the initial vector value Ⅳ in the initialization unit is 00400c0400000100, the required number of update rounds b is 8, and the generated hash value H is 256 bits long;

[0041] In the chip implementation device of the Ascon-Xof algorithm, the initial vector value Ⅳ in the initialization unit is 00400c0000000000, the required number of update rounds b is 12, and the generated hash value H is l bits long for any l;

[0042] In the chip implementation device of the Ascon-XofA algorithm, the initial vector value Ⅳ in the initialization unit is 00400c0400000000, the required number of update rounds b is 8, and the generated hash value H is l bits long for any l.

[0043] In a further technical solution, the update result in the initialization stage is solidified into the chip, and the four algorithms correspond to four internal states S after the initialization ends; under the pre-computation optimization method, the initialization stage is omitted, and the internal state after 12 rounds of update of the permutation network in the initialization stage is directly obtained according to the algorithm, and the information absorption stage is directly entered.

[0044] In a further technical solution, for the permutation network processing, under the throughput optimization algorithm, the permutation network is copied w times, and w updates are performed per clock cycle;

[0045] Under the throughput optimization algorithm, the internal state update is expressed as:

[0046] S←(p w ) 12 / w (IV||0 256 )

[0047]

[0048] where, ‖ represents concatenating two binary numbers; ⊕ represents exclusive OR; p w represents performing w rounds of update; M represents the input information, M i represents the i-th information block of the input information; S x~y represents the x-th to y-th bits of the internal state S; s represents the number of information blocks after the input information is filled; t represents the number of output hash blocks.

[0049] Through the above technical solutions, a chip implementation device and method of an ultra-lightweight Ascon hash algorithm provided by the present invention have the following beneficial effects:

[0050] The Ascon hash algorithm fully considers lightweight and security in its design. Its sponge-structure-based design enables it to perform excellently in resource-constrained environments and is suitable for scenarios such as the Internet of Things and embedded systems. This invention designs a chip implementation solution based on the Ascon hash algorithm, covering the chip implementations of four algorithms: Ascon-Hash, Ascon-HashA, Ascon-Xof, and Ascon-XofA, and supports 256-bit and variable hash-length outputs.

[0051] This invention optimizes the design of the permutation network and adopts a combination of round constant addition, substitution layer, and linear diffusion layer to ensure efficient data processing and security. In addition, this invention also provides pre-computation optimization and throughput optimization schemes, further enhancing the efficiency and performance of the chip implementation. By solidifying the internal state in the initialization stage, the computational overhead in the initialization stage is reduced, while the throughput optimization algorithm significantly reduces the number of processing cycles and improves the overall throughput by parallelly processing multiple update operations.

[0052] The hardware implementation device and method of this invention not only meet the requirements of resource-constrained devices for efficient and secure hash algorithms but also enhance the flexibility and applicability of the hardware implementation through optimization schemes, and can be widely applied to the hash calculation requirements in fields such as the Internet of Things and embedded systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of this invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0054] Figure 1 Schematic diagram of a chip implementation device for a ultra-lightweight Ascon hash algorithm disclosed in an embodiment of this invention;

[0055] Figure 2 Flowchart of a chip implementation method for a ultra-lightweight Ascon hash algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of this invention with reference to the drawings in the embodiments of this invention.

[0057] This invention provides a chip implementation device for a ultra-lightweight Ascon hash algorithm, as Figure 1 shown, which includes a control unit, a permutation network, an initialization unit, an information absorption unit, and a hash generation unit.

[0058] I. Control Unit

[0059] The control unit controls the hash algorithm to go through the initialization stage, the information absorption stage, and the hash generation stage in sequence.

[0060] The control unit includes a counter and a state machine. The counter has 4 bits, with a counting range from 0 to 15. It is cleared after reset or completion of an update, and increments by 1 each time an update start signal is received. The counter records the number of update rounds to control state transitions. The state machine has a total of 4 states: idle state, initialization state, information absorption state, and hash generation state, which are used to indicate different stages.

[0061] II. Permutation Network

[0062] The permutation network is used to update the internal state S. The permutation network is denoted as p and consists of round constant addition, a substitution layer, and a linear diffusion layer, which is used to update the 320-bit internal state S. In the permutation network, the internal state S is divided into 5 data blocks of 64 bits each. In the round constant addition stage, the third 64-bit data block is XORed with the round constant. In the substitution layer, NOT, AND, and XOR operations are performed on the data blocks. In the linear diffusion layer, shift and XOR operations are performed on each data block.

[0063] In the permutation network, the 320-bit internal state is evenly divided into 5 data blocks (x0, x1, x2, x3, x4) of 64 bits each. The internal state S = x0||x1||x2||x3||x4;

[0064] The operation of round constant addition is expressed as:

[0065] x2 ← x2 ⊕ c r

[0066] If a total of 12 rounds of updates are required, c r = 8'hf0 - (count value - 1) * 15;

[0067] If a total of 8 rounds of updates are required, c r = 8'hb4 - (count value - 1) * 15;

[0068] The operation of the substitution layer is expressed as:

[0069]

[0070] The operation of the linear diffusion layer is expressed as:

[0071] x0 ← x0 ⊕ (x0 >> 19) ⊕ (x0 >> 28)

[0072] x1 ← x1 ⊕ (x1 >> 61) ⊕ (x1 >> 39)

[0073] x2 ← x2 ⊕ (x2 >> 1) ⊕ (x2 >> 6)

[0074] x3 ← x3 ⊕ (x3 >> 10) ⊕ (x3 >> 17)

[0075] x4 ← x4 ⊕ (x4 >> 7) ⊕ (x4 >> 41)

[0076] Among them, ‖ represents concatenating two binary numbers, ⊕ represents exclusive OR, represents taking the bitwise complement of x, >> represents circular right shift, c r is the round constant.

[0077] III. Initialization Unit

[0078] The initialization unit is used to fix the connection of the initial vector Ⅳ and 256 bits of 0 in the initialization state, initialize the internal state S, and send it to the permutation network for 12 rounds of update to update the internal state S.

[0079] IV. Information Absorption Unit

[0080] The information absorption unit is used to absorb the input information M into the updated internal state S, and update the internal state S through the permutation network during the absorption process. After all the input information M is absorbed, it is sent to the permutation network for update.

[0081] Specifically, the information absorption unit divides the input information into several information blocks of r bits. If the last information block is less than r bits, it is padded with a bit string starting with 1 followed by several 0s to make its length equal to r; the first r bits of the internal state S are taken and XORed with the first information block, then sent to the permutation network for b rounds of update. After the update, the first r bits of the internal state S are taken and XORed with the next information block, and so on, until after XORing with the last information block, the internal state S is sent to the permutation network for 12 rounds of update.

[0082] V. Hash Generation Unit

[0083] The hash generation unit generates and outputs a hash value using the internal state S that has absorbed the input information M, and updates the internal state S through the permutation network during the process of generating the hash value.

[0084] Specifically, the hash generation unit first obtains a 320-bit internal state S from the permutation network, and copies the first r bits of the 320-bit internal state S as the first hash block; then the 320-bit internal state S is sent to the permutation network for b rounds of update. After the update, the first r bits of the 320-bit internal state S are copied as the second hash block, and the generation process of the second hash block is repeated until t hash blocks of length r bits are generated; t is the smallest integer greater than l / r. The t-th hash block is truncated by l mod r bits, and then concatenated with the previous t - 1 hash blocks in order to obtain a hash value H of length l bits and output it.

[0085] The ultra-lightweight Ascon hash algorithm includes the algorithms Ascon-Hash, Ascon-HashA, Ascon-Xof, and Ascon-XofA, where r = 64 in the information absorption unit and the hash generation unit;

[0086] In the chip implementation device of the algorithm Ascon-Hash, the initial vector value Ⅳ in the initialization unit is 00400c0000000100, the required number of update rounds b = 12, and the generated hash value H is 256 bits long;

[0087] In the chip implementation device of the algorithm Ascon-HashA, the initial vector value Ⅳ in the initialization unit is 00400c0400000100, the required number of update rounds b = 8, and the generated hash value H is 256 bits long;

[0088] In the chip implementation device of the algorithm Ascon-Xof, the initial vector value Ⅳ in the initialization unit is 00400c0000000000, the required number of update rounds b = 12, and the generated hash value H is any l bits long;

[0089] In the chip implementation device of the algorithm Ascon-XofA, the initial vector value Ⅳ in the initialization unit is 00400c0400000000, the required number of update rounds b = 8, and the generated hash value H is any l bits long.

[0090] A chip implementation method of an ultra-lightweight Ascon hash algorithm, using the chip implementation device of an ultra-lightweight Ascon hash algorithm as described above, as Figure 2 shown, includes the following steps:

[0091] Step 1, initialization stage:

[0092] Under the control of the control unit, after the device receives the hash value generation start signal, it enters the initialization state. The initialization unit concatenates the initial vector IV with 256 bits of 0 as the initial value of the internal state S, and then sends the internal state S to the permutation network for 12 rounds of update. Each round of update passes through round constant addition, substitution layer, and linear diffusion layer in sequence. The counter records the number of update rounds. When the count value reaches 11, the initialization state ends, the counter is cleared, and the internal state S is sent to the information absorption unit;

[0093] Step 2, information absorption stage:

[0094] In the information absorption unit, the input information M is divided into several information blocks of r bits. If the last information block is less than r bits, it is padded with a bit string starting with 1 followed by several 0s to make its length equal to r. After taking the first r bits of the internal state S and performing an exclusive OR operation with the first information block, the internal state S is sent into the permutation network for b rounds of update. Each round of update successively passes through round constant addition, substitution layer, and linear diffusion layer. The counter records the number of update rounds. When the count value is b - 1, it indicates that the processing of this information block is completed, and the counter is cleared. Then, take the first r bits of the internal state S and perform an exclusive OR operation with the next information block. After completion, send the internal state S into the permutation network for b rounds of update. Repeat this process until after performing the exclusive OR operation with the last information block, send the internal state S into the permutation network for 12 rounds of update. After the update is completed, the information absorption stage is completed.

[0095] Step 3, hash generation stage:

[0096] The hash generation unit first obtains a 320-bit internal state from the permutation network, copies the first r bits of the 320-bit internal state as the first hash block, and then sends the 320-bit internal state into the permutation network for b rounds of update. After the update is completed, copy the first r bits of the 320-bit internal state as the second hash block. Repeat the process of generating the second hash block until t hash blocks of length r bits are generated. t is the smallest integer greater than l / r. The t-th hash block intercepts l mod r bits. After concatenating all the hash blocks, a hash value H of length l bits is obtained. After outputting the hash value H, the hash generation stage is completed.

[0097] In a further technical solution, since the four algorithms correspond to four initial vectors IV, the initialization stage can be simplified. The update result of the initialization stage is solidified in the chip. The four algorithms correspond to four internal states S after initialization. Under the pre-computation optimization method, the initialization stage is omitted, and the internal state after 12 rounds of update in the permutation network in the initialization stage is directly obtained according to the algorithm, and directly enter the information absorption stage.

[0098] The internal state S directly entering the information absorption stage = x0||x1||x2||x3||x4;

[0099] Algorithm Ascon-Hash:

[0100] x0 = ee9398aadb67f03d

[0101] x1 = 8bb21831c60f1002

[0102] x2 = b48a92db98d5da62

[0103] x3 = 43189921b8f8e3e8

[0104] x4 = 348fa5c9d525e140

[0105] Algorithm Ascon - HashA:

[0106] x0 = 01470194fc6528a6

[0107] x1 = 738ec38ac0adffa7

[0108] x2 = 2ec8e3296c76384c

[0109] x3 = d6f6a54d7f52377d

[0110] x4 = a13c42a223be8d87

[0111] Algorithm Ascon - HashXof:

[0112] x0 = b57e273b814cd416

[0113] x1 = 2b51042562ae2420

[0114] x2 = 66a3a7768ddf2218

[0115] x3 = 5aad0a7a8153650c

[0116] x4 = 4f3e0e32539493b6

[0117] Algorithm Ascon - HashXofA:

[0118] x0 = 44906568b77b9832

[0119] x1 = cd8d6cae53455532

[0120] x2 = f7b5212756422129

[0121] x3 = 246885e1de0d225b

[0122] x4 = a8cb5ce33449973f

[0123] In a further technical solution, for the permutation network processing, under the throughput optimization algorithm, the permutation network is replicated w times, and w updates are performed per clock cycle, denoted as p w , thus reducing the number of cycles in the processing process.

[0124] Under the throughput optimization algorithm, the internal state update is denoted as:

[0125] S ← (p w ) 12 / w (IV || 0 256 )

[0126]

[0127] Where, ‖ represents concatenating two binary numbers; ⊕ represents exclusive or; p w represents performing w rounds of updates; M represents the input message, M i represents the i-th message block of the input message; S x~y represents the x-th to y-th bits of the internal state S; s represents the number of message blocks after filling the input message; t represents the number of output hash blocks.

[0128] Through experiments, the maximum throughput of the Ascon-Hash chip implementation device can reach 182 Mbps on the Zynq-7020 model FPGA, occupying 1182 LUTs and 987 FFs; on the Zynq-UltraScale model FPGA, the maximum throughput can reach 630 Mbps, occupying 1175 LUTs and 985 FFs. The maximum throughput of the Ascon-HashA chip implementation device can reach 186 Mbps on the Zynq-7020 model FPGA, occupying 1063 LUTs and 983 FFs; on the Zynq-UltraScale model FPGA, the maximum throughput can reach 663 Mbps, occupying 1061 LUTs and 983 FFs. The maximum throughput of the Ascon-HashXof chip implementation device can reach 267 Mbps on the Zynq-7020 model FPGA, occupying 1223 LUTs and 1247 FFs; on the Zynq-UltraScale model FPGA, the maximum throughput can reach 706 Mbps, occupying 1220 LUTs and 1247 FFs. The maximum throughput of the Ascon-HashXofA chip implementation device can reach 288 Mbps on the Zynq-7020 model FPGA, occupying 1077 LUTs and 1240 FFs, and on the Zynq-UltraScale model FPGA, the maximum throughput can reach 746 Mbps, occupying 1074 LUTs and 1240 FFs.

[0129] The throughput optimization method can optimize one round of update per clock cycle into multiple rounds of update per clock cycle. This design method improves the amount of data processed by the permutation network per clock cycle. Under this method, the maximum throughput of the Ascon-Hash chip implementation device can reach 312 Mbps on the Zynq-7020 FPGA, occupying 3783 LUTs and 1623 FFs; on the Zynq-UltraScale FPGA, the maximum throughput can reach 638 Mbps, occupying 3788 LUTs and 1631 FFs. The maximum throughput of the Ascon-HashA chip implementation device can reach 216 Mbps on the Zynq-7020 FPGA, occupying 3628 LUTs and 1625 FFs; on the Zynq-UltraScale FPGA, the maximum throughput can reach 697 Mbps, occupying 3635 LUTs and 1633 FFs. The maximum throughput of the Ascon-HashXof chip implementation device can reach 376 Mbps on the Zynq-7020 FPGA, occupying 3190 LUTs and 1247 FFs; on the Zynq-UltraScale FPGA, the maximum throughput can reach 881 Mbps, occupying 3125 LUTs and 1247 FFs. The maximum throughput of the Ascon-HashXofA chip implementation device can reach 405 Mbps on the Zynq-7020 FPGA, occupying 3236 LUTs and 1241 FFs, and on the Zynq-UltraScale FPGA, the maximum throughput can reach 750 Mbps, occupying 3229 LUTs and 1240 FFs.

[0130] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A chip implementation device for an ultra-lightweight Ascon hash algorithm, characterized in that: It includes a control unit, a permutation network, an initialization unit, an information absorption unit and a hash generation unit; The control unit controls the hash algorithm to go through an initialization phase, an information absorption phase, and a hash generation phase in sequence; The permutation network is used to update the internal state S; The initialization unit is used to connect and fix the initial vector IV and 256 bits of 0 in the initialization state to achieve the initialization of the internal state S, and send it to the permutation network for 12 rounds of update to update the internal state S; The information absorption unit is used to absorb the input information M into the updated internal state S, and in the absorption process, the internal state S is updated through the replacement network, and all the input information M is absorbed and sent to the replacement network for updating; The hash generation unit generates a hash value by using the internal state S that absorbs the input information M and outputs it. In the process of generating the hash value, the internal state S is updated through the permutation network.

2. According to claim 1, a chip implementation device of an ultra-lightweight Ascon hash algorithm is characterized in that: The control unit includes a counter and a state machine, wherein the counter has 4 bits, the counting range is from 0 to 15, is cleared after reset or update is completed, and is incremented by 1 each time an update start signal is received. The counter records the number of update rounds to control state switching; the state machine has 4 states: idle state, initialization state, information absorption state, and hash generation state, which are used to indicate different stages.

3. According to claim 1, a chip implementation device of an ultra-lightweight Ascon hash algorithm is characterized in that: The permutation network is denoted as p, and is composed of a round constant addition, a replacement layer, and a linear diffusion layer, and is used to update the 320-bit internal state S; in the permutation network, the internal state S is divided into five 64-bit data blocks; In the round constant addition phase, the third 64-bit data block is XORed with the round constant; in the replacement layer, the data block is NOT, AND, and XORed; In the linear diffusion layer, each data block is shifted and XORed.

4. According to claim 3, a chip implementation device of an ultra-lightweight Ascon hash algorithm is characterized in that: In the permutation network, the 320-bit internal state is divided into five 64-bit data blocks (x0, x1, x2, x3, x4), and the internal state S = x0||x1||x2||x3||x4; The operation of adding a round constant is expressed as: If a total of 12 rounds of updates are required, c r =8'hf0-(count value-1)*15; If a total of 8 rounds of updates are required, c r =8'hb4-(count value-1)*15; The operation of replacing the layer is expressed as: The operation of the linear diffusion layer is expressed as: Among them, ‖ represents the connection of two binary numbers, represents XOR, Indicates bitwise inversion of x, >> indicates circular right shift, c r is the round constant.

5. According to claim 1, a chip implementation device of an ultra-lightweight Ascon hash algorithm is characterized in that: The information absorption unit divides the input information into several information blocks of r bits. If the last information block is less than r bits, it is filled with a bit string starting with 1 and followed by several 0s to make its length equal to r; the first r bits of the internal state S are taken and XORed with the first information block, and then sent to the permutation network for b rounds of update. After the update, the first r bits of the internal state S are taken and XORed with the next information block, and this is repeated until the XOR is completed with the last information block, and the internal state S is sent to the permutation network for 12 rounds of update.

6. The chip implementation device of the ultra-lightweight Ascon hash algorithm according to claim 1 is characterized in that: The hash generation unit first obtains a 320-bit internal state S from a permutation network, and copies the first r bits of the 320-bit internal state S as the first hash block; then the 320-bit internal state S is sent to the permutation network for b rounds of update, and after the update is completed, the first r bits of the 320-bit internal state S are copied as the second hash block, and the generation process of the second hash block is repeated until t hash blocks with a length of r bits are generated; t is the smallest integer greater than l / r, and the tth hash block is truncated by l mod r bits, and is sequentially concatenated with the previous t-1 hash blocks to obtain a hash value H with a length of l bits, which is output.

7. A chip implementation method of an ultra-lightweight Ascon hash algorithm, using an ultra-lightweight Ascon hash algorithm chip implementation device as claimed in any one of claims 1 to 6, characterized in that: The steps include: Step 1, initialization phase: Under the control of the control unit, the device enters the initialization state after receiving the hash value generation start signal. The initialization unit connects the initial vector IV with 256 bits of 0 as the initial value of the internal state S, and then sends the internal state S to the permutation network for 12 rounds of update. Each round of update passes through the round constant addition, the replacement layer and the linear diffusion layer in sequence. The counter records the number of update rounds until the count value is 11, the initialization state ends, the counter is reset, and the internal state S is sent to the information absorption unit; Step 2, information absorption stage: In the information absorption unit, the input information M is divided into several information blocks of r bits. If the last information is less than r bits, it is filled with a bit string starting with 1 and followed by several 0s to make its length equal to r; the first r bits of the internal state S are taken to be XORed with the first information block, and after completion, the internal state S is sent to the permutation network for b rounds of update. Each round of update passes through the round constant addition, replacement layer and linear diffusion layer in turn. The counter records the number of update rounds. When the count value is b-1, it means that the information block has been processed and the counter is reset; then the first r bits of the internal state S are taken to be XORed with the next information block. After completion, the internal state S is sent to the permutation network for b rounds of update, and this is repeated until the internal state S is XORed with the last information block, and then sent to the permutation network for 12 rounds of update. After the update is completed, the information absorption stage is completed; Step 3, hash generation phase: The hash generation unit first obtains a 320-bit internal state from the permutation network, copies the first r bits of the 320-bit internal state as the first hash block, and then sends the 320-bit internal state to the permutation network for b rounds of update. After the update is completed, the first r bits of the 320-bit internal state are copied as the second hash block, and the process of generating the second hash block is repeated until t hash blocks of length r bits are generated; t is the smallest integer greater than l / r, the tth hash block intercepts l mod r bits, and all hash blocks are concatenated to obtain a hash value H of length l bits. After outputting the hash value H, the hash generation stage is completed.

8. The chip implementation method of the ultra-lightweight Ascon hash algorithm according to claim 7 is characterized in that: The ultra-lightweight Ascon hash algorithm includes algorithm Ascon-Hash, algorithm Ascon-HashA, algorithm Ascon-Xof and algorithm Ascon-XofA, wherein r in the information absorption unit and the hash generation unit is 64; In the chip implementation device of the algorithm Ascon-Hash, the initial vector value IV in the initialization unit is 00400c0000000100, the required update round number b is 12, and the generated hash value H is 256 bits long; In the chip implementation device of the algorithm Ascon-HashA, the initial vector value IV in the initialization unit is 00400c0400000100, the required update round number b is 8, and the generated hash value H is 256 bits long; In the chip implementation device of the algorithm Ascon-Xof, the initial vector value IV in the initialization unit is 00400c0000000000, the required update round number b is 12, and the generated hash value H is arbitrarily l bits long; In the chip implementation device of the algorithm Ascon-XofA, the initial vector value IV in the initialization unit is 00400c0400000000, the required number of update rounds b is 8, and the generated hash value H is arbitrarily l bits long.

9. The chip implementation method of the ultra-lightweight Ascon hash algorithm according to claim 8 is characterized in that: The update results of the initialization phase are solidified into the chip. The four algorithms correspond to the four internal states S after the initialization. Under the pre-calculation optimization method, the initialization phase is omitted, and the internal state of the permutation network after 12 rounds of updates in the initialization phase is directly obtained according to the algorithm, and the information absorption phase is directly entered.

10. The chip implementation method of the ultra-lightweight Ascon hash algorithm according to claim 7, characterized in that: For permutation network processing, under the throughput optimization algorithm, the permutation network is replicated w times and updated w times per clock cycle; Under the throughput optimization algorithm, the internal state update is expressed as: S←(p w ) 12 / w (IV||0 256 ) Among them, ‖ means connecting two binary numbers; Indicates exclusive OR; p w indicates w rounds of updates; M indicates input information, M i represents the i-th information block of the input information; S x~y Represents the xth to yth bits of the internal state S; s represents the number of information blocks after the input information is filled; t represents the number of output hash blocks.