A data encryption implementation method and device for SM3 algorithm

By parallelly optimizing the key paths of CSA and CPA adders in the SM3 algorithm hardware structure, and using multiple sets of extended compression modules to run in parallel, the problems of high circuit resource occupancy and large operation power consumption in the existing SM3 algorithm hardware structure are solved, and lower power consumption and higher applicability are achieved.

CN116155481BActive Publication Date: 2025-05-13CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310164866.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-24
Publication Date
2025-05-13
Estimated Expiration
2043-02-24

AI Technical Summary

Technical Problem

When performing 64 rounds of iterative compression of function of the existing SM3 algorithm hardware structure, the circuit resource occupies a large amount of power consumption, and the operational power consumption is high, and the applicability is low.

Method used

By setting up a CSA carry-reserved adder and a CPA carry-transfer adder in parallel in the critical path, the message word is iteratively compressed and calculated, and runs in parallel through multiple sets of extended compression modules to reduce addition delay and circuit resource usage.

Benefits of technology

It reduces the operating power consumption and circuit resource occupancy of the SM3 algorithm, improves the single encryption speed, and has higher applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116155481B_ABST
    Figure CN116155481B_ABST
Patent Text Reader

Abstract

The present invention discloses a data encryption implementation method and device of an SM3 algorithm, the method comprising: filling a plaintext message through a state machine; grouping the filled messages, respectively expanding them, and outputting message words; optimizing the iterative compression calculation of the message words in parallel through a CSA carry-preserving adder and a CPA carry-propagating adder to obtain encrypted data. The device comprises a message filling module, a plurality of groups of expansion and compression modules, and a control module, wherein the control module is respectively connected to the message filling module and the expansion and compression module, and the expansion and compression module comprises an expansion module and a compression module. The present invention solves the problems of large circuit resource occupancy and high operating power consumption of the existing SM3 algorithm implementation structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of encryption technology, and in particular to a method and device for implementing data encryption of an SM3 algorithm. Background Art

[0002] The SM3 algorithm is a cryptographic hash algorithm issued by the State Cryptography Administration. It is used for digital signatures and verification of commercial cryptography, generation and verification of message authentication codes, and generation of random numbers. Its security and efficiency are equivalent to SHA-256. However, at this stage, the hardware structure design for the SM3 algorithm generally adopts a multi-stage pipeline structure when performing 64 rounds of function iterative compression, that is, the iterative compression in the SM3 algorithm is cyclically expanded, and the output of each round of calculation is used as the input of the next round. The calculation continues until there is no remaining content of the current hash value that needs to be continued to be calculated, and the final hash value of the calculation is output. However, when implementing in hardware, it often takes up a large amount of circuit resources and has high operating power consumption, so it is not suitable for low-power application scenarios and has low applicability. Summary of the invention

[0003] 1. Technical issues to be resolved

[0004] Based on the above problems, the present invention provides a data encryption implementation method and device of SM3 algorithm, which solves the problems of large circuit resource occupancy and high operating power consumption of the existing SM3 algorithm implementation structure.

[0005] (II) Technical solution

[0006] Based on the above technical problems, the present invention provides a data encryption implementation method of the SM3 algorithm, comprising:

[0007] S1, fill the plaintext message through the state machine;

[0008] S2, group the filled messages, expand them separately, and output the message words;

[0009] S3. Optimize the iterative compression calculation of the message word in parallel through the CSA carry-save adder and the CPA carry-propagation adder to obtain encrypted data.

[0010] Furthermore, in step S3, the CSA carry-save adder and CPA carry-propagate adder set in the critical path are used to perform iterative compression calculation on the message word, and the critical path is the calculation path with the largest amount of calculation in the iterative compression calculation.

[0011] Furthermore, the critical path is a calculation path of word registers E and A where intermediate variables SS1, SS2, TT1, and TT2 are located in the iterative compression function.

[0012] Further, in step S3, the iterative compression calculation of the message word is performed by parallel optimization of the three CSAs and the two CPAs.

[0013] Furthermore, the first CSA calculates SS1 and the second CSA calculates GG j (E,F,G)+H+W j , the third CSA calculates FF j (A,B,C)+D+W j ′ , three CSAs are calculated in parallel, and then the first CPA calculates TT2 according to the calculation results of the first CSA and the second CSA, and the second CPA calculates TT1 according to SS2 obtained by the XOR operation of the calculation result of the first CSA and the calculation result of the third CSA, and the two CPAs are calculated in parallel, and then the iterative compression calculation is completed according to the calculation results of the two CPAs, wherein A, B, C, D, E, F, G, H are 8 word registers, GG j and FF j represents a Boolean function, which takes different expressions as j changes. W j and W j ′ In step S1, the state machine includes 8 states: idle, normal message, last word, add 1, add 0, fill high, fill low and complete state, and the state jump method of the state machine includes:

[0014] Enters idle state after reset or power-on;

[0015] In idle state, if the message valid signal and the last word signal are received, the last word state is entered; if only the message valid signal is received, the message normal state is entered;

[0016] When the message is in normal state, if the message valid signal and the last word signal are received, it will enter the last word state;

[0017] In the last word state, if the counter cnt=13, that is, the remaining padding area is exactly 64 bits, it jumps to the filling high state, otherwise it automatically jumps to the plus 1 state;

[0018] In the plus 1 state, after adding the bit "1" to the end of the message, it automatically jumps to the plus 0 state;

[0019] In the add 0 state, k "0"s are added to the end of the message, and then the state of filling high bits is automatically jumped to. k is the smallest non-negative integer that satisfies l+1+k≡448mod512, l(l<2 64 ) is the bit length of the plaintext message;

[0020] When filling the high-bit state, fill the high 32 bits of the binary representation representing the message length, and automatically jump to the filling low-bit state after filling is completed;

[0021] When filling the low-bit state, fill the lower 32 bits of the binary representation representing the message length, and automatically jump to the completed state after filling is completed;

[0022] When in the completion state, the filling completion signal is set to 1, and then automatically jumps to the idle state.

[0023] Furthermore, steps S2 and S3 are executed in parallel by multiple expansion and compression modules.

[0024] The present invention also discloses a data encryption implementation device of the SM3 algorithm, comprising a message filling module, a plurality of groups of expansion and compression modules and a control module, wherein the control module is respectively connected to the message filling module and the expansion and compression module, and the expansion and compression module comprises an expansion module and a compression module;

[0025] The message filling module runs step S1 of the data encryption implementation method of the SM3 algorithm;

[0026] The extension module runs step S2 of the data encryption implementation method of the SM3 algorithm;

[0027] The compression module runs step S3 in the data encryption implementation method of the SM3 algorithm.

[0028] Furthermore, it also includes a data cache module, and the data cache module is respectively connected to the multiple groups of expansion and compression modules.

[0029] Furthermore, there are three groups of expansion and compression modules.

[0030] (III) Beneficial effects

[0031] The above technical solution of the present invention has the following advantages:

[0032] (1) The method of the present invention optimizes the critical path of iterative operations by parallelizing CSA and CPA adders, and reduces the addition delay when multiple inputs are used through parallel operations, thereby reducing the operating power consumption of SM3 and reducing the circuit resource occupancy rate; and further reduces the addition delay and operating power consumption by taking advantage of the fact that the calculation delay of CSA and CPA adders themselves is relatively small when high inputs are used;

[0033] (2) The present invention sets multiple expansion and compression modules in parallel, so the running power consumption and circuit resources occupied are significantly lower than those of the pipeline structure design, and it is also beneficial to improve the single encryption speed;

[0034] (3) The device of the present invention integrates a message filling module, which is implemented by an improved state machine, making software control simpler. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The features and advantages of the present invention will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present invention in any way. In the accompanying drawings:

[0036] Figure 1 Flow chart of a method for implementing SM3 algorithm data encryption according to an embodiment of the present invention;

[0037] Figure 2 A schematic diagram of the jump operation principle of the state machine of an embodiment of the present invention;

[0038] Figure 3 A schematic diagram of message filling in an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of message expansion according to an embodiment of the present invention;

[0040] Figure 5 A schematic diagram of compression logic according to an embodiment of the present invention;

[0041] Figure 6 A schematic diagram of parallel operation of CSA and CPA adders according to an embodiment of the present invention;

[0042] Figure 7 It is a structural diagram of the SM3 algorithm data encryption implementation device of an embodiment of the present invention. DETAILED DESCRIPTION

[0043] The specific implementation of the present invention is further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0044] Embodiment 1 of the present invention is a data encryption implementation method of the SM3 algorithm, such as Figure 1 As shown, including:

[0045] S1. Fill the plaintext message through the state machine to obtain a message with a length that is a multiple of 512 bits;

[0046] The state machine includes 8 states: idle, normal message, last word, add 1, add 0, fill high, fill low and completion state. Figure 2 As shown, the 8 states and the jump conditions of each state are as follows:

[0047] 1. Idle state: enter the idle state after reset / power-on. In the idle state, if the message valid msg_valid and the last word last_word signal are received, it means that the message input is complete and enters the last_word state; if only the message valid msg_valid signal is received, it means that the message input begins and enters the normal_msg state;

[0048] 2. When the message is normal, that is, in the normal_msg state: receiving the message valid msg_valid and the last word last_word signal, it means that the message has been fully input and enters the last_word state;

[0049] 3. The last word, i.e., last_word state: This state is the last state of receiving a plaintext message. If the counter cnt=13, i.e., the remaining padding area is exactly 64 bits, then it jumps to the add_lenh state, otherwise the state machine automatically jumps to the add_80 state;

[0050] 4. Add 1, i.e. add_80 state: This state will add the bit "1" to the end of the message. After that, the state machine automatically jumps to the add_00 state;

[0051] 5. Add 0, i.e. add_00 state: This state will add k "0" to the end of the message. 64 ) bits of message, k is the smallest non-negative integer that satisfies l+1+k≡448mod512. Then the state machine jumps to the add_lenh state;

[0052] 6. When filling the high bits, i.e., add_lenh state: this state will fill the high 32 bits of the binary representation representing the message length. After filling, it will automatically jump to the add_lenl state;

[0053] 7. When filling the lower bits, i.e., add_lenl state: this state will fill the lower 32 bits of the binary representation representing the message length. After the filling is completed, it will automatically jump to the finish state;

[0054] 8. When the state is completed, it means that the message filling is completed, the filling completion signal is set to 1, and then it automatically jumps to the idle state;

[0055] That is to say, the state machine is in the idle state. When receiving the reset signal or power-on, the state machine enters the idle state; when receiving the msg_vaild signal, it means that the message starts to be input, and the state machine jumps to the normal_msg state; when receiving the msg_vaild and last_word signals, it means that the message is fully input, and the state machine jumps to the last_word state; then the bit "1" is added to the end of the message, and then k "0"s are added. Then fill the high bit of the bit string representing the length of the message, and fill the low bit of the bit string representing the length of the message after the next clock cycle; when the filling is completed, the state machine enters the Finish state, sends a filling completion signal to the control module, and then the state machine enters the idle state.

[0056] The state machine adds a bit "l" to the end of the message m, then adds k "0"s to satisfy the minimum non-negative integer l+1+k≡448mod512, and then adds a 64-bit bit string, which is the binary representation of the length l, such as Figure 3 As shown in FIG. 1 , P1 represents the permutation function in the message expansion. The bit length of the padded message m′ is a multiple of 512.

[0057] S2, group the padded messages into 512-bit lengths, expand them separately through 16-word shift registers, and output message words W j and W j ′;

[0058] The padded message m′ is grouped and numbered according to the length of 512 bits, and the obtained message B (0) B (1) ···B (n-1) , n represents the total number of groups, n = (l+1+k) / 512.

[0059] Through the 16-word shift register W0...W 15 It is implemented as a buffer for generating 132 message words. (i) Perform message expansion to generate 132 words W0, W1, .., W 67 , W′0, W′1,···,W′ 63 The first 16 words of the shift register are transmitted through message group B. (i) Directly divide it, no calculation is needed, and then shift one word to the left each time the calculation is extended. Figure 4 The value of the register is extracted at the position shown in the figure for logical operation, and the result is used as W 15 The updated value of . Extract W0 as the current round W j The output of W0 and W4 is extracted and XORed as the current round W j ′ output. That is:

[0060] FOR j=16 TO 67

[0061]

[0062]

[0063] ENDFOR

[0064] FOR j=0 TO 63

[0065]

[0066] ENDFOR

[0067] S3. Optimize the iterative compression calculation of the message word in parallel through the CSA carry-save adder and the CPA carry-propagation adder to obtain encrypted data.

[0068] The running efficiency of the compression function is mainly determined by the delay of the key operation path, and the delay of the key path is mainly caused by the delay of the addition operation. Therefore, the CSA carry-save adder and the CPA carry-propagation adder set on the key path are used to parallel optimize the iterative compression calculation of the message word. The key path is the calculation path with the largest amount of calculation in the iterative compression calculation, which can be divided into SS1, SS2, TT1, TT2 variable generation circuit and GG j , FF j Boolean operation function circuit.

[0069] The iteratively compressed logic design diagram is as follows Figure 5 As shown in the figure, A, B, C, D, E, F, G, H are 8 word registers or the series connection of their values, GG j Represents a Boolean function, which takes different expressions as j changes. FF j represents a Boolean function, which takes different expressions as j changes. SS1, SS2, TT1, and TT2 are intermediate variables. P0 represents the permutation function in the compression function. <<< represents a bit operation of a 32-bit circular left shift. ⊕ represents a 32-bit XOR operation. It has the following iterative compression function:

[0070] SS1←((A<<<12)+E+(T j <<<j))<<<7

[0071]

[0072] TT1←FF j (A,B,C)+D+SS2+W j '

[0073] TT2←GG j (E,F,G)+H+SS1+W j

[0074] D←C

[0075] C←B<<<9

[0076] B←A

[0077] A←TT1

[0078] H←G

[0079] G←F<<<19

[0080] F←E

[0081] E←P0(TT2)

[0082] In this embodiment, among the critical paths, the longest critical path is the part of calculating the word register E:

[0083] SS1←((A<<<12)+E+(T j <<<j))<<<7

[0084] TT2←GG j (E,F,G)+H+SS1+W j

[0085] E←P0(TT2)

[0086] For the longest critical path, two CSA structures are used to parallelize the calculation of SS1 and GG. j (E,F,G)+H+W j , and then calculate TT2 through the CPA structure according to the calculation results of the two CSA structures, and then calculate the word register E according to P0(TT2). In the critical path, except for the longest critical path, the operation of directly generating SS2 is an XOR operation, and its delay can be ignored, so no parallel calculation is performed, but the calculation of TT1 and TT2 is parallel calculation, so three CSA structures and two CPA structures are used for parallel calculation.

[0087] The process of parallel operation using three CSA structures and two CPA structures in this embodiment is as follows: Figure 6 As shown, including: the first CSA structure calculation SS1, the second CSA structure calculation GG j (E,F,G)+H+W j , the third CSA structure calculates FF j (A,B,C)+D+W j′, the three CSA structures are calculated in parallel, and then the first CPA structure calculates TT2 according to the calculation results of the first CSA structure and the second CSA structure, and the second CPA structure calculates TT1 according to SS2 obtained by the XOR operation of the calculation result of the first CSA structure and the calculation result of the third CSA structure. The two CPA structures are calculated in parallel, and then the operation of the iterative compression function is completed according to the calculation results of the two CPA structures.

[0088] By introducing CSA (Carry Save Adder) and CPA (Carry Propagateadder) structures, the addition delay in case of multiple inputs is reduced, and more sequential addition operations are performed in parallel, thereby reducing the delay of a single iteration calculation. Through optimization, the original 5 addition delays are shortened to 2 addition delays. The calculation delay of the compression module is further reduced by taking advantage of the fact that the calculation delay of CSA and CPA adders is less than that of conventional traveling wave adders in the case of high input.

[0089] After all message packets are processed, the output of the last 512-bit packet is the algorithm hash value.

[0090] The SM3 algorithm uses a message word processing method that combines message double words to achieve rapid diffusion and confusion of messages in a local area. 64 ) bits of plaintext message, the SM3 hash algorithm performs message padding and iterative compression to finally generate a hash value of 256 bits in length.

[0091] Embodiment 2 of the present invention is a data encryption implementation device of the SM3 algorithm, such as Figure 7 As shown, it includes a message filling module, multiple groups of extension and compression modules, a control module and a data cache module. The control module is respectively connected to the message filling module and the extension and compression module, and the data cache module is respectively connected to the extension and compression modules. The multiple groups of extension and compression modules are the same, including extension modules and compression modules. Three groups of extension and compression modules are used in this embodiment.

[0092] The message filling module runs step S1 in Example 1 to fill the input message according to the SM3 algorithm to obtain a message with a bit length that is a multiple of 512; the expansion module runs step S2 in Example 1 to perform word expansion on the 512-bit message after the filling grouping; the compression module runs step S3 in Example 1 to iteratively compress and obtain encrypted data.

[0093] The data encryption implementation device of the SM3 algorithm is used to make the message m be grouped into 512 bits after being filled by the message filling module, and the grouped message m is ′Input to multiple groups of expansion and compression modules for parallel calculation, and the expansion modules in each group output W under the control of the control module j and W j ′ To the corresponding compression module, after 64 rounds of iterative compression, the data is finally stored in the data cache module. Since the SM3 algorithm needs to perform 64 rounds of calculations during the iterative compression process, this embodiment adopts a three-way expansion compression module to perform parallel calculations at the same time, which improves the throughput of encryption operations and reduces latency without increasing a large amount of resources. The circuit resources occupied are 550 slices, and the power consumption is 0.027W, which is obviously less circuit resources occupied and consumes less power than the pipeline structure design.

[0094] Moreover, since the expansion and compression modules are three-way parallel, each clk clock can calculate 3 W j and W j ′ Value, respectively W j and W j ′ Under the control of the control module, it is sent to the compression module for compression. Compared with the design scheme using multi-stage pipeline structure and unipolar structure, each clk clock can only calculate 1W j and W j ′ The numerical value shows that the computational efficiency of message expansion in this embodiment is improved by 3 times, so that the time required for a single encryption calculation is shorter.

[0095] In summary, the above-mentioned SM3 algorithm data encryption implementation method and device have the following beneficial effects:

[0096] (1) The method of the present invention optimizes the critical path of iterative operations by parallelizing CSA and CPA adders, and reduces the addition delay when multiple inputs are used through parallel operations, thereby reducing the operating power consumption of SM3 and reducing the circuit resource occupancy rate; and further reduces the addition delay and operating power consumption by taking advantage of the fact that the calculation delay of CSA and CPA adders themselves is relatively small when high inputs are used;

[0097] (2) The present invention sets multiple expansion and compression modules in parallel, so the running power consumption and circuit resources occupied are significantly lower than those of the pipeline structure design, and it is also beneficial to improve the single encryption speed;

[0098] (3) The device of the present invention integrates a message filling module, which is implemented by an improved state machine, making software control simpler.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations shall fall within the scope defined by the appended claims.

Claims

1. A data encryption implementation method of the SM3 algorithm, characterized in that: include: S1, fill the plaintext message through the state machine; S2, group the filled messages, expand them separately, and output the message words; S3, performing iterative compression calculation on the message word through parallel optimization of CSA carry-save adder and CPA carry-propagation adder to obtain encrypted data: using the CSA carry-save adder and CPA carry-propagation adder set in the critical path to perform iterative compression calculation on the message word, the critical path is the calculation path with the largest amount of calculation in the iterative compression calculation, the critical path is the calculation path of word registers E and A where intermediate variables SS1, SS2, TT1, and TT2 are located in the iterative compression function, and performing iterative compression calculation on the message word through parallel optimization of three CSAs and two CPAs, including: the first CSA calculates SS1, the second CSA calculates , the third CSA calculation , three CSAs are calculated in parallel, and then the first CPA calculates TT2 according to the calculation results of the first CSA and the second CSA, and the second CPA calculates TT1 according to SS2 obtained by the XOR operation of the calculation result of the first CSA and the calculation result of the third CSA, and the two CPAs are calculated in parallel, and then the iterative compression calculation is completed according to the calculation results of the two CPAs, wherein A, B, C, D, E, F, G, and H are 8 word registers, and represents a Boolean function, which takes different expressions as j changes. and Indicates the message word.

2. The data encryption implementation method of the SM3 algorithm according to claim 1 is characterized in that: In step S1, the state machine includes 8 states: idle, normal message, last word, add 1, add 0, fill high bit, fill low bit and completion state, and the state jump method of the state machine includes: Enters idle state after reset or power-on; In idle state, if the message valid signal and the last word signal are received, the last word state is entered; if only the message valid signal is received, the message normal state is entered; When the message is in normal state, if the message valid signal and the last word signal are received, it will enter the last word state; In the last word state, if the counter cnt=13, that is, the remaining padding area is exactly 64 bits, it jumps to the filling high state, otherwise it automatically jumps to the plus 1 state; In the plus 1 state, after adding the bit "1" to the end of the message, it automatically jumps to the plus 0 state; When the state is 0, After a "0" is added to the end of the message, it automatically jumps to the filling high state; To satisfy The smallest non-negative integer, is the bit length of the plaintext message, ; When filling the high-bit state, fill the high 32 bits of the binary representation representing the message length, and automatically jump to the filling low-bit state after filling is completed; When filling the low-bit state, fill the lower 32 bits of the binary representation representing the message length, and automatically jump to the completed state after filling is completed; When in the completion state, the filling completion signal is set to 1, and then automatically jumps to the idle state.

3. The data encryption implementation method of the SM3 algorithm according to claim 1 is characterized in that: The steps S2 and S3 are executed in parallel by multiple expansion and compression modules.

4. A data encryption implementation device of the SM3 algorithm according to any one of claims 1 to 3, characterized in that: It includes a message filling module, multiple groups of expansion and compression modules and a control module, wherein the control module is connected to the message filling module and the expansion and compression module respectively, and the expansion and compression module includes an expansion module and a compression module; The message filling module runs step S1 of the data encryption implementation method of the SM3 algorithm; The extension module runs step S2 of the data encryption implementation method of the SM3 algorithm; The compression module runs step S3 in the data encryption implementation method of the SM3 algorithm.

5. The data encryption implementation device of the SM3 algorithm according to claim 4 is characterized in that: It also includes a data cache module, and the data cache module is respectively connected to the multiple groups of expansion and compression modules.

6. The data encryption implementation device of the SM3 algorithm according to claim 5, characterized in that: There are three groups of expansion and compression modules.

Citation Information

Patent Citations

  • Arithmetic processing unit that performs multiply and multiply-add operations with saturation and method therefor

    CN102804128A

  • Implementation method and device of hash algorithm

    CN112084534A