Fast Implementation Method of SM2 Encryption and Decryption Based on SIMD

By adopting SIMD parallel processing technology in the SM2 algorithm, the problem of low encryption efficiency when processing long data is solved, and significant encryption speed and computing efficiency are achieved.

CN115174038BActive Publication Date: 2025-06-24SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210846869.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-19
Publication Date
2025-06-24
Estimated Expiration
2042-07-19

AI Technical Summary

Technical Problem

When the existing SM2 public key encryption algorithm processes longer data, the KDF part takes a long time, resulting in low encryption efficiency and affecting the encryption rate of the SM2 algorithm.

Method used

The parallel processing method based on SIMD is adopted to improve local computing efficiency and improve the overall computing efficiency of the SM2 algorithm, thereby improving data encryption efficiency. The specific implementation includes using the VPGATHERDD instruction and the PSHUFB instruction for message expansion and data rearrangement, and instead of traditional pointers to implement byte reverse operation.

Benefits of technology

Through SIMD parallelization processing, the speed of SM2 encryption and decryption is significantly improved, the overall computing efficiency of the algorithm is improved, the amount of code is reduced, and the process of loading data is omitted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115174038B_ABST
    Figure CN115174038B_ABST
Patent Text Reader

Abstract

The present invention provides a fast implementation method for SM2 encryption and decryption based on SIMD. For multiple pieces of data after adding CT values, a preset number of bits are taken out from each piece and integrated into one piece and stored in the message array. The pre-computed index values and the VPGATHERDD instruction are used to implement the first step of message expansion, and the PSHUFB instruction is used to rearrange the expanded data, replacing the pointer to implement the byte reverse order function, improving the overall operation efficiency of the algorithm; when the first step of message expansion is completed, the data has been loaded into the register through the VPGATHERDD instruction, omitting the subsequent separate loading process; the second and third steps of message expansion are implemented using automated loop instructions, reducing the code volume.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data encryption processing, and particularly relates to a fast implementation method for SM2 encryption and decryption based on SIMD. Background Art

[0002] The statements in this part only provide the background art related to the present invention and do not necessarily constitute the prior art.

[0003] The SM2 algorithm is a public-key cryptography algorithm for data encryption. The KDF part in the SM2 algorithm uses the SM3 algorithm for hash operation, and the execution efficiency of KDF is the key to affecting the speed of SM2 encryption and decryption.

[0004] The inventors found that in the KDF part of the existing SM2 public-key encryption algorithm, when processing long data, it takes a long time and the encryption efficiency is low, which greatly affects the encryption rate of the SM2 public-key cryptography algorithm. Summary of the Invention

[0005] In order to solve the deficiencies of the prior art, the present invention provides a fast implementation method for SM2 encryption and decryption based on SIMD. After using SIMD parallel processing, by improving the local operation efficiency, the overall operation efficiency of the SM2 algorithm is improved, and thus the data encryption efficiency is increased.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The first aspect of the present invention provides a method for improving the SM2 data encryption speed based on SIMD.

[0008] A method for improving the SM2 data encryption speed based on SIMD includes the following processes:

[0009] Use a random number generator to generate a random number k, use k to calculate the elliptic curve point C1(x1, y1), and convert the data type of C1 into a bit string. Calculate the elliptic curve point (x2, y2) according to the random number k;

[0010] Convert the data type of the elliptic curve point (x2, y2) into a bit string, merge them into one piece of data, and copy it into multiple identical pieces of data, and add the corresponding 32-bit counter CT value in sequence at the end of each piece of data;

[0011] Define an unsigned integer array index32 and load it into the register index. For each of the multiple pieces of data after adding the CT value, extract a preset number of bits and integrate them into one piece and store them in the message array;

[0012] Use the VPGATHERDD instruction and the index to put the data in the message into 16 variables of type __m256i, and use the PSHUFB instruction to perform byte reverse order operation on the data in the 16 variables of type __m256i;

[0013] Put the 16 variables of type __m256i after the byte reverse order operation into the group of type __m256i array; perform the first step of message expansion on the 16 data in the group, and finally expand to 68; perform the second step of message expansion on the 68 data in the group, and finally expand to 64 and put them into the group2 array of type __m256i;

[0014] Define 8 unsigned integer arrays, each array stores 8 hexadecimal initial values of the SM3 hash function and the contents of the 8 arrays are the same; define 8 word registers of type __m256i, and put the same data in the 8 unsigned integer arrays into the same register;

[0015] Define temporary data SS1, SS2, TT2, all of type __m256i; define boolean functions FF1, FF2, GG1 and GG2; according to group, group2, SS1, SS2, TT2, FF1, FF2, GG1 and GG2, execute the 16-round round function RoundFunction1 and the 48-round round function RoundFunction2;

[0016] After 64 rounds of round function operations, reload the 8 word registers of type __m256i back into the array, and perform exclusive OR with the 8 unsigned integer arrays to obtain the ciphertext of this round of operation;

[0017] Take out the remaining part after the preset number of bits of each data, add the bit string representing the data length at the end, and integrate them into one data in order and put it into the message array, and execute the steps after using the VPGATHERDD instruction to obtain the final ciphertext;

[0018] According to the plaintext input length L, repeat the above process. After executing several times, put the ciphertext output in each round into the unsigned character array t;

[0019] Perform exclusive OR on t and message to obtain C2;

[0020] Merge the coordinate point x2, the message message, and the coordinate point y2 in order (x2||message||y2), and perform an operation on it using the SM3 hash function to obtain C3;

[0021] Merge C1, C3, and C2 in sequence to obtain C = C1||C3||C2 and output it. C is the encryption result.

[0022] The second aspect of the present invention provides a method for improving the SM2 data decryption speed based on SIMD.

[0023] A method for improving the SM2 data decryption speed based on SIMD includes the following processes:

[0024] Assume the encrypted ciphertext is C, and C = C1||C3||C2;

[0025] Extract C1 from the ciphertext and convert the data type of C1 into a point on the elliptic curve;

[0026] Calculate the elliptic curve points x2 and y2, then convert the data types of the coordinates x2 and y2 into bit strings and merge x2 and y2 into one (x2||y2);

[0027] Use the parallelized KDF strategy involved in the encryption process to calculate x2 and y2, and obtain the operation result t;

[0028] Extract the bit string C2 from C, and perform an exclusive OR operation on t and C2 to obtain the plaintext M;

[0029] Merge the coordinates x2, the plaintext M, and the coordinates y2 (x2||M||y2), and perform an operation on it using the SM3 hashing algorithm to obtain the result u;

[0030] Extract the bit string C3 from C, determine whether u is equal to C3. If they are not equal, report an error and exit. Otherwise, output the plaintext M;

[0031] Among them, the parallelized KDF strategy includes the following processes:

[0032] Obtain multiple input data, and add the corresponding 32-bit counter CT value to the end of each data in sequence;

[0033] Define an unsigned integer array index32 and load it into the register index. For each of the multiple data after adding the CT value, extract a preset number of bits and integrate them into one and store them in the message array;

[0034] Use the VPGATHERDD instruction and the index index to put the data in message into 16 __m256i type variables, and use the PSHUFB instruction to perform a byte reverse operation on the data in the 16 __m256i type variables;

[0035] Put the 16 __m256i type variables after byte reverse operation into the group of __m256i array type; perform the first step of message expansion on the 16 data in the group, and finally expand them to 68; perform the second step of message expansion on the 68 data in the group, and finally expand them to 64 and put them into the __m256i type group2 array;

[0036] Define 8 unsigned integer arrays, each array stores 8 hexadecimal initial values of the SM3 hash function and the contents of the 8 arrays are the same; define 8 __m256i type word registers, and put the same data in the 8 unsigned integer arrays into the same register;

[0037] Define temporary data SS1, SS2, TT2, all of which are of __m256i type; define boolean functions FF1, FF2, GG1 and GG2; execute the 16-round round function RoundFunction1 and the 48-round round function RoundFunction2 according to group, group2, SS1, SS2, TT2, FF1, FF2, GG1 and GG2;

[0038] After 64 rounds of round function operations, reload the 8 __m256i type word registers back into the array and perform exclusive OR with the 8 unsigned integer arrays to obtain the ciphertext of this round of operation;

[0039] Take out the remaining part after the preset number of bits from each piece of data, add the bit string representing the data length at the end, and integrate them into one piece of data in order and put it into the message array, and execute the steps after using the VPGATHERDD instruction to obtain the final ciphertext;

[0040] According to the plaintext input length L, repeat the above process. After executing several times, put the ciphertext output in each round into the unsigned character type array t.

[0041] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in the first aspect of the present invention or the steps in the method for improving the SM2 data decryption speed based on SIMD described in the second aspect of the present invention.

[0042] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in the first aspect of the present invention or the steps in the method for improving the SM2 data decryption speed based on SIMD described in the second aspect of the present invention.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] 1. For the method for fast implementation of SM2 encryption and decryption based on SIMD of the present invention, multiple pieces of data after adding CT values are each taken out a preset number of bits, integrated into one piece and stored in the message array. The first step of message expansion is implemented using the pre-calculated index values and the VPGATHERDD instruction, and the expanded data is rearranged using the PSHUFB instruction, replacing the pointer to implement the byte reverse order function, thereby improving the overall operation efficiency of the algorithm.

[0045] 2. For the method for fast implementation of SM2 encryption and decryption based on SIMD of the present invention, when the first step of message expansion is completed, the data has been loaded into the register through the VPGATHERDD instruction, omitting the subsequent separate loading process; the second and third steps of message expansion are implemented using automated loop instructions, reducing the amount of code.

[0046] Advantages of additional aspects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0048] Figure 1 It is a schematic diagram of the method for improving the SM2 data encryption speed based on SIMD provided in Embodiment 1 of the present invention.

[0049] Figure 2 It is an effect diagram of putting the same data in hash0 to hash7 into the same register provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The present invention will be further described below in conjunction with the drawings and embodiments.

[0051] It should be noted that the following detailed descriptions are all exemplary and are intended to provide a further description of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0052] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0053] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0054] Embodiment 1:

[0055] As Figure 1 shown, Embodiment 1 of the present invention provides a method for improving the SM2 data encryption speed based on SIMD, including the following processes:

[0056] S1: Use a random number generator to generate a random number k, use k to calculate the elliptic curve point C1(x1, y1), and convert the data type of C1 into a bit string. Calculate the elliptic curve point (x2, y2) according to the random number k, and convert the data type of the elliptic curve point (x2, y2) into a bit string and merge them into a piece of data z as the input for the next step;

[0057] S2: According to the length L bits of the input data z, pre-calculate all 32-bit counter CT values and put them into an array for waiting to be used. The value of CT is from 1 to L / 32.

[0058] S3: Copy the data z into eight identical ones, each with a fixed length of 64 bytes, and add the corresponding 32-bit counter CT value at the end of each piece of data in sequence.

[0059] S4: Define temporary variables participating in the operation.

[0060] S4.1: Define an unsigned integer array index32 and load it into the register index;

[0061] index32[8] = {0, 64, 128, 192, 256, 320, 384, 448};

[0062] index = _mm256_loadu_si256((__m256i*)index32_0);

[0063] S5: Take the 8 pieces of incoming data input[0] to input[7], take out 512 bits from each and integrate them into one piece and store them in the message array;

[0064] memcpy(message, input[0], 64);

[0065] memcpy(message + 64, input[1], 64);

[0066] memcpy(message + 128, input[2], 64);

[0067] memcpy(message + 192, input[3], 64);

[0068] memcpy(message + 256, input[4], 64);

[0069] memcpy(message + 320, input[5], 64);

[0070] memcpy(message + 384, input[6], 64);

[0071] memcpy(message + 448, input[7], 64).

[0072] S6: Use the VPGATHERDD instruction and the index defined in S4.1 to put the data in message into 16 __m256i - type variables from w0 to w15.

[0073] w0 = _mm256_i32gather_epi32(message, index, 1);

[0074] w1 = _mm256_i32gather_epi32(message + 4, index, 1);

[0075] w2 = _mm256_i32gather_epi32(message + 8, index, 1);

[0076] w3 = _mm256_i32gather_epi32(message + 12, index, 1);

[0077] w4 = _mm256_i32gather_epi32(message + 16, index, 1);

[0078] w5 = _mm256_i32gather_epi32(message + 20, index, 1);

[0079] w6 = _mm256_i32gather_epi32(message + 24, index, 1);

[0080] w7 = _mm256_i32gather_epi32(message + 28, index, 1);

[0081] w8 = _mm256_i32gather_epi32(message + 32, index, 1);

[0082] w9 = _mm256_i32gather_epi32(message + 36, index, 1);

[0083] w10 = _mm256_i32gather_epi32(message + 40, index, 1);

[0084] w11 = _mm256_i32gather_epi32(message + 44, index, 1);

[0085] w12 = _mm256_i32gather_epi32(message + 48, index, 1);

[0086] w13 = _mm256_i32gather_epi32(message + 52, index, 1);

[0087] w14g = _mm256_i32gather_epi32(message + 56, index, 1);

[0088] w15g = _mm256_i32gather_epi32(message + 60, index, 1).

[0089] S7: Use the PSHUFB instruction to perform byte reverse order operation on the data in w0 to w15.

[0090] w0 = _mm256_shuffle_epi8(w0, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3));

[0091] w1 = _mm256_shuffle_epi8(w1, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3));

[0092] ···

[0093] w15 = _mm256_shuffle_epi8(w15, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3));

[0094] S8: Put the processed w0 to w15 into group of __m256i array type.

[0095] group

[68] = {w0, w1, w2, w3, w4, w5, w6, w7, w8, w9, w10, w11, w12, w13, w14, w15};

[0096] S9: Perform the first - step message expansion on the 16 data in group, and finally expand them to 68.

[0097] FOR i = 16 to 67;

[0098] group[i] = _mm256_xor_si256(_mm256_xor_si256(P1(_mm256_xor_si256(_mm256_xor_si256(group[i - 16], group[i - 9]), MoveLeft(group[i - 3], 15))), MoveLeft(group[i - 13], 7)), group[i - 6]);

[0099] ENDFOR

[0100] S9.1: The P1 function is:

[0101] P1(X) = X ⊕ (X <<< 15) ⊕ (X <<< 23);

[0102] S9.2: The MoveLeft function is:

[0103] MoveLeft(data, length) = _mm256_xor_si256(_mm256_slli_epi32(data, length), _mm256_srli_epi32(data, 32 - length));

[0104] S10: Perform the second - step message expansion on 68 data in group, and finally expand them into 64 data, and put them into the group2 array of type __m256i.

[0105] FOR i = 0 to 63;

[0106] group2[i] = _mm256_xor_si256(group[i], group[i + 4]);

[0107] ENDFOR

[0108] S11: Define 8 unsigned integer arrays hash0 to hash7, and store the initial values of the SM3 hash function in 16 - hexadecimal format in each array, and the contents of the 8 arrays are the same;

[0109] Hash = {0x7380166f, 0x4914b2b9, 0x172442d7, 0xda8a0600, 0xa96f30bc, 0x163138aa, 0xe38dee4d, 0xb0fb0e4e}.

[0110] S12: Define 8 __m256i - type word registers: A, B, C, D, E, F, G, H, and put the same data in hash0 to hash7 into the same register, as Figure 2 shown.

[0111] S13: Define temporary data SS1, SS2, TT2, all of type __m256i.

[0112] Define an unsigned integer array T, and its content is:

[0113] 0x79cc4519, 0xf3988a32, 0xe7311465, 0xce6228cb, 0x9cc45197, 0x3988a32f, 0x7311465e, 0xe6228cbc, 0xcc451979, 0x988a32f3, 0x311465e7, 0x6228cbce, 0xc451979c, 0x88a32f39, 0x11465e73, 0x228cbce6, 0x9d8a7a87, 0x3b14f50f, 0x7629ea1e, 0xec53d43c, 0xd8a7a879, 0xb14f50f3, 0x629ea1e7, 0xc53d43ce, 0x8a7a879d, 0x14f50f3b, 0x29ea1e76, 0x53d43cec, 0xa7a879d8, 0x4f50f3b1, 0x9ea1e762, 0x3d43cec5, 0x7a879d8a, 0xf50f3b14, 0xea1e7629, 0xd43cec53, 0xa879d8a7, 0x50f3b14f, 0xa1e7629e, 0x43cec53d, 0x879d8a7a, 0xf3b14f5, 0x1e7629ea, 0x3cec53d4, 0x79d8a7a8, 0xf3b14f50, 0xe7629ea1, 0xcec53d43, 0x9d8a7a87, 0x3b14f50f, 0x7629ea1e, 0xec53d43c, 0xd8a7a879, 0xb14f50f3, 0x629ea1e7, 0xc53d43ce, 0x8a7a879d, 0x14f50f3b, 0x29ea1e76, 0x53d43cec, 0xa7a879d8, 0x4f50f3b1, 0x9ea1e762, 0x3d43cec5。

[0114] S14: Define Boolean functions FF1, FF2, GG1, and GG2.

[0115] FF1(X, Y, Z) = X ⊕ Y ⊕ Z;

[0116] FF2(X, Y, Z) = (X ∧ Y) ∨ (X ∧ Z) ∨ (Y ∧ Z);

[0117] GG1(X, Y, Z) = X ⊕ Y ⊕ Z;

[0118]

[0119] S15: Execute the round function RoundFunction1 for 16 rounds. The data involved in the operation is of type __m256i, and the AVX2 instruction is used to replace the original logical operation. Here, temp is a temporary variable.

[0120] FOR i = 0 TO 15;

[0121] temp = A <<< 12;

[0122] SS1 = (temp + E + T[i]) <<< 7;

[0123] SS2 = SS1 ⊕ temp;

[0124] D = FF1(A, B, C) + D + SS2 + group2[i];

[0125] TT2 = GG1(E, F, G) + H + SS1 + group[i];

[0126] B = B <<< 9;

[0127] F = F <<< 19;

[0128] H = P0(TT2);

[0129] ENDFOR;

[0130] Where P0 is X ⊕ (X <<< 9) ⊕ (X <<< 17).

[0131] S16: Execute the round function RoundFunction2 for 48 rounds. The data involved in the operation is of type __m256i, and the AVX2 instruction is used to replace the original logical operation. Here, temp is a temporary variable.

[0132] FOR i = 16 TO 63;

[0133] temp = A <<< 12;

[0134] SS1 = (temp + E + T[i]) <<< 7;

[0135] SS2 = SS1 ⊕ temp;

[0136] D = FF2(A, B, C) + D + SS2 + group2[i];

[0137] TT2 = GG2(E, F, G) + H + SS1 + group[i];

[0138] B = B <<< 9;

[0139] F = F <<< 19;

[0140] H = P0(TT2);

[0141] ENDFOR.

[0142] S17: After performing 64 rounds of round function operations, reload A, B, C, D, E, F, G, H back into the array, and perform an exclusive OR operation with hash0 to hash7 to obtain the ciphertext of this round of operation.

[0143] S18: Append the bit string representing the data length to the end of the remaining part of each piece of data in S4, and integrate them into one piece of data in order and put it into the message array, then execute S6 to S17 to obtain the final ciphertext.

[0144] S19: According to the plaintext input length L, repeatedly execute S3 - S18. After executing several times, put the ciphertext output in each round into the unsigned character type array t;

[0145] S20: Output the result t of the KDF operation, perform an exclusive OR operation on t and message to obtain C2;

[0146] S21: Combine the coordinate point x2, the message message, and the coordinate point y2 in order (x2||message||y2), and perform an operation on it using the SM3 hash function to obtain C3;

[0147] S22: Combine C1, C3, and C2 in order as C (C = C1||C3||C2) and output it. C is the encryption result.

[0148] Example 2:

[0149] The second embodiment of the present invention provides a method for improving the SM2 data decryption speed based on SIMD, including:

[0150] Assume the encrypted ciphertext is C, C = C1||C3||C2;

[0151] B1: Take out C1 from the ciphertext and convert the data type of C1 into a point on the elliptic curve;

[0152] B2: Calculate the elliptic curve points x2 and y2, then convert the data types of the coordinates x2 and y2 into bit strings and combine x2 and y2 into one (x2||y2).

[0153] B3: Use the parallelized KDF strategy involved in the encryption process to calculate x2 and y2 to obtain the operation result t.

[0154] B4: Take out the bit string C2 from C, and perform an exclusive OR operation on t and C2 to obtain the plaintext M.

[0155] B5: Combine the coordinates x2, the plaintext M, and the coordinate y2 (x2||M||y2), and perform an operation on it using the SM3 hashing algorithm to obtain the result u.

[0156] B6: Take out the bit string C3 from C, and determine whether u is equal to C3. If they are not equal, report an error and exit. Otherwise, output the plaintext M.

[0157] Among them, the parallelized KDF strategy includes:

[0158] F1: According to the input data length of L bits, pre-compute all 32-bit counter CT values and put them into an array for waiting to be used. The value of CT is from 1 to L / 32.

[0159] F2: Obtain 8 pieces of input data, each with a fixed length of 64 bytes, and add the corresponding 32-bit counter CT value at the end of each piece of data in sequence.

[0160] F3: Define temporary variables participating in the operation.

[0161] F3.1: Define an unsigned integer array index32 and load it into the register index;

[0162] index32[8] = {0, 64, 128, 192, 256, 320, 384, 448};

[0163] index = _mm256_loadu_si256((__m256i*)index32_0);

[0164] F4: Take the 8 pieces of incoming data input[0] to input[7], extract 512 bits from each and integrate them into one piece and store it in the message array;

[0165] memcpy(message, input[0], 64);

[0166] memcpy(message + 64, input[1], 64);

[0167] memcpy(message + 128, input[2], 64);

[0168] memcpy(message + 192, input[3], 64);

[0169] memcpy(message + 256, input[4], 64);

[0170] memcpy(message + 320, input[5], 64);

[0171] memcpy(message + 384, input[6], 64);

[0172] memcpy(message + 448, input[7], 64).

[0173] F5: Use the VPGATHERDD instruction and the index defined in F3.1 to put the data in message into 16 __m256i - type variables from w0 to w15.

[0174] w0 = _mm256_i32gather_epi32(message, index, 1);

[0175] w1 = _mm256_i32gather_epi32(message + 4, index, 1);

[0176] w2 = _mm256_i32gather_epi32(message + 8, index, 1);

[0177] w3 = _mm256_i32gather_epi32(message + 12, index, 1);

[0178] w4 = _mm256_i32gather_epi32(message + 16, index, 1);

[0179] w5 = _mm256_i32gather_epi32(message + 20, index, 1);

[0180] w6 = _mm256_i32gather_epi32(message + 24, index, 1);

[0181] w7 = _mm256_i32gather_epi32(message + 28, index, 1);

[0182] w8 = _mm256_i32gather_epi32(message + 32, index, 1);

[0183] w9 = _mm256_i32gather_epi32(message + 36, index, 1);

[0184] w10 = _mm256_i32gather_epi32(message + 40, index, 1);

[0185] w11 = _mm256_i32gather_epi32(message + 44, index, 1);

[0186] w12 = _mm256_i32gather_epi32(message + 48, index, 1);

[0187] w13 = _mm256_i32gather_epi32(message + 52, index, 1);

[0188] w14g = _mm256_i32gather_epi32(message + 56, index, 1);

[0189] w15g = _mm256_i32gather_epi32(message + 60, index, 1).

[0190] F6: Use the PSHUFB instruction to perform byte reverse order operation on the data in w0 to w15.

[0191] w0 = _mm256_shuffle_epi8(w0, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3));

[0192] w1 = _mm256_shuffle_epi8(w1, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3));

[0193] ···

[0194] w15 = _mm256_shuffle_epi8(w15, _mm256_set_epi8(28, 29, 30, 31, 24, 25, 26, 27, 20, 21, 22, 23, 16, 17, 18, 19, 12, 13, 14, 15, 8, 9, 10, 11, 4, 5, 6, 7, 0, 1, 2, 3)).

[0195] F7: Put the processed w0 to w15 into the group of __m256i array type.

[0196] group

[68] = {w0, w1, w2, w3, w4, w5, w6, w7, w8, w9, w10, w11, w12, w13, w14, w15};

[0197] F8: Perform the first - step message expansion on the 16 data in group, and finally expand it to 68.

[0198] FOR i = 16 to 67;

[0199] group[i] = _mm256_xor_si256(_mm256_xor_si256(P1(_mm256_xor_si256(_mm256_xor_si256(group[i - 16], group[i - 9]), MoveLeft(group[i - 3], 15))), MoveLeft(group[i - 13], 7)), group[i - 6]);

[0200] ENDFOR

[0201] F8.1: Where the P1 function is:

[0202] P1(X) = X ⊕ (X <<< 15) ⊕ (X <<< 23);

[0203] F8.2: The MoveLeft function is:

[0204] MoveLeft(data, length) = _mm256_xor_si256(_mm256_slli_epi32(data, length), _mm256_srli_epi32(data, 32 - length));

[0205] F9: Perform the second - step message expansion on the 68 data in group, and finally expand it to 64, and store them in the group2 array of type __m256i.

[0206] FOR i = 0 to 63;

[0207] group2[i] = _mm256_xor_si256(group[i], group[i + 4]);

[0208] ENDFOR

[0209] F10: Define 8 unsigned integer arrays hash0 to hash7, each array stores 8 hexadecimal initial values of the SM3 hash function and the contents of the 8 arrays are the same;

[0210] Hash = {0x7380166f, 0x4914b2b9, 0x172442d7, 0xda8a0600, 0xa96f30bc, 0x163138aa, 0xe38dee4d, 0xb0fb0e4e}。

[0211] F11: Define 8 __m256i - type word registers: A, B, C, D, E, F, G, H, and put the same data among hash0 to hash7 into the same register, such as Figure 2 shown.

[0212] F12: Define temporary data SS1, SS2, TT2, all of which are of __m256i type.

[0213] Define an unsigned integer array T, whose content is:

[0214] 0x79cc4519, 0xf3988a32, 0xe7311465, 0xce6228cb, 0x9cc45197, 0x3988a32f, 0x7311465e, 0xe6228cbc, 0xcc451979, 0x988a32f3, 0x311465e7, 0x6228cbce, 0xc451979c, 0x88a32f39, 0x11465e73, 0x228cbce6, 0x9d8a7a87, 0x3b14f50f, 0x7629ea1e, 0xec53d43c, 0xd8a7a879, 0xb14f50f3, 0x629ea1e7, 0xc53d43ce, 0x8a7a879d, 0x14f50f3b, 0x29ea1e76, 0x53d43cec, 0xa7a879d8, 0x4f50f3b1, 0x9ea1e762, 0x3d43cec5, 0x7a879d8a, 0xf50f3b14, 0xea1e7629, 0xd43cec53, 0xa879d8a7, 0x50f3b14f, 0xa1e7629e, 0x43cec53d, 0x879d8a7a, 0xf3b14f5, 0x1e7629ea, 0x3cec53d4, 0x79d8a7a8, 0xf3b14f50, 0xe7629ea1, 0xcec53d43, 0x9d8a7a87, 0x3b14f50f, 0x7629ea1e, 0xec53d43c, 0xd8a7a879, 0xb14f50f3, 0x629ea1e7, 0xc53d43ce, 0x8a7a879d, 0x14f50f3b, 0x29ea1e76, 0x53d43cec, 0xa7a879d8, 0x4f50f3b1, 0x9ea1e762, 0x3d43cec5。

[0215] F13: Define the Boolean functions FF1, FF2, GG1, and GG2.

[0216] FF1(X, Y, Z) = X ⊕ Y ⊕ Z;

[0217] FF2(X, Y, Z) = (X ∧ Y) ∨ (X ∧ Z) ∨ (Y ∧ Z);

[0218] GG1(X, Y, Z) = X ⊕ Y ⊕ Z;

[0219]

[0220] F14: Execute the round function RoundFunction1 for 16 rounds. The data involved in the operation is of type __m256i, and AVX2 instructions are used to replace the original logical operations. Here, temp is a temporary variable.

[0221] FOR i = 0 TO 15;

[0222] temp = A <<< 12;

[0223] SS1 = (temp + E + T[i]) <<< 7;

[0224] SS2 = SS1 ⊕ temp;

[0225] D = FF1(A, B, C) + D + SS2 + group2[i];

[0226] TT2 = GG1(E, F, G) + H + SS1 + group[i];

[0227] B = B <<< 9;

[0228] F = F <<< 19;

[0229] H = P0(TT2);

[0230] ENDFOR;

[0231] Where P0 is X ⊕ (X <<< 9) ⊕ (X <<< 17).

[0232] F15: Execute the round function RoundFunction2 for 48 rounds. The data involved in the operation is of type __m256i, and AVX2 instructions are used to replace the original logical operations. Here, temp is a temporary variable.

[0233] FOR i = 16 TO 63;

[0234] temp = A <<< 12;

[0235] SS1 = (temp + E + T[i]) <<< 7;

[0236] SS2 = SS1 ⊕ temp;

[0237] D = FF2(A, B, C) + D + SS2 + group2[i];

[0238] TT2 = GG2(E, F, G) + H + SS1 + group[i];

[0239] B = B <<< 9;

[0240] F = F <<< 19;

[0241] H = P0(TT2);

[0242] ENDFOR.

[0243] F16: After performing 64 rounds of round function operations, reload A, B, C, D, E, F, G, and H back into the array, and perform an exclusive OR operation with hash0 to hash7 to obtain the ciphertext for this round of operation.

[0244] F17: Append the bit string representing the data length to the end of the remaining part of each piece of data in S4, and integrate them into one piece of data in order and then put it into the message array, and execute F5 to F16 to obtain the final ciphertext.

[0245] F18: According to the plaintext input length L, repeatedly execute F2 - F17. After executing several times, put the ciphertext output in each round into the unsigned character type array t.

[0246] F19: Output the KDF operation result t.

[0247] Example 3:

[0248] Embodiment 3 of the present invention provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in Embodiment 1 of the present invention or the steps in the method for improving the SM2 data decryption speed based on SIMD described in Embodiment 2.

[0249] Example 4:

[0250] Embodiment 4 of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in Embodiment 1 of the present invention or the steps in the method for improving the SM2 data decryption speed based on SIMD described in Embodiment 2.

[0251] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.

[0252] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0253] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0254] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.

[0255] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0256] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for improving the data encryption speed of SM2 based on SIMD, characterized in that: It includes the following processes: Use a random number generator to generate a random number k, calculate the elliptic curve point C1(x1, y1) using k, convert the data type of C1 to a bit string, and calculate the elliptic curve point (x2, y2) according to the random number k; Convert the data type of the elliptic curve point (x2, y2) to a bit string, merge them into one piece of data, copy it into multiple identical pieces of data, and add the corresponding 32-bit counter CT value at the end of each piece of data in sequence; Define an unsigned integer array index32 and load it into the register index. For each of the multiple pieces of data with the CT value added, extract a preset number of bits from each and integrate them into one piece and store it in the message array; Use the VPGATHERDD instruction and the index to put the data in message into 16 variables of type __m256i, and use the PSHUFB instruction to perform a byte reverse order operation on the data in the 16 variables of type __m256i; Put the 16 variables of type __m256i after the byte reverse order operation into the group of type __m256i array; perform the first step of message expansion on the 16 data in group, and finally expand it to 68; perform the second step of message expansion on the 68 data in group, and finally expand it to 64, and put them into the group2 array of type __m256i; Define 8 unsigned integer arrays, each array stores 8 hexadecimal initial values of the SM3 hash function and the contents of the 8 arrays are the same; define 8 word registers of type __m256i, and put the same data in the 8 unsigned integer arrays into the same register; Define temporary data SS1, SS2, TT2, all of type __m256i; define boolean functions FF1, FF2, GG1 and GG2; according to group, group2, SS1, SS2, TT2, FF1, FF2, GG1 and GG2, execute the 16-round round function RoundFunction1 and the 48-round round function RoundFunction2; After performing 64 rounds of round function operations, reload the 8 word registers of type __m256i back into the array, and perform an exclusive OR with the 8 unsigned integer arrays to obtain the ciphertext of this round of operation; Add a bit string representing the data length to the end of the remaining part after extracting the preset number of bits from each piece of data, integrate them into one piece of data in sequence and put it into the message array, and execute the steps after using the VPGATHERDD instruction to obtain the final ciphertext; According to the plaintext input length L, repeat the above process. After executing several times, put the ciphertext output in each round into the unsigned character array t; Perform an exclusive OR on t and message to obtain C2; Merge the coordinate points x2, the message message, and the coordinate point y2 in sequence (x2||message||y2), and perform an operation on it using the SM3 hash function to obtain C3; Merge C1, C3, and C2 in sequence to obtain C = C1||C3||C2 and output it. C is the encryption result.

2. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Obtain 8 input data, each with a fixed length of 64 bytes; Take out 512 bits from each and integrate them into one and store them in the message array.

3. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Calculate the elliptic curve points x2 and y2, convert the data types of x2 and y2 into bit strings and merge them into one data as the plaintext data. According to the length L bits of the plaintext data, calculate the values of all 32-bit counters CT and put them into an array for waiting to be used. The value of CT is from 1 to L / 32.

4. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Perform the first-step message expansion on 16 data in group, and finally expand them to 68, including: i takes values from 16 to 67 in sequence; group[i] = _mm256_xor_si256(_mm256_xor_si256(P1(_mm256_xor_si256(_mm256_xor_si256(group[i - 16], group[i - 9]), MoveLeft(group[i - 3], 15))), MoveLeft(group[i - 13], 7)), group[i - 6]); Among them, P1(X) = X ⊕ (X <<< 15) ⊕ (X <<< 23); MoveLeft(data, length) = _mm256_xor_si256(_mm256_slli_epi32(data, length), _mm256_srli_epi32(data, 32 - length)).

5. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Perform the second-step message expansion on 68 data in group, and finally expand them to 64, and put them into the group2 array of __m256i type, including: i takes values from 0 to 63 in sequence; group2[i] = _mm256_xor_si256(group[i], group[i + 4]).

6. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Set 8 __m256i type word registers to be: A, B, C, D, E, F, G, H; Then the 16-round round function RoundFunction1 includes: i takes values from 0 to 15 in sequence; temp = A <<< 12; SS1 = (temp + E + T[i]) <<< 7; SS2 = SS1 ⊕ temp; D = FF1(A, B, C) + D + SS2 + group2[i]; TT2 = GG1(E, F, G) + H + SS1 + group[i]; B = B <<< 9; F = F <<< 19; H = P0(TT2); P0 = X ⊕ (X <<< 9) ⊕ (X <<< 17).

7. The method for improving the SM2 data encryption speed based on SIMD according to claim 1, characterized in that: Set 8 __m256i type word registers as: A, B, C, D, E, F, G, H; Then the 48-round round function RoundFunction2 includes: i takes values from 16 to 63 in sequence; temp = A <<< 12; SS1 = (temp + E + T[i]) <<< 7; SS2 = SS1 ⊕ temp; D = FF2(A, B, C) + D + SS2 + group2[i]; TT2 = GG2(E, F, G) + H + SS1 + group[i]; B = B <<< 9; F = F <<< 19; H = P0(TT2).

8. A method for improving the SM2 data decryption speed based on SIMD, characterized in that: Let the encrypted ciphertext be C, C = C1||C3||C2; Extract C1 from the ciphertext and convert the data type of C1 into a point on the elliptic curve; Calculate the elliptic curve points x2, y2, and then convert the data types of the coordinates x2, y2 into bit strings and merge x2, y2 into one (x2||y2); Use the parallelized KDF strategy involved in the encryption process to calculate x2, y2, and obtain the operation result t; Extract the bit string C2 from C, and perform an exclusive OR operation on t and C2 to obtain the plaintext M; Merge the coordinates x2, the plaintext M, and the coordinates y2 (x2||M||y2), and perform an operation on it using the SM3 hashing algorithm to obtain the result u; Extract the bit string C3 from C, determine whether u is equal to C3, if not, report an error and exit, otherwise output the plaintext M; Among them, the parallelized KDF strategy includes the following processes: Obtain multiple input data, and sequentially add the corresponding 32-bit counter CT value at the end of each data; Define an unsigned integer array index32 and load it into the register index, and for each of the multiple data after adding the CT value, extract a preset number of bits and integrate them into one and store them in the message array; Use the VPGATHERDD instruction and the index index to put the data in the message into 16 __m256i type variables, and use the PSHUFB instruction to perform a byte reverse operation on the data in the 16 __m256i type variables; Put the 16 __m256i - type variables after byte - reverse operation into the __m256i array - type group; perform the first - step message expansion on the 16 data in the group, and finally expand them to 68; perform the second - step message expansion on the 68 data in the group, and finally expand them to 64, and put them into the __m256i - type group2 array; Define 8 unsigned - integer arrays, each array stores 8 hexadecimal initial values of the SM3 hash function and the contents of the 8 arrays are the same; define 8 __m256i - type word registers, and put the same data in the 8 unsigned - integer arrays into the same register; Define temporary data SS1, SS2, TT2, all of which are of __m256i type; define boolean functions FF1, FF2, GG1 and GG2; execute the 16 - round round function RoundFunction1 and the 48 - round round function RoundFunction2 according to group, group2, SS1, SS2, TT2, FF1, FF2, GG1 and GG2; After 64 - round round - function operations, reload the 8 __m256i - type word registers back into the array, and perform an exclusive - OR operation with the 8 unsigned - integer arrays to obtain the ciphertext of this round of operation; Take the remaining part after removing the preset number of bits from each piece of data, append the bit string representing the data length at the end, and integrate them into a piece of data in order and put it into the message array, and execute the steps after using the VPGATHERDD instruction to obtain the final ciphertext; According to the plain - text input length L, repeat the above process. After executing several times, put the ciphertext output in each round into the unsigned - character - type array t.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in any one of claims 1 - 7 or the steps in the method for improving the SM2 data decryption speed based on SIMD described in claim 8.

10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for improving the SM2 data encryption speed based on SIMD described in any one of claims 1 - 7 or the steps in the method for improving the SM2 data decryption speed based on SIMD described in claim 8.

Citation Information

Patent Citations

  • Realization method of SM2 encryption algorithm on binary extension field

    CN107147495A

  • SM3 parallel data encryption operation method and system based on SIMD

    CN113794552A