Method and system for quickly realizing cryptographic algorithm encryption and decryption based on key derivation function

Optimizing the key derived function of the SM2 public key encryption algorithm through pre-computation and SIMD instructions, the problem of high computational complexity of the key derived function is solved, the encryption and decryption speed and efficiency are improved, and the security and memory consumption are achieved.

CN120433928APending Publication Date: 2025-08-05SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510675340.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-13
Filing Date
2025-05-23
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the prior art, the key derivation function of the SM2 public key encryption algorithm has a high computational complexity, resulting in slow encryption and decryption speed, occupying computer memory, and affecting the performance of the elliptic curve algorithm.

Method used

Optimize key derivation functions through pre-calculation and SIMD instructions, including message expansion and iterative compression of the standard SM3 hash algorithm for the first 512 bits, and then batch processing of the remaining 32 bits to improve KDF speed.

Benefits of technology

It greatly improves the computing speed of key-derived functions, reduces the computational complexity, improves the encryption and decryption efficiency, and realizes the advantages of security and low memory footprint.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433928A_ABST
    Figure CN120433928A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cryptographic algorithm encryption and decryption, in particular to a cryptographic algorithm encryption and decryption method and system based on a key derivation function, and the method comprises the steps of encrypting a message and decrypting a ciphertext. A key derivation function in the step of encrypting a message and the step of decrypting a ciphertext is accelerated and comprises the following steps of: firstly, executing message expansion and iterative compression of a standard SM3 Hash algorithm on the first 512 bits Z to obtain a 256-bit Hash value H; and then, carrying out batch processing on the remaining 32 bits to obtain a calculation result of the key derivation function. The calculation complexity is reduced through pre-calculation and SIMD instructions, the KDF speed is greatly improved, and the method has the advantages of being universal, small in memory occupation, high in efficiency, safe to achieve and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of encryption and decryption of national secret algorithms, and in particular to a method and system for quickly implementing encryption and decryption of national secret algorithms based on a key derivation function. Background Art

[0002] The SM2 public key encryption algorithm is a public key encryption algorithm based on elliptic curve cryptography and is widely used in operations such as data encryption, decryption, and digital signatures.

[0003] The SM2 public key encryption algorithm includes executing a key derivation function. The execution process of the key derivation function is a key factor affecting the encryption and decryption performance of the elliptic curve algorithm. Although the existing technology has accelerated the SM3 algorithm based on SIMD parallel methods, which has greatly improved the speed compared to the standard implementation, its speed is still far behind that of the symmetric encryption algorithm.

[0004] During the encryption and decryption process, the key derivation function in the elliptic curve has high computational complexity and slow operation speed, which occupies computer memory and affects the speed of encryption and decryption. Summary of the Invention

[0005] In order to address the shortcomings of the existing technology, the present invention provides a national secret algorithm encryption and decryption method and system based on the rapid implementation of the key derivation function; by pre-calculation and SIMD instructions, the computational complexity is reduced and the KDF speed is greatly improved. The method of the present invention has the advantages of versatility, low memory usage, high efficiency, and security.

[0006] On the one hand, a national secret algorithm encryption and decryption method based on a fast implementation of a key derivation function is provided, including: a step of encrypting a message and a step of decrypting a ciphertext; the key derivation function in the step of encrypting a message and the step of decrypting a ciphertext is accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; then, batch processing the remaining 32 bits to obtain the calculation result of the key derivation function.

[0007] On the other hand, a national secret algorithm encryption and decryption system based on a fast implementation of a key derivation function is provided, including: a module for encrypting messages and a module for decrypting ciphertext; the key derivation function in the module for encrypting messages and the module for decrypting ciphertext is accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; then, batch processing the remaining 32 bits to obtain the calculation result of the key derivation function.

[0008] The above technical solution has the following advantages or beneficial effects:

[0009] The key derivation function in the steps of encrypting a message and decrypting a ciphertext is accelerated. First, the standard SM3 hash algorithm is used to expand and iteratively compress the first 512 bits Z to obtain a 256-bit hash value H. Then, the remaining 32 bits are processed in batches to obtain the calculation result of the key derivation function. Compared with previous technologies, the method of the present invention reduces computational complexity through pre-computation and SIMD instructions, greatly improving the KDF speed, and has the advantages of versatility, low memory usage, high efficiency, and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0011] Figure 1 This is a flow chart of the method of embodiment 1. DETAILED DESCRIPTION

[0012] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.

[0013] Example 1, as Figure 1 As shown, this embodiment provides a national secret algorithm encryption and decryption method based on a fast implementation of a key derivation function, including: a step of encrypting a message and a step of decrypting a ciphertext; the key derivation function in the step of encrypting the message and the step of decrypting the ciphertext is accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; then, batch processing the remaining 32 bits to obtain the calculation result of the key derivation function.

[0014] Furthermore, the step of encrypting the message includes:

[0015] Suppose the message to be sent is a bit string M, and klen is the bit length of M;

[0016] To encrypt plaintext M, user A, as the encryptor, performs the following operations:

[0017] A1: Use a random number generator to generate a random number k∈[1,N-1]; N represents the order of SM2;

[0018] A2: Calculate the elliptic curve point C1 = [k]G = (x1, y1) and convert the data type of C1 to a bit string.

[0019] A3: Calculate the elliptic curve point S = [h]P B , if S is a point at infinity, report an error and exit;

[0020] A4: Calculate elliptic curve point [k]P B =(x2,y2), convert the data type of x2,y2 to bit string;

[0021] A5: Calculate t = KDF(x²||y²,klen). If t is a string of all 0 bits, return to A1.

[0022] A6: Calculation

[0023] A7: Calculate C3 = Hash(x2||M||y2);

[0024] A8: Output ciphertext C=C1||C2||C3.

[0025] Furthermore, the step of decrypting the ciphertext includes:

[0026] Assume klen is the bit length of C2 in the ciphertext; to decrypt the ciphertext C = C1||C2||C3, user B, as the decryptor, should perform the following operations:

[0027] B1: Extracts the bit string C1 from the ciphertext C, converts the data type of C1 to a point on the elliptic curve, and verifies whether C1 satisfies the elliptic curve equation. If not, an error is reported and the process exits.

[0028] B2: Calculate the elliptic curve point S = [h]C1. If S is a point at infinity, report an error and exit.

[0029] B3: Calculate [d B ]C1=(x2,y2), convert the data type of x2,y2 to bit string;

[0030] B4: Calculate t = KDF(x²||y²,klen). If t is a string of all 0 bits, report an error and exit.

[0031] B5: Extract the bit string C2 from the ciphertext C and calculate

[0032] B6: Calculate u = Hash(x2||M′||y2), extract the bit string C3 from the ciphertext C, and report an error and exit if u≠C3.

[0033] B7: Output plaintext M′.

[0034] Furthermore, the key derivation function includes:

[0035] Let the cryptographic hash function be H v (·), whose output is a hash value of length v bits.

[0036] Key derivation function KDF(Z,klen):

[0037] Input: bit string Z, integer klen; klen represents the bit length of the key data to be obtained;

[0038] Output: key data bit string K of length klen;

[0039] (a) Initialize a 32-bit counter ct = 0x00000001;

[0040] (b) For i from 1 to implement:

[0041] (b1) Calculate H ai =H v (Z||ct);

[0042] (b2)ct plus 1;

[0043] (c) If klen / v is an integer, then let Otherwise, let for The leftmost Bit;

[0044] (d) Order

[0045] Furthermore, before accelerating the execution of the key derivation function in the steps of encrypting the message and decrypting the ciphertext, the method further includes:

[0046] Precompute and save parameter tables:

[0047] SM3_Tj

[64] ={0x79cc4519,0xf3988a32,0xe7311465,0xce6228cb,0x9cc45197,0x3988a32f,0x7311465e, 0xe6228cbc,0xcc451979,0x988a32f3,0x311465e7,0x6228cbce,0xc451979c,0x88a32f39,0x11465e73,0x 228cbce6,0x9d8a7a87,0x3b14f50f,0x7629ea1e,0xec53d43c,0xd8a7a879,0xb14f50f3,0x629ea1e7,0xc 53d43ce,0x8a7a879d,0x14f50f3b,0x29ea1e76,0x53d43cec,0xa7a879d8,0x4f50f3b1,0x9ea1e762,0x3d4 3cec5,0x7a879d8a,0xf50f3b14,0xea1e7629,0xd43cec53,0xa879d8a7,0x50f3b14f,0xa1e7629e,0x43ce c53d,0x879d8a7a,0x0f3b14f5,0x1e7629ea,0x3cec53d4,0x79d8a7a8,0xf3b14f50,0xe7629ea1,0xcec53d 43,0x9d8a7a87,0x3b14f50f,0x7629ea1e,0xec53d43c,0xd8a7a879,0xb14f50f3,0x629ea1e7,0xc53d43ce ,0x8a7a879d,0x14f50f3b,0x29ea1e76,0x53d43cec,0xa7a879d8,0x4f50f3b1,0x9ea1e762,0x3d43cec5}.

[0048] Because the input message length is fixed at 544 bits, and only the first 32 bits of the last 512 bits are different, and the rest are fixed values, some message extension values can be precalculated.

[0049] Furthermore, the key derivation function in the step of encrypting the message and the step of decrypting the ciphertext is accelerated, including:

[0050] Assume that the width of the SIMD register is w bits (usually 128, 256, or 512), and the step size is n = w / 32. Each register can be divided into n 32-bit blocks. Set the contents of each block in the SIMD register m0 to 1, 2, ... n-1, H = SM3Update(Z), and split H into 8 32-bit data blocks B0, B1....B7. Fill the SIMD array SW[8], and fill SW[index] with Bindex from 0 to 7, i is equal to 0; i represents the position of the currently processed data; index represents the number;

[0051] (1): A=SW[0]; B=SW[1]; C=SW[2]; D=SW[3]; E=SW[4]; F=SW[5]; G=SW[6]; H=SW[7];

[0052] Among them, A, B, C, D, E, F, G, H represent 8 SIMD registers;

[0053] Execute the p1(x) function on register m0 to obtain array SW_0; fill SW_1 with 0x80404000; fill array SW_2 with 0x01108888; fill SW_4 with 0x10005000; fill SW_5 with 0x0022028a; fill SW_7 with 0xac545c04; and fill SW_8 with 0x99cda89a; the p1(x) function is specifically implemented as follows: perform a 15-bit circular left shift on x in 32-bit groups, perform a 23-bit circular left shift on x in 32-bit groups, and then perform an XOR operation on x, and then output to x;

[0054] Execute a 15-bit circular left shift on SW_0 in 32-bit groups, and then execute the p1(x) function to obtain SW_3; execute a 15-bit circular left shift on SW_3 in 32-bit groups to obtain t0; execute the p1(x) function on t0, and then perform an XOR operation with SW_0 to obtain SW_6; execute a 15-bit circular left shift on SW_6 in 32-bit groups, and then perform an XOR operation with SW_0 to obtain t0; execute p1(x) on t0, and then perform an XOR operation with SW_3 to obtain SW_9; fill SW_10 with 0xa0003000; fill SW_11 with 0x002202a8; fill SW_12 with 0x11000; fill SW_13 with 0xa4505804; SW_0 to SW_15 are all SIMD registers;

[0055] Execute a cyclic left shift of 15 bits on SW_9 in 32-bit groups, and then perform an XOR operation with SW_3 to obtain t0. Execute p1(x) on t0, and then perform an XOR operation with SW_3 and SW_12 to obtain SW_12.

[0056] SW_0 is grouped into 32 bits and circularly shifted left by 7 bits. Then, it is XORed with SW_13 to obtain SW_13. SW_14 is filled with 0xf45691fb. T4 is filled with 0x00000220.

[0057] Call the status word generation instruction sequence Gen_W(T4,SW_6,SW_12,SW_2,SW_15,SW_9); t0=T4;

[0058] Call the first round function instruction sequence OneRoundA(A,E,B,F,C,G,D,H,m0,m0,SM3_Tj[0]); fill t1 with 0x80000000;

[0059] Call the first round function instruction sequence OneRoundA(D,H,A,E,B,F,C,G,t1,t1,SM3_Tj[1]);

[0060] Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj[2]);

[0061] Call the second round function instruction sequence OneRoundB(B,F,C,G,D,H,A,E,SM3_Tj[3]);

[0062] Call the second round function instruction sequence OneRoundB(A,E,B,F,C,G,D,H,SM3_Tj[4]);

[0063] Call the second round function instruction sequence OneRoundB(D,H,A,E,B,F,C,G,SM3_Tj[5]);

[0064] Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj[6]);

[0065] Call the second round function instruction sequence OneRoundB(B,F,C,G,D,H,A,E,SM3_Tj[7]);

[0066] Call the second round function instruction sequence OneRoundB(A,E,B,F,C,G,D,H,SM3_Tj[8]);

[0067] Call the second round function instruction sequence OneRoundB(D,H,A,E,B,F,C,G,SM3_Tj[9]);

[0068] Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj

[10] );

[0069] Call the third round function instruction sequence OneRounds(B,F,C,G,D,H,A,E,t0,SM3_Tj

[11] );

[0070] Call the third round function instruction sequence OneRounds(A,E,B,F,C,G,D,H,SW_0,SM3_Tj

[12] );

[0071] Call the third round function instruction sequence OneRounds(D,H,A,E,B,F,C,G,SW_1,SM3_Tj

[13] );

[0072] Call the third round function instruction sequence OneRounds(C,G,D,H,A,E,B,F,SW_2,SM3_Tj

[14] );

[0073] Perform an XOR operation on t0 and SW_3 to obtain T3;

[0074] Call the first round function instruction sequence OneRoundA(B,F,C,G,D,H,A,E,t0,T3,SM3_Tj

[15] );

[0075] Perform an XOR operation on SW_0 and SW_4 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_0,T3,SM3_Tj

[16] );

[0076] Call the status word generation instruction sequence Gen_W(SW_0,SW_7,SW_13,SW_3,SW_0,SW_10);

[0077] Perform an XOR operation on SW_1 and SW_5 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_1,T3,SM3_Tj

[17] ); call the status word generation instruction sequence Gen_W(SW_1,SW_8,SW_14,SW_4,SW_1,SW_11);

[0078] Perform XOR operation on SW_2 and SW_6 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_2,T3,SM3_Tj

[18] );

[0079] Call the status word generation instruction sequence Gen_W(SW_2,SW_9,SW_15,SW_5,SW_2,SW_12)

[0080] Perform XOR operation on SW_3 and SW_7 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_3,T3,SM3_Tj

[19] );

[0081] Call the status word generation instruction sequence Gen_W (SW_3, SW_10, SW_0, SW_6, SW_3, SW_13)

[0082] Perform XOR operation on SW_4 and SW_8 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_4,T3,SM3_Tj

[20] ); call the status word generation instruction sequence Gen_W(SW_4,SW_11,SW_1,SW_7,SW_4,SW_14);

[0083] Perform XOR operation on SW_5 and SW_9 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_5,T3,SM3_Tj

[21] ); call the status word generation instruction sequence Gen_W(SW_5,SW_12,SW_2,SW_8,SW_5,SW_15);

[0084] Perform XOR operation on SW_6 and SW_10 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_6,T3,SM3_Tj

[22] );

[0085] Call the status word generation instruction sequence Gen_W (SW_6, SW_13, SW_3, SW_9, SW_6, SW_0)

[0086] Perform XOR operation on SW_7 and SW_11 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_7,T3,SM3_Tj

[23] ); call the status word generation instruction sequence Gen_W(SW_7,SW_14,SW_4,SW_10,SW_7,SW_1);

[0087] Perform XOR operation on SW_8 and SW_12 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_8,T3,SM3_Tj

[24] ); call the status word generation instruction sequence Gen_W(SW_8,SW_15,SW_5,SW_11,SW_8,SW_2);

[0088] Perform XOR operation on SW_9 and SW_13 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_9,T3,SM3_Tj

[25] ); call the status word generation instruction sequence Gen_W(SW_9,SW_0,SW_6,SW_12,SW_9,SW_3);

[0089] Perform XOR operation on SW_10 and SW_14 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_10, T3, SM3_Tj

[26] ); call the status word generation instruction sequence Gen_W(SW_10, SW_1, SW_7, SW_13, SW_10, SW_4);

[0090] Perform XOR operation on SW_11 and SW_15 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_11,T3,SM3_Tj

[27] ); call the status word generation instruction sequence Gen_W(SW_11,SW_2,SW_8,SW_14,SW_11,SW_5);

[0091] Perform XOR operation on SW_12 and SW_0 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_12,T3,SM3_Tj

[28] ); call the status word generation instruction sequence Gen_W(SW_12,SW_3,SW_9,SW_15,SW_12,SW_6);

[0092] Perform XOR operation on SW_13 and SW_1 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_13,T3,SM3_Tj

[29] ); call the status word generation instruction sequence Gen_W(SW_13,SW_4,SW_10,SW_0,SW_13,SW_7);

[0093] Perform XOR operation on SW_14 and SW_2 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_14, T3, SM3_Tj

[30] ); call the status word generation instruction sequence Gen_W(SW_14, SW_5, SW_11, SW_1, SW_14, SW_8);

[0094] Perform XOR operation on SW_15 and SW_3 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_15,T3,SM3_Tj

[31] ); call the status word generation instruction sequence Gen_W(SW_15,SW_6,SW_12,SW_2,SW_15,SW_9);

[0095] Perform XOR operation on SW_0 and SW_4 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_0,T3,SM3_Tj

[32] ); call the status word generation instruction sequence Gen_W(SW_0,SW_7,SW_13,SW_3,SW_0,SW_10);

[0096] Perform XOR operation on SW_1 and SW_5 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_1,T3,SM3_Tj

[33] ); call the status word generation instruction sequence Gen_W(SW_1,SW_8,SW_14,SW_4,SW_1,SW_11);

[0097] Perform XOR operation on SW_2 and SW_6 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_2, T3, SM3_Tj

[34] ); call the status word generation instruction sequence Gen_W(SW_2, SW_9, SW_15, SW_5, SW_2, SW_12);

[0098] Perform XOR operation on SW_3 and SW_7 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_3,T3,SM3_Tj

[35] ); call the status word generation instruction sequence Gen_W(SW_3,SW_10,SW_0,SW_6,SW_3,SW_13);

[0099] Perform XOR operation on SW_4 and SW_8 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_4,T3,SM3_Tj

[36] ); call the status word generation instruction sequence Gen_W(SW_4,SW_11,SW_1,SW_7,SW_4,SW_14);

[0100] Perform XOR operation on SW_5 and SW_9 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_5,T3,SM3_Tj

[37] ); call the status word generation instruction sequence Gen_W(SW_5,SW_12,SW_2,SW_8,SW_5,SW_15);

[0101] Perform XOR operation on SW_6 and SW_10 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_6, T3, SM3_Tj

[38] ); call the status word generation instruction sequence Gen_W(SW_6, SW_13, SW_3, SW_9, SW_6, SW_0);

[0102] Perform XOR operation on SW_7 and SW_11 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_7,T3,SM3_Tj

[39] ); call the status word generation instruction sequence Gen_W(SW_7,SW_14,SW_4,SW_10,SW_7,SW_1);

[0103] Perform XOR operation on SW_8 and SW_12 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_8,T3,SM3_Tj

[40] ); call the status word generation instruction sequence Gen_W(SW_8,SW_15,SW_5,SW_11,SW_8,SW_2);

[0104] Perform XOR operation on SW_9 and SW_13 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_9,T3,SM3_Tj

[41] ); call the status word generation instruction sequence Gen_W(SW_9,SW_0,SW_6,SW_12,SW_9,SW_3);

[0105] Perform XOR operation on SW_10 and SW_14 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_10, T3, SM3_Tj

[42] ); call the status word generation instruction sequence Gen_W(SW_10, SW_1, SW_7, SW_13, SW_10, SW_4);

[0106] Perform XOR operation on SW_11 and SW_15 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_11,T3,SM3_Tj

[43] ); call the status word generation instruction sequence Gen_W(SW_11,SW_2,SW_8,SW_14,SW_11,SW_5);

[0107] Perform an XOR operation on SW_12 and SW_0 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_12,T3,SM3_Tj

[44] ); call the status word generation instruction sequence Gen_W(SW_12,SW_3,SW_9,SW_15,SW_12,SW_6);

[0108] Perform XOR operation on SW_13 and SW_1 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_13,T3,SM3_Tj

[45] ); call the status word generation instruction sequence Gen_W(SW_13,SW_4,SW_10,SW_0,SW_13,SW_7);

[0109] Perform XOR operation on SW_14 and SW_2 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_14, T3, SM3_Tj

[46] ); call the status word generation instruction sequence Gen_W(SW_14, SW_5, SW_11, SW_1, SW_14, SW_8);

[0110] Perform XOR operation on SW_15 and SW_3 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_15,T3,SM3_Tj

[47] ); call the status word generation instruction sequence Gen_W(SW_15,SW_6,SW_12,SW_2,SW_15,SW_9);

[0111] Perform XOR operation on SW_0 and SW_4 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_0,T3,SM3_Tj

[48] ); call the status word generation instruction sequence Gen_W(SW_0,SW_7,SW_13,SW_3,SW_0,SW_10);

[0112] Perform XOR operation on SW_1 and SW_5 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_1,T3,SM3_Tj

[49] ); call the status word generation instruction sequence Gen_W(SW_1,SW_8,SW_14,SW_4,SW_1,SW_11);

[0113] Perform XOR operation on SW_2 and SW_6 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_2, T3, SM3_Tj

[50] ); call the status word generation instruction sequence Gen_W(SW_2, SW_9, SW_15, SW_5, SW_2, SW_12);

[0114] Perform XOR operation on SW_3 and SW_7 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_3,T3,SM3_Tj

[51] ); call the status word generation instruction sequence Gen_W(SW_3,SW_10,SW_0,SW_6,SW_3,SW_13);

[0115] Perform an XOR operation on SW_4 and SW_8 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_4,T3,SM3_Tj

[52] );

[0116] Perform an XOR operation on SW_5 and SW_9 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_5,T3,SM3_Tj

[53] );

[0117] Perform an XOR operation on SW_6 and SW_10 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_6,T3,SM3_Tj

[54] );

[0118] Perform an XOR operation on SW_7 and SW_11 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_7,T3,SM3_Tj

[55] );

[0119] Perform an XOR operation on SW_8 and SW_12 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_8,T3,SM3_Tj

[56] );

[0120] Perform an XOR operation on SW_9 and SW_13 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_9,T3,SM3_Tj

[57] );

[0121] Perform an XOR operation on SW_10 and SW_14 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_10,T3,SM3_Tj

[58] );

[0122] Perform an XOR operation on SW_11 and SW_15 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_11,T3,SM3_Tj

[59] );

[0123] Perform an XOR operation on SW_12 and SW_0 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_12,T3,SM3_Tj

[60] );

[0124] Perform an XOR operation on SW_13 and SW_1 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_13,T3,SM3_Tj

[61] );

[0125] Perform an XOR operation on SW_14 and SW_2 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_14,T3,SM3_Tj

[62] );

[0126] Perform an XOR operation on SW_15 and SW_3 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_15,T3,SM3_Tj

[63] );

[0127] Perform an XOR operation on A and SW[2] to obtain A; perform an XOR operation on B and SW[3] to obtain B; perform an XOR operation on C and SW[4] to obtain C; perform an XOR operation on D and SW[5] to obtain D; perform an XOR operation on E and SW[6] to obtain E; perform an XOR operation on F and SW[7] to obtain F; perform an XOR operation on G and SW[8] to obtain G; perform an XOR operation on H and SW[9] to obtain H; SW[6] represents the sixth element of array SW, SW[8] represents the eighth element of array SW, and SW[9] represents the ninth element of array SW;

[0128] (2): Group A, B, C, D, E, F, G, and H into 32-bit blocks, each containing n 32-bit data blocks.

[0129] (3): Let sequence number j be executed from 0 to 7 to extract the jth 32-bit data block from A, B, C, D, E, F, G, H, and output it to the position m0 starting at the w+j*32th byte of K, grouped into 32 bits, each group + n; where K represents the key data bit string of length klen;

[0130] (4): Add w to the value of i. If i is less than klen / v, go to step (1); otherwise, go to step (5);

[0131] Wherein, klen represents the bit length of the key data to be obtained;

[0132] v represents the cryptographic hash function H v (·) The output is a hash value of length v bits;

[0133] (5): SW[0] = m0.

[0134] Furthermore, the p0(x) is specifically implemented as follows: performing a 9-bit cyclic left shift on x in 32-bit groups, performing a 17-bit cyclic left shift on x in 32-bit groups, performing an XOR operation on the result, and outputting the result to x.

[0135] Furthermore, the p1(x) is specifically implemented as follows: performing a 15-bit cyclic left shift on x in 32-bit groups, performing a 23-bit cyclic left shift on x in 32-bit groups, performing an XOR operation on the result, and outputting the result to x.

[0136] Furthermore, the Ff(A, B, C) = A XOR B XOR C. Furthermore, the FF(A, B, C) = (A and B) or (A and C) or (B and C).

[0137] Furthermore, GG(A, B, C)=C XOR(A&(B XOR C)).

[0138] Furthermore, the instructions for the status word generation instruction sequence Gen_W (SW_0, SW_7, SW_13, SW_3, SW_out, SW_10) are: SW_13 is grouped into 32 bits and circularly shifted left by 15 bits, and then XORed with SW_7 and SW_0, and then output to t0; SW_3 is grouped into 32 bits and circularly shifted left by 7 bits, and then output to t1; t0 is grouped into t1 and p1(x) is executed, and then XORed with SW_10 and t1, and then output to SW_out.

[0139] Furthermore, OneRoundA(A,E,B,F,C,G,D,H,w0smd,w4smd,Ti) is implemented as follows:

[0140] A is grouped into 32-bit groups and circularly shifted left by 12 bits before being output to the extended message word sm3ff. A, E, B, F, C, G, D, H, w0smd, w4smd, and Ti all represent formal parameters. During execution, the names of the corresponding positions will be replaced by the actual corresponding registers.

[0141] After executing Ff(E,F,G), perform unsigned addition on the output of the Ff(E,F,G) function, w0smd, and H in 32-bit groups, and output the result to H; Ff(E,F,G) = E XOR F XOR G;

[0142] Tismd is filled with Ti, Ti is grouped with sm3ff and E to perform unsigned addition operation, and then a circular left shift of 7 bits is performed and output to T2;

[0143] After performing an unsigned addition operation on H and T2 in 32-bit groups, the p0(x) function is executed and output to H; the p0(x) is specifically implemented as follows: performing a 9-bit cyclic left shift on x in 32-bit groups, performing a 17-bit cyclic left shift on x in 32-bit groups, and performing an XOR operation on x, and then outputting to x;

[0144] After performing XOR operation on T2 and sm3ff, the output is sent to T2;

[0145] Perform unsigned addition on D and the extended message word w4smd in 32-bit groups and output to D;

[0146] After executing Ff(A,B,C), perform unsigned addition operation on the output result of Ff(A,B,C) function and T2 in 32-bit groups and output the result to T2.

[0147] Perform unsigned addition on D and T2 in 32-bit groups and output to D;

[0148] Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F;

[0149] Perform a 9-bit cyclic left shift on B in 32-bit groups and output it to B.

[0150] Furthermore, OneRoundB(A,E,B,F,C,G,D,H,Ti) is implemented as follows:

[0151] A is grouped into 32-bit groups and then cyclically shifted left by 12 bits and output to sm3ff;

[0152] After executing Ff(E,F,G), perform unsigned addition operation on the output result of Ff(E,F,G) and H in 32-bit groups and output them to H;

[0153] Fill Tismd with Ti, perform unsigned addition on Ti, sm3ff, and E in 32-bit groups, perform a circular left shift by 7 bits, and output to T2;

[0154] After performing unsigned addition on H and T2 in 32-bit groups, execute the p0(x) function and output it to H;

[0155] After performing XOR operation on T2 and sm3ff, the output is sent to T2;

[0156] After executing Ff(A,B,C), perform unsigned addition operation on the output value of Ff(A,B,C) and T2 in 32-bit groups and output the result to T2.

[0157] Perform unsigned addition on D and T2 in 32-bit groups and output to D;

[0158] Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F;

[0159] Perform a 9-bit cyclic left shift on B in 32-bit groups and output it to B.

[0160] Furthermore, OneRounds(A,E,B,F,C,G,D,H,w4smd,Ti) is implemented as follows:

[0161] A is grouped into 32-bit groups and then circularly shifted left by 12 bits and output to register sm3ff;

[0162] After executing Ff(E,F,G), perform unsigned addition on the output value of the Ff(E,F,G) function and H in 32-bit groups and output the result to H; Ff(E,F,G) = E XOR F XOR G;

[0163] Fill Tismd with Ti, perform unsigned addition on Ti, sm3ff, and E in 32-bit groups, perform a circular left shift by 7 bits, and output to T2;

[0164] Perform unsigned addition on H and T2 in 32-bit groups, then execute the p0(x) function and output to H.

[0165] Perform XOR operation on T2 and sm3ff and output to T2;

[0166] Perform unsigned addition on D and w4simd in 32-bit groups and output to D.

[0167] After executing Ff(A,B,C), perform unsigned addition operation on the output value of Ff(A,B,C) and T2 in 32-bit groups and output them to T2;

[0168] Perform unsigned addition on D and T2 in 32-bit groups and output to D;

[0169] Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F;

[0170] Perform a 9-bit cyclic left shift on B in 32-bit groups and output it to B.

[0171] Furthermore, OneRoundC(A,E,B,F,C,G,D,H,w0smd,w4smd,Ti) is implemented as follows:

[0172] A is grouped into 32-bit groups and then cyclically shifted left by 12 bits and output to sm3ff;

[0173] After executing GG(E, F, G), the output value of GG(E, F, G) is subjected to unsigned addition operation with w0smd and H in 32-bit groups and output to H; GG(E, F, G) = G XOR(E&(F XOR G));

[0174] Tismd is filled with Ti, Ti is grouped with sm3ff and E to perform unsigned addition operation, and then it is rotated left by 7 bits and output to T2;

[0175] Perform unsigned addition on H and T2 in 32-bit groups, then execute the p0(x) function of the SM3 algorithm and output it to H; the p0(x) is specifically implemented as follows: perform a 9-bit cyclic left shift on x in 32-bit groups, perform a 17-bit cyclic left shift on x in 32-bit groups, and then XOR the result with x and output it to x;

[0176] After performing XOR operation on T2 and sm3ff, the output is sent to T2;

[0177] Perform unsigned addition on D and w4simd in 32-bit groups and output to D.

[0178] After executing FF(A,B,C), perform unsigned addition on the output values of FF(A,B,C) and T2 in 32-bit groups and output them to T2; FF(A,B,C) = (A and B) or (A and C) or (B and C);

[0179] Perform unsigned addition on D and T2 in 32-bit groups and output to D;

[0180] For F, perform a 19-bit cyclic left shift on the 32-bit group and output it to F;

[0181] For B, perform a 9-bit circular left shift in 32-bit groups and output it to B.

[0182] Example 2: This embodiment provides a national secret algorithm encryption and decryption system based on a fast implementation of a key derivation function, including: a module for encrypting messages and a module for decrypting ciphertext; the key derivation function in the module for encrypting messages and the module for decrypting ciphertext is accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; then, batch processing the remaining 32 bits to obtain the calculation result of the key derivation function. The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention can be modified and varied in various ways. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. The encryption and decryption method of the national secret algorithm based on the key derivation function is characterized by: include: The steps to encrypt the message and the steps to decrypt the ciphertext; The key derivation function in the step of encrypting the message and the step of decrypting the ciphertext is accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; Then, the remaining 32 bits are processed in batches to obtain the calculation result of the key derivation function.

2. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 1 is characterized in that: Before the key derivation function in the steps of encrypting the message and decrypting the ciphertext is accelerated, the method further includes: precalculating and saving a parameter table: SM3_Tj[64]={0x79cc4519,0xf3988a32,0xe7311465,0xce6228cb,0x9cc45197,0x3988a32f,0x7311465e, 0xe6228cbc,0xcc451979,0x988a32f3,0x311465e7,0x6228cbce,0xc451979c,0x88a32f39,0x11465e73,0x 228cbce6,0x9d8a7a87,0x3b14f50f,0x7629ea1e,0xec53d43c,0xd8a7a879,0xb14f50f3,0x629ea1e7,0xc 53d43ce,0x8a7a879d,0x14f50f3b,0x29ea1e76,0x53d43cec,0xa7a879d8,0x4f50f3b1,0x9ea1e762,0x3d4 3cec5,0x7a879d8a,0xf50f3b14,0xea1e7629,0xd43cec53,0xa879d8a7,0x50f3b14f,0xa1e7629e,0x43ce c53d,0x879d8a7a,0x0f3b14f5,0x1e7629ea,0x3cec53d4,0x79d8a7a8,0xf3b14f50,0xe7629ea1,0xcec53d 43,0x9d8a7a87,0x3b14f50f,0x7629ea1e,0xec53d43c,0xd8a7a879,0xb14f50f3,0x629ea1e7,0xc53d43ce ,0x8a7a879d,0x14f50f3b,0x29ea1e76,0x53d43cec,0xa7a879d8,0x4f50f3b1,0x9ea1e762,0x3d43cec5}.

3. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 1 is characterized in that: Accelerating the execution of key derivation functions in the steps of encrypting a message and decrypting a ciphertext, including: Assume that the width of the SIMD register is w bits, and the step size is n=w / 32. Each register can be divided into n 32-bit blocks. Set the contents of each block in the SIMD register m0 to 1, 2, ...n-1, H=SM3Update(Z), and split H into 8 32-bit data blocks B0, B1....B7. Fill the SIMD array SW[8] with index from 0 to 7, SW[index] is filled with Bindex, i is equal to 0; i represents the position of the currently processed data; index represents the number; (1): A=SW[0];B=SW[1];C=SW[2];D=SW[3];E=SW[4];F=SW[5];G=SW[6];H=SW[7];where A, B, C, D, E, F, G, H represent 8 SIMD registers; Execute the p1(x) function on register m0 to obtain array SW_0; fill SW_1 with 0x80404000; fill SW_2 with 0x01108888; fill SW_4 with 0x10005000; fill SW_5 with 0x0022028a; fill SW_7 with 0xac545c04; and fill SW_8 with 0x99cda89a; the p1(x) function is specifically implemented as follows: perform a 15-bit circular left shift on x in 32-bit groups, perform a 23-bit circular left shift on x in 32-bit groups, perform an XOR operation on x, and output the result to x; Execute a 15-bit circular left shift on SW_0 in 32-bit groups, and then execute the p1(x) function to obtain SW_3; execute a 15-bit circular left shift on SW_3 in 32-bit groups to obtain t0; execute the p1(x) function on t0, and then perform an XOR operation with SW_0 to obtain SW_6; execute a 15-bit circular left shift on SW_6 in 32-bit groups, and then perform an XOR operation with SW_0 to obtain t0; execute p1(x) on t0, and then perform an XOR operation with SW_3 to obtain SW_9; fill SW_10 with 0xa0003000; fill SW_11 with 0x002202a8; fill SW_12 with 0x11000; fill SW_13 with 0xa4505804; SW_0 to SW_15 are all SIMD registers; Execute a cyclic left shift of 15 bits on SW_9 in 32-bit groups, and then perform an XOR operation with SW_3 to obtain t0. Execute p1(x) on t0, and then perform an XOR operation with SW_3 and SW_12 to obtain SW_12. Shift SW_0 7 bits to the left in 32-bit groups, then perform an XOR operation with SW_13 to obtain SW_13; fill SW_14 with 0xf45691fb; fill T4 with 0x00000220; Calling status word generation instruction sequence and calling function instruction sequence; Perform an XOR operation on A and SW[2] to obtain A; perform an XOR operation on B and SW[3] to obtain B; perform an XOR operation on C and SW[4] to obtain C; perform an XOR operation on D and SW[5] to obtain D; perform an XOR operation on E and SW[6] to obtain E; perform an XOR operation on F and SW[7] to obtain F; perform an XOR operation on G and SW[8] to obtain G; perform an XOR operation on H and SW[9] to obtain H; SW[6] represents the sixth element of the SIMD array SW, SW[8] represents the eighth element of the SIMD array SW, and SW[9] represents the ninth element of the SIMD array SW; (2): Group A, B, C, D, E, F, G, and H into 32-bit blocks, each containing n 32-bit data blocks. (3): Let sequence number j be executed from 0 to 7, extract the jth 32-bit data block from A, B, C, D, E, F, G, H, and output it to the position m0 starting from the w+j*32th byte of K, grouped by 32 bits, each group + n; Where K represents a key data bit string of length klen; (4): Add w to the value of i. If i is less than klen / v, go to step (1); otherwise, go to step (5); Wherein, klen represents the bit length of the key data to be obtained; v represents the cryptographic hash function H v (·) The output is a hash value of length v bits; (5): SW[0] = m0.

4. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 3 is characterized in that: The calling status word generation instruction sequence and the calling function instruction sequence include: Call the status word generation instruction sequence Gen_W(T4,SW_6,SW_12,SW_2,SW_15,SW_9); t0=T4; Call the first round function instruction sequence OneRoundA(A,E,B,F,C,G,D,H,m0,m0,SM3_Tj[0]); fill t1 with 0x80000000; Call the first round function instruction sequence OneRoundA(D,H,A,E,B,F,C,G,t1,t1,SM3_Tj[1]); Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj[2]); Call the second round function instruction sequence OneRoundB(B,F,C,G,D,H,A,E,SM3_Tj[3]); Call the second round function instruction sequence OneRoundB(A,E,B,F,C,G,D,H,SM3_Tj[4]); Call the second round function instruction sequence OneRoundB(D,H,A,E,B,F,C,G,SM3_Tj[5]); Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj[6]); Call the second round function instruction sequence OneRoundB(B,F,C,G,D,H,A,E,SM3_Tj[7]); Call the second round function instruction sequence OneRoundB(A,E,B,F,C,G,D,H,SM3_Tj[8]); Call the second round function instruction sequence OneRoundB(D,H,A,E,B,F,C,G,SM3_Tj[9]); Call the second round function instruction sequence OneRoundB(C,G,D,H,A,E,B,F,SM3_Tj[10]); Call the third round function instruction sequence OneRounds(B,F,C,G,D,H,A,E,t0,SM3_Tj[11]); Call the third round function instruction sequence OneRounds(A,E,B,F,C,G,D,H,SW_0,SM3_Tj[12]); Call the third round function instruction sequence OneRounds(D,H,A,E,B,F,C,G,SW_1,SM3_Tj[13]); Call the third round function instruction sequence OneRounds(C,G,D,H,A,E,B,F,SW_2,SM3_Tj[14]); Perform an XOR operation on t0 and SW_3 to obtain T3.

5. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 4 is characterized in that: After performing an XOR operation on t0 and SW_3 to obtain T3, it also includes: Call the first round function instruction sequence OneRoundA(B,F,C,G,D,H,A,E,t0,T3,SM3_Tj[15]); Perform an XOR operation on SW_0 and SW_4 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_0,T3,SM3_Tj[16]); Call the status word generation instruction sequence Gen_W(SW_0,SW_7,SW_13,SW_3,SW_0,SW_10); Perform an XOR operation on SW_1 and SW_5 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_1,T3,SM3_Tj[17]); call the status word generation instruction sequence Gen_W(SW_1,SW_8,SW_14,SW_4,SW_1,SW_11); Perform XOR operation on SW_2 and SW_6 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_2,T3,SM3_Tj[18]); Call the status word generation instruction sequence Gen_W(SW_2,SW_9,SW_15,SW_5,SW_2,SW_12) Perform XOR operation on SW_3 and SW_7 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_3,T3,SM3_Tj[19]); Call the status word generation instruction sequence Gen_W(SW_3,SW_10,SW_0,SW_6,SW_3,SW_13) to perform an XOR operation on SW_4 and SW_8 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_4,T3,SM3_Tj[20])); call the status word generation instruction sequence Gen_W(SW_4,SW_11,SW_1,SW_7,SW_4,SW_14); Perform XOR operation on SW_5 and SW_9 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_5,T3,SM3_Tj[21]); call the status word generation instruction sequence Gen_W(SW_5,SW_12,SW_2,SW_8,SW_5,SW_15); Perform XOR operation on SW_6 and SW_10 to obtain T3; call the fourth round function instruction sequence OneRoundC(C,G,D,H,A,E,B,F,SW_6,T3,SM3_Tj[22]); Call the status word generation instruction sequence Gen_W (SW_6, SW_13, SW_3, SW_9, SW_6, SW_0) Perform XOR operation on SW_7 and SW_11 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_7,T3,SM3_Tj[23]); call the status word generation instruction sequence Gen_W(SW_7,SW_14,SW_4,SW_10,SW_7,SW_1).

6. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 5 is characterized in that: After calling the status word generation instruction sequence Gen_W (SW_7, SW_14, SW_4, SW_10, SW_7, SW_1), it also includes: Perform XOR operation on SW_8 and SW_12 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_8,T3,SM3_Tj[24]); call the status word generation instruction sequence Gen_W(SW_8,SW_15,SW_5,SW_11,SW_8,SW_2); Perform XOR operation on SW_9 and SW_13 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_9,T3,SM3_Tj[25]); call the status word generation instruction sequence Gen_W(SW_9,SW_0,SW_6,SW_12,SW_9,SW_3); Perform XOR operation on SW_10 and SW_14 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_10, T3, SM3_Tj[26]); call the status word generation instruction sequence Gen_W(SW_10, SW_1, SW_7, SW_13, SW_10, SW_4); Perform XOR operation on SW_11 and SW_15 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_11,T3,SM3_Tj[27]); call the status word generation instruction sequence Gen_W(SW_11,SW_2,SW_8,SW_14,SW_11,SW_5); Perform XOR operation on SW_12 and SW_0 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_12,T3,SM3_Tj[28]); call the status word generation instruction sequence Gen_W(SW_12,SW_3,SW_9,SW_15,SW_12,SW_6); Perform XOR operation on SW_13 and SW_1 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_13,T3,SM3_Tj[29]); call the status word generation instruction sequence Gen_W(SW_13,SW_4,SW_10,SW_0,SW_13,SW_7); Perform XOR operation on SW_14 and SW_2 to obtain T3; call the fourth round function instruction sequence OneRoundC(C, G, D, H, A, E, B, F, SW_14, T3, SM3_Tj[30]); call the status word generation instruction sequence Gen_W(SW_14, SW_5, SW_11, SW_1, SW_14, SW_8); Perform XOR operation on SW_15 and SW_3 to obtain T3; call the fourth round function instruction sequence OneRoundC(B,F,C,G,D,H,A,E,SW_15,T3,SM3_Tj[31]); call the status word generation instruction sequence Gen_W(SW_15,SW_6,SW_12,SW_2,SW_15,SW_9); Perform XOR operation on SW_0 and SW_4 to obtain T3; call the fourth round function instruction sequence OneRoundC(A,E,B,F,C,G,D,H,SW_0,T3,SM3_Tj[32]); call the status word generation instruction sequence Gen_W(SW_0,SW_7,SW_13,SW_3,SW_0,SW_10); Perform XOR operation on SW_1 and SW_5 to obtain T3; call the fourth round function instruction sequence OneRoundC(D,H,A,E,B,F,C,G,SW_1,T3,SM3_Tj[33]); call the status word generation instruction sequence Gen_W(SW_1,SW_8,SW_14,SW_4,SW_1,SW_11).

7. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 4 is characterized in that: OneRoundA(A,E,B,F,C,G,D,H,w0smd,w4smd,Ti) is implemented as follows: A is grouped into 32-bit groups and circularly shifted left by 12 bits before being output to the extended message word sm3ff. A, E, B, F, C, G, D, H, w0smd, w4smd, and Ti all represent formal parameters. During execution, the names of the corresponding positions will be replaced by the actual corresponding registers. After executing Ff(E,F,G), perform unsigned addition on the output of the Ff(E,F,G) function, w0smd, and H in 32-bit groups, and output the result to H; Ff(E,F,G) = E XOR F XOR G; Tismd is filled with Ti, Ti is grouped with sm3ff and E to perform unsigned addition operation, and then a circular left shift of 7 bits is performed and output to T2; After performing an unsigned addition operation on H and T2 in 32-bit groups, the p0(x) function is executed and output to H; the p0(x) is specifically implemented as follows: performing a 9-bit cyclic left shift on x in 32-bit groups, performing a 17-bit cyclic left shift on x in 32-bit groups, and performing an XOR operation on x, and then outputting to x; After performing XOR operation on T2 and sm3ff, the output is sent to T2; Perform unsigned addition on D and the extended message word w4smd in 32-bit groups and output to D; After executing Ff(A,B,C), perform unsigned addition operation on the output result of Ff(A,B,C) function and T2 in 32-bit groups and output the result to T2. Perform unsigned addition on D and T2 in 32-bit groups and output to D; Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F; Perform a 9-bit cyclic left shift on B according to the 32-bit group and output it to B; OneRoundB(A,E,B,F,C,G,D,H,Ti) is implemented as follows: A is grouped into 32-bit groups and then cyclically shifted left by 12 bits and output to sm3ff; After executing Ff(E,F,G), perform unsigned addition operation on the output result of Ff(E,F,G) and H in 32-bit groups and output them to H; Fill Tismd with Ti, perform unsigned addition on Ti, sm3ff, and E in 32-bit groups, perform a circular left shift by 7 bits, and output to T2; Perform unsigned addition on H and T2 in 32-bit groups, then execute the p0(x) function and output to H. After performing XOR operation on T2 and sm3ff, the output is sent to T2; After executing Ff(A,B,C), perform unsigned addition operation on the output value of Ff(A,B,C) and T2 in 32-bit groups and output the result to T2. Perform unsigned addition on D and T2 in 32-bit groups and output to D; Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F; Perform a 9-bit cyclic left shift on B in 32-bit groups and output it to B.

8. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 4 is characterized in that: OneRounds(A,E,B,F,C,G,D,H,w4smd,Ti) is implemented as follows: A is grouped into 32-bit groups and then circularly shifted left by 12 bits and output to register sm3ff; After executing Ff(E,F,G), perform unsigned addition on the output value of the Ff(E,F,G) function and H in 32-bit groups and output the result to H. Fill Tismd with Ti, perform unsigned addition on Ti, sm3ff, and E in 32-bit groups, perform a circular left shift by 7 bits, and output to T2; Perform unsigned addition on H and T2 in 32-bit groups, then execute the p0(x) function and output to H. Perform XOR operation on T2 and sm3ff and output to T2; Perform unsigned addition on D and w4simd in 32-bit groups and output to D. After executing Ff(A,B,C), perform unsigned addition operation on the output value of Ff(A,B,C) and T2 in 32-bit groups and output them to T2; Perform unsigned addition on D and T2 in 32-bit groups and output to D; Perform a 19-bit cyclic left shift on F in 32-bit groups and output to F; Perform a 9-bit cyclic left shift on B in 32-bit groups and output it to B.

9. The encryption and decryption method of the national secret algorithm based on the rapid implementation of the key derivation function as claimed in claim 5 is characterized in that: OneRoundC(A,E,B,F,C,G,D,H,w0smd,w4smd,Ti) is implemented as follows: A is grouped into 32-bit groups and then cyclically shifted left by 12 bits and output to sm3ff; After executing GG(E, F, G), the output value of GG(E, F, G) is subjected to unsigned addition operation with w0smd and H in 32-bit groups and output to H; GG(E, F, G) = G XOR(E&(F XOR G)); Tismd is filled with Ti, and Ti is grouped with sm3ff and E to perform unsigned addition operation, and then a circular left shift of 7 bits is performed and output to T2; Perform unsigned addition on H and T2 in 32-bit groups, then execute the p0(x) function of the SM3 algorithm and output it to H; the p0(x) is specifically implemented as follows: perform a 9-bit cyclic left shift on x in 32-bit groups, perform a 17-bit cyclic left shift on x in 32-bit groups, and then XOR the result with x and output it to x; After performing XOR operation on T2 and sm3ff, the output is sent to T2; Perform unsigned addition on D and w4simd in 32-bit groups and output to D. After executing FF(A,B,C), perform unsigned addition on the output values of FF(A,B,C) and T2 in 32-bit groups and output them to T2; FF(A,B,C) = (A and B) or (A and C) or (B and C); Perform unsigned addition on D and T2 in 32-bit groups and output to D; For F, perform a 19-bit cyclic left shift on the 32-bit group and output it to F; For B, perform a 9-bit circular left shift in 32-bit groups and output it to B.

10. The encryption and decryption system of the national secret algorithm based on the key derivation function is characterized by: include: A module for encrypting messages and a module for decrypting ciphertext; The key derivation functions in the message encryption module and the ciphertext decryption module are accelerated, including: first, performing message expansion and iterative compression of the standard SM3 hash algorithm on the first 512 bits Z to obtain a 256-bit hash value H; Then, the remaining 32 bits are processed in batches to obtain the calculation result of the key derivation function.