Method and system compatible with AES and SM4 block cipher algorithms

By converting the unified round of operation of SM4 and AES into global pipelined operations and optimizing the timing isolation segment series, the problem of poor performance compatibility with AES and SM4 in the prior art is solved, and a unified round of operation with high performance and low resource overhead is achieved.

CN114374507BActive Publication Date: 2025-05-16SHENZHEN STATE MICRO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210028032.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-11
Publication Date
2025-05-16
Estimated Expiration
2042-01-11

AI Technical Summary

Technical Problem

The existing methods and system compatible with two packet cipher algorithms, AES and SM4, have poor performance, cannot meet the throughput requirements, and have a large resource overhead.

Method used

By uniformly converting the unified operation process of SM4 and AES into global pipelined operation processes of front linear transformation, S-box replacement transformation and post-linear transformation, and optimizing processing performance and resource overhead by designing timing isolation segment series.

Benefits of technology

It realizes high-performance unified round of operations compatible with two packet cipher algorithms AES and SM4, which reduces resource overhead and improves processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114374507B_ABST
    Figure CN114374507B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system compatible with two block cipher algorithms, AES and SM4. The method compatible with two block cipher algorithms, AES and SM4, comprises: converting the operation process of a single general iterative round in the unified round operation of SM4 and the unified round operation of AES from the first round to the eighth round into a global pipeline operation process of a front linear transformation, an S-box substitution transformation and a back linear transformation; the front linear transformation comprises at least one level of timing isolation segment, and any level of the timing isolation segment is composed of at least one XOR gate or is a linear transformation timing isolation segment for realizing permutation processing. The present invention can achieve a balance between high performance and low resource consumption compatible with two block cipher algorithms, AES and SM4.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of block encryption algorithms, and in particular to a method and system compatible with two block encryption algorithms, AES and SM4. Background Art

[0002] The SM4 symmetric encryption algorithm, developed by the China Information Security Standardization Technical Committee, is a commercial block cipher algorithm. It was released as a cryptographic industry standard in 2012, converted into a Chinese national standard in 2016, and became an ISO / IEC international standard in June 2021. This marks the continuous improvement of my country's commercial encryption technology level and international standardization capabilities.

[0003] The AES symmetric encryption algorithm, developed by the National Institute of Standards and Technology (NIST) of the United States, was officially released in November 2001 and has long become the de facto standard adopted by many applications based on symmetric key encryption internationally.

[0004] In recent years, with the advancement of SM4 standardization and its widespread adoption on many application platforms, there have been frequent application requirements for compatible support of these two mature algorithms on the same platform. In particular, when formulating the high-speed interface transmission protection system standard for China's ultra-high-definition television, there has been an application requirement for compatible encryption / decryption of program code streams in SM4 or AES standards for 8K definition image frame data.

[0005] However, the performance of existing methods and systems that are compatible with both SM4 and AES algorithms is relatively poor, and cannot meet the corresponding throughput requirements and save resource overhead at the same time. Summary of the invention

[0006] In order to solve the technical problem of poor performance of methods and systems compatible with both SM4 and AES algorithms in the prior art, the present invention proposes a method and system compatible with both AES and SM4 block cipher algorithms.

[0007] The method proposed by the present invention is compatible with two block cipher algorithms, AES and SM4, comprising: uniformly converting the operation process of a single general iterative round in the first to eighth rounds of the unified round operation of SM4 and the unified round operation of AES into a global pipeline operation process of a front linear transformation, an S-box substitution transformation and a back linear transformation;

[0008] The front linear transformation includes at least one level of timing isolation segment, and the timing isolation segment at any level is composed of at least one XOR gate or is a linear transformation timing isolation segment for implementing replacement processing.

[0009] Further, it also includes: uniformly converting the S-box substitution transformation of AES and SM4 into a local pipeline operation process of front mapping transformation, composite domain inversion transformation, and back mapping transformation, so that the S-box substitution transformation includes multiple levels of timing isolation segments, and the timing isolation segment at any level is composed of at least one XOR gate and / or at least one AND gate. In one embodiment, the XOR gate can be a two-input XOR gate, and the AND gate can be a two-input AND gate;

[0010] The forward mapping transformation converts the elements on the GF(2^8) fields that are different between AES and SM4 in the S-box substitution transformation into the same target composite field that is linearly isomorphic to the respective original GF(2^8) fields.

[0011] Furthermore, the preset standards of processing performance and storage resource overhead are achieved by designing the sum of the number of sequential isolation segments of the front linear transformation required by SM4 in a single general iteration round and the number of sequential isolation segments of the S-box substitution transformation.

[0012] Furthermore, the single iterative round operation process from the 9th round to the last round in the unified round operation of the AES is also uniformly converted into a global pipeline operation process of the front linear transformation, S-box substitution transformation and the back linear transformation.

[0013] Further, when the algorithm being processed is AES, the first iteration round has only a round key sub-transformation, the front linear transformations of other iteration rounds include a row shift operation of AES, the back linear transformations include an XOR network transformation formed by combining a column obfuscation of AES and a round key operation, and the back linear transformation of the last iteration round does not include a column obfuscation of AES;

[0014] When the processing algorithm is SM4, the front linear transformation includes the round key addition operation of SM4, the rear linear transformation includes the XOR network transformation formed by merging the L transformation and XOR combination operation of SM4, and the FP replacement is performed after the last round of rear linear transformation is completed.

[0015] Further, when the algorithm being processed is an AES encryption algorithm, the post-mapping transformation performs an AES standard affine transformation after performing an AES corresponding GF(2^8) isomorphic inverse mapping; when the algorithm being processed is an AES decryption algorithm, the pre-mapping transformation performs an AES standard inverse affine transformation before performing an AES corresponding GF(2^8) isomorphic mapping;

[0016] When the processing algorithm is SM4, the front mapping transformation performs the first affine transformation of the SM4 standard before converting GF(2^8) into the target composite domain. At the same time, the back mapping transformation of SM4 performs the second affine transformation of the SM4 standard after performing GF(2^8) isomorphic inverse mapping on the target composite domain.

[0017] Furthermore, the target composite domain is a GF((2^4)^2) domain.

[0018] Furthermore, the principle of dividing the timing isolation segments is that the combined path delay of each isolation segment does not exceed the sum of the circuit delays of N levels of basic gate units, and the value of N is determined according to preset standards of processing performance and storage resource overhead.

[0019] The system proposed by the present invention for implementing the method for implementing the two block cipher algorithms compatible with AES and SM4 described in the above technical solution includes multiple iterative round implementation circuits, wherein the multiple iterative round implementation circuits include a general iterative round implementation circuit, and a single general iterative round implementation circuit includes:

[0020] A front linear transformation circuit, which is used to implement the row shift operation of AES and the round key addition operation of SM4;

[0021] S-box substitution transformation circuit, which is used to realize front linear mapping, composite domain inverse transformation, and back mapping transformation;

[0022] The post-linear transformation circuit is used to implement the XOR network transformation (XNWA) formed by combining the column confusion and round key addition operations of AES, and the XOR network transformation (XNWS) formed by combining the L transformation and XOR combination operations of SM4.

[0023] Furthermore, the front-end linear transformation circuit includes at least one XOR gate and a multiplexer.

[0024] Based on the above technical solution, the present invention realizes a high-performance implementation of a unified round operation compatible with and supporting both AES and SM4 block cipher algorithms, which is realized based on a design method of a pipeline mechanism at both global and local levels.

[0025] It is used to achieve high-performance unified round operations that are compatible with both AES and SM4 block cipher algorithms. The S-box substitution transformation logic is implemented by combining composite domain conversion with DACSE algorithm.

[0026] A low-resource-overhead implementation of unified round operations for compatible block cipher algorithms, AES and SM4, is implemented based on the design idea of ​​maximally reusing S-box substitution transformation logic and suboptimally reusing other linear transformation logic.

[0027] The unified S-box search and replacement logic is used to be compatible with the two block cipher algorithms AES and SM4. Its search and replacement operation is implemented by a circuit structure that sequentially includes a front mapping transformation, a composite domain inverse transformation, and a back mapping transformation.

[0028] A unified round operation processing logic (RND) is used to support both AES and SM4 block cipher algorithms. A single general iterative round operation is implemented using a circuit structure that follows the sequence of "front linear transformation, S-box substitution transformation, and rear linear transformation".

[0029] The independent version of round key expansion logic (KEP) used to support two block cipher algorithms, AES and SM4, can be implemented based on local hardware circuits or external hardware circuits for the overall expansion operation. Each round key obtained by the expansion operation will be stored in a non-volatile storage element (commonly known as a D flip-flop) so that a large amount of each round key data required by the round operation in the pipeline stage can be obtained simultaneously and quickly.

[0030] The switching control logic for supporting "autonomous hardware selection of multiple sets of keys for each round" is implemented based on the circuit structure of the multiplexer.

[0031] In order to balance the two implementation indicators of high processing performance and low resource overhead, the SM4 algorithm will weigh the number of integrated storage elements required for multi-level shifting of state variables in its round operation pipeline stage (which is proportional to the hardware resource overhead of this part of the sub-circuit) and the number of timing isolation segments of the related linear transformation logic in the unified round operation (which is proportional to the maximum operating frequency of the entire hardware module) according to the specific application requirements expected by the entire cryptographic system.

[0032] The above-mentioned global-level pipeline mechanism is defined as: the overall processing of the round operation logic of the block cipher algorithm will be implemented using a pipeline design architecture, that is, the round operation processing of each iterative round will be implemented separately using physically independent different groups of circuit logic.

[0033] The above-mentioned local-level pipeline mechanism is defined as: a single iterative round processing of the round operation logic that implements the block cipher algorithm will be implemented using a pipelined design architecture, that is, a single round of round operator transformation processing will be split into multiple segments in terms of timing after relevant trade-offs, and each segment will be implemented separately using a group of circuit logics that should be isolated in timing.

[0034] The S-box substitution transformation logic in the unified round operation that is compatible with the two block cipher algorithms will be implemented using a design method that combines composite domain conversion + DACSE algorithm, and a circuit structure that is "front mapping transformation, composite domain inverse transformation, and back mapping transformation" in sequence.

[0035] The pre-mapping transformation and / or post-mapping transformation will include the conversion part of the composite domain transformation, while the composite domain inverse transformation corresponds to the substitution part of the composite domain. For the pre-mapping transformation and post-mapping transformation, which are essentially linear transformations, the optimization algorithm referred to as DACSE can be further used to analyze and infer the minimum number or shortest delay of the basic gate unit circuits required for the pre-mapping (post-mapping) transformation based on the specific physical implementation environment of the target algorithm circuit (including but not limited to the underlying platform (ASIC or FPGA), the process & technology of the foundry, and the EDA tool implemented by the back-end), so as to make a trade-off between the two implementation indicators of high processing performance and low resource overhead.

[0036] The above SBOX logic (i.e. S-box substitution transformation logic) always exists in the round operations of the two algorithms (Note: except for round #0 of AES), and as the only nonlinear transformation logic in the block cipher algorithm, it is the most critical in terms of both timing performance and resource overhead, so the reuse implementation of the SBOX logic should be considered to the greatest extent. As for other linear transformation logics of the round operations of the algorithm, due to the natural differences in the definitions of the AES and SM4 algorithms themselves, their reuse implementation may not be worth the cost in terms of resource overhead, so they should not be considered to a high degree.

[0037] Rounds #1 to #8 of the AES round operation will implement the multiplexing implementation of the corresponding transformation logic with all rounds (i.e., rounds #0 to #31) of the SM4 round operation.

[0038] There is no need to consider multiplexing the AES round #9 to #N (N=10 / 12 / 14 for AES-128 / 192 / 256) with the SM4 round, that is, its corresponding transformation logic (especially the S-box substitution transformation logic) can be an independent version that only supports the AES algorithm, thus saving the corresponding hardware resource overhead.

[0039] The above-mentioned target composite domain includes but is not limited to GF((2^4)^2).

[0040] The linear transformation of the unified round operation includes: the row shift operation of AES and the round key addition operation of SM4.

[0041] The S-box replacement transformation of the unified round operation is: a byte replacement operation carried out with bytes (8-bits) as the granularity for the whole or part of the state variable.

[0042] The post-linear transformation of the unified round operation includes: an XOR network transformation operation formed by combining the column confusion and round key addition operations of AES, and an XOR network transformation operation formed by combining the L transformation and XOR transformation operations of SM4.

[0043] All linear property transformation operations can be implemented using basic gate unit circuits such as "two-input XOR gate, two-input AND gate" to obtain the weighted measurement of the hardware implementation under the two implementation indicators of high processing performance and low resource overhead as mentioned above.

[0044] The S-box lookup operation of the unified round operation mentioned above will use a multiplexed implementation version that is compatible with various SBOX logics that support AES and SM4. Note: The S-box lookup operation dedicated to the later rounds of AES can be another multiplexed implementation version of the SBOX logic that does not include SM4 (that is, it is just a multiplex of the forward and reverse versions of the AES SBOX logic).

[0045] The post-linear transformation of the unified round operation will implement two independent sets of different XOR network transformation circuits for the AES and SM4 algorithms, and the reasons have been explained in the above description.

[0046] The decryption direction of the AES algorithm round operation will adopt the sub-transformation order reconstruction method well known in the industry to achieve process consistency with the encryption direction. Therefore, these two processing directions no longer have essential differences in hardware implementation process control, and they will no longer be distinguished in this article.

[0047] In application scenarios where high performance processing is a higher priority (note: in this scenario, the update frequency of the initial key is assumed to be very low), the KEP logic responsible for providing the 'round key' data for the round operation will be implemented separately from the round operation logic.

[0048] Generally, the KEP logic can be integrated with the round operation logic and implemented by the current hardware circuit module.

[0049] Optionally, the KEP logic is integrated and implemented by other circuit modules outside the scope of the current hardware circuit module.

[0050] Regardless of the above options, each round key obtained by the expansion operation will be stored using a non-volatile storage element (commonly known as a D flip-flop) so that a large amount of each round key data required by the round operation in the pipeline stage can be obtained simultaneously and quickly.

[0051] Alternatively, under the above-mentioned "generally" implementation premise, the SBOX logic required for the round key expansion process may reuse the corresponding logic circuit in the round operation to save some hardware resource overhead.

[0052] Under the premise of the same block cipher algorithm, in order to support cryptographic processing for application scenarios such as "ciphertext data encrypted based on different initial keys with unpredictable arrival order will be received within a continuous period of time", the current hardware circuit module will implement a storage circuit that can store an appropriate number of sets of round keys, and pre-define the corresponding initial key identifier (KID). Then, in the global hierarchical pipeline of the round operation as described above, the hardware will automatically select the correct one of the round keys in real time based on the above KID identifier and the circuit structure of the multiplexer for use.

[0053] Based on the definition characteristics of the SM4 algorithm itself and the above-mentioned pipeline design mechanism, the present invention requires multi-level pipeline shifting and storage of the corresponding state logic variables of the SM4 round operation in the pipeline stage.

[0054] The number of integration levels of the required non-volatile storage elements (such as D flip-flops) will depend on the sum of the number of timing isolation segments of the two transformation logics of the front linear transformation and S-box substitution transformation of the SM4 round operation. The more the number of timing isolation segments, the simpler the calculation tasks to be completed in each working clock cycle, and the higher the clock frequency at which the entire hardware circuit can run stably, and then the processing performance of the entire hardware circuit will be raised to a higher level; but the cost of such implementation is that the storage resource overhead required for pipeline transfer will become larger due to the larger number of timing isolation segments.

[0055] Therefore, if the entire hardware circuit module does not need to run at its highest clock frequency in actual applications, the designer can significantly reduce the storage resource overhead required for the pipeline shift by reducing the number of timing isolation stages of the above two parts of the transformation logic. Moreover, when the working clock frequency of the entire hardware circuit module is constrained to a lower level, the relevant EDA tools can also select the basic gate unit circuit with a larger delay during physical implementation, which will also bring a certain degree of benefit to the saving of hardware resource overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The present invention is described in detail below with reference to the embodiments and accompanying drawings, wherein:

[0057] Figure 1 The present invention is compatible with the design architecture of a single iterative round supporting unified round operations of both AES and SM4 algorithms.

[0058] Figure 2 The present invention provides a local pipeline architecture design for various SBOX logics for the AES and SM4 algorithms.

[0059] Figure 3The present invention is compatible with the circuit implementation of the entire grouping operation supporting both AES and SM4 algorithms under the global pipeline mechanism.

[0060] Figure 4 A circuit implementation architecture of a basic pipeline logic of a unified round operation under a global pipeline design architecture according to an embodiment of the present invention.

[0061] Figure 5 A circuit implementation of the multiplexed SBOX logic of the unified round operation in a local pipeline design architecture according to an embodiment of the present invention.

[0062] Figure 6 A circuit implementation of non-multiplication inversion logic in multiplexed SBOX logic in an embodiment of the present invention under a local pipeline design architecture.

[0063] Figure 7 The invention provides a circuit implementation of the multiplication inversion logic on the target complex domain of the multiplexed SBOX logic in a local pipeline design architecture according to an embodiment of the present invention.

[0064] Figure 8 The invention provides an overall circuit implementation of a unified round operation under a global pipeline mechanism according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0066] Thus, a feature indicated in this specification will be used to illustrate one of the features of an embodiment of the present invention, rather than implying that each embodiment of the present invention must have the described feature. In addition, it should be noted that this specification describes many features. Although some features can be combined together to illustrate possible system designs, these features can also be used in other combinations that are not explicitly described. Thus, unless otherwise stated, the described combinations are not intended to be limiting.

[0067] For general symmetric key block cipher algorithms, whether encryption or decryption, it is essentially an iterative process for "data variables with a bit width of one block (group) size". This iteratively processed data block variable is usually called a state (also called a "state matrix"). For AES and SM4, both of which have a 16-byte wide state, the hardware circuit implementation of the algorithm, the implementation architecture that does not prioritize performance indicators is usually: physically only one set of sequential logic elements (such as D flip-flops) is integrated to store the state results after each iteration of the round operation, and each sub-transformation of each iteration is implemented by the corresponding combinational logic element. Obviously, this implementation architecture that prioritizes resource overhead indicators cannot meet the implementation indicators of ultra-high throughput, so it is generally considered to use an implementation architecture based on a pipeline mechanism to consider the hardware circuit implementation of the algorithm.

[0068] The present invention implements the method of the present invention compatible with the two block cipher algorithms of AES and SM4 through a pipeline mechanism design framework. The method of the present invention is to convert the unified round operation of SM4 and the single general iterative round operation process from the first round to the eighth round in the unified round operation of AES into a global pipeline operation process of front linear transformation, S-box substitution transformation and back linear transformation. Among them, the front linear transformation includes at least one level of timing isolation segment, and any one level of timing isolation segment is composed of at least one XOR gate or is a linear transformation timing isolation segment for realizing permutation processing.

[0069] The AES algorithm has N+1 iteration rounds, N=10, or N=12, or N=14. Except for the first iteration round which can be recorded as "Round #0", the other iteration rounds can be called general iteration rounds. The general iteration rounds of the AES algorithm except the last general iteration round contain 4 seed transformations. The last general iteration round only contains 3 seed transformations and is a special general iteration round. The SM4 algorithm has 32 iteration rounds, all of which are general iteration rounds. A single general iteration round and other single iteration rounds are both single iteration rounds.

[0070] In a preferred embodiment, the single iterative round operation process of the unified round operation of AES and SM4 can be uniformly converted into a global pipeline operation process of a front linear transformation, an S-box substitution transformation and a back linear transformation, that is, the single iterative round operation process from the 9th round to the last round in the unified round operation of AES is also uniformly converted into a global pipeline operation process of a front linear transformation, an S-box substitution transformation and a back linear transformation.

[0071] When the processing algorithm is AES, the first iteration round only has the round key sub-transformation, and the front linear transformations of other iteration rounds include the row shift operation of AES. In one embodiment, the row shift operation specifically adopts the linear transformation timing isolation segment for realizing the permutation processing. In other embodiments, the row shift operation can also be implemented without the timing isolation segment. The back linear transformation includes the XOR network transformation formed by combining the column confusion of AES and the round key sub-transformation, and the back linear transformation of the last iteration round does not contain the column confusion of AES.

[0072] When the processing algorithm is SM4, the front linear transformation includes the round key addition operation of SM4, the rear linear transformation includes the XOR network transformation formed by merging the L transformation and XOR combination operation of SM4, and the FP permutation (Final Permutation) is performed after the last round of rear linear transformation.

[0073] The present invention abstractly presents the design architecture of a single iteration round compatible with the unified round operation of the two algorithms as Figure 1 The architecture shown is implemented. The implementation architecture is centered on the "S-box substitution transformation" in a single iteration round, and the sub-transformation operation before it is defined as the "pre-linear transformation (MLpre)", and the sub-transformation operation after it is defined as the "post-linear transformation (MLaft)". With this abstract division, it is convenient to further carry out the architecture design and implementation of the pipeline mechanism at the local level by dividing and conquering the three parts. That is, the circuit part of the pre-linear transformation of a single iteration round of unified round operation includes: the row shift operation of AES, and the round key addition operation of SM4. The circuit of the S-box substitution transformation of a single iteration round of unified round operation is a byte replacement operation carried out on the whole or part of the state variable with the byte (8-bit) as the granularity. The circuit of the post-linear transformation of a single iteration round of unified round operation includes: the XOR network transformation (XNWA) formed by merging the column confusion and round key addition operations of AES, and the XOR network transformation (XNWS) formed by merging the L transformation and XOR combination operations of SM4.

[0074] On the basis of the above, the present invention adopts the basic idea of ​​combining composite domain conversion with multiplication inversion on the same target composite domain, and uniformly converts the S-box substitution transformation of AES and SM4 into a local pipeline operation process of front mapping transformation, composite domain inversion transformation, and back mapping transformation, so that the S-box substitution transformation includes multiple levels of timing isolation segments, and any level of timing isolation segment is composed of at least one XOR gate and / or at least one AND gate, wherein the XOR gate can be a two-input XOR gate, and the AND gate can be a two-input AND gate. The front mapping transformation converts the different GF(2^8) domain elements of AES and SM4 in the S-box substitution transformation into the same target composite domain that is linearly isomorphic to their respective original GF(2^8) domains.

[0075] The previous solution for a dedicated hardware accelerator (i.e., the hardware circuit module mentioned above in this article) that is compatible with implementing SM4 and AES block cipher algorithms is generally to use the same physical memory logic (e.g., SRAM) inside the hardware circuit module to time-division multiplex the different SBOX logics of the two algorithms, i.e., to load the memory logic with the S-box lookup result corresponding to a certain algorithm during the period when one algorithm is activated.

[0076] Under the premise of this general solution, SBOX logic is usually implemented using a lookup table (LUT). In addition, the access performance of the corresponding memory logic is upper bounded within the framework of the established integrated circuit process technology, which often results in it failing to meet the performance requirements for higher data throughput in some specific application scenarios.

[0077] If the above-mentioned LUT+SRAM solution is not adopted, and one wants to save hardware resource overhead as much as possible while meeting the ultra-high throughput requirements, the difficulty and key point lies in: the efficient reuse of the different SBOX logics of the two algorithms, and the appropriate reuse of other sub-transformation logics of the round operation under the pipeline mechanism.

[0078] The present invention adopts the general idea of ​​composite domain conversion + multiplication inversion for efficient reuse of different SBOX logics in round operations of the two algorithms, so that the nonlinear part of the SBOX logic itself (i.e., multiplication inversion on the GF(2^8) domain) can find the same reusable part in different definitions of the two algorithms. In addition, for the linear part of the SBOX logic itself (i.e., (inverse) affine transformation) and the isomorphic (inverse) mapping transformation involved in the mutual conversion of elements between the GF(2^8) domain and the target composite domain (note: both are linear transformations), under the realization index of ultra-high throughput / performance, they can optionally weigh and consider the specific circuit implementation based on the DACSE optimization algorithm.

[0079] Symmetric key cryptographic algorithms such as SM4 and AES involve the concept of finite fields (also known as Galois fields, abbreviated as GF) and arithmetic operations such as addition, multiplication and inversion on the corresponding finite fields in the various seed transformation operations of their round operations. In the context of modern computer applications and CMOS digital logic design, due to the relevant integrated implementation limitations, the corresponding data is usually processed / stored / transmitted in binary form. Therefore, symmetric key cryptographic algorithms such as AES and SM4 took this feature of the implementation platform into consideration when they were first formulated.

[0080] A field is a set of elements on which operations such as addition, subtraction, multiplication, and division can be performed without exceeding the value of the field. A field containing a finite number of elements is called a finite field, and the number of its elements is called the "order of the finite field". The order of every finite field must be a power of a prime number, that is, the order of a finite field can be expressed as p^n (p is a prime number and n is a positive integer). Finite fields are often also called Galois Fields, denoted as GF(p^n).

[0081] When n=1 and p is a prime number, there exists a finite field GF(p) which is also called the prime number field.

[0082] When n=1 and p=2, it corresponds to the smallest prime number field GF(2). In cryptography, its elements are 0 and 1 (1-bit binary numbers), and addition and multiplication on this field are equivalent to logical exclusive OR (XOR) and logical AND (AND) operations, respectively.

[0083] When n>1 and p=2, it corresponds to the GF(2^n) field. In cryptography, its elements are generally not represented by integers, but are represented as a polynomial whose "highest term is x^(n-1) and the coefficient of each term is an element in the GF(2) field."

[0084] Finite fields are widely used in cryptography, and the most commonly used are prime number fields GF(p) and GF(2^n).

[0085] For the two symmetric key cryptographic algorithms, SM4 and AES, the S-box Chautor transforms of their round operations specifically involve the GF(2^8) field, and an 8-bit byte data can be mapped to "one of 256 possible 7th-order polynomials". Based on this, the addition and subtraction operations on the GF(2^8) field in cryptography are defined as: the XOR operation of each coefficient of the polynomial corresponding to the two operands according to the order, which is equivalent to the "bitwise XOR operation of the two bytes to be added / subtracted". Similarly, the definition basis of multiplication and division operations on the GF(2^8) field in cryptography includes: specifying an 8th-order irreducible polynomial and the modulo operation based on it.

[0086] Therefore, for multiplication, it is defined as: the two operands are first multiplied by a polynomial, and then modulo the irreducible polynomial. The remainder of the modulo is the multiplication result, and the multiplication result must be an element in the GF(2^8) domain. For multiplication inversion, it is equivalent to "division operation with the dividend restricted to element 1 in the GF(2^8) domain". In other words, the multiplication inversion of element X is to find the element Y such that the remainder after multiplying X and Y modulo the irreducible polynomial is the element 1.

[0087] The SBOX logic of AES and SM4 algorithms is obviously different. In addition to the different definitions of affine transformation, the more critical reason is the difference in the irreducible polynomials that the multiplication inversion over the GF(2^8) field depends on.

[0088] The 8th-order irreducible polynomial of AES is SA(x)=(x^8+x^4+x^3+x+1), while that of SM4 is SS(x)=(x^8+x^7+x^6+x^5+x^4+x^2+1).

[0089] In order to find the common points of the multiplication inversion operation of the SBOX logic of the AES and SM4 algorithms on the GF(2^8) domain, it is necessary to use the decomposition operation to convert the two different GF(2^8) domains into the same target composite domain that is linearly isomorphic to their respective original GF(2^8) domains. Based on this, the multiplication inversion operation on the GF(2^8) domain can be converted into a corresponding isomorphic operation on the target composite domain that is more suitable for the implementation of digital circuits on the binary basis. In addition, because the GF(2^8) domain and the target composite domain are linearly isomorphic, the element mapping conversion between the two types of domains can be completed through the isomorphic (inverse) mapping transformation of "the implementation method is matrix multiplication".

[0090] In one embodiment, the target composite domain can be a GF((2^4)^2) domain, which is generated from a base domain GF(2^4) based on a specified 2nd-order irreducible polynomial (P(y)=y^2+y+ν, ν={0010}2), and the GF(2^4) base domain is generated from a GF(2) domain by another specified 4th-order irreducible polynomial (Q(x)=x^4+x^3+x^2+x+1). Thus, the multiplication inversion operation on the GF(2^8) domain can be isomorphically mapped to the multiplication inversion operation on the target composite domain. In order to achieve the isomorphic (inverse) mapping of elements in two GF domains, the corresponding isomorphic (inverse) mapping transformation specific to a single algorithm will be adopted. In terms of hardware circuit implementation, the multiplication inversion on the target composite domain will be implemented by using a circuit structure combination of the following sequence: "element isomorphism mapping of GF(2^8) domain → GF((2^4)^2) domain, multiplication inversion on GF((2^4)^2) domain, and element isomorphism inverse mapping of GF((2^4)^2) domain → GF(2^8) domain".

[0091] In order to be compatible with and support ultra-high performance implementation of both algorithms, a preferred embodiment of the present invention designs the above-mentioned global and local pipeline mechanisms to implement a method compatible with both AES and SM4 block cipher algorithms.

[0092] Figure 2The architecture design of the local pipeline of various SBOX logics for the two algorithms AES and SM4 is shown. The AES algorithm uses two versions of SBOX logic for encryption and decryption, respectively, and the SM4 algorithm uses the same version of SBOX logic for encryption / decryption. Regardless of the algorithm, or the forward / reverse version of AES SBOX, these SBOX logics can be split into a combination of sub-transformation circuits that are sequentially "front affine transformation, multiplication inversion on the GF(2^8) field, and back affine transformation" (Note: This split corresponds to the original definition of the algorithm standard).

[0093] Specifically, the "front affine transformations" of the three SBOX logics, namely, the forward SBOX of AES, the reverse SBOX of AES, and the SBOX of SM4, are: the empty transformation, the inverse affine transformation of the AES standard, and the first affine transformation of the SM4 standard. The "back affine transformations" of these three SBOX logics are: the affine transformation of the AES standard, the empty transformation, and the second affine transformation of the SM4 standard.

[0094] Under the general idea of ​​"composite domain conversion + multiplication inversion" described in this specification, the above standard can be split and adjusted as follows: Figure 2 The leftmost column shows the optimization split. That is, when the algorithm being processed is the AES encryption algorithm, the post-mapping transformation performs an AES standard affine transformation after performing an isomorphic inverse mapping of the target composite domain to GF(2^8); when the algorithm being processed is the AES decryption algorithm, the pre-mapping transformation performs an AES standard inverse affine transformation before converting GF(2^8) to the target composite domain; when the algorithm being processed is SM4, the pre-mapping transformation performs an SM4 standard first affine transformation before converting GF(2^8) to the target composite domain, and at the same time, the SM4 post-mapping transformation performs an SM4 standard second affine transformation after performing an isomorphic inverse mapping of the target composite domain to GF(2^8).

[0095] Based on the above description of the SBOX logic optimization splitting and the corresponding description of the "isomorphic implementation based on composite domain transformation" regarding the inversion of multiplication on the Galois field mentioned above, the present invention abstractly summarizes the circuit implementation architecture of the SBOX logic that is compatible with the unified round operation of the two algorithms as: a sequential combination of the front mapping transformation, the composite domain inversion transformation, and the back mapping transformation.

[0096] The present invention also protects a system for implementing the above-mentioned technical solution that is compatible with both AES and SM4 block cipher algorithms, the system comprising multiple iterative round implementation circuits, wherein a single iterative round implementation circuit comprises a front linear transformation circuit, an S-box substitution transformation circuit, and a back linear transformation circuit.

[0097] The front linear transformation circuit is used to implement the row shift operation of AES and the round key addition operation of SM4.

[0098] The S-box substitution transform circuit is used to implement the front linear mapping, the composite domain inverse transformation, and the back mapping transformation.

[0099] The post-linear transformation circuit is used to implement the XOR network transformation (XNWA) formed by combining the column confusion and round key addition operations of AES, and the XOR network transformation (XNWS) formed by combining the L transformation and XOR combination operations of SM4.

[0100] Figure 3 The circuit implementation of the entire group operation under the global pipeline mechanism that is compatible with the two algorithms is shown. As mentioned above, the general iteration round of the unified round operation can be implemented in sequence by the circuit combination of "front linear transformation, S-box substitution transformation, and back linear transformation". In addition to this general iteration round, the following points should also be noted:

[0101] (1) For the SM4 algorithm, after State (31) is obtained, "FP permutation" must be performed to obtain the output grouping result. That is, SM4 performs FP permutation to obtain the output grouping result after the unified round of operations is completed.

[0102] (2) Round #0 of the AES algorithm only has the "add round key" sub-transformation.

[0103] (3) The MLaft of round #N (i.e., the last round, N=10 / 12 / 14) of the AES algorithm does not contain column confusion (Note: To distinguish, Figure 3 The middle mark is MLaft').

[0104] (4) The resource allocation ratio in the S-box lookup logic box of the two algorithms shown in the figure is different, the specific ratio is 4:1 (AES:SM4). The resource allocation ratio here refers to the number of times the two algorithms call the S-box logic (Note: in hardware circuit implementation, it is also the number of single S-box logic circuits).

[0105] Figure 4 In one embodiment, the circuit implementation architecture of a basic pipeline logic (circuit submodule) of the unified round operation under the global pipeline design architecture is shown, that is, the implementation circuit of a general single iteration round.

[0106] K and A / B / C / D represent the round key input and State stimulus input of the submodule respectively, and PO and NA / NB / NC represent the State result output of the submodule. PO = Primary Output, NA = New A, and the bit width of these signals is one word (32-bit).

[0107] The light grey box marked with the word XOR on the top corresponds to the circuit implementation of the "front linear transformation" of the SM4 algorithm (Note: the front linear transformation of the row shift of the AES algorithm is not included in the basic pipeline logic). The XOR circuit can be implemented by a set of four-input XOR gates or multiple sets of two-input XOR gates.

[0108] It can be seen from this specific example that the front-end linear transformation circuit includes at least one XOR gate and a multiplexer.

[0109] The light grey box in the middle marked with SB4X corresponds to the circuit implementation of the "S-box substitution transformation" reused by the two algorithms (Note: 4X means substitution processing for 4 bytes).

[0110] The light grey box marked with XNWA / XNWS at the bottom corresponds to the circuit implementation of the "post-linear transformation", that is, the XOR network sub-transformation of the AES / SM4 algorithm.

[0111] The is_AES signal is 1 / 0, indicating that the AES / SM4 algorithm is currently selected and activated.

[0112] The light gray box marked with DFFuN (N=1,2...) corresponds to the multi-stage shift circuit implementation of the SM4 algorithm for the input A / B / C / D signals under the local pipeline design architecture. The specific number of shift stages required should be equal to the sum of the number of timing isolation segments of the two sub-circuits of XOR and SB4X in this figure. For the trade-off between the two implementation indicators of high processing performance and low resource overhead, please refer to the relevant description of claim 3.

[0113] Figure 5 In one embodiment, a circuit implementation architecture of the multiplexed SBOX logic of unified round operations under a local pipeline design architecture is shown.

[0114] Aiso and Siso respectively represent the isomorphic mapping transformation logic of the two algorithms, which are responsible for isomorphic mapping of an element on the GF(2^8) field to the corresponding element on the GF((2^4)^2) composite field, that is, Aiso is the isomorphic mapping transformation logic of AES, and Siso is the isomorphic mapping transformation logic of SM4.

[0115] Aiso -1 and Siso -1 They respectively represent the inverse transformations of the above two isomorphic mappings, that is, the isomorphic inverse mapping transformation logic of each algorithm.

[0116] Aaf and Aaf -1 They represent the affine transformation logic and inverse affine transformation logic of the AES algorithm respectively.

[0117] Saf1 and Saf2 represent the first affine transformation logic and the second affine transformation logic of the SM4 algorithm respectively.

[0118] MI on GF((2^4)^2) represents the unified multiplication inversion logic on the target composite field.

[0119] The three rows from top to bottom in this figure correspond to the SBOX logic of AES reverse version, AES forward version, and SM4 respectively.

[0120] Figure 6 In one embodiment, a circuit implementation of the non-multiplicative inversion logic in the multiplexed SBOX logic (the figure only takes the affine transformation logic Aaf of the AES algorithm as an example) under a local pipeline design architecture is shown.

[0121] Isomorphic (inverse) mapping transformation logic is equivalent to a certain matrix multiplication logic. Under the premise that the coefficients of the polynomials are all elements on the GF(2) field (i.e., 0 or 1), their circuit implementation can be abstracted as follows: the input data is X with a bit width of 8, and each bit is recorded as x7, x6..., x0; the output data is Y with a bit width of 8, and each bit is recorded as y7, y6..., y0; the calculation function of a certain output bit yi (i∈[0,7]) can be expressed as the formula yi=fi(x7,x6...,x0). In short, an output bit yi is the XOR result of several input bits. For example, in one embodiment, the SM4 algorithm is based on the isomorphic mapping logic of the corresponding target composite domain (i.e. Figure 5 The implementation expression of Siso shown in FIG. 1 can be expressed as follows (the "^" symbol here represents a single-bit XOR process):

[0122] y7=x6^x5^x4^x3^x2,

[0123] y6=x7^x3^x2^x1,

[0124] y5=x7^x5^x3^x2,

[0125] y4=x5^x3^x2,

[0126] y3=x7^x6^x5^x1,

[0127] y2=x6^x5^x4^x2,

[0128] y1=x6^x5^x2^x1,

[0129] y0=x6^x5^x1^x0,

[0130] From the above implementation expression, it can be observed that each yi output bit corresponds to its own xi combination XOR calculation function, that is, the fi(x7,x6...,x0) function. Note that since there are at most 8 xi, the XOR processing objects for calculating a certain yi will not exceed 8.

[0131] When the above implementation expression is RTL-coded based on the HDL language, it can be written in a style similar to pseudocode or in a manually-intervened style based on the DACSE optimization algorithm. In one embodiment, the manually-intervened writing style can use the most basic XOR2 / AND2 (two-input XOR / AND) gate unit circuit to implement RTL coding. More broadly, for different circuit implementation underlying platforms (such as ASIC or FPGA), different foundry process technologies (such as SMIC 40LL or UMC 55ULP), or back-end EDA tools from different manufacturers, the two styles of RTL code may have unpredictable differences in both performance and resources. Therefore, for a certain given implementation scenario combination, a comparative analysis of the actual implementation results of the two coding styles is the appropriate strategy to quickly obtain better results.

[0132] Similar to the implementation of isomorphic mapping transformation, (inverse) affine transformation logic is also equivalent to some kind of matrix multiplication logic, but its input data may have an additional constant 1 or 0 (Note: the original affine transformation is equivalent to multiplying the input by a specified constant matrix, and then adding a specified 8-bit constant vector). Its calculation function can be expressed as the formula "yi = fi(x7,x6...,x0,c); (c = 0, 1)". In short, an output bit yi is the XOR result of several input bits and a single-bit constant that is either 0 or 1. For example, the implementation expression of the first affine transformation logic (Saf1) of the SM4 algorithm can be expressed as (here the "^" symbol represents the single-bit XOR processing; remember c1 = single-bit constant 1)

[0133] y7=x7^x6^x4^x1^x0^c1,

[0134] y6=x7^x6^x5^x3^x0^c1,

[0135] y5=x7^x6^x5^x4^x2,

[0136] y4=x6^x5^x4^x3^x1^c1,

[0137] y3=x5^x4^x3^x2^x0,

[0138] y2=x7^x4^x3^x2^x1,

[0139] y1=x6^x3^x2^x1^x0^c1,

[0140] y0=x7^x5^x2^x1^x0^c1,

[0141] From the above implementation expression, it can be observed that each yi output bit corresponds to its own xi&c1 combination XOR calculation function, that is, the fi(x7,x6...,x0,c) function. Note that since xi will not appear 8 times in the same expression, the XOR processing objects for calculating a certain yi will not exceed 8.

[0142] In summary, whether it is isomorphic (inverse) mapping transformation or (inverse) affine transformation, both types of logic can be implemented using logic circuits containing only XOR gate units. Specifically, under a certain combination of implementation scenarios, the implementation result circuits obtained based on the two RTL coding styles of the pseudocode and manual intervention often have certain performance differences in terms of processing performance and resource overhead.

[0143] For example, if a certain manufacturer's back-end EDA tool has powerful performance and can make multiple iterative attempts in a short time, then it is more suitable to adopt a pseudocode coding style, so that the tool can obtain a higher automatic selection right, thereby realizing a result circuit that better meets the implementation constraints. On the contrary, if it is known that the performance of the EDA manufacturer's back-end tool is not very stable, you can consider the coding style of manual intervention and directly adopt a certain XORn / ANDn gate unit (n represents the number of inputs) to obtain a "path delay predictable" circuit implementation. Combined with gate unit call branches based on various possible implementation scenario combinations, you can obtain a reusable design of "relaxing circuit clock frequency constraints based on a specified performance indicator to reduce resource overhead" with minimal modification work.

[0144] Figure 7 In one embodiment, a circuit implementation of the multiplication inversion logic on the target complex domain of the multiplexed SBOX logic under the local pipeline design architecture is shown. We can assume that the input and output of the multiplication inversion are Xi and Yo respectively; the sum of the circuit delays of several levels of AND2 gate units is na, and the sum of the circuit delays of several levels of XOR2 gate units is nx (n = 1, 2... represents the number of levels, and a / x represents the circuit delay of a single-level gate unit).

[0145] The gray italic characters in the figure represent the principle mark of the intermediate variables of the inversion algorithm, the regular characters represent the signal name mark in the circuit implementation, and the italic underlined font represents the corresponding multiplication and multiplication inversion logic subcircuit on the GF(2^4) basis field.

[0146] The rectangular box with a small triangle on a gray background in the figure represents a sequential logic element such as a D flip-flop, and the elliptical box with a white background represents the corresponding gate unit combinational logic element (the words 1x, 2a, etc. inside are the sum of the delays of the above-mentioned gate unit circuits). Figure 7 Each D flip-flop is a separation point of the timing isolation segments on its left and right sides.

[0147] In the figure, H represents the upper 4 bits of input data, L represents the lower 4 bits of input data, and M1 / M2 / M3 are the output results of the operation circuit in the dotted box on the left (Note: the letter M indicates that the dotted box circuit is the corresponding multiplication logic).

[0148] In the current embodiment, as shown in this figure, the division principle of the timing isolation segment is: the combination path delay of each isolation segment does not exceed the sum of the circuit delays of N levels of basic gate units, and the value of N is determined according to the preset standards of processing performance and storage resource overhead. For example, in a specific embodiment, the combination path delay of each isolation segment does not exceed the sum of the circuit delays of 3 levels of basic gate units. For example, in the U_mip subcircuit, the sum of the circuit delays before its first-level D flip-flop is "1x+2a" (Note: its "P transformation from H_ff2 to HHv" belongs to the position-based permutation type processing, so it can be considered that no circuit component delay will be introduced in the hardware circuit implementation), and the sum of the circuit delays before its second-level D flip-flop is "3x". Similarly, for Figure 7 For the illustrated embodiment, the multiplication inversion logic circuit on its target complex domain is designed to have 6 levels of sequential isolation segments.

[0149] By extension, the hardware implementation method of the present invention, that is, the circuit implementation architecture, can easily achieve the circuit structure optimization of the key sub-circuit of "multiplication inversion logic on the target composite domain of the reused SBOX logic" under different implementation scenario combinations (including but not limited to the underlying platform, foundry process technology, back-end implementation tools and other dimensions) by splitting or merging relevant combinational logic paths and adding or deleting corresponding sequential logic elements, thereby obtaining priority preference or trade-off considerations on the two implementation indicators of high processing performance and / or low resource overhead.

[0150] Figure 8 FIG. 1 shows an overall circuit implementation of a unified round operation under a global pipeline mechanism in one embodiment. We can assume that Figure 4 A basic pipeline logic is shown in this figure by a gray rectangle with the word 1pp.

[0151] The square boxes with diagonal stripes in this figure represent the sequential logic elements of the 32-bit wide register state variable. The square boxes with white background represent the part of the single round key with 32 bits wide (such as k0, k1, k2, k3), which are also implemented as corresponding sequential logic elements. The is_AES signal is 1 / 0, which indicates that the AES / SM4 algorithm is currently selected and activated. Note: Due to the particularity of the AES row shift sub-transformation itself and the application characteristics of the global pipeline architecture, in this embodiment, the "row shift" is independent of the 1pp sub-circuit.

[0152] For the sake of simplicity, the multiplexer logic of "autonomously selecting the expected round key for multiple sets of round keys in real time by hardware" is not included in this diagram. Similarly, the differences in the number of different iterations of the round operation due to the three initial key length modes of AES are also not included in this diagram.

[0153] In order to emphasize the pipeline mechanism, the output signals that are actually integrated in the 1pp subcircuit and implemented in the form of sequential logic elements are moved outside the 1pp frame in this figure. For example: for the #1 round of the AES algorithm, the PO outputs of the four 1pp subcircuits are illustrated as s0, s1, s2, and s3 in sequence; and for the #0 round of the SM4 algorithm, the PO, NA, NB, and NC outputs of the single 1pp subcircuit are illustrated as i1, i2, i3, and s0 drawn together in sequence.

[0154] According to the original definition of the AES algorithm standard, a single iteration round of AES (except round #0) requires 16 S-box substitution transformation calculation operations. In this embodiment, they are calculated and processed in parallel on the time axis by four 1pp sub-circuits (which integrate a total of 16 multiplexed SBOX logics).

[0155] According to the original definition of the SM4 algorithm standard, a single iteration round of SM4 only requires four S-box lookup transformation calculation operations. In this embodiment, they are calculated and processed in parallel on the time axis by a 1pp subcircuit (which integrates a total of four multiplexed SBOX logics).

[0156] Combining the original definitions of the two algorithm standards, and considering that the operations of rounds #0 to #8 in the three initial key length modes of AES are the same in each mode, it can be deduced that rounds #1 to #8 of the AES algorithm and rounds #0 to #31 of the SM4 algorithm require the same number of 128 SBOX logics in the design architecture of the pipeline mechanism. Therefore, in this embodiment, these SBOX logics can be integrated into a multiplexed version compatible with different algorithms and different encryption / decryption directions. For the other several iterations of the AES algorithm after round #8, in order to save resource overhead, the SBOX logic will adopt a simplified multiplexed version that removes the SM4 content, contains only the AES content, but is compatible with the encryption / decryption direction.

[0157] The following is a detailed explanation of the English abbreviations used in this patent. Such as AFISO, CFINV, AFIIA; ABIAI, ABII; SAISO, SIIA, etc. Among them, A=AES, S=SM4; F=Forward, B=Backward; CF=Composite Field, INV=Inversion; ISO=Isomorphic mapping, IIA=InverseIsomorphic&Affine mapping, II=Inverse Isomorphic mapping, IA=Inverse Affinemapping.

[0158] The forward version of the AES algorithm, the SBOX logic's forward mapping transformation, is "the isomorphic mapping transformation that maps elements on the GF(2^8) field to the target composite field" - AFISO.

[0159] The finite field inverse transformation of the forward version of the AES algorithm SBOX logic is the "multiplication inverse transformation on the target composite field" - CFINV.

[0160] The post-mapping transformation of the forward version of the SBOX logic of the AES algorithm includes "the isomorphic inverse mapping transformation that maps the elements on the target composite field back to the GF(2^8) field, and the affine transformation in the AES standard" - AFIIA.

[0161] The forward mapping transformation of the reverse version of the SBOX logic of the AES algorithm includes "the inverse affine transformation in the AES standard, and the isomorphic mapping transformation that maps elements on the GF(2^8) field to the target composite field" - ABIAI.

[0162] The finite field inverse transformation of the reverse version of the AES algorithm, SBOX logic, is the "multiplication inverse transformation on the target composite field" - CFINV.

[0163] The post-mapping transformation of the reverse version of the SBOX logic of the AES algorithm is "the isomorphic inverse mapping transformation that maps the elements on the target composite field back to the GF(2^8) field" - ABII.

[0164] The pre-mapping transformation of the SBOX logic of the SM4 algorithm includes "the first affine transformation in the SM4 standard, and the isomorphic mapping transformation that maps elements over the GF(2^8) field to the target composite field" - SAISO.

[0165] The finite field inversion transform of the SBOX logic of the SM4 algorithm is the "multiplicative inversion transform on the target composite field" - CFINV.

[0166] The post-mapping transformation of the SM4 algorithm's SBOX logic includes "an isomorphic inverse mapping transformation that maps elements on the target composite field back to the GF(2^8) field, and a second affine transformation in the SM4 standard" - SIIA.

[0167] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method compatible with both AES and SM4 block cipher algorithms, characterized in that: include: The unified round operation of SM4 and the single general iterative round operation process from the 1st round to the 8th round in the unified round operation of AES are uniformly converted into a global pipeline operation process of front linear transformation, S-box substitution transformation and back linear transformation; The front linear transformation includes at least one level of timing isolation segment, and any level of the timing isolation segment is composed of at least one XOR gate or is a linear transformation timing isolation segment for implementing a replacement process; The S-box substitution transformation of AES and SM4 is uniformly converted into a local pipeline operation process of front mapping transformation, composite domain inversion transformation, and back mapping transformation, so that the S-box substitution transformation includes multiple levels of timing isolation segments, and the timing isolation segment at any level is composed of at least one XOR gate and / or at least one AND gate; The forward mapping transformation converts the elements on the GF(2^8) fields that are different between AES and SM4 in the S-box substitution transformation into the same target composite field that is linearly isomorphic to the respective original GF(2^8) fields; The preset standards of processing performance and storage resource overhead are achieved by designing the sum of the number of temporal isolation segments of the front linear transformation and the number of temporal isolation segments of the S-box substitution transformation required by SM4 in a single general iteration round; The principle of dividing the timing isolation segments is that the combined path delay of each isolation segment does not exceed the sum of the circuit delays of N levels of basic gate units, and the value of N is determined according to preset standards of processing performance and storage resource overhead.

2. The method for being compatible with both AES and SM4 block cipher algorithms as claimed in claim 1, characterized in that: The single iterative round operation process from the 9th round to the last round in the unified round operation of the AES is also uniformly converted into a global pipeline operation process of the front linear transformation, S-box substitution transformation and the back linear transformation.

3. The method for being compatible with both AES and SM4 block cipher algorithms as claimed in claim 1, characterized in that: When the algorithm being processed is AES, the first iteration round has only a round key sub-transformation, the front linear transformations of other iteration rounds include an AES row shift operation, the back linear transformations include an XOR network transformation formed by combining an AES column obfuscation and a round key operation, and the back linear transformation of the last iteration round does not include an AES column obfuscation; When the processing algorithm is SM4, the front linear transformation includes the round key addition operation of SM4, the rear linear transformation includes the XOR network transformation formed by merging the L transformation and XOR combination operation of SM4, and the FP replacement is performed after the last round of rear linear transformation is completed.

4. The method for being compatible with both AES and SM4 block cipher algorithms as claimed in claim 1, characterized in that: When the algorithm being processed is the AES encryption algorithm, the post-mapping transformation performs the GF(2^8) isomorphic inverse mapping corresponding to AES and then performs the AES standard affine transformation; when the algorithm being processed is the AES decryption algorithm, the pre-mapping transformation performs the AES standard inverse affine transformation and then performs the AES standard GF(2^8) isomorphic mapping; When the processing algorithm is SM4, the front mapping transformation performs the first affine transformation of the SM4 standard before converting the GF(2^8) domain to the target composite domain. At the same time, the back mapping transformation of the SM4 performs the second affine transformation of the SM4 standard after converting the target composite domain back to the corresponding GF(2^8) domain through the GF(2^8) isomorphic inverse mapping.

5. The method for being compatible with both AES and SM4 block cipher algorithms as claimed in claim 1, characterized in that: The target composite domain is the same GF((2^4)^2) domain.

6. A system for implementing the method of being compatible with both AES and SM4 block cipher algorithms as claimed in any one of claims 1 to 5, comprising a plurality of iterative round implementation circuits, wherein the plurality of iterative round implementation circuits include a general iterative round implementation circuit, characterized in that: A single general iteration round implementation circuit includes: A front linear transformation circuit, which is used to implement the row shift operation of AES and the round key addition operation of SM4; S-box substitution transformation circuit, which is used to realize front linear mapping, composite domain inverse transformation, and back mapping transformation; The post-linear transformation circuit is used to implement the XOR network transformation (XNWA) formed by combining the column confusion and round key addition operations of AES, and the XOR network transformation (XNWS) formed by combining the L transformation and XOR combination operations of SM4.

7. The system for the method of being compatible with both AES and SM4 block cipher algorithms as claimed in claim 6, characterized in that: The front-end linear conversion circuit includes at least one XOR gate and a multiplexer.