A Hardware AES-GCM Method and System Based on FPGA
By designing a hardware encryption system based on AES-GCM on FPGA, using multi-core structures and parallel processing arrays, the problems of low security and low efficiency of data transmission in the prior art are solved, and efficient and secure data encryption processing are achieved.
Patent Information
- Application Number
- CN202411733585.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-29
AI Technical Summary
The existing hardware data encryption technology is not very secure during data transmission, and the encryption time is long when large data transmission is transmitted, and the transmission efficiency is low.
A hardware AES-GCM system based on FPGA is designed, and a secondary flow structure of AES core with the characteristics of master and slave cores, AES-GCM parallel processing array, and AES and GCM is realized to realize real-time online encryption of data flow.
Real-time online encryption of data streams is realized to meet the network throughput requirements of 10Gbps. The test results show that the module can achieve a processing speed of 12Gbps, improving the security and efficiency of data transmission.
Smart Images

Figure CN119210696B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hardware data encryption, and in particular to a hardware AES-GCM method and system based on FPGA. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] A Field-Programmable Gate Array (FPGA) is a programmable semiconductor device that allows designers to configure and reconfigure the hardware logic after manufacturing. FPGA technology is very popular in electronic design and digital system design because it provides flexibility and customizability, enabling engineers to quickly develop and modify hardware designs according to needs.
[0004] An FPGA can be reprogrammed by loading different configuration files (usually called bitstreams), which allows them to be used in multiple applications without having to design and manufacture new hardware each time. Multiple operations can also be performed simultaneously, making them very suitable for applications that require high-speed parallel processing, such as digital signal processing, image processing, and communication systems. Due to the hardware parallelism of FPGAs, they can achieve very low data processing latency, which is crucial for real-time systems and high-speed communication. FPGAs can be configured to perform specific tasks as needed, making them very useful in applications that need to quickly adapt to new technologies or standard changes. At the same time, its design can be easily extended to meet more complex system requirements.
[0005] FPGA design process:
[0006] 1. Conceptual design: Define the requirements and specifications of the system;
[0007] 2. Hardware Description Language (HDL) coding: Write the logic of the FPGA using hardware description languages such as Verilog or VHDL;
[0008] 3. Simulation: Simulate the design on a computer to verify its functionality and performance;
[0009] 4. Synthesis: Convert the HDL code into a logic netlist that the FPGA can understand;
[0010] 5. Placement and routing: Map the logic netlist to the physical resources of the FPGA;
[0011] 6. Programming: Download the generated bitstream file into the FPGA to configure its logic;
[0012] 7. Testing: Test the FPGA on actual hardware to ensure it works as expected.
[0013] FPGA is widely used in various fields, including but not limited to: Communications: For high-speed data transmission and signal processing. Military and Aerospace: For radar systems, navigation systems, and communication equipment. Consumer Electronics: For video game consoles, digital TVs, and audio equipment. Industrial Control: For automated control systems and robotics. Automotive: For Advanced Driver Assistance Systems (ADAS) and vehicle networks.
[0014] The Advanced Encryption Standard (AES) is a symmetric-key encryption algorithm used to protect the security of electronic data. AES is widely used in various scenarios, including wireless networks, online transactions, data storage, and cloud computing, etc. It is commonly used to protect sensitive information and ensure the security and integrity of data during transmission and storage.
[0015] Galois Counter Mode (GCM) is a block cipher mode of operation that provides authenticated encryption by hashing over the binary Galois field of order 2 128 represented as GF(2 128 ). In recent years, GCM has been incorporated into multiple standards, such as in IETF RFC 4106 for secure encapsulation of secure payloads and in IEEE P1619.1 for authenticated encryption of storage devices.
[0016] The AES-GCM encryption technology, which combines AES and GCM, is an encryption method using key lengths of 128, 192, and 256 bits. It processes data in 128-bit block sizes. This encryption mode not only ensures data confidentiality but also provides authentication of data integrity and authenticity through the GCM mode. Due to its high efficiency, security, and protection of data integrity, the AES-GCM encryption mode has been widely used in communication and data storage fields that require a high level of security, such as the security layer of the TLS / SSL protocol, Virtual Private Networks (VPNs), data encryption for mobile devices, and various secure communication protocols.
[0017] During the data transmission process in hardware, common solutions use an encryption method such as AES or GCM to encrypt data. This encryption method is limited by the limitations of the algorithm itself, and the data transmission security is not high. In addition, when the data volume is large, it takes a long time to encrypt the data, and the transmission efficiency is low. Summary of the Invention
[0018] To solve the technical problems existing in the above-mentioned background technology, the present invention provides a hardware AES-GCM method and system based on FPGA. The present invention realizes real-time online encryption of data streams through the AES-GCM module. The module includes an AES core with the characteristics of a master core and a slave core, an AES-GCM parallel processing array, and a two-stage pipelined structure of AES and GCM. To meet the network throughput requirement of 10 Gbps. The test results show that the module can reach a processing speed of 12 Gbps.
[0019] To achieve the above object, the present invention adopts the following technical solutions:
[0020] The first aspect of the present invention provides a hardware AES-GCM system based on FPGA.
[0021] A hardware AES-GCM system based on FPGA deploys a number of AES-GCM modules on the FPGA hardware chip. The operations of each AES-GCM module are independent of each other and no data transmission is performed;
[0022] Among them, each of the AES-GCM modules includes a GCTR module and a GHASH module. At least two AES modules are integrated inside the GCTR module, and the GCM module is implemented by the GHASH module; the GCTR module performs an exclusive OR operation on the generated ciphertext and Hash_key and inputs them into the GHASH module for authentication processing to generate a TAG label for verifying whether the data has been tampered with during the encryption process.
[0023] Further, the operations of the AES module and the GCM module are independent of each other, and an array design structure is adopted in terms of structure.
[0024] Further, the array design structure is: when the AES module transmits the output result to the GCM module for processing, the AES module continues to process the next batch of data.
[0025] Further, each of the AES modules is responsible for encrypting or decrypting a part of the data.
[0026] Further, each AES module is expanded from a single channel to four channels.
[0027] Further, the AES modules work in parallel and share the same key, initialization vector, and counter.
[0028] Further, each of the GCM modules includes a plurality of GF multipliers.
[0029] Further, each of the AES modules is used to perform byte substitution, row shift, column mixing, and round key addition operations.
[0030] Furthermore, each of the AES-GCM modules has independent encryption and decryption functions.
[0031] The second aspect of the present invention provides a method for a hardware AES-GCM system based on FPGA as described in the first aspect.
[0032] The input set of data is divided into blocks, and different data blocks enter different AES-GCM modules of the array for processing respectively;
[0033] In one AES-GCM module, the two four-channel AES cores included in the GCTR function are used to process the AES encryption of data, and the GHASH function is used to perform encryption authentication on the AES-encrypted data; at the same time, the GCTR and GHASH functions form a two-stage pipeline to process the data entering this AES-GCM module.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] Based on the Xilinx Kintex UltraScale series FPGA, the present invention designs an efficient AES-GCM encryption module, realizing real-time online encryption of data streams. The module includes an AES core with the characteristics of a master core and a slave core, an AES-GCM parallel processing array, and a two-stage pipeline structure of AES and GCM to meet the network throughput requirement of 10 Gbps. The test results show that the module can reach a processing speed of 12 Gbps.
[0036] The present invention designs a multi-AES core structure with the characteristics of a master core and a slave core. This structure can not only efficiently execute AES encryption tasks, but also achieve higher data processing capabilities and security through the collaborative work between the master and slave cores. The design of this multi-core structure enables each core to work independently while being able to cooperate with each other, thus performing excellently when processing a large amount of data.
[0037] The present invention designs an AES-GCM parallel processing array, which is composed of multiple independent AES-GCM modules. These modules can simultaneously perform data encryption and authentication work, greatly improving the processing speed. Each module can independently complete AES encryption and GCM authentication tasks, thus realizing true parallel processing. This design not only improves the overall performance of the system, but also can effectively cope with high-concurrency scenarios, ensuring the security and integrity of data.
[0038] The present invention designs a two-stage pipelined structure of AES and GCM, which separates the AES encryption process from the GCM authentication process. After AES finishes processing the first batch of data, while GCM calculates the authentication result, AES continues to process the next batch of data. This pipelined design enables the two processes to run in parallel, greatly improving the overall processing efficiency. In this way, the system can achieve higher data throughput while ensuring data security, meeting the requirements of high-performance application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention, and the schematic embodiments and descriptions thereof are used to explain the present invention without unduly limiting the present invention.
[0040] Figure 1 is the main flowchart of AES encryption and decryption shown in the present invention;
[0041] Figure 2 is the flowchart of matrix transformation shown in the present invention;
[0042] Figure 3 is the schematic diagram of key expansion shown in the present invention;
[0043] Figure 4 is the flowchart of key expansion shown in the present invention;
[0044] Figure 5 is the flowchart of byte substitution shown in the present invention;
[0045] Figure 6 is the schematic diagram of row shift shown in the present invention;
[0046] Figure 7 is the schematic diagram of reverse row shift shown in the present invention;
[0047] Figure 8 is the schematic diagram of the GCTR module shown in the present invention;
[0048] Figure 9 is the flowchart of the GCM encryption process shown in the present invention;
[0049] Figure 10 is the flowchart of the GCM decryption process shown in the present invention;
[0050] Figure 11 is the schematic diagram of AES-GCM expansion shown in the present invention;
[0051] Figure 12 is the schematic diagram of the relationship between U1 and U2 shown in the present invention;
[0052] Figure 13It is a schematic diagram of the AES-GCM processing array shown in the present invention;
[0053] Figure 14 It is a schematic diagram of the parallelized structure of GCTR and GHASH calculations shown in the present invention;
[0054] Figure 15 It is a schematic diagram of the two-stage pipeline structure of GCTR and GHASH shown in the present invention. Detailed implementation manners
[0055] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0056] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0057] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0058] Embodiment 1
[0059] As Figure 1 shown, this embodiment provides a hardware AES-GCM system based on FPGA.
[0060] A hardware AES-GCM system based on FPGA deploys a number of AES-GCM modules on the FPGA hardware chip. The operations of each AES-GCM module are independent of each other and no data transmission is performed;
[0061] Among them, each of the AES-GCM modules includes a GCTR module and a GHASH module. At least two AES modules are integrated inside the GCTR module, and the GCM module is implemented by the GHASH module; the GCTR module performs an exclusive OR operation on the ciphertext and Hash_key it generates and then inputs them into the GHASH module for authentication processing to generate a TAG tag for verifying whether the data has been tampered with during the encryption process.
[0062] In some embodiments, the operations of the AES module and the GCM module are independent of each other, and an array design structure is adopted in terms of structure.
[0063] In some embodiments, the array design structure is as follows: when the AES module transmits the output result to the GCM module for processing, the AES module continues to process the next batch of data.
[0064] In some embodiments, each AES module is responsible for the encryption or decryption task of a part of the data.
[0065] In some embodiments, each AES module is expanded from a single channel to four channels.
[0066] In some embodiments, the AES modules work in parallel and share the same key, initialization vector, and counter.
[0067] In some embodiments, each GCM module includes a plurality of GF multipliers.
[0068] In some embodiments, each AES module is used to perform byte substitution, row shift, column mixing, and add-round-key operations.
[0069] In some embodiments, each AES-GCM module has independent encryption and decryption functions
[0070] Specifically, the AES algorithm, GCM algorithm, and parallel AES-GCM structure are described in detail below:
[0071] 1. AES Algorithm
[0072] AES is an algorithm that uses block encryption technology. It divides the plaintext data into data blocks of a fixed length, and then encrypts each data block in turn until the entire plaintext is fully encrypted. There are three choices for the key length supported by AES: 128 bits, 192 bits, or 256 bits. Different key lengths are recommended to use different numbers of encryption rounds. Generally speaking, the more encryption rounds, the stronger the security provided, but at the same time, it will also increase the time required for encryption processing. Table 1 shows the correspondence between the key length and the corresponding recommended number of encryption rounds.
[0073] Table 1 Correspondence between Key - Block - Rounds
[0074]
[0075] The AES encryption algorithm mainly consists of four operations: byte substitution, row shift, column mixing, and add-round-key. In addition, the original key needs to be expanded. Taking the 10-round encryption process as an example: First, the plaintext performs an add-round-key operation; then, byte substitution, row shift, column mixing, and add-round-key are cycled nine times; finally, the column mixing step is omitted in the tenth round. The main processes of AES encryption and decryption are as Figure 3As shown. Among them, W[0,3] refers to the 128-bit key composed of W[0], W[1], W[2], and W[3] in series. The round functions from the 1st round to the 9th round of encryption are the same, including 4 operations: substitution bytes, shift rows, mix columns, and add round key. The column mixing is not performed in the last round of iteration. In addition, before the first round of iteration, an XOR encryption operation is first performed on the plaintext and the original key.
[0076] At the same time, Figure 1 The AES decryption process is also shown. The decryption process is still 10 rounds, and the operation of each round is the inverse operation of the encryption operation. Since the 4 round operations of AES are all reversible, therefore, one round of the decryption operation is to sequentially perform inverse shift rows, inverse substitution bytes, add round key, and inverse mix columns. Similar to the encryption operation, the inverse mix columns is not performed in the last round. Before the first round of decryption, a key addition operation is performed once.
[0077] The 128-bit input plaintext block P is divided into 16 bytes, respectively represented as P = P0, P1 … P15. Usually, the plaintext block is described by a square matrix in bytes, and this matrix is called the state matrix. In each round of iteration of the algorithm, the content of the state matrix is continuously updated, and the final result is the ciphertext output. In this matrix, the arrangement order of the bytes follows the order from top to bottom and from left to right, and the specific arrangement is as Figure 2 shown.
[0078] Assume that a 128-bit key is used. The input key K is also divided into 16 bytes, respectively represented as K = K0, K1… K15, and is also described by a key matrix in bytes. Through the key scheduling function, this key matrix is expanded into a sequence of 44 words W[0], W[1], …, W
[43] . The first 4 elements W[0], W[1], W[2], W[3] of this sequence are the original key and are used as the initial key addition in the encryption operation. The following 40 words are divided into 10 groups, and each group of 4 words (128 bits) is respectively used for the add round key in 10 rounds of encryption operations, as Figure 3 shown.
[0079] The specific process of key expansion is as follows: First, convert the 16-byte initial key into 4 32-bit words by column, that is, W[0], W[1], W[2], W[3]. 10 groups of expanded keys, W[i], i = 4......43, are calculated through W[0], W[1], W[2], W[3]. As Figure 4 shown, the specific expansion rules are as follows:
[0080] 1) If i is not a multiple of 4, then the i-th column is determined by the following equation: w[i] = w[i - 4] ⊕ w[i - 1];
[0081] 2) If i is a multiple of 4, then the i-th column is determined by the following equation: w[i] = w[i - 4] ⊕ T(w[i - 1]).
[0082] The algorithm T consists of three parts: word rotation, byte substitution, and round constant XOR. Word rotation means rotating the 4 bytes in a word to the left by 1 byte. Byte substitution performs byte substitution on the result of word rotation using the S-box. Round constant XOR is to XOR the result of the previous two steps with the round constant Rcon[j], where j represents the round number, and Rcon is a fixed one-dimensional array with values Rcon = {00, 01, 02, 04, 08, 10, 20, 40, 80, 1B, 36}.
[0083] In the AES encryption algorithm, the byte substitution process is actually a lookup table operation. The algorithm defines an S-box (substitution box) and its inverse counterpart, the inverse S-box. Each element in the state matrix is mapped to a new byte by the following method: using the high 4 bits of the byte as the row index and the low 4 bits as the column index, and then selecting the corresponding element from the S-box or inverse S-box as the output result.
[0084] For example, during encryption, if the output byte is 0x2, then look up the 0x02 row and 0x06 column of the S-box to get the value 0x3f, and then replace the original 0x26 with 0x3f. After byte substitution of the state matrix Figure 5 as shown.
[0085] As Figure 6 shown, the shift is a left circular shift operation. When the key length is 128 bits, the 0th row of the state matrix is shifted left by 0 bytes, the 1st row is shifted left by 1 byte, the 2nd row is shifted left by 2 bytes, and the 3rd row is shifted left by 3 bytes.
[0086] As Figure 7 shown, the inverse shift transformation is to perform the opposite shift operation on each row in the state matrix. The 0th row of the state matrix is shifted right by 0 bytes, the 1st row is shifted right by 1 byte, the 2nd row is shifted right by 2 bytes, and the 3rd row is shifted right by 3 bytes.
[0087] The column mixing transformation is implemented by matrix multiplication. The state matrix after row shift is multiplied by a fixed matrix to obtain the confused state matrix, and its definition is as follows:
[0088]
[0089] The inverse column mixing transformation can be defined by the matrix multiplication of the following formula:
[0090]
[0091] Among them, 0 ≤ c ≤ Nb, where Nb is the length of the grouped data. It can be verified that the product of the inverse transformation matrix and the forward transformation matrix is exactly the identity matrix; the column mixing input is a 4x4 state matrix - , where each element is a byte. This matrix is actually another representation of a word (128 bits). - represents the transformed state matrix.
[0092] 2. GCM Algorithm
[0093] Let n and u denote the unique pair of positive integers such that the total number of bits of the plaintext is (n - 1)*128 + u, where 1 ≤ u ≤ 128. The plaintext consists of a sequence of n bit strings, where the last bit string has a bit length of u and the other bit strings have a bit length of 128. The sequence is denoted as P1, P2,..., Pn - 1, Pn. Similarly, the ciphertext is denoted as C1, C2,..., Cn - 1. The additional Au in the authentication data A is denoted as A1, A2,..., Am - 1, A, where the last bit string A can be a partial block of length v, and m and v denote the unique pair of positive integers such that the total number of bits in A is (m - 1)128 + v and 1 ≤ v ≤ 128. The authenticated encryption operation is defined as follows:
[0094]
[0095] This part is divided into an encryption part and data integrity verification. The two main functions used in GCM are block cipher encryption and multiplication over the field GF(2 128 ). Encrypting the value X with the key K using the block cipher is denoted as E(K, X), multiplying two elements X, Y ∈ GF(2 128 ) is denoted as X·Y, and adding X and Y is denoted as X⊕Y. Addition in this field is equivalent to a bitwise exclusive - or operation. The function len() returns a 64 - bit string that contains the non - negative integer describing the number of bits in its argument, with the least significant bit on the right. The expression 0 t denotes a string consisting of t 0 - bits, A||B denotes the concatenation of two bit strings A and B. The function MSB t (S) returns a bit string that contains only the left - most t bits of S, and the symbol {} denotes a bit string of length zero. The incr() function is used to generate consecutive counter values. This function treats the right - most 32 bits of its argument as a non - negative integer, with the least significant bit on the right, and increments that value modulo 2 32 . More formally, the value of incr(F||I) is F ||(i + 1 mod 2 32) Authentication tag T, with a length that can be any value between 0 and 128. The length of the tag is denoted as t.
[0096] The AES-GCM module relies on the GCTR encryption module and the GHASH implementation module. The GCTR encryption module is a high-performance encryption module that implements the GCTR mode based on the AES algorithm and can process the encryption of 8 data blocks simultaneously. The module includes control signal processing, data input and output, an internal state machine, AES core instantiation, and input / output multiplexing logic, supports the input of initialization vector (IV), key and its length, as well as the processing of special hash keys and Y0 values, and it controls the data flow and encryption process through a finite state machine.
[0097] The GHASH function receives 8 ciphertext blocks of 128 bits and the corresponding Y values, combines them with the hash key through XOR operations, then calls the GF multiplier module to perform finite field multiplication, and finally outputs 8 Y values of 128 bits and a valid signal for the authentication step in the GCM encryption algorithm.
[0098] The encryption part is the GCTR encryption module. The data integrity verification part is the GHASH module, as Figure 8 shown.
[0099] The data encryption module mainly realizes the encryption operation of the input plaintext data through the AES symmetric encryption algorithm, and is responsible for encrypting the plaintext data into ciphertext to ensure the security of the data. The GHASH module multiplies the ciphertext encrypted by the GCTR encryption algorithm by the Hash Key in GF(2 128 ) Finally, a tag TAG is generated to check whether the data has been tampered with during the encryption process. The process is as Figure 9 shown, where P-1 and P-2 are plaintexts, and C-1 and C-2 are ciphertexts.
[0100] The authentication decryption operation is similar to the encryption operation, but the order of the hash step and the encryption step is reversed. It is defined as follows:
[0101]
[0102] As Figure 10 shown, the tag T' calculated by the decryption operation is compared with the tag T related to the ciphertext C. If the two tags match (length and value), the ciphertext is returned. In the figure, P-1 and P-2 are plaintexts, and C-1 and C-2 are ciphertexts.
[0103] In this specific context, each element is defined as a vector consisting of 128 bits. To describe this concept in more detail, each element can be represented as a sequence containing 128 binary bits. Specifically, element X can be described as a sequence consisting of 128 bits, that is, X0 , X 1 , X 2 , ..., X 126 , X 127 。Here, X 0 represents the leftmost bit, while X 127 represents the rightmost bit.
[0104] Next, the special element R used in the multiplication operation is explained in detail. This element is defined as a 16-character string, specifically represented as "11100001||0120". The "||" symbol here indicates that this is an element consisting of two parts, where "11100001" is the first 8 bits and "0120" is the last 8 bits. Therefore, the entire element R can be understood as a binary sequence consisting of 16 characters.
[0105] In addition, the function of the function rightshift() also needs to be described in detail. The function of this function is to shift the bits of its argument one bit to the right. Specifically, if there is an argument V and the function rightshift(V) is called, then the result W will be a new 128-bit vector. In this new vector, each bit W i (where 1 ≤ i ≤ 127) will be equal to the bit V in the original vector V i−1 . In other words, each bit is shifted one bit to the right. The rightmost bit W 0 will be set to 0 to fill the empty space created by the right shift operation. In this way, the function rightshift() realizes the function of shifting each bit of a 128-bit vector one bit to the right. The multiplication process in GF(2 128 ) is as follows:
[0106] (1) Initialize Z to 0 and V as a copy of X.
[0107] (2) For each i from 0 to 127 (a total of 128 iterations):
[0108] a. Check whether the current lowest bit of Y (i.e., the i-th bit) is 1.
[0109] b. If the i-th bit of Y is 1, then perform an exclusive OR (XOR) operation on Z and V to update the value of Z.
[0110] c. Check whether the highest bit of V (i.e., the 128th bit) is 0.
[0111] d. If the highest bit of V is 0, then shift V one bit to the right (equivalent to dividing by 2).
[0112] e. If the most significant bit of V is not 0, then shift V one bit to the right and perform an exclusive - OR operation with a predefined "R" value to update the value of V.
[0113] (3)After the loop ends, return the value of Z, which is the product of X and Y in the GF(2 128 ) field.
[0114] 3. Parallel AES - GCM Structure
[0115] Since there is a performance bottleneck in a single AES core on Kintex UltraScale series FPGA devices, especially when performing GF(2 128 ) module calculations, it takes 128 clock cycles to complete. This situation is particularly obvious when dealing with high - speed infinite - bandwidth networks because these networks usually need to support a high throughput of 10 Gbps. To meet this requirement, parallelization techniques are adopted in AES encryption and GHASH hash functions. Through parallel processing, multiple computing tasks can be executed simultaneously, thus significantly improving the overall performance. Specifically, multiple AES cores can be deployed on the FPGA, and each core is responsible for encrypting or decrypting a part of the data. In this way, the GF(2 128 ) module calculations can be distributed to multiple cores for parallel execution, and multiple parts of data can be processed simultaneously in 128 clock cycles, thus greatly shortening the total processing time. In addition, by parallelizing the GHASH function, higher efficiency can also be obtained when dealing with data integrity verification.
[0116] The AES core is extended to 4 channels, enabling it to process four 128 - bit data blocks simultaneously. Through this extension, as Figure 11 shown, the entire AES - GCM module is upgraded from the original single - channel single - AES core to an efficient structure with 8 channels and 2 AES cores. Such an improvement significantly increases the theoretical processing speed from the original 0.2 Gbps to 1.6 Gbps. In addition, to further improve the processing capacity, 2 AES cores are integrated inside a GCTR module, thus ensuring higher data processing efficiency and stronger encryption performance.
[0117] In addition, these two AES cores U1 and U2 work in parallel logically. They share the key, initialization vector (IV), and counter, but process different data blocks respectively. This design enables the system to encrypt or decrypt multiple data blocks simultaneously, thus significantly improving the processing efficiency. The relationship between the two is as Figure 12As shown. Specifically, when the AES cores U1 and U2 perform encryption or decryption tasks, they work in cooperation but independently process the data blocks allocated to them. Since they share the same key, IV, and counter, they can maintain data consistency and security. This parallel processing mechanism enables the system to process more data at the same time, thus greatly improving the overall processing speed and efficiency. In this way, the system can better meet the encryption or decryption requirements of a large amount of data, ensuring the efficiency and real-time nature of data processing.
[0118] The key_len signal received by the top-level module from the outside world passes the key length key_len to the U2 instance. In AES encryption, the key length determines the behavior of the key scheduling algorithm. The key_ready signal is passed from the U1 instance to the U2 instance to indicate whether the key is ready. The keymem_sboxw signal passes the result of the S-box (Substitution-Box) write operation to the U2 instance. In the AES key scheduling process, the S-box is used for non-linear transformation of the key bytes, and it is calculated during the key expansion process and used to generate round keys. The round_key signal passes the generated round key round_key to the U2 instance, and this signal indicates whether the key scheduling has been completed and the AES core is ready for encryption or decryption operations.
[0119] Among them, the value of keymem_sboxw comes from the output of the key expansion module. The key expansion is responsible for generating round keys according to the input key, key length, and initialization signal, and outputs these round keys and related S-box values. Specifically, keymem_sboxw is a 32-bit wide signal output by the key expansion module, representing one of the S-box values used in the AES algorithm. These S-box values are used for subsequent byte substitution operations, which is one of the key steps in the AES algorithm. round_key is a 128-bit wide signal output by the key expansion module, representing the current round key, which is used for the round key addition operation in the AES algorithm.
[0120] To further improve weak scalability, the present invention designs an efficient AES-GCM processing array. This array consists of 8 independent AES-GCM modules, each of which can process data independently. Through this design, the speed and efficiency of data processing can be significantly improved. Specifically, each AES-GCM module can perform AES encryption and GCM authentication operations, thus achieving high-speed data processing. The theoretical processing speed of the entire array can reach approximately 12.5 - 12.8 Gbps. This means that in practical applications, the present invention can process a large amount of data without bottlenecks. This design not only improves the speed of data processing but also has good scalability. Since each module is independent, the present invention can further improve the processing speed by adding more modules. This weak scalability enables our system to flexibly handle different application scenarios and requirements.
[0121] The present invention uses AES-GCM modules to specifically demonstrate Figure 13 the encryption array in. This encryption array contains eight mutually independent AES-GCM encryption and decryption modules, and these modules have two independent encryption and decryption functions. In the figure, the present invention uses a control signal Enc or Dec to indicate whether the next switch or node should decrypt or re-encrypt the data. In the program, this control signal is represented by a 4-bit control signal Ctrl, where the third bit is used to indicate decryption (1) or encryption (0). In addition, the first bit of Ctrl is used to control whether the AES-GCM module is enabled.
[0122] Since the GHASH function of GCM uses a 128-bit finite field multiplier, and through in-depth research on the finite field multiplier, it can be known that the unit with the largest resource consumption in the GCM system is the finite field multiplier, and its performance also becomes one of the limiting factors of the overall system performance. Given the progress of elliptic curve cryptosystems and the requirements of various cryptographic algorithms, the finite field multiplier has always been the focus and challenge in the field of cryptographic hardware research. To adapt to resource-constrained environments, it is particularly necessary to solve the problem of excessive resource consumption of the finite field multiplier.
[0123] Previously, the prior art has tried to optimize the GF(2 128 ) module. For example, by using the matrix method to combine the multiplication and modulo reduction steps of finite field multiplication, a new reduction matrix formula in the PB form is derived, and based on this, a finite field multiplier scheme with a low-complexity bit-parallel architecture based on PB is proposed. Although this scheme can complete the finite field multiplication of two 128-bit inputs in a single cycle, it consumes 7,261 lookup tables (LUTs). Another example is to use a GF(2 128) A multiplier was proposed, and multiplier schemes of single-cycle parallelization and four-cycle serialization were proposed simultaneously to reduce hardware overhead. After synthesis in Vivado, the multiplier of this structure still consumes 8,319 lookup tables (LUTs).
[0124] Although the above scheme can significantly reduce the calculation cycles of the GF(2 128 ) multiplier, its resource consumption is large, and it is difficult to meet the requirements of the lightweight AES-GCM encryption module for on-chip area. Therefore, in the present invention, the operations of AES and GCM are made independent of each other, and an array design is still adopted in the structure. As Figure 14 shown, the ciphertext C generated by GCTR is XORed with Hash_key and then input into the GHASH function. The calculated intermediate value ghash_result is output again and XORed with the next batch of ciphertext until all data is processed.
[0125] As Figure 15 shown, in terms of the process, a two-stage pipeline method is adopted to achieve parallel data processing of AES and GCM, so as to reduce the overall time cost as much as possible. The GCTR module transfers the encrypted ciphertext to the GHASH function. During the process of the GHASH function calculating and generating the TAG, the GCTR module continues to process the subsequent data; and after all data is processed, the GHASH function returns a TAG for the receiving party to detect whether the data is correct.
[0126] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hardware AES-GCM system based on FPGA, characterized in that: Several AES-GCM modules are deployed on the FPGA hardware chip. The operation of each AES-GCM module is independent of each other and no data transmission is performed. Among them, each AES-GCM module includes a GCTR module and a GHASH module. The GCTR module integrates at least two AES modules inside, and the GCM module is implemented by the GHASH module. The GCTR module performs an XOR operation on the generated ciphertext and Hash_key and then inputs it into the GHASH module for authentication processing, generating a TAG tag to verify whether the data has been tampered with during the encryption process. The operations of the AES module and the GCM module are independent of each other, and the structure adopts an array design structure; The array design structure is: when the AES module transmits the output result to the GCM module for processing, the AES module continues to process the next batch of data; The process adopts a two-level pipeline method; the GCTR module passes the encrypted ciphertext to the GHASH function. While the GHASH function is calculating and generating TAG, the GCTR module continues to process the next batch of data; and after all data are processed, the GHASH function returns TAG so that the receiver can check whether the data is correct; The AES-GCM processing array consists of 8 independent AES-GCM modules, each of which can process data independently. A multi-AES core structure with master and slave core characteristics is designed. AES cores U1 and U2 work together when performing encryption or decryption tasks, but each processes the data blocks assigned to them independently. They share the same key, IV and counter to maintain data consistency and security. In this way, the system can cope with the encryption or decryption needs of large amounts of data; The key_len signal received by the top-level module from the outside world passes the key length key_len to the U2 instance; in AES encryption, the key length determines the behavior of the key scheduling algorithm. The key_ready signal is passed from the U1 instance to the U2 instance to indicate whether the key is ready. The keymem_sboxw signal passes the result of the S-box write operation to the U2 instance. In the AES key scheduling process, the S-box is used to perform nonlinear transformations on the key bytes. It is calculated during the key expansion process and used to generate round keys; the round_key signal passes the generated round key round_key to the U2 instance. This signal indicates whether the key scheduling has been completed and the AES core is ready for encryption or decryption operations.
2. The FPGA-based hardware AES-GCM system according to claim 1, characterized in that: Each of the AES modules is responsible for the encryption or decryption of a portion of data.
3. The FPGA-based hardware AES-GCM system according to claim 1, characterized in that: Expand each AES module from single channel to four channels.
4. The FPGA-based hardware AES-GCM system according to claim 1, characterized in that: Each of the GCM modules includes a plurality of GF multipliers.
5. The FPGA-based hardware AES-GCM system according to claim 1, characterized in that: Each of the AES modules is used to perform byte substitution, row shift, column mixing and round key addition operations.
6. The FPGA-based hardware AES-GCM system according to claim 1, characterized in that: Each of the AES-GCM modules has independent encryption and decryption functions.
7. A method for a hardware AES-GCM system based on FPGA as claimed in any one of claims 1 to 6, characterized in that: include: A set of input data is divided into blocks, and different data blocks enter different AES-GCM modules of the array for processing; In an AES-GCM module, the two four-channel AES cores contained in the GCTR function are used to process the AES encryption of data, and the GHASH function is used to encrypt and authenticate the AES-encrypted data; at the same time, the GCTR and GHASH functions form a two-level pipeline to process the data entering the AES-GCM module.
Citation Information
Patent Citations
IEEE802.1AE protocol-based GCM high-speed encryption and decryption equipment
CN101827107A