FPGA (Field Programmable Gate Array)-based national cryptographic algorithm acceleration device and electronic equipment

Through the design of FPGA hardware acceleration platform and PCIE interface, the problem of insufficient computing performance of Guomi algorithm on the CPU is solved, efficient data interaction and algorithm acceleration are achieved, and it is suitable for blockchain, finance and data security fields.

CN120498669APending Publication Date: 2025-08-15SHANGHAI JIAOTONG UNIV

Patent Information

Application Number
CN202510648068.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the Guomi algorithm has insufficient computing performance when implemented on the CPU, especially in large-scale data processing and high concurrent tasks, and how to seamlessly integrate with existing blockchain systems is a difficult problem.

Method used

The FPGA hardware acceleration platform is adopted, combined with the PCIE interface and the AXI bus protocol, and the Guomi algorithm acceleration device is designed, including the PCIE bus processing module, the AXI bus processing module, the storage module and the Guomi algorithm control module to achieve efficient data interaction and algorithm acceleration.

Benefits of technology

Significantly improve the computing performance of Guoxin algorithm, reduce latency, improve resource utilization, and seamlessly integrate with the host system through PCIE interface, suitable for blockchain, finance and data security fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498669A_ABST
    Figure CN120498669A_ABST
Patent Text Reader

Abstract

The invention provides an FPGA (Field Programmable Gate Array)-based national secret algorithm acceleration device and electronic equipment, the device comprises a PCIE (Peripheral Component Interface Express) bus processing module, the PCIE bus processing module is used for completing data packaging and unpackaging work of each layer of a PCIE protocol, and the PCIE bus processing module further provides an AXI4 interface and supports AXIFULL and AXILITE protocols; the AXI bus processing module comprises an AXIFULL protocol processing module and an AXILITE protocol processing module; the storage module adopts a high-capacity full-dual-port BRAM (Border Random Access Memory) in the FPGA and is used for storing intermediate data and input and output data; the national cryptographic algorithm control module is responsible for managing and scheduling operation of three national cryptographic algorithm sub-modules SM2, SM3 and SM4 and processing data interaction between the sub-modules and the host; and the hardware acceleration unit comprises three independent hardware acceleration modules SM2, SM3 and SM4, each hardware acceleration module is responsible for a calculation task of a corresponding algorithm, and each hardware acceleration module comprises a control interface and a BRAM read-write interface for instruction and data interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of algorithm operation processing, and in particular to an FPGA-based national secret algorithm acceleration device and electronic equipment. Background Art

[0002] With the rapid growth of blockchain technology, the financial sector, and data security needs, encryption technology has become increasingly important in protecting data privacy and ensuring information security. However, these sectors have long relied on internationally accepted cryptographic algorithms (such as RSA, SHA-1, and 3DES), the core technologies of which are often controlled by foreign countries. To this end, my country has independently developed national cryptographic algorithms (abbreviated as "national cryptographic algorithms"), which include asymmetric encryption algorithms SM2, symmetric encryption algorithms SM4, and cryptographic hash algorithms SM3. These algorithms incorporate the characteristics of specific Chinese application scenarios, offering high security and computational efficiency. They have become core standards for various secure communications and data protection applications and have been incorporated into the ISO / IEC international standards system.

[0003] In traditional software implementations, encryption algorithms typically run on general-purpose central processing units (CPUs). While this approach offers high flexibility and facilitates development, iteration, and maintenance, its computational performance is highly dependent on the computing power of general-purpose processors. Due to a lack of architectural designs optimized for encryption algorithms, CPUs often suffer from high computational latency and insufficient throughput when processing large amounts of data or highly concurrent tasks. Therefore, hardware acceleration has become an effective solution to this problem. Field-programmable gate arrays (FPGAs), in particular, as a highly customizable hardware platform with powerful parallel processing capabilities, have become an ideal choice for accelerating cryptographic algorithm computations. FPGA hardware acceleration can significantly improve the computational speed of encryption algorithms, reduce latency, and achieve low power consumption and efficient resource utilization.

[0004] However, in the process of implementing hardware acceleration for national cryptographic algorithms, a crucial issue is how to efficiently integrate the acceleration module into existing blockchain systems and ensure seamless integration with upper-layer applications. To this end, selecting an efficient communication interface is particularly important. The PCIE (Peripheral Component Interconnect Express) protocol, as a high-bandwidth, low-latency communication protocol, is particularly well-suited for high-performance computing requirements and data-intensive applications, and is widely supported by almost all modern computer motherboards and servers. Thanks to the standardized nature of the PCIE interface, hardware acceleration modules can be seamlessly integrated into existing systems, significantly reducing the complexity of system modifications while maintaining the integrity of the host architecture. Summary of the Invention

[0005] In response to the defects in the prior art, the purpose of the present invention is to provide a national secret algorithm acceleration device and electronic equipment based on FPGA, which significantly improves the computing performance of the national secret algorithms SM2, SM3 and SM4 by optimizing the hardware architecture and interface design, reduces latency, improves resource utilization, and conducts efficient data interaction with the host through a high-speed PCIE interface.

[0006] In order to solve the above problems, the technical solution of the present invention is:

[0007] An FPGA-based national secret algorithm acceleration device, comprising:

[0008] The PCIE bus processing module is used to complete the data packaging and unpacking work of each layer of the PCIE protocol and integrate the DMA engine. The PCIE bus processing module also provides an AXI4 interface to realize the conversion between the PCIE bus and the AXI bus protocol inside the FPGA, supporting AXI_FULL and AXI_LITE protocols;

[0009] AXI bus processing module, including AXI_FULL and AXI_LITE protocol processing modules. The AXI_FULL protocol is used for large-scale data exchange and high-speed memory connection. It dynamically adjusts the data transmission path to achieve concurrent communication and load balancing between multiple master devices and slave devices. The AXI_LITE protocol is used to control the logic modules within the FPGA;

[0010] The storage module uses the large-capacity full-dual-port BRAM inside the FPGA to store intermediate data and input and output data;

[0011] The national secret algorithm control module is responsible for managing and scheduling the operations of the three national secret algorithm sub-modules SM2, SM3, and SM4, and processing the data interaction between the sub-modules and the host;

[0012] The hardware acceleration unit includes three independent hardware acceleration modules: SM2, SM3, and SM4. Each hardware acceleration module is responsible for the calculation tasks of the corresponding algorithm. Each hardware acceleration module includes a control interface and a BRAM read and write interface for instruction and data interaction.

[0013] Preferably, the large-capacity full-dual-port BRAM has one group of ports connected to the AXI BRAM Controller module, converting the AXI_FULL protocol into BRAM read and write instructions through the AXI protocol converter, providing data storage and exchange channels, and supporting data transmission in DMA mode; the other group of ports uses a bare BRAM interface for user logic to directly access the BRAM.

[0014] Preferably, the national secret algorithm control module includes two groups of external interfaces, AXI_LITE and BRAM, register processing logic, register space, memory access control logic and sub-algorithm control logic;

[0015] The AXI_LITE interface is used to exchange control signals with the host and configure the acceleration device through read and write operations on the control register, including selecting the algorithm to be executed, configuring input and output parameters, and monitoring the execution status;

[0016] The BRAM interface is used for large-scale data exchange to ensure the order of algorithm execution and the correctness of data flow;

[0017] The register space design is divided into two categories: global registers and sub-module registers. The global register is responsible for managing the overall configuration and operating status of the national secret algorithm module, supporting the collaborative work of different algorithm modules, and providing global operating information to facilitate the host's operation and management of the module; the sub-module registers are designed for the independent needs of the three algorithm modules SM2, SM3 and SM4, respectively, to support their respective core functions;

[0018] The memory access control logic is responsible for coordinating BRAM read and write requests from different submodules to ensure the orderliness of data access and efficient use of resources;

[0019] The control logic of the sub-algorithm can dynamically schedule and manage the execution process of each algorithm module according to the control signal generated by the register processing logic.

[0020] Preferably, the SM2 algorithm hardware acceleration module includes a BRAM read-write interface and a control interface. The BRAM read-write interface is used for data interaction with the host system; the control interface realizes the control and status monitoring of the SM2 algorithm by the upper-layer control system; the overall control logic of the SM2 algorithm is designed in the form of a two-stage finite state machine to complete the protocol layer function, key generation and management function, point multiplication operator module and finite field operator module.

[0021] Preferably, the external interface of the SM3 algorithm hardware acceleration module includes a control interface and a BRAM read-write interface, and the internal interface includes message filling, message expansion and iterative compression processes. The function of message filling is to fill the input message to meet the requirements of the SM3 algorithm for the data block format. The iterative compression process is responsible for managing each round of the compression process, ensuring that the output of the message expansion module can be correctly passed to the iterative calculation unit, and completing 64 rounds of iterations in the order defined by the algorithm.

[0022] Preferably, the external interface of the SM4 algorithm hardware acceleration module includes a control interface and a BRAM read-write interface; the internal key expansion module is responsible for generating 32 round keys based on the 128-bit master key. Its design goal is to ensure that the expanded round keys have strong security and randomness, while efficiently supporting the operation of the hardware pipeline; the task of the internal round function module is to perform 32 rounds of iterative operations on the input data. In order to improve the data throughput, 32 round function sub-modules are instantiated, each module is responsible for one round of SM4 operation, and the modules are cascaded through registers to achieve pipeline processing.

[0023] Furthermore, the present invention also provides an electronic device, comprising the FPGA-based national encryption algorithm acceleration device as described above.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] 1. The present invention adopts FPGA as the hardware acceleration platform. By optimizing the hardware architecture and interface design, it realizes the efficient acceleration of the national secret algorithms SM2, SM3, and SM4, and significantly improves the computing performance of the national secret algorithms.

[0026] 2. The present invention realizes efficient data interaction with the host system through the PCIE interface, is easy to be integrated into the existing system, and can meet the needs of different application scenarios.

[0027] 3. This invention can be widely used in blockchain, finance, data security and other fields, providing strong hardware support for the promotion and application of domestic cryptographic algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0029] Figure 1 This is the architectural design diagram of the FPGA-based national secret algorithm acceleration device of the present invention;

[0030] Figure 2 This is the structural diagram of the SM2 algorithm hardware acceleration module;

[0031] Figure 3 This is the flow chart of the point multiplication algorithm in the SM2 algorithm;

[0032] Figure 4 This is the structural diagram of the SM3 algorithm hardware acceleration module;

[0033] Figure 5 This is the structural diagram of the SM4 algorithm hardware acceleration module. DETAILED DESCRIPTION

[0034] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.

[0035] Specifically, the present invention provides a national secret algorithm acceleration device based on FPGA, such as Figure 1 As shown, the device includes: a PCIE bus processing module for efficient data exchange and protocol conversion; an AXI bus processing module for achieving high-speed data transmission and load balancing; a storage module using a large-capacity full-dual-port BRAM to provide efficient data storage and access; a national secret algorithm control module responsible for algorithm scheduling, data interaction and error handling; a hardware acceleration unit including independent SM2, SM3, and SM4 acceleration modules to achieve parallel algorithm computing.

[0036] The PCIE bus processing module is used to complete the data packaging and unpacking work of each layer of the PCIE protocol, and integrates a DMA engine for efficient large-scale data exchange; the PCIE bus processing module also provides an AXI4 interface, which can realize the conversion between the PCIE bus and the AXI bus protocol inside the FPGA, and supports AXI_FULL and AXI_LITE protocols.

[0037] The AXI bus processing module includes AXI_FULL and AXI_LITE protocol processing modules. The AXI_FULL protocol is used for large-scale data exchange and high-speed memory connection. It realizes concurrent communication and load balancing between multiple master devices and slave devices by dynamically adjusting the data transmission path, thereby improving the data throughput of the system; the AXI_LITE protocol is used to accurately control the logic module inside the FPGA, and is suitable for efficient interaction of small data volumes such as control instructions and status queries.

[0038] The storage module uses the FPGA's internal large-capacity full-dual-port BRAM to store intermediate data and input and output data. The FPGA's internal acceleration module exchanges data with the host through the BRAM, supporting efficient parallel read and write operations and avoiding access bottlenecks caused by external memory.

[0039] Furthermore, one set of ports in the large-capacity, full-dual-port BRAM is connected to the AXI BRAM Controller module. Using an AXI protocol converter, it converts AXI_FULL protocol into BRAM read and write instructions, providing efficient data storage and exchange channels and supporting data transfer in DMA mode. The other set of ports uses a raw BRAM interface, allowing user logic to directly access the BRAM, providing lower latency access.

[0040] The national secret algorithm control module is responsible for managing and scheduling the operations of the three national secret algorithm sub-modules SM2, SM3, and SM4, and processing data interaction between the sub-modules and the host; the national secret algorithm control module is also responsible for the error detection and recovery mechanism to ensure the stable operation of the accelerator card under high load conditions.

[0041] Furthermore, the national secret algorithm control module includes two sets of external interfaces: AXI_LITE and BRAM; register processing logic; register space; memory access control logic; and sub-algorithm control logic. The AXI_LITE interface is used to exchange control signals with the host and configure the accelerator card through read and write operations on the control registers, including selecting the algorithm to be executed, configuring input and output parameters, and monitoring execution status. The BRAM interface is used for large-scale data exchange, ensuring the order of algorithm execution and the correctness of data flow. The register space is designed into two categories: global registers and sub-module registers. Global registers are responsible for managing the overall configuration and operating status of the national secret algorithm module, supporting the collaborative operation of different algorithm modules and providing global operating information to facilitate host operation and management of the module. Sub-module registers are designed for the independent requirements of the three algorithm modules SM2, SM3, and SM4, supporting their respective core functions. The memory access control logic is responsible for coordinating BRAM read and write requests from different sub-modules, ensuring orderly data access and efficient resource utilization. The sub-algorithm control logic dynamically schedules and manages the execution of each algorithm module based on the control signals generated by the register processing logic, effectively avoiding resource conflicts and performance bottlenecks between modules. In addition, the system should be able to dynamically select the execution algorithm according to different needs and adjust related parameters to meet the requirements of various application scenarios.

[0042] The hardware acceleration unit specifically includes three independent hardware acceleration modules: SM2, SM3, and SM4. Each hardware acceleration module is responsible for the calculation task of the corresponding algorithm. The three modules are independent of each other and have their own external control interface and BRAM read and write interface for instruction and data exchange.

[0043] Specifically, the SM2 algorithm hardware acceleration module includes a BRAM read-write interface and a control interface. The BRAM read-write interface is used to interact with the host system for data; the control interface enables the upper control system to control and monitor the status of the SM2 algorithm.

[0044] like Figure 2 As shown, the SM2 algorithm hardware acceleration module adopts a layered architecture design, including a protocol layer, key generation and management functions, a point multiplication operator module and a finite field operator module.

[0045] The protocol layer includes digital signatures, digital signature verification, key exchange, and asymmetric encryption and decryption. Digital signatures are based on the SM2-1 standard, implementing signature generation and verification logic. Signature generation: Calculates the hash value e, generates a random number k, and performs a point multiplication k·G to obtain the elliptic curve point (x1, y1), ultimately outputting the signature (r, s). Signature verification: Verifies that r ≡ (e + x1) mod n holds through a point multiplication. Key exchange: Utilizes the ECMQV protocol, integrating temporary key pair generation and shared key calculation logic.

[0046] like Figure 3 As shown, the point multiplication algorithm optimization adopts the binary expansion method to realize the point multiplication operation, and reduces the number of modular inverse operations through Jacobian coordinate transformation. The specific process is as follows:

[0047] (1) Initialize point P to Jacobian coordinates (X, Y, Z);

[0048] (2) Traverse each bit of the key k and select the point addition or point doubling operation based on the conditional judgment;

[0049] (3) The final result is converted to affine coordinates (x, y) through modular inverse operation, and the output is Q = k·P;

[0050] The external interface of the SM3 algorithm hardware acceleration module also includes a control interface and a BRAM read and write interface, and the internal process mainly includes message filling, message expansion and iterative compression, such as Figure 4 As shown in the figure, the core process of the SM3 algorithm hardware acceleration module is as follows:

[0051] (1) Message padding: The input message is padded to an integer multiple of 512 bits to meet the data block format requirements of the SM3 algorithm. Padding rules include appending bits of 1, 0, and appending 64 bits of message length. A pipelined padding unit is used to support parallel processing of consecutive message blocks, avoiding additional delays caused by padding operations.

[0052] (2) Message expansion: The 512-bit message group is divided into 132 32-bit words (W0~W67, W'0~W'63), and the message expansion is achieved through nonlinear functions.

[0053] (3) Iterative compression: Each round of iteration includes bitwise operations, modular addition operations, and register updates, totaling 64 rounds. The final hash value is generated by concatenating registers A to H after 64 rounds of iteration, and the output is a 256-bit hash result.

[0054] like Figure 5As shown, the external interface of the SM4 algorithm hardware acceleration module also includes a control interface and a BRAM read-write interface; the internal key expansion module is responsible for generating 32 round keys based on the 128-bit master key. Its design goal is to ensure that the expanded round keys have strong security and randomness, while efficiently supporting the operation of the hardware pipeline; the task of the internal round function module is to perform 32 rounds of iterative operations on the input data. In order to improve the data throughput, an optimized architecture based on the pipeline is designed, that is, 32 round function sub-modules are instantiated, each module is responsible for one round of SM4 operation, and the modules are cascaded through registers to achieve pipeline processing.

[0055] According to another embodiment of the present invention, the present invention also provides an electronic device, including the FPGA-based national encryption algorithm acceleration device as described above.

[0056] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A national secret algorithm acceleration device based on FPGA, characterized in that: The device comprises: The PCIE bus processing module is used to complete the data packaging and unpacking work of each layer of the PCIE protocol and integrate the DMA engine. The PCIE bus processing module also provides an AXI4 interface to realize the conversion between the PCIE bus and the AXI bus protocol inside the FPGA, supporting AXI_FULL and AXI_LITE protocols; AXI bus processing module, including AXI_FULL and AXI_LITE protocol processing modules. The AXI_FULL protocol is used for large-scale data exchange and high-speed memory connection. It dynamically adjusts the data transmission path to achieve concurrent communication and load balancing between multiple master devices and slave devices. The AXI_LITE protocol is used to control the logic modules within the FPGA; The storage module uses the large-capacity full-dual-port BRAM inside the FPGA to store intermediate data and input and output data; The national secret algorithm control module is responsible for managing and scheduling the operations of the three national secret algorithm sub-modules SM2, SM3, and SM4, and processing the data interaction between the sub-modules and the host; The hardware acceleration unit includes three independent hardware acceleration modules: SM2, SM3, and SM4. Each hardware acceleration module is responsible for the calculation tasks of the corresponding algorithm. Each hardware acceleration module includes a control interface and a BRAM read and write interface for instruction and data interaction.

2. The FPGA-based national secret algorithm acceleration device according to claim 1, characterized in that: The large-capacity full-dual-port BRAM has one group of ports connected to the AXI BRAM Controller module, converting the AXI_FULL protocol into BRAM read and write instructions through the AXI protocol converter, providing data storage and exchange channels, and supporting data transmission in DMA mode; the other group of ports uses a bare BRAM interface for user logic to directly access the BRAM.

3. The FPGA-based national secret algorithm acceleration device according to claim 1, characterized in that: The national secret algorithm control module includes two groups of external interfaces, AXI_LITE and BRAM, register processing logic, register space, memory access control logic and sub-algorithm control logic; The AXI_LITE interface is used to exchange control signals with the host and configure the acceleration device through read and write operations on the control register, including selecting the algorithm to be executed, configuring input and output parameters, and monitoring the execution status; The BRAM interface is used for large-scale data exchange to ensure the order of algorithm execution and the correctness of data flow; The register space design is divided into two categories: global registers and sub-module registers. The global register is responsible for managing the overall configuration and operating status of the national secret algorithm module, supporting the collaborative work of different algorithm modules, and providing global operating information to facilitate the host's operation and management of the module; the sub-module registers are designed for the independent needs of the three algorithm modules SM2, SM3 and SM4, respectively, to support their respective core functions; The memory access control logic is responsible for coordinating BRAM read and write requests from different submodules to ensure the orderliness of data access and efficient use of resources; The control logic of the sub-algorithm can dynamically schedule and manage the execution process of each algorithm module according to the control signal generated by the register processing logic.

4. The FPGA-based national secret algorithm acceleration device according to claim 1, characterized in that: The SM2 algorithm hardware acceleration module includes a BRAM read-write interface and a control interface. The BRAM read-write interface is used for data interaction with the host system; the control interface realizes the control and status monitoring of the SM2 algorithm by the upper-level control system; the overall control logic of the SM2 algorithm is designed in the form of a two-stage finite state machine to complete the protocol layer functions, key generation and management functions, point multiplication operator module and finite field operator module.

5. The FPGA-based national secret algorithm acceleration device according to claim 1 is characterized in that: The external interfaces of the SM3 algorithm hardware acceleration module include a control interface and a BRAM read-write interface, and the internal interfaces include message filling, message expansion, and iterative compression processes. The function of message filling is to fill the input message to meet the requirements of the SM3 algorithm for the data block format. The iterative compression process is responsible for managing each round of the compression process, ensuring that the output of the message expansion module can be correctly passed to the iterative calculation unit and completing 64 rounds of iterations in the order defined by the algorithm.

6. The FPGA-based national secret algorithm acceleration device according to claim 1, characterized in that: The external interfaces of the SM4 algorithm hardware acceleration module include a control interface and a BRAM read-write interface; the internal key expansion module is responsible for generating 32 round keys based on the 128-bit master key. Its design goal is to ensure that the expanded round keys have strong security and randomness, while efficiently supporting the operation of the hardware pipeline; the task of the internal round function module is to perform 32 rounds of iterative operations on the input data. In order to improve the data throughput, 32 round function sub-modules are instantiated, each module is responsible for one round of SM4 operations, and the modules are cascaded through registers to achieve pipeline processing.

7. An electronic device, characterized in that: Including the FPGA-based national secret algorithm acceleration device as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Verification system for cryptographic algorithms

    CN106650411A

  • A national cryptographic algorithm acceleration processing system based on an FPGA

    CN109902043A

  • National cryptographic algorithm SM4 acceleration processing method and system based on high-level integration

    CN111914307A

  • FPGA optimization implementation method and system for SM4 cryptographic algorithm and application

    CN113078996A

  • Configurable secret key SM4 encryption and decryption system based on FPGA

    CN116506106A

Cited By

  • Encrypted storage device and NVMe over TCP write-in and read request processing method

    CN121578941A