SM4 hardware acceleration and anti-side channel protection method and system for domestic FPGA

Through the pre-built PCIe bus connection and dynamic resource configuration in FPGA, combined with the comprehensive use of acceleration blocks and guard blocks, the configuration problem of SM4 algorithm in FPGA is solved, efficient encryption processing and anti-side channel protection are achieved, and the performance and security of the system are improved.

CN120358028AActive Publication Date: 2025-07-22JINAN UNIVERSITY

Patent Information

Application Number
CN202510845836.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the prior art, the configuration of SM4 and anti-side channel protection in FPGAs usually relies on subjective adjustment by technicians, resulting in waste of performance and poor dynamic protection capabilities, making it difficult to meet high throughput and security needs.

Method used

The acceleration protection FPGA is adapted to the target computer through the pre-constructed PCIe bus connection, and the acceleration block is used to identify data types and configure dynamic resources, and combine the protection block to perform physical layer suppression leakage, logical layer obfuscation association and system layer dynamic response, so as to realize the acceleration and anti-side channel protection of the SM4 algorithm.

Benefits of technology

Improves encryption processing performance and security, meets high throughput requirements, and effectively protects side channel attacks to ensure overall system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358028A_ABST
    Figure CN120358028A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of FPGA architecture, in particular to an SM4 hardware acceleration and anti-side-channel protection method and system of a domestic FPGA, which comprises the following steps of: performing data type identification operation on bright-ciphertext batch data by utilizing an acceleration block to obtain a data type, and according to a pre-constructed parallelization strategy and the data type, performing data type identification operation on the SM4 hardware acceleration and anti-side-channel protection on the SM4 hardware acceleration and anti-side-channel protection on the SM4 hardware acceleration and anti-side-channel protection system; performing dynamic resource configuration on the plaintext-ciphertext batch data to obtain resource configuration information, and performing configuration according to the resource configuration information, the key-mode parameter and the dynamic control instruction to obtain a configured SM4 algorithm; according to the configured SM4 algorithm, performing encryption processing on the plaintext-ciphertext batch data to obtain to-be-protected encrypted data; and performing protection operation based on physical layer leakage suppression, logic layer confusion association and system layer dynamic response on key-mode parameters, dynamic control instructions, plaintext-ciphertext batch data and encryption processing by using the protection block, and converting to-be-protected encrypted data into protected encrypted data. According to the invention, the encryption processing performance and security can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of FPGA architectures, and in particular to a method and system for SM4 hardware acceleration and side-channel attack protection of domestic FPGAs. Background Art

[0002] As a commercial cryptographic standard algorithm in China, the hardware acceleration and side-channel attack protection of SM4 are the core requirements for improving the security and efficiency of cryptographic systems in FPGAs. Hardware acceleration implements algorithm logic through dedicated circuits, which can increase the operation speed by dozens of times compared to software execution, meet the high-throughput requirements of real-time sensitive scenarios such as the Internet of Things and financial payments, and at the same time reduce the overall power consumption of the system. Among them, side-channel attack protection targets side-channel attack means such as power consumption analysis and electromagnetic radiation, and adopts protection mechanisms such as masking technology and randomized execution paths to prevent key information from leaking through physical side channels.

[0003] However, in specific usage scenarios, the configuration of SM4 and side-channel attack protection is usually used by default. To adjust the FPGA speed-up configuration, it depends on the subjective configuration of technicians and has a high usage threshold, resulting in performance waste and poor dynamic protection capabilities in the process of using SM4 and side-channel attack protection. Summary of the Invention

[0004] The present invention provides a method for SM4 hardware acceleration and side-channel attack protection of domestic FPGAs, and its main purpose is to improve the encryption processing performance and security.

[0005] To achieve the above object, a method for SM4 hardware acceleration and side-channel attack protection of domestic FPGAs provided by the present invention includes: Using a pre-built PCIe bus to connect a pre-built acceleration and protection FPGA to a pre-built target computer, where the acceleration and protection FPGA includes an acceleration block and a protection block, and the target computer includes a CPU and a DDR; Obtaining the key-mode parameters and dynamic control instructions generated in the CPU, and obtaining the plain-ciphertext batch data in the DDR, where the key-mode parameters include the main key, encryption mode parameters, and function switching signals, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; Using the acceleration block to perform data type identification operations on the plain-ciphertext batch data to obtain the data type, and dynamically configuring the plain-ciphertext batch data according to a pre-built parallelization strategy and the data type to obtain resource configuration information, and configuring the configured SM4 algorithm according to the resource configuration information, key-mode parameters, and dynamic control instructions; According to the configured SM4 algorithm, encrypt the plaintext-ciphertext batch data to obtain the encrypted data to be protected; Using the protection block, perform protection operations on the key-mode parameters, dynamic control instructions, plaintext-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transform the encrypted data to be protected into the encrypted data that has been protected.

[0006] Optionally, obtaining the key-mode parameters and dynamic control instructions generated in the CPU, and obtaining the plaintext-ciphertext batch data in the DDR includes: Using a pre-built AXI-Lite bus, send the key-mode parameters and dynamic control instructions generated in the CPU to the pre-built control register in the acceleration block; Using a pre-built AXI-Stream bus, send the plaintext-ciphertext batch data pre-fetched from the DDR by the CPU to the pre-built input buffer module in the acceleration block.

[0007] Optionally, performing a data type recognition operation on the plaintext-ciphertext batch data to obtain a data type includes: Perform data block meta-information analysis on the plaintext-ciphertext batch data to obtain the data block size distribution and transmission interval characteristics; Use a pre-built sliding time window to perform throughput and latency correlation analysis on the plaintext-ciphertext batch data to obtain the data processing volume and average response latency; According to a preset weight coefficient sequence, perform feature weighting calculation on the data block size distribution, transmission interval characteristics, data processing volume, and average response latency to obtain a data type score; Obtain the size relationship between the data type score and a pre-built classification threshold, and obtain the data type according to the size relationship, where the data type includes a large data block type and a small data block type.

[0008] Optionally, performing dynamic resource allocation on the plaintext-ciphertext batch data according to a pre-built parallelization strategy and the data type to obtain resource allocation information includes: According to a pre-built parallelization strategy, determine whether the data type is a large data block type or a small data block type; When the data type is a large data block type, determine that the resource allocation information is the first type of configuration information; When the data type is a small data block type, determine that the resource allocation information is the second type of configuration information; Among them, the first type of configuration information includes using the CPU to obtain task-related data, and using the PCIe bus to pre-load the task-related data into the DDR for storage, and using a pre-built FPGA multi-core parallel architecture and 32-level pipeline design to process the plaintext-ciphertext batch data; Among them, the second type of configuration information includes using a pre-built bit-slice service to reorganize the plaintext-ciphertext batch data into a bit width supported by a preset SIMD instruction, obtaining a data grouping result, and performing batch processing on multiple groups in the data grouping result.

[0009] Optionally, configuring the configured SM4 algorithm according to the resource configuration information, key-mode parameter, and dynamic control instruction includes: Performing an acceleration architecture selection operation on the pre-built SM4 algorithm based on the resource configuration information; Parsing the key-mode parameter to obtain a master key, an encryption mode parameter, and a function switching signal, and performing a key expansion operation on the master key based on an overlay random mask according to a pre-built SM4 key expansion instruction to obtain round keys, and storing the round keys in a pre-built secure storage area; Configuring the encryption process in the SM4 algorithm according to the encryption mode parameter, and configuring the working logic in the SM4 algorithm according to the function switching signal; Configuring the combination mode of the pipeline stage number and the parallel core number in the SM4 algorithm according to the dynamic control instruction to obtain the configured SM4 algorithm.

[0010] Optionally, using the protection block to perform protection operations on the key-mode parameter, dynamic control instruction, plaintext-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, including: Using the protection block to perform a signal weakening operation based on power consumption and electromagnetic radiation characteristics, and a correlation perturbation operation based on timing and clock jitter on the encryption process to complete the physical layer leakage suppression process; Performing an intermediate value randomization operation on the encryption process to complete the logical layer confusion association process; Performing a system layer dynamic response process on the AXI-Lite bus and AXI-Stream bus based on abnormal event detection and response.

[0011] Optionally, using the protection block to perform a signal weakening operation based on power consumption and electromagnetic radiation characteristics, and a correlation perturbation operation based on timing and clock jitter on the encryption process, including: Use the pre-built adjustable load circuit in the protection block to perform power consumption difference balancing operation in the operation stage of the encryption process; Use the pre-built metal shielding layer to cover the preset high-leakage risk area in the configured SM4 algorithm; Use the pre-built true random number to perform timing perturbation operation on the data transmission in the AXI-Lite bus and AXI-Stream bus.

[0012] Optionally, the intermediate value randomization operation on the encryption process is performed to complete the logical layer confusion association process, including: Perform threshold segmentation on the preset intermediate value in the encryption process, where the intermediate value includes the S-box output and the round function intermediate value; Use the pre-built dual-rail redundant calculation module in the protection block to perform real-time comparison of output consistency on the encryption process to complete the logical layer confusion association process.

[0013] Optionally, the system layer dynamic response process for the AXI-Lite bus and AXI-Stream bus based on abnormal event detection and response includes: Use the pre-built CNN attack recognition model in the protection block to perform attack mode recognition monitoring based on timing characteristics on the AXI-Lite bus and AXI-Stream bus to obtain the target attack mode; According to the preset dynamic switching strategy, protect the target attack mode type to complete the system layer dynamic response process.

[0014] To achieve the above object, the present invention also provides a domestic FPGA-based SM4 hardware acceleration and side-channel attack protection system, including: A data acquisition module, configured to use the pre-built PCIe bus to connect the pre-built adaptive acceleration protection FPGA to the pre-built target computer, where the adaptive acceleration protection FPGA includes an acceleration block and a protection block, the target computer includes a CPU and a DDR, and obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plain-ciphertext batch data in the DDR, where the key-mode parameters include the main key, encryption mode parameters, and function switching signals, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; The SM4 algorithm acceleration configuration module is used to utilize the acceleration block to perform data type identification operations on the plaintext-ciphertext batch data to obtain the data type, and perform dynamic resource allocation on the plaintext-ciphertext batch data according to the pre-constructed parallelization strategy and the data type to obtain resource allocation information, and configure the configured SM4 algorithm according to the resource allocation information, key-mode parameters, and dynamic control instructions; The side-channel resistance protection module is used to perform encryption processing on the plaintext-ciphertext batch data according to the configured SM4 algorithm to obtain the encrypted data to be protected, and utilize the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plaintext-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transform the encrypted data to be protected into the encrypted data that has been protected.

[0015] To solve the above problems, the present invention also provides an electronic device, which includes: A memory that stores at least one instruction; A processor that executes the instructions stored in the memory to implement the above-mentioned SM4 hardware acceleration and side-channel resistance protection method for domestic FPGAs.

[0016] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned SM4 hardware acceleration and side-channel resistance protection method for domestic FPGAs.

[0017] To solve the problems described in the background art, the SM4 algorithm in the FPGA of the present invention obtains data from two aspects: the CPU and the DDR. Among them, the batch data of plaintext-ciphertext sent by the DDR is the main data to be processed. Through the cooperation between the DDR memory and the FPGA, a large amount of data can be quickly processed, thereby improving the data processing efficiency. The key-mode parameters and dynamic control instructions sent by the CPU have a small amount of data, but they play a controlling role in the SM4 algorithm in the FPGA. In this solution, the CPU will pre-adjust the pipeline depth in the FPGA according to the task requirements. In addition, the present invention also configures a parallelization strategy in the acceleration block in the FPGA. Through the mutual coordination of the pipeline depth adjustment and the parallelization strategy, the pipeline depth optimizes the single-core timing efficiency, and the parallelization expands the multi-core processing ability. The two goals are complementary to improve the processing efficiency of the acceleration block. In addition, due to the changes caused by the encryption processing process of the SM4 algorithm in the acceleration block and the changes in the data acquisition paths from the CPU and the DDR, the anti-side-channel protection scheme should also change. The present invention adapts to the changes in the protection block for anti-side-channel protection from three aspects: suppressing leakage at the physical layer, confusing associations at the logical layer, and dynamically responding at the system layer, so as to ensure the overall security of the system while ensuring SM4 hardware acceleration. Therefore, the present invention can improve the encryption processing performance and security. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 FIG. is a schematic flow chart of a method for SM4 hardware acceleration and anti-side-channel protection of a domestic FPGA provided by an embodiment of the present invention; Figure 2 FIG. is a functional module diagram of a system for SM4 hardware acceleration and anti-side-channel protection of a domestic FPGA provided by an embodiment of the present invention; Figure 3 FIG. is a schematic structural diagram of an electronic device for implementing the method for SM4 hardware acceleration and anti-side-channel protection of the domestic FPGA provided by an embodiment of the present invention.

[0019] DESCRIPTION OF THE REFERENCE NUMERALS: 1. Electronic device; 10. Processor; 11. Memory; 12. Bus.

[0020] The implementation, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0022] The embodiments of the present application provide a method for SM4 hardware acceleration and side-channel attack resistance protection of domestic FPGAs. The execution subjects of the method for SM4 hardware acceleration and side-channel attack resistance protection of domestic FPGAs include, but are not limited to, at least one of electronic devices such as servers, terminals, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the method for SM4 hardware acceleration and side-channel attack resistance protection of domestic FPGAs can be executed by software or hardware installed on terminal devices or server devices, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0023] Referring to Figure 1 As shown, it is a schematic flowchart of the method for SM4 hardware acceleration and side-channel attack resistance protection of domestic FPGAs provided by an embodiment of the present invention. In this embodiment, the method for SM4 hardware acceleration and side-channel attack resistance protection of domestic FPGAs includes: S1. Use a pre-built PCIe bus to connect a pre-built acceleration and protection FPGA adapted to the target computer, where the acceleration and protection FPGA adapted includes an acceleration block and a protection block, and the target computer includes a CPU and a DDR.

[0024] Among them, the PCIe bus refers to a high-speed serial computer expansion bus standard used to connect a computer motherboard to external devices (such as graphics cards, solid-state drives, network cards, FPGAs, etc.).

[0025] Among them, FPGA refers to a field-programmable gate array, which belongs to a programmable logic device and can reconfigure its internal structure to achieve different functions. It can be understood as editable hardware, and the hardware function changes with the editing.

[0026] Among them, the acceleration and protection FPGA adapted refers to an FPGA that adaptively configures the processing process of the SM4 algorithm according to the task requirements of the target computer. The design purpose of the present invention is to improve the execution speed and security of the SM4 algorithm. The part responsible for calling and processing SM4 encrypted data is called the acceleration block. And the part that performs side-channel attack resistance protection on the data source and data processing process during the encryption process of the acceleration block is called the protection block.

[0027] Specifically, in the embodiments of the present invention, the acceleration block in the adaptive acceleration protection FPGA aims to improve the execution efficiency of the SM4 symmetric encryption algorithm through a dedicated hardware architecture or instruction set optimization, so as to meet the encryption requirements of high throughput and low latency. The anti-side-channel protection refers to the defense technology against side-channel attacks (Side-Channel Attack), which aims to suppress or confuse the physical or logical leakage information (such as power consumption, electromagnetic radiation, timing differences, etc.) generated by the encryption device during operation, and prevent attackers from stealing sensitive data (such as keys, passwords, etc.) through non-direct means. In the following, the adaptive acceleration protection FPGA will be directly referred to as FPGA in the present invention.

[0028] Specifically, in the embodiments of the present invention, the target computer refers to a computer that needs to process a large amount of video data and financial data, and at least includes a CPU and a DDR. The CPU refers to the central processing unit, and the DDR refers to the memory module.

[0029] In the embodiments of the present invention, the adaptive acceleration protection FPGA is connected to the target computer through a pre-built PCIe bus, so as to accelerate the process of the target computer's inability to encrypt and process massive data.

[0030] Specifically, the FPGA is connected to the host through the PCIe bus, supporting DMA transmission, realizing low-latency data interaction between the CPU and the FPGA, and the maximum bandwidth can reach 64 Gbps (PCIe Gen4 x16). Then, through the AXI interconnection protocol: the AXI4 bus is used to realize data scheduling between internal modules of the FPGA, supporting multi-channel concurrent transmission to ensure the efficiency of task distribution and result feedback.

[0031] S2. Obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plaintext-ciphertext batch data in the DDR, where the key-mode parameters include the master key, encryption mode parameters, and function switching signals, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plaintext-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information.

[0032] Among them, the key-mode parameters include the master key, encryption mode parameters, and function switching signals. Among them, the master key (MK) is used to generate round keys (rk0-rk31) and is the core input of the SM4 algorithm. The encryption mode parameters (such as the initialization vector IV of CBC and the counter value of CTR) are used to control the context of the encryption process. The function switching signal (encryption / decryption, fixed key mode / transformed key mode) is used to determine the working logic of the algorithm core.

[0033] Among them, the dynamic control instruction refers to an instruction used to control the data processing progress and mode. The start / suspend data transmission is an instruction for managing the start and suspension during the data processing of the FPGA. Adjusting the pipeline depth refers to an instruction for managing the relationship between the pipeline and the number of parallel cores during the data processing of the FPGA.

[0034] Among them, the plaintext-ciphertext batch data refers to some data that needs to be processed in the CPU task. During the process of storing these data in the DDR, some of the data will be encrypted. Therefore, it has partial plaintext and partial ciphertext, that is, plaintext / ciphertext data blocks. In addition, the intermediate state cache information refers to the intermediate results (such as the rk value after round key expansion) during the pipeline encryption process, and is used to save the context during the multi-core task switching through the DDR.

[0035] Specifically, in the embodiment of the present invention, obtaining the key-mode parameters and dynamic control instructions generated in the CPU, and obtaining the plaintext-ciphertext batch data in the DDR includes: Using a pre-built AXI-Lite bus, sending the key-mode parameters and dynamic control instructions generated in the CPU to the pre-built control register in the acceleration block; Using a pre-built AXI-Stream bus, sending the plaintext-ciphertext batch data pre-fetched by the CPU in the DDR to the pre-built input cache module in the acceleration block.

[0036] Among them, the AXI-Lite bus is a reduced version of the memory-mapped bus, which is specifically used for low-throughput data transmission between a host (such as a CPU) and a slave (such as a peripheral register, which is the control register in the present invention). The control register is a key component for configuring the hardware logic behavior, and the functions or parameters of each module in the FPGA can be dynamically adjusted through programming.

[0037] Among them, the AXI-Stream bus is a transmission line that does not require address management, supports burst transmission of unlimited length, and transmits data in a continuous stream form. The input cache module is a component of the input / output block (IOB) in the FPGA, and is responsible for receiving external signals and performing level conversion, noise filtering, and timing synchronization to ensure stable transmission of signals to the internal logic unit of the FPGA.

[0038] Specifically, in the embodiments of the present invention, the AXI-Lite bus does not support burst transactions and can only read and write data of a single address (such as 32 bits or 64 bits) each time. However, the channel structure includes independent read address, read data, write address, write data, and write response channels, and each channel implements a handshaking mechanism through VALID and READY signals, which is suitable for effectively controlling the FPGA. The AXI-Lite bus sends the key-mode parameters and dynamic control instructions generated in the CPU to the control registers pre-built in the acceleration block, so as to control the FPGA by using the control registers.

[0039] Specifically, in the embodiments of the present invention, the AXI-Stream bus supports high-speed channel design and only includes data-related signals (such as TDATA, TLAST, TREADY), cancels the address line, and reduces protocol overhead. The AXI-Stream bus sends the plaintext-ciphertext batch data to the input buffer module pre-built in the acceleration block, so as to input the data into the acceleration block by using the input buffer module.

[0040] In the embodiments of the present invention, the AXI-Lite bus focuses on low-speed and precise register-level control and is suitable for simple peripheral interactions. The AXI-Stream bus is oriented towards high-speed and continuous data stream transmission and adapts to scenarios with high real-time requirements. The two are interconnected through AXI Interconnect to achieve an efficient architecture with separated control and data.

[0041] In addition, it should be known that the core operations of the SM4 algorithm (such as S-box transformation and round function iteration) can be implemented through the general computing capabilities of the CPU, especially optimized for parallel processing with the SIMD instruction set (such as AVX, SSE). For example, the vectorized instructions of the CPU can be used to achieve parallel encryption of multiple packets. However, limited by the serial execution architecture of the CPU, it is easy to become a performance bottleneck when processing massive data. Among them, the SIMD (Single Instruction, Multiple Data) is a parallel computing technology that allows the processor to perform the same operation on multiple data elements through a single instruction. Its core goal is to improve the operation efficiency of data-intensive tasks, especially in scenarios that require batch processing of a large number of similar data, with remarkable effects.

[0042] Specifically, in the embodiments of the present invention, the CPU can generate an instruction for adjusting the pipeline depth in the dynamic control instruction according to the task requirements through a pipeline optimization algorithm. By splitting the long path into multiple short paths (inserting registers), the pipeline shortens the critical path delay, thereby increasing the clock frequency (i.e., increasing F). For example, a three-stage pipeline design with an original path delay of t can shorten the clock cycle to approximately t / 3, and the system rate can be increased to three times the original. This mechanism improves the throughput by optimizing the timing efficiency of a single task (trading area for speed).

[0043] S3. Use the acceleration block to perform a data type identification operation on the plaintext-ciphertext batch data to obtain the data type, and perform dynamic resource allocation on the plaintext-ciphertext batch data according to the pre-constructed parallelization strategy and the data type to obtain resource allocation information, and configure the configured SM4 algorithm according to the resource allocation information, the key-mode parameter, and the dynamic control instruction.

[0044] Among them, the data type identification operation refers to an operation of analyzing whether the data processed by the acceleration block is a large data stream scenario for video encryption or a small data stream scenario similar to Internet of Things instructions. The data type refers to the preset data types corresponding to the two data stream scenarios, including the large data block type and the small data block type.

[0045] Among them, the essence of the parallelization strategy is: relying on the independence of the FPGA programmable logic unit, increasing the number of operations per cycle (N) by simultaneously processing multiple tasks (such as multi-core and multi-channel designs), and directly increasing the throughput. For example, by designing 4 parallel SM4 cores, each core independently processes the packet data, and theoretically the throughput can be increased by 4 times.

[0046] Among them, the dynamic resource allocation refers to the configuration information of the data source, as well as the data retrieval and processing methods. The resource allocation information refers to the two types of results of the dynamic resource allocation. One is to directly obtain data from the DDR and then perform parallel processing. The other is single-core processing, grouping the data and processing multiple groups of data at a time.

[0047] Among them, the process of configuring the configured SM4 algorithm refers to optimizing the dedicated hardware architecture or instruction set in the SM4 algorithm, aiming to improve the execution efficiency of the SM4 symmetric encryption algorithm, so as to meet the encryption requirements of high throughput and low latency. The configured SM4 algorithm represents the FPGA structure information with an architecture and instruction configuration completed and can be used to implement the SM4 algorithm function.

[0048] Specifically, in the embodiments of the present invention, the performing a data type identification operation on the plaintext-ciphertext batch data to obtain the data type includes: Perform data block meta - information analysis on the plain - ciphertext batch data to obtain the data block size distribution and transmission interval characteristics; Use a pre - constructed sliding time window to perform throughput - latency correlation analysis on the plain - ciphertext batch data to obtain the data processing volume and average response latency; According to a preset weight coefficient sequence, perform feature - weighted calculation on the data block size distribution, transmission interval characteristics, data processing volume, and average response latency to obtain a data type score; Obtain the size relationship between the data type score and a pre - constructed classification threshold, and based on this size relationship, obtain the data type, where the data type includes a large data block type and a small data block type.

[0049] Among them, the data block meta - information analysis refers to the process of analyzing the size, speed, and frequency of data in the plain - ciphertext batch data. The data block size distribution refers to the statistical average data volume of a single transmission. In a large data stream scenario, it usually shows continuous and uniform block sizes (such as 1MB - 10MB / block), while in a small data block scenario, it is mostly discretely distributed (such as in KB level). The transmission interval characteristic means that a large data stream has a high - frequency continuous transmission characteristic (such as micro - second - level interval), and the small data block type shows low - frequency burst transmission (such as second - level interval).

[0050] Among them, the sliding time window is a dynamic statistical or control mechanism. By dividing time into multiple consecutive or overlapping sub - windows, it realizes flexible analysis of data or resource management in different time periods. Its core goal is to dynamically adjust the statistical range according to time movement, so as to more accurately reflect the system state or process streaming data.

[0051] Among them, the throughput - latency correlation analysis refers to the process of statistically calculating the data processing volume (TPS) and average response latency per unit time through a sliding time window. The data processing volume and the average response latency mean that in a large data stream scenario, TPS and latency are linearly negatively correlated (the latency increases suddenly when resources are saturated), and in a small data block scenario, TPS fluctuates greatly and the latency is relatively stable (low resource utilization rate).

[0052] Among them, the weight coefficient sequence refers to the influence degree scores of each index on the judgment of the data type size. For example, the data block size distribution is 0.3, the transmission interval characteristic is 0.3, the data processing volume is 0.2, and the average response latency is 0.2.

[0053] Among them, the feature - weighted calculation refers to the process of performing weighted calculation on the probability scores of each feature for separately judging the data type. The data type score refers to the weighted result of the feature - weighted calculation.

[0054] Among them, the classification threshold is configured as 50%.

[0055] Specifically, in the embodiments of the present invention, through two methods of data block meta-information analysis and throughput-latency correlation analysis, four indicators for distinguishing data flow types, namely data block size distribution, transmission interval characteristics, data processing volume, and average response latency, are obtained. Then, through a pre-constructed weight coefficient sequence, weights are assigned to each indicator, and finally, feature weighted calculation is performed to obtain a data type score. When the data type score is less than 50%, it is determined as the small data block type, and when the data type score is greater than 50%, it is determined as the large data block type.

[0056] In detail, in the embodiments of the present invention, the dynamic resource allocation for the plaintext-ciphertext batch data according to the pre-constructed parallelization strategy and the data type to obtain resource allocation information includes: Judging whether the data type is a large data block type or a small data block type according to the pre-constructed parallelization strategy; When the data type is a large data block type, determining that the resource allocation information is the first type of configuration information; When the data type is a small data block type, determining that the resource allocation information is the second type of configuration information; Among them, the first type of configuration information includes using the CPU to obtain task-related data, and using the PCIe bus to pre-load the task-related data into the DDR for storage, and using the pre-constructed FPGA multi-core parallel architecture and 32-level pipeline design to process the plaintext-ciphertext batch data; Among them, the second type of configuration information includes using the pre-constructed bit slicing service to reorganize the plaintext-ciphertext batch data into the bit width supported by the preset SIMD instruction to obtain a data grouping result, and performing batch processing on multiple groups in the data grouping result.

[0057] Among them, the first type of configuration information and the second type of configuration information represent the two ways of the above-mentioned resource allocation information in the expression of "dynamic resource allocation", which will not be elaborated here.

[0058] Among them, the task-related data refers to a large amount of data to be processed, such as traffic video data, financial trend data, etc.

[0059] Among them, the FPGA multi-core parallel architecture and 32-level pipeline design are the initialized FPGA operation modes. The processing of the plaintext-ciphertext batch data refers to encrypting the plaintext-ciphertext batch data using the SM4 algorithm.

[0060] Among them, the Bitslicing service is an optimized implementation technology for cryptographic algorithms. Its core goal is to improve encryption efficiency through data reorganization and parallel processing, while enhancing the ability to resist side-channel attacks.

[0061] Among them, the bit width supported by the SIMD instruction refers to the fact that SIMD packs multiple data elements (such as pixel values, floating-point numbers, integers, etc.) into a wide register (such as 128 bits, 256 bits, or even 512 bits) and uses a single instruction to complete batch operations. The data grouping result refers to the grouped data after packing. The same-batch processing refers to encrypting multiple groups of grouped data simultaneously using the SM4 algorithm at one time.

[0062] Specifically, in the embodiments of the present invention, the resource configuration information is configured according to the data type. When the data type is a large data block type, it is determined that the resource configuration information is the first type of configuration information. When the data type is a small data block type, it is determined that the resource configuration information is the second type of configuration information.

[0063] In detail, in the embodiments of the present invention, the configured SM4 algorithm is configured according to the resource configuration information, the key-mode parameter, and the dynamic control instruction, including: Performing an acceleration architecture selection operation on the pre-constructed SM4 algorithm based on the resource configuration information; Parsing the key-mode parameter to obtain the master key, the encryption mode parameter, and the function switching signal, and performing a key expansion operation on the master key based on the superimposed random mask according to the pre-constructed SM4 key expansion instruction to obtain the round keys, and storing the round keys in the pre-constructed secure storage area; Configuring the encryption process in the SM4 algorithm according to the encryption mode parameter, and configuring the working logic in the SM4 algorithm according to the function switching signal; Configuring the combination mode of the pipeline stages and the parallel cores in the SM4 algorithm according to the dynamic control instruction to obtain the configured SM4 algorithm.

[0064] Among them, the SM4 algorithm refers to the commercial cryptographic algorithm released by the China National Cryptography Administration, which belongs to block cipher and is commonly used for data encryption.

[0065] Among them, the acceleration architecture selection operation refers to configuring the data retrieval and processing methods in the SM4 algorithm as the first type of configuration information or the second type of configuration information in the resource configuration information.

[0066] Among them, the parsing refers to the process of regularizing and normalizing the content of the key-mode parameter. The SM4 key expansion instruction refers to the normalized process instruction in the SM4 algorithm.

[0067] Among them, the key expansion operation based on superimposed random masks refers to adding true random numbers (TRNG) each time during the key expansion process. The key expansion process refers to the process in which the initial key generates round keys through the SM4 key expansion instruction and stores them in the secure storage area. The key operands contain multiple round keys and support dynamic switching (such as changing keys for each batch of data). The random masks generated by superimposing TRNG during the key expansion process can destroy the correlation of side-channel attacks.

[0068] Among them, the round key refers to the result of the key expansion of the master key in a certain round. The secure storage area refers to a storage area such as a TEE or an encrypted BRAM.

[0069] Among them, the encryption mode parameters can support modes such as ECB, CBC, and CTR. For example, the CBC mode requires configuring the initialization vector and writing it to the register through AXI-Lite.

[0070] Among them, configuring the encryption process in the SM4 algorithm refers to selecting a certain type among the control encryption mode parameters, such as ECB, CBC, CTR, etc.

[0071] Among them, configuring the working logic in the SM4 algorithm refers to configuring the start and pause functions.

[0072] Among them, configuring the combined mode of the pipeline stage number and the number of parallel cores in the SM4 algorithm refers to monitoring the depth of the data queue and dynamically adjusting the pipeline stage number or the number of cores. For example, when the queue depth > 8KB, 4-core parallelism is enabled, otherwise it switches to the single-core high-frequency mode, etc.

[0073] Specifically, in the embodiments of the present invention, the above content is a configuration scheme. An electronic design automation (EDA) tool can be executed. According to the structure of the SM4 algorithm, and then according to the resource configuration information, key-mode parameters, master key, encryption mode parameters, function switching signals, and dynamic control instructions, the FPGA structure is programmed to obtain the configured SM4 algorithm.

[0074] S4. Encrypt the plain-ciphertext batch data according to the configured SM4 algorithm to obtain the encrypted data to be protected.

[0075] In the embodiments of the present invention, the configured SM4 algorithm is still the calculation process of the national standard SM4 algorithm, but the data retrieval efficiency is improved. Therefore, the process of encrypting the plain-ciphertext batch data is not described in detail here. The encrypted data to be protected refers to the encryption result of the configured SM4 algorithm for the plain-ciphertext batch data.

[0076] S5. Use the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transform the to-be-protected encrypted data into protected encrypted data.

[0077] Among them, the protection operation refers to the specific implementation means to improve the security of the encryption result.

[0078] Among them, the physical layer leakage suppression refers to the method of protecting the hardware through physical methods. The logical layer confusion association refers to the protection method of randomizing the intermediate values in the calculation process to ensure that the key information cannot be restored from a single point of leakage. The system layer dynamic response refers to the method of extracting the timing characteristics during the system processing process and data transmission process through AI to determine whether there is an attacker interfering with the data encryption process.

[0079] Among them, the protected encrypted data refers to the to-be-protected encrypted data that has passed the security authentication of each protection process.

[0080] Specifically, in the embodiment of the present invention, the use of the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response includes: Use the protection block to perform signal weakening operations based on power consumption and electromagnetic radiation characteristics on the encryption process, and correlation perturbation operations based on timing and clock jitter to complete the physical layer leakage suppression process; Perform intermediate value randomization operations on the encryption process to complete the logical layer confusion association process; Perform system layer dynamic response processes on the AXI-Lite bus and AXI-Stream bus based on abnormal event detection and response.

[0081] Among them, the signal weakening operation based on power consumption and electromagnetic radiation characteristics means that information may be exposed during data transmission due to the power consumption and electromagnetic radiation generated by each unit and route. Therefore, weakening this information, such as balancing the power consumption in each process, or covering the electromagnetic radiation through a layer, can effectively reduce information leakage.

[0082] Among them, the correlation perturbation operation based on timing and clock jitter means that an attacker can perform a timing attack through the correlation between timing and clock jitter. By disturbing the correlation between timing and clock jitter, the attacker's correlation analysis of read and write timing can be disrupted.

[0083] Among them, the intermediate value randomization operation refers to an operation used to ensure that key information cannot be restored from single-point leakage. For example, a dual-rail redundancy calculation module is inserted into the encryption pipeline to compare the output consistency in real time, detect and block fault injection attacks.

[0084] Among them, the system-level dynamic response process based on abnormal event detection and response refers to the process of detecting abnormal events in the data transmission process and processing process through an AI model, and can perform autonomous processing according to the detection results. For example, when a low risk is detected, redundant verification is turned off to improve performance; when an attack is detected, a 3rd-order threshold mask is enabled, etc.

[0085] Specifically, in the embodiments of the present invention, using the protection block to perform signal weakening operations based on power consumption and electromagnetic radiation characteristics, and correlation perturbation operations based on timing and clock jitter on the encryption processing process includes: Using the pre-built adjustable load circuit in the protection block to perform power consumption difference balancing operations in the operation stage of the encryption processing process; Using the pre-built metal shielding layer to cover the preset high-leakage risk area in the configured SM4 algorithm; Using the pre-built true random number to perform timing perturbation operations on the data transmission in the AXI-Lite bus and AXI-Stream bus.

[0086] Among them, the adjustable load circuit is an electronic system that can dynamically adjust equivalent resistance or load parameters, mainly used to simulate different load conditions and test the performance of power supplies, batteries or other power supply devices.

[0087] Among them, the power consumption difference balancing operation in the operation stage refers to embedding an adjustable load circuit in the FPGA logic unit to balance the power consumption differences in different operation stages of the SM4 algorithm (such as S-box substitution, round key XOR), and weaken the effectiveness of differential power analysis (DPA).

[0088] Among them, the metal shielding layer is a material that can shield or suppress electromagnetic waves. The covering is used to shield electromagnetic waves.

[0089] Among them, the true random number refers to randomly generated numbers or characters. The timing perturbation operation refers to introducing timing perturbation based on true random numbers (TRNG) in the AXI bus transmission path to disrupt the attacker's correlation analysis of read and write timing.

[0090] Specifically, in the embodiments of the present invention, the adjustable load circuit pre-constructed in the protection block is used to perform power consumption difference balancing operation on the operation stage of the encryption process, which is beneficial to weakening the effectiveness of differential power analysis; the pre-constructed metal shielding layer is used to cover the preset high-leakage risk area in the configured SM4 algorithm to suppress the electromagnetic radiation signal during the encryption process; the pre-constructed true random number is used to perform timing perturbation operation on the data transmission in the AXI-Lite bus and the AXI-Stream bus to disrupt the attacker's correlation analysis of the read-write timing. Thus, the leakage suppression process at the physical layer is completed.

[0091] Specifically, in the embodiments of the present invention, the intermediate value randomization operation is performed on the encryption process to complete the logical layer confusion association process, including: Performing threshold segmentation on the preset intermediate value in the encryption process, where the intermediate value includes the S-box output and the round function intermediate value; Using the dual-rail redundant calculation module pre-constructed in the protection block to perform real-time comparison of output consistency on the encryption process to complete the logical layer confusion association process.

[0092] Among them, the intermediate value refers to the S-box output and the round function intermediate value.

[0093] Among them, the threshold segmentation is a cryptographic technique aimed at decomposing sensitive information (such as keys or data) into multiple independent parts (called "shares" or "shadows") and distributing them to different participants. Only when a preset number of shares (i.e., the "threshold value") are combined can the original secret be restored.

[0094] Among them, the dual-rail redundant calculation module is a hardware redundancy technique that improves the reliability of the system through parallel calculation and real-time verification. Its core principle is to synchronously execute the same task through two independent calculation units and compare the output results to detect or correct errors.

[0095] Among them, the real-time comparison of output consistency means that in the scenario of dynamic data changes, the system ensures that the content, status, or logical relationship of the output results of different data sources or processing flows is always consistent by real-time monitoring and comparison. Its core goal is to quickly discover and locate inconsistent nodes during the continuous update of data, thereby ensuring the correctness and reliability of the business.

[0096] Specifically, in the embodiments of the present invention, it can be achieved through multiplication masks and thresholds: perform threshold segmentation (such as 3rd-order threshold masking) on the S-box output and the intermediate values of the round function to ensure that the key information cannot be restored from a single-point leakage, and implement it through redundant check logic: insert a dual-rail redundant calculation module in the encryption pipeline to compare the output consistency in real time, detect and block fault injection attacks, thereby completing the intermediate value randomization operation and realizing the logical layer confusion association process.

[0097] Specifically, in the embodiments of the present invention, the system layer dynamic response process for performing anomaly event detection and response on the AXI-Lite bus and the AXI-Stream bus includes: Using the pre-built CNN attack recognition model in the protection block, perform attack pattern recognition monitoring based on timing characteristics on the AXI-Lite bus and the AXI-Stream bus to obtain the target attack pattern; According to the preset dynamic switching strategy, protect the type of the target attack pattern to complete the system layer dynamic response process.

[0098] Among them, the CNN attack recognition model refers to a simply trained lightweight CNN neural network, which is used to detect attacks and identify attack patterns according to the abnormal timing characteristics during data transmission and processing. The attack pattern recognition monitoring refers to the execution process of the CNN attack recognition model. The target attack pattern refers to the recognition result of the CNN attack recognition model, which can be no attack, DoS attack, fault injection attack, etc.

[0099] Among them, the dynamic switching strategy refers to adjusting the protection intensity according to the threat level. The protection refers to the execution process of the dynamic switching strategy. For example, when the risk is low, the redundant check is turned off to improve performance, and when an attack is detected, a 3rd-order threshold mask is enabled.

[0100] Specifically, in the embodiments of the present invention, the CNN attack recognition model is used to detect attacks on the data transmission and processing timing information in the AXI-Lite bus and the AXI-Stream bus. Once an attacker is detected, automatic protection is performed according to the dynamic switching strategy, thereby completing the system layer dynamic response process.

[0101] Specifically, in the embodiments of the present invention, a full-stack protection system for physical layer leakage suppression, logical layer confusion association, and system layer dynamic response is constructed. If the encrypted data to be protected passes the authentication of the above protection process, the encrypted data to be protected is transformed into protected encrypted data.

[0102] To solve the problems described in the background art, the SM4 algorithm in the FPGA of the present invention obtains data from two aspects: the CPU and the DDR. Among them, the batch data of plaintext-ciphertext sent by the DDR is the main data to be processed. Through the cooperation between the DDR memory and the FPGA, a large amount of data can be processed quickly, thereby improving the data processing efficiency; while the key-mode parameters and dynamic control instructions sent by the CPU have a small data volume, but play a controlling role in the SM4 algorithm in the FPGA. In this solution, the CPU will pre-adjust the pipeline depth in the FPGA according to the task requirements. In addition, the present invention also configures a parallelization strategy in the acceleration block in the FPGA. Through the mutual cooperation of the pipeline depth adjustment and the parallelization strategy, the pipeline depth optimizes the single-core timing efficiency, and the parallelization expands the multi-core processing ability. The two goals complement each other to improve the processing efficiency of the acceleration block; in addition, due to the changes caused by the encryption processing process of the SM4 algorithm in the acceleration block and the changes in the data acquisition paths from the CPU and the DDR, the anti-side-channel protection scheme should also change. The present invention adapts to the changes in the protection block from three aspects: physical layer leakage suppression, logical layer confusion association, and system layer dynamic response for anti-side-channel protection, so as to ensure the overall security of the system while ensuring SM4 hardware acceleration. Therefore, the present invention can improve the encryption processing performance and security.

[0103] As Figure 2 shown, it is a functional block diagram of the SM4 hardware acceleration and anti-side-channel protection system of the domestic FPGA provided by an embodiment of the present invention.

[0104] The SM4 hardware acceleration and anti-side-channel protection system 100 of the domestic FPGA described in the present invention can be installed in an electronic device. According to the implemented functions, the SM4 hardware acceleration and anti-side-channel protection system 100 of the domestic FPGA can include a data acquisition module 101, an SM4 algorithm acceleration configuration module 102, and an anti-side-channel protection module 103. The modules described in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0105] The data acquisition module 101 is used to connect a pre-built adaptive acceleration and protection FPGA to a pre-built target computer by using a pre-built PCIe bus. The adaptive acceleration and protection FPGA includes an acceleration block and a protection block. The target computer includes a CPU and a DDR. The data acquisition module 101 is also used to obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plain-ciphertext batch data in the DDR. The key-mode parameters include a master key, an encryption mode parameter, and a function switching signal. The dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth. The plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information. The SM4 algorithm acceleration configuration module 102 is used to perform a data type recognition operation on the plain-ciphertext batch data by using the acceleration block to obtain a data type. According to a pre-built parallelization strategy and the data type, the SM4 algorithm acceleration configuration module 102 performs dynamic resource allocation on the plain-ciphertext batch data to obtain resource allocation information. According to the resource allocation information, key-mode parameters, and dynamic control instructions, the SM4 algorithm acceleration configuration module 102 configures a configured SM4 algorithm. The side-channel attack protection module 103 is used to encrypt the plain-ciphertext batch data according to the configured SM4 algorithm to obtain encrypted data to be protected. The side-channel attack protection module 103 is also used to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response by using the protection block, and convert the encrypted data to be protected into protected encrypted data.

[0106] Specifically, each module in the SM4 hardware acceleration and side-channel attack protection system 100 of the domestic FPGA in the embodiment of the present invention adopts the same technical means as those in the Figure 1 SM4 hardware acceleration and side-channel attack protection method of the domestic FPGA described above, and can produce the same technical effects, which will not be elaborated here.

[0107] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the SM4 hardware acceleration and side-channel attack protection method of the domestic FPGA provided by an embodiment of the present invention.

[0108] The electronic device 1 may include a processor 10, a memory 11, and a bus 12. The electronic device 1 may further include a computer program stored in the memory 11 and executable on the processor 10, such as a program for the SM4 hardware acceleration and side-channel attack protection method of the domestic FPGA.

[0109] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 further includes the internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can not only be used to store application software installed on the electronic device 1 and various types of data, such as the code of the SM4 hardware acceleration and anti-side-channel protection method program of domestic FPGA, but also be used to temporarily store data that has been output or will be output.

[0110] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the SM4 hardware acceleration and anti-side-channel protection method program of domestic FPGA), and calling data stored in the memory 11, to execute various functions of the electronic device 1 and process data.

[0111] The bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0112] Figure 3 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3The structures shown do not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0113] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0114] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0115] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0116] The program of the SM4 hardware acceleration and anti-side-channel protection method of the domestic FPGA stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve: Connect a pre-built adaptive acceleration protection FPGA to a pre-built target computer by using a pre-built PCIe bus, where the adaptive acceleration protection FPGA includes an acceleration block and a protection block, and the target computer includes a CPU and a DDR; Obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plain-ciphertext batch data in the DDR, where the key-mode parameters include the master key, encryption mode parameters, and function switching signals, and the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; Use the acceleration block to perform a data type recognition operation on the plain-ciphertext batch data to obtain the data type, and perform dynamic resource allocation on the plain-ciphertext batch data according to the pre-constructed parallelization strategy and the data type to obtain resource allocation information, and configure the configured SM4 algorithm according to the resource allocation information, key-mode parameters, and dynamic control instructions; Encrypt the plain-ciphertext batch data according to the configured SM4 algorithm to obtain the encrypted data to be protected; Use the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transform the encrypted data to be protected into the encrypted data that has been protected.

[0117] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figures 1 to 3 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0118] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0119] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, and when the computer program is executed by the processor of the electronic device, it can implement: Connect the pre-constructed adaptive acceleration protection FPGA to the pre-constructed target computer by using the pre-constructed PCIe bus, where the adaptive acceleration protection FPGA includes an acceleration block and a protection block, and the target computer includes a CPU and a DDR; Obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plain-ciphertext batch data in the DDR, where the key-mode parameters include the master key, encryption mode parameters, and function switching signals, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; Use the acceleration block to perform a data type recognition operation on the plain-ciphertext batch data to obtain the data type, and perform dynamic resource allocation on the plain-ciphertext batch data according to the pre-constructed parallelization strategy and the data type to obtain resource allocation information, and configure the configured SM4 algorithm according to the resource allocation information, key-mode parameters, and dynamic control instructions; Encrypt the plain-ciphertext batch data according to the configured SM4 algorithm to obtain the encrypted data to be protected; Use the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transform the encrypted data to be protected into the encrypted data that has been protected.

[0120] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative, and there can be other division methods in actual implementation.

[0121] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, the functional modules in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0123] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for SM4 hardware acceleration and anti-side-channel protection of a domestic FPGA, characterized in that, The method includes: Connecting a pre-built adaptive acceleration protection FPGA to a pre-built target computer using a pre-built PCIe bus, where the adaptive acceleration protection FPGA includes an acceleration block and a protection block, and the target computer includes a CPU and a DDR; Obtaining the key-mode parameters and dynamic control instructions generated in the CPU, and obtaining the plain-ciphertext batch data in the DDR, where the key-mode parameters include a master key, an encryption mode parameter, and a function switching signal, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plain-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; Using the acceleration block to perform a data type identification operation on the plain-ciphertext batch data to obtain a data type, and dynamically configuring resources for the plain-ciphertext batch data according to a pre-built parallelization strategy and the data type to obtain resource configuration information, and configuring a configured SM4 algorithm according to the resource configuration information, key-mode parameters, and dynamic control instructions; Performing an encryption process on the plain-ciphertext batch data according to the configured SM4 algorithm to obtain encrypted data to be protected; Using the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plain-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and transforming the encrypted data to be protected into encrypted data that has been protected.

2. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 1, characterized in that, The obtaining the key-mode parameters and dynamic control instructions generated in the CPU, and obtaining the plain-ciphertext batch data in the DDR, includes: Sending the key-mode parameters and dynamic control instructions generated in the CPU to a pre-built control register in the acceleration block using a pre-built AXI-Lite bus; Sending the plain-ciphertext batch data pre-fetched from the DDR by the CPU to a pre-built input cache module in the acceleration block using a pre-built AXI-Stream bus.

3. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 2, characterized in that, The performing a data type identification operation on the plain-ciphertext batch data to obtain a data type, includes: Analyzing the metadata of the data blocks in the plain-ciphertext batch data to obtain the data block size distribution and transmission interval characteristics; Performing throughput and latency correlation analysis on the plain-ciphertext batch data using a pre-built sliding time window to obtain the data processing volume and average response latency; Performing feature weighted calculation on the data block size distribution, transmission interval characteristics, data processing volume, and average response latency according to a preset weight coefficient sequence to obtain a data type score; Obtaining the size relationship between the data type score and a pre-built classification threshold, and obtaining the data type according to the size relationship, where the data type includes a large data block type and a small data block type.

4. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGA according to claim 3, characterized in that, The dynamically configuring resources for the plain-ciphertext batch data according to a pre-built parallelization strategy and the data type to obtain resource configuration information, includes: According to the pre-built parallelization strategy, determine whether the data type is a large data block type or a small data block type; When the data type is a large data block type, determine that the resource configuration information is the first type of configuration information; When the data type is a small data block type, determine that the resource configuration information is the second type of configuration information; Among them, the first type of configuration information includes using the CPU to obtain task-related data, and using the PCIe bus to pre-load the task-related data into the DDR for storage, and using the pre-built FPGA multi-core parallel architecture and 32-level pipeline design to process the plaintext-ciphertext batch data; Among them, the second type of configuration information includes using the pre-built bit-slice service to reorganize the plaintext-ciphertext batch data into the bit width supported by the preset SIMD instruction to obtain a data grouping result, and performing batch processing on multiple groups in the data grouping result.

5. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 4, characterized in that The configured SM4 algorithm configured according to the resource configuration information, key-mode parameter, and dynamic control instruction includes: Perform an acceleration architecture selection operation on the pre-built SM4 algorithm based on the resource configuration information; Parse the key-mode parameter to obtain the master key, encryption mode parameter, and function switching signal, and perform a key expansion operation on the master key based on the superimposed random mask according to the pre-built SM4 key expansion instruction to obtain the round key, and store the round key in the pre-built secure storage area; Configure the encryption process in the SM4 algorithm according to the encryption mode parameter, and configure the working logic in the SM4 algorithm according to the function switching signal; Configure the combination mode of the pipeline stage number and the number of parallel cores in the SM4 algorithm according to the dynamic control instruction to obtain the configured SM4 algorithm.

6. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 5, characterized in that, The protection operation of the key-mode parameter, dynamic control instruction, plaintext-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response using the protection block includes: Use the protection block to perform a signal weakening operation based on the power consumption and electromagnetic radiation characteristics of the encryption process, and a correlation perturbation operation based on the timing and clock jitter to complete the physical layer leakage suppression process; Perform an intermediate value randomization operation on the encryption process to complete the logical layer confusion association process; Perform a system layer dynamic response process on the AXI-Lite bus and AXI-Stream bus based on abnormal event detection and response.

7. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 6, characterized in that, The use of the protection block to perform a signal weakening operation based on the power consumption and electromagnetic radiation characteristics of the encryption process, and a correlation perturbation operation based on the timing and clock jitter includes: Use the pre-built adjustable load circuit in the protection block to perform an operation stage power consumption difference balancing operation on the encryption process; Use the pre-built metal shielding layer to cover the preset high leakage risk area in the configured SM4 algorithm; Perform a timing perturbation operation on the data transmission in the AXI-Lite bus and the AXI-Stream bus by using pre-built true random numbers.

8. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGA according to claim 7, characterized in that Perform an intermediate value randomization operation on the encryption process to complete the logical layer confusion association process, including: Perform threshold segmentation on the intermediate values preset in the encryption process, where the intermediate values include S-box outputs and round function intermediate values; Use the pre-built dual-rail redundancy calculation module in the protection block to perform real-time comparison of output consistency on the encryption process to complete the logical layer confusion association process.

9. The SM4 hardware acceleration and anti-side-channel protection method for domestic FPGAs according to claim 8, wherein, Perform a system layer dynamic response process for the AXI-Lite bus and the AXI-Stream bus based on abnormal event detection and response, including: Use the pre-built CNN attack recognition model in the protection block to perform attack mode recognition monitoring based on timing characteristics on the AXI-Lite bus and the AXI-Stream bus to obtain the target attack mode; According to the preset dynamic switching strategy, protect the target attack mode type to complete the system layer dynamic response process.

10. A SM4 hardware acceleration and anti-side-channel protection system for domestic FPGA, characterized in that, The system includes: A data acquisition module, which is used to connect a pre-built adaptive acceleration protection FPGA to a pre-built target computer by using a pre-built PCIe bus, where the adaptive acceleration protection FPGA includes an acceleration block and a protection block, the target computer includes a CPU and a DDR, and obtain the key-mode parameters and dynamic control instructions generated in the CPU, and obtain the plaintext-ciphertext batch data in the DDR, where the key-mode parameters include the main key, encryption mode parameters, and function switching signals, the dynamic control instructions include instructions for starting / suspending data transmission and adjusting the pipeline depth, and the plaintext-ciphertext batch data includes plaintext / ciphertext data blocks and intermediate state cache information; An SM4 algorithm acceleration configuration module, which is used to use the acceleration block to perform data type recognition operations on the plaintext-ciphertext batch data to obtain the data type, and perform dynamic resource configuration on the plaintext-ciphertext batch data according to the pre-built parallelization strategy and the data type to obtain resource configuration information, and configure the configured SM4 algorithm according to the resource configuration information, key-mode parameters, and dynamic control instructions; A side-channel attack protection module, which is used to encrypt the plaintext-ciphertext batch data according to the configured SM4 algorithm to obtain the encrypted data to be protected, and use the protection block to perform protection operations on the key-mode parameters, dynamic control instructions, plaintext-ciphertext batch data, and the encryption process based on physical layer leakage suppression, logical layer confusion association, and system layer dynamic response, and convert the encrypted data to be protected into the encrypted data that has been protected.

Citation Information

Patent Citations

  • Privacy computing heterogeneous acceleration method and device based on fully homomorphic encryption

    CN115622684A

  • SM4 encryption method based on FPGA

    CN119201832A

  • Communication encryption acceleration device and method, Internet of Things equipment and system

    CN120165866A

  • Categorizing encrypted data files

    US20210263904A1

Cited By

  • National cryptographic algorithm acceleration system based on Numba just-in-time compiling technology

    CN120560665A

  • SM4 symmetric cryptographic algorithm low-delay hardware acceleration method and system

    CN120729509A

  • Terminal security access and data protection method based on virtual power plant

    CN122204401A