An SM4 Encryption Method Based on FPGA

By operating multiple SM4 encryption cores in parallel on the FPGA chip, combining pipeline architecture and random key generation, the shortcomings of existing SM4 algorithm hardware implementation in anti-side channel attacks are solved, and higher anti-side channel attack capabilities and key randomness are achieved.

CN119201832BActive Publication Date: 2025-06-17青岛青软晶尊微电子科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411064828.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-05
Publication Date
2025-06-17
Estimated Expiration
2044-08-05

AI Technical Summary

Technical Problem

The existing SM4 algorithm hardware implementation has shortcomings in resisting side channel attacks, lacks flexibility and randomness, and is easily analyzed and predicted by attackers.

Method used

By implementing parallel operations of multiple SM4 encryption cores on the FPGA chip, using pipeline architecture and random key generation, combining polynomial operations and logical operations, increasing the uncertainty of operations and anti-side channel attack capabilities.

Benefits of technology

It improves the anti-side channel attack capability of the SM4 encryption method, enhances the randomness and unpredictability of the key, and effectively resists advanced attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119201832B_ABST
    Figure CN119201832B_ABST
Patent Text Reader

Abstract

The present application discloses an SM4 encryption method based on FPGA, which relates to the technical field of data encryption and includes: obtaining plaintext data to be encrypted; configuring multiple SM4 encryption cores on an FPGA chip to perform parallel encryption operations on the plaintext data to be encrypted; wherein each SM4 encryption core adopts a pipeline architecture, and the pipeline architecture includes a data fetcher, a round function arithmetic unit, and a key adder; using an extraction circuit based on physical noise integrated on the FPGA chip to generate an N-bit random number as the encryption key for each SM4 encryption core; collecting handshake signals to calculate the operation efficiency of each SM4 encryption core; according to the operation efficiency, adopting a priority scheduling algorithm to adjust the FPGA resource allocation among the SM4 encryption cores; during the encryption operation, using a mixed calculation of polynomial and logical operations for power consumption balancing, and using random waiting cycles for timing randomization. Aiming at the problem of weak anti-differential sensitivity in the prior art, the present application improves the anti-side-channel attack ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data encryption technology, and particularly to an SM4 encryption method based on FPGA. Background Art

[0002] With the rapid development of information technology and the wide popularization of network applications, the problem of data security has become increasingly prominent. In various information systems and network platforms, protecting critical data and user privacy from illegal access and malicious attacks has become an important and urgent task. As the core means to ensure information security, cryptographic technology plays a key role in data encryption and identity authentication. The SM4 algorithm encrypts data using a 128-bit key, has a high security strength and good performance, and can effectively resist various conventional cryptographic attacks. However, the SM4 algorithm still faces the threat of side-channel attacks in the implementation process.

[0003] Side-channel attack is an attack method based on the physical implementation information of cryptographic devices. By analyzing side-channel information such as power consumption, electromagnetic radiation, and execution time during the operation of cryptographic devices, the key is speculated or the ciphertext is cracked. Common side-channel attacks include power analysis attacks, electromagnetic analysis attacks, time analysis attacks, etc. These attacks utilize the correlation between the physical characteristics in the practice of cryptographic algorithms and the key. Through statistical analysis and related calculations, the key or plaintext information can be restored, thus destroying the security of SM4 encryption.

[0004] The existing hardware implementation schemes of the SM4 algorithm mainly use devices such as ASIC (Application Specific Integrated Circuit) and FPGA (Field Programmable Gate Array). Although these schemes can provide high encryption performance, there are still some deficiencies in resisting side-channel attacks. First, most of the existing schemes adopt fixed encryption structures and operation methods, lacking flexibility and randomness, and are easily analyzed and predicted by attackers. Second, the operation timing and energy consumption during the encryption process are relatively regular and have a certain correlation with the key, which may leak key information. In addition, limited by hardware resources and costs, it is difficult for existing schemes to implement complex anti-side-channel attack measures, and the ability to resist advanced attacks is limited. Summary of the Invention

[0005] Aiming at the problem of limited anti-side-channel attack ability in the encryption of the SM4 algorithm in the prior art, this application provides an SM4 encryption method based on FPGA. By performing parallel encryption operations on plaintext data through multiple SM4 encryption cores with a pipeline architecture, etc., the anti-side-channel attack ability is improved.

[0006] The purpose of this application is achieved through the following technical solutions.

[0007] This application provides an SM4 encryption method based on FPGA, which realizes parallel operation of multiple SM4 encryption cores on the FPGA chip, including: Input of plaintext data: Use high-speed IO interfaces of FPGA, such as LVDS, SerDes, etc., to achieve high-speed input of plaintext data. Design a data input buffer to temporarily store the input plaintext data and balance the difference between the data input rate and the encryption core processing rate. Use a data distribution circuit to distribute the input plaintext data to each SM4 encryption core to achieve parallel processing of data. Multi-core parallel encryption: Determine the number of SM4 encryption cores that can be implemented according to the resources of the FPGA chip, such as lookup tables (LUTs), flip-flops (FFs), block RAMs (BRAMs), etc. Adopt a modular design method to design the SM4 encryption core as an independent functional module for easy replication and instantiation. Use a parallel control circuit to synchronously coordinate the operating states of each SM4 encryption core to achieve parallel encryption processing.

[0008] Pipeline architecture: Split the round function of the SM4 encryption algorithm into multiple sub-functions, such as non-linear transformation, linear transformation, etc., to achieve a fine-grained pipeline design. Adopt multi-stage pipeline registers and insert pipeline registers between the fetcher, round function arithmetic unit, and key adder to improve the parallelism of data processing. Optimize the scheduling and control logic of the pipeline to minimize data dependencies and conflicts between pipeline stages and improve the efficiency of the pipeline. Random key generation: Integrate physical noise sources, such as resistor thermal noise sources, ring oscillators, etc., on the FPGA chip to generate random noise signals. Design a random number extraction circuit to sample, quantize, and post-process the analog noise signals generated by the physical noise source to extract a random bit sequence. Adopt cryptographically secure random number post-processing algorithms, such as Von Neumann correction, XOR mixing, etc., to improve the statistical characteristics and unpredictability of random numbers.

[0009] Handshake signal acquisition: Set handshake signal lines between each stage of the pipeline to indicate the validity of data and the completion of processing. Adopt an asynchronous handshake protocol, such as the request - response protocol, to achieve reliable communication and synchronization between pipeline stages. Design a handshake signal acquisition circuit to monitor the state changes of handshake signals in real time and record the timestamps and durations of handshake signals. Computational efficiency evaluation: Design an efficiency evaluation circuit to calculate the computational time and data throughput rate of each stage of the pipeline based on the acquired handshake signal information. Adopt a weighted average algorithm to comprehensively consider the computational time and data throughput rate of each stage of the pipeline to obtain a comprehensive computational efficiency index. Through threshold comparison and sorting algorithms, evaluate and sort the computational efficiency of each SM4 encryption core. Dynamic resource scheduling: Design a resource scheduling controller to dynamically adjust the allocation of FPGA resources among each SM4 encryption core according to the computational efficiency evaluation results. Adopt priority scheduling algorithms, such as efficiency - based priority scheduling, fairness - based round - robin scheduling, etc., to achieve dynamic resource allocation. By configuring the reconfigurable logic units of the FPGA, dynamically change the number and location of SM4 encryption cores to adapt to different task requirements and resource constraints.

[0010] Specifically, after each SM4 encryption core completes a round of encryption operations, it packages the computational efficiency data and sends it to the central controller through a dedicated efficiency data channel. The computational efficiency data includes key information such as the number of the SM4 encryption core, computational time, resource utilization rate, etc. After receiving the computational efficiency data, the central controller parses and stores the data for subsequent priority scheduling. The central controller performs priority sorting on each SM4 encryption core according to the collected computational efficiency data. The sorting algorithm can adopt hierarchical sorting based on efficiency thresholds or weighted scoring sorting, etc., and select a suitable algorithm according to actual needs. The sorting result generates a priority list, arranging the SM4 encryption cores in descending order of computational efficiency. The central controller generates resource allocation control instructions according to the priority list. For the SM4 encryption core with the highest computational efficiency, generate a resource priority guarantee instruction to ensure that it obtains the optimal resource support. For the SM4 encryption core with the second - highest computational efficiency, generate a resource enhancement instruction to allocate more FPGA resources to it through dynamic partial reconfiguration technology. For the SM4 encryption core with relatively low computational efficiency, generate a resource recovery instruction to recover some of its FPGA resources through dynamic partial reconfiguration technology.

[0011] For an SM4 encryption core with extremely low or abnormal operation efficiency, generate a resource suspension instruction, suspend its operation through gating technology, and release resources. The central controller packages the generated resource allocation control instructions and distributes them to each SM4 encryption core through a dedicated control instruction channel. After each SM4 encryption core receives the resource allocation control instruction, it performs corresponding resource adjustment operations according to the instruction type. For the resource priority guarantee instruction, the SM4 encryption core maintains the current resource configuration and continues to operate efficiently. For the resource enhancement instruction, the SM4 encryption core loads a new encryption module in the idle area of the FPGA chip through dynamic partial reconfiguration technology to expand the parallelism and throughput. For the resource recycling instruction, the SM4 encryption core unloads some encryption modules through dynamic partial reconfiguration technology to release FPGA resources and reports the released resource information to the central controller. For the resource suspension instruction, the SM4 encryption core suspends the clock signal through gating technology, enters the low-power state, and reports the suspension state to the central controller. After the central controller receives the resource information released by the SM4 encryption core, it incorporates these resources into the allocable resource pool. According to the priority list and the situation of the allocable resource pool, the central controller reallocates resources. The released resources are preferentially allocated to the SM4 encryption core with higher operation efficiency to improve its parallelism and throughput. Through dynamic partial reconfiguration technology, a new encryption module is loaded into the SM4 encryption core that obtains resource support.

[0012] Hybrid operation mode: Design a polynomial operation circuit to implement operations such as addition and multiplication of polynomials, and introduce non-linear transformation. Use a random number generator to randomly select polynomial operations or logical operations during the encryption operation to increase the uncertainty of the operation. By dynamically adjusting the coefficients and orders of the polynomials, change the complexity of the polynomial operations and increase the difficulty of power analysis attacks. Random operation timing: Design a random delay circuit and insert random delay units, such as LUT lookup tables and flip-flops, on the critical path of the encryption operation. Use a pseudo-random number generator to control the insertion position and duration of the random delay to introduce time randomness. By dynamically adjusting the parameters of the random delay, such as delay granularity and delay probability, adapt to different security requirements and performance requirements.

[0013] Furthermore, write the SM4 encryption IP core using a hardware description language: Use Verilog HDL or VHDL hardware description language to write the RTL (Register Transfer Level) - level description code of the SM4 encryption IP core according to the specifications and processes of the SM4 encryption algorithm. In the RTL - level description, define parameters such as the input and output interfaces, data bit - widths, and encryption rounds of the SM4 encryption, and implement each functional module of the SM4 encryption, such as round key generation, non - linear transformation, and linear transformation, through sequential logic and combinational logic circuits. Adopt a structured design method to divide the SM4 encryption IP core into multiple sub - modules to improve the readability and maintainability of the code. Through simulation testing and verification, ensure that the RTL - level description of the SM4 encryption IP core meets the functional and performance requirements of the SM4 encryption algorithm.

[0014] Logic synthesis to generate a gate - level netlist: Input the RTL - level description code of the SM4 encryption IP core into a logic synthesis tool, such as Synopsys Design Compiler, Cadence RTL Compiler, etc. In the logic synthesis tool, set the technology library and constraint files of the target FPGA device, and specify design metrics such as timing constraints and area constraints. Run the logic synthesis tool to perform logic optimization and mapping on the RTL - level description of the SM4 encryption IP core, and convert the RTL - level description into a gate - level netlist. The gate - level netlist consists of basic logic units (such as AND gates, OR gates, NOT gates, flip - flops, etc.) and describes the logical functions and connection relationships of the SM4 encryption IP core. Perform functional verification and timing analysis on the generated gate - level netlist to ensure that the synthesized circuit is functionally consistent with the original RTL - level description and meets the timing constraints.

[0015] Placement and routing to map to FPGA physical resources: Input the gate - level netlist file of the SM4 encryption IP core into an FPGA placement and routing tool, such as Xilinx Vivado, Intel Quartus Prime, etc. In the placement and routing tool, select the target FPGA device model and set physical constraint conditions such as pin constraints and timing constraints. Run the placement and routing tool to map the logic units of the SM4 encryption IP core to the physical resources of the FPGA device, such as lookup tables (LUTs), flip - flops (FFs), etc. The placement and routing tool uses heuristic algorithms and optimization strategies to perform placement and routing on the mapped physical resources, determine the positions and connection methods of the physical resources, and generate a bit - stream file after placement and routing. Perform timing analysis and power consumption analysis on the circuit after placement and routing to verify whether the timing performance and power consumption characteristics of the circuit meet the design requirements.

[0016] Clock Tree Balancing and Asynchronous Reset: Before performing FPGA placement and routing, balance and optimize the clock tree on the FPGA chip to ensure that the clock arrival times of all physical resources of the SM4 encryption IP core are consistent, reducing clock skew and jitter. Adopt clock tree synthesis (CTS) technology to automatically generate a balanced clock tree structure, insert clock buffers and delay units, and adjust the propagation path and driving strength of the clock signal. In the top-level module of the SM4 encryption IP core, add an asynchronous reset signal to place the SM4 encryption IP core in a preset initial state. The asynchronous reset signal is connected to the key registers of the SM4 encryption IP core. When the reset signal is valid, clear or set the registers to the preset value to ensure that the SM4 encryption IP core starts working from a known state.

[0017] Port Mapping Connection to Form the SM4 Encryption Core: Connect the placed and routed SM4 encryption IP core to the physical resources of the FPGA chip through port mapping to form a complete SM4 encryption core. In the top-level design file of the FPGA device, instantiate multiple SM4 encryption IP cores and connect their input and output ports to the physical pins or internal interconnection resources of the FPGA. Through port mapping, connect the data input, key input, encryption enable, reset, and other control signals of the SM4 encryption IP core to the corresponding ports of the FPGA. At the same time, connect the data output of the SM4 encryption IP core to the output port of the FPGA or the input port of the next-level module to form a complete data encryption path.

[0018] Multiple SM4 encryption cores are configured and connected in parallel to form a parallel SM4 encryption array, improving the overall encryption efficiency and throughput. Multiple SM4 encryption cores are configured and implemented on the FPGA chip, covering the complete design process from high-level language description to physical resource mapping. Utilize the reconfigurable characteristics and parallel processing capabilities of the FPGA to achieve efficient SM4 encryption hardware acceleration, meeting the encryption requirements of high throughput and low latency. At the same time, through technical means such as clock tree balancing and asynchronous reset, ensure the reliability and stability of the SM4 encryption core, providing a solid hardware foundation for realizing secure and trustworthy data encryption.

[0019] Furthermore, for the selection and application of the logic synthesis tool and the placement and routing synthesis tool, the selection of the logic synthesis tool: Select at least one logic synthesis tool from Synopsys Design Compiler, Cadence RTL Compiler, and Mentor Graphics Precision to convert the RTL-level description of the SM4 encryption IP core into a logic gate-level netlist. Synopsys Design Compiler: A widely used logic synthesis tool in the industry, providing high-quality synthesis results and optimization capabilities. Supports multiple hardware description languages such as Verilog, VHDL, and SystemVerilog, etc. Provides rich constraint setting options such as timing constraints, area constraints, power consumption constraints, etc., which can be optimized according to design requirements. Integrates functions such as design rule check (DRC) and floorplanning, which helps to detect design problems early.

[0020] Cadence RTL Compiler: A logic synthesis tool from Cadence, seamlessly integrated with Cadence's digital IC design flow. Adopts leading synthesis algorithms and optimization strategies to generate high-quality logic gate-level netlists. Supports multiple hardware description languages and mixed-language designs, providing flexible design input methods. Provides comprehensive design constraint management and design space exploration functions to help designers quickly converge the design. Mentor Graphics Precision: A logic synthesis tool of Mentor Graphics (now Siemens EDA), widely used in FPGA and ASIC designs. Adopts mature synthesis engines and optimization technologies to generate area-optimized and performance-optimized logic gate-level netlists. Supports multiple hardware description languages, providing flexible design input and constraint setting options. Integrates function verification and timing analysis functions to ensure the correctness and performance of the synthesized circuit. The selection of the placement and routing synthesis tool: Select at least one placement and routing synthesis tool from Xilinx Vivado, Intel Quartus, and Lattice Diamond to map the logic gate-level netlist of the SM4 encryption IP core to the physical resources of the FPGA device and complete the placement and routing.

[0021] Xilinx Vivado: The FPGA design suite of Xilinx, which supports the design and implementation of Xilinx's full range of FPGA devices. It provides integrated synthesis, placement and routing, and verification tools to achieve a complete design process from RTL-level description to bitstream generation. It adopts advanced placement and routing algorithms and optimization strategies, such as clock tree synthesis, congestion-aware placement, etc., to generate high-quality placement and routing results. It provides rich design constraints and optimization options, such as timing constraints, pin constraints, power consumption optimization, etc., to meet different design requirements. Intel Quartus: The FPGA design suite of Intel (formerly Altera), which supports the design and implementation of Intel FPGA devices. It provides a complete design process, including synthesis, placement and routing, timing analysis, and power consumption analysis, etc. It adopts an intelligent placement and routing engine and optimization algorithms to automatically complete the mapping and interconnection of physical resources and generate efficient placement and routing results. It provides rich design constraints and optimization options, such as timing optimization, logic optimization, power consumption optimization, etc., to meet the design requirements of high performance and low power consumption. Lattice Diamond: The FPGA design suite of Lattice Semiconductor, which supports the design and implementation of Lattice FPGA devices. It provides an integrated design environment, including tools such as synthesis, placement and routing, simulation, and debugging. It adopts optimized placement and routing algorithms and strategies to automatically complete the mapping and routing of logic cells and generate compact and efficient placement and routing results. It provides flexible design constraints and optimization options, such as timing constraints, pin constraints, power consumption optimization, etc., to meet different design requirements.

[0022] Application process of logic synthesis tool and placement and routing synthesis tool: In the selected logic synthesis tool, use the RTL-level description of the SM4 encryption IP core as the input, set appropriate design constraints and optimization options, run the synthesis process, and generate a logic gate-level netlist. Perform functional verification and timing analysis on the generated logic gate-level netlist to ensure the correctness and performance of the synthesized circuit meet the requirements. Import the synthesized logic gate-level netlist into the selected placement and routing synthesis tool, and set the target FPGA device model and physical constraints, such as pin assignment, clock constraint, etc. In the placement and routing synthesis tool, run the placement and routing process, map the logic gate-level netlist to the physical resources of the FPGA device, such as lookup tables (LUTs), flip-flops (FFs), etc., and complete the placement and routing of the physical resources. Perform timing analysis, power consumption analysis, and design rule checking on the placed and routed circuit to verify whether the performance, power consumption, and manufacturability of the circuit meet the design requirements. If the verification results meet the requirements, generate the final FPGA bitstream file for the configuration and implementation of the FPGA device. By selecting appropriate logic synthesis tools and placement and routing synthesis tools and following the standard design process, the implementation of the SM4 encryption IP core on the FPGA can be efficiently completed. These tools provide powerful synthesis, optimization, and placement and routing functions, helping designers quickly and reliably map the SM4 encryption algorithm to the FPGA hardware platform to achieve high-performance and low-power encryption operations. At the same time, through reasonable design constraints and optimization strategies, the parallel processing ability and hardware acceleration advantages of the FPGA device can be fully utilized to meet the high-throughput and low-latency encryption requirements.

[0023] Furthermore, for the design of the fetcher, round function calculator, and key adder in the pipeline architecture, the fetcher adopts an architecture that supports variable-length packet fetching. According to the block length of the SM4 encryption algorithm, the input plaintext data to be encrypted is divided into packets, the plaintext data is divided into multiple data packets of equal length, and the divided data packets are sequentially output to the next stage of the pipeline; the round function calculator adopts a programmable architecture that supports dynamically configuring the number of encryption rounds. The number of rounds of SM4 encryption is set through configurable registers, and the number of round function operations is controlled according to the configured number of rounds. In each round of round function operation, non-linear transformation and linear transformation are performed on the input data packet, and the round function operation result is output to the next stage of the pipeline; the key adder adopts an architecture that supports key packet division and packet key prediction. The input encryption key is divided into multiple key packets, the key packets are extended through a lookup table method, the next key packet is predicted based on the current key packet, and the round function operation result is XORed with the predicted key packet to generate the encryption result for the current round.

[0024] Further, the data fetcher includes a grouped storage area: The grouped storage area is a storage unit for caching the plaintext data groups to be fetched. According to the block length (128 bits) of the SM4 encryption algorithm, the bit width of the grouped storage area is set to 128 bits to store a complete plaintext data group. The depth of the grouped storage area is appropriately set according to the design requirements and throughput requirements to meet the needs of data caching and reading. The grouped storage area is implemented using dual-port RAM (DPRAM), with one port for writing the plaintext data group and the other port for reading the plaintext data group. Circular buffer: The circular buffer manages the plaintext data groups using a linked list structure, and realizes data caching and reading through the insertion and deletion operations of the linked list. The linked list structure contains multiple linked list nodes, and each linked list node is used to record the start address and block length of a plaintext data group.

[0025] The data structure of the linked list node is as follows:

[0026] struct ListNode{uint32_t startAddr; / / The start address of the plaintext data group

[0027] uint32_t length; / / The length of the plaintext data group

[0028] ListNode*next; / / Pointer to the next linked list node};

[0029] The head pointer (head) of the linked list points to the first node in the linked list and is used to identify the start position of the linked list. The linked list operation interface provides management operations for inserting and deleting nodes in the linked list, including: insertNode(ListNode*node): Inserts a new linked list node at the end of the linked list. deleteNode(ListNode*node): Deletes the specified linked list node from the linked list. The circular buffer is implemented using SRAM or distributed RAM (DRAM), and stores and reads the plaintext data groups according to the start address and block length recorded by the linked list nodes. Data fetching controller: The data fetching controller traverses the linked list through the linked list operation interface according to the block length of the SM4 encryption algorithm, and sequentially reads the plaintext data groups recorded in the linked list.

[0030] The working process of the data fetching controller is as follows: Initialize the linked list head pointer (head) to be empty. When a new plaintext data packet arrives, insert a new linked list node into the tail of the linked list through the linked list operation interface insertNode(). According to the block length of the SM4 encryption algorithm, traverse the linked list through the linked list operation interface and sequentially read the plaintext data packets recorded in the linked list nodes. For each read plaintext data packet, read the corresponding data from the circular buffer according to the starting address and block length recorded in the linked list node. After reading a plaintext data packet, delete the corresponding linked list node through the linked list operation interface deleteNode() to perform variable-length block data reading. Repeat until the linked list is empty, indicating that all plaintext data packets have been read. The data fetching controller manages the caching and reading of plaintext data packets in the circular buffer by controlling the write pointer and read pointer of the circular buffer. The write pointer (writePtr) points to the next writable position in the circular buffer. The read pointer (readPtr) points to the next position to be read in the circular buffer. When the write pointer catches up with the read pointer, it means the circular buffer is full and new data packets cannot be written until the read pointer advances. When the read pointer catches up with the write pointer, it means the circular buffer is empty and new data packets need to be written by the write pointer before further reading can occur.

[0031] Furthermore, the round function arithmetic unit includes: Round number configuration register: The round number configuration register is used to store the round number configuration value of the SM4 encryption algorithm, which is set according to security and performance requirements. The bit width of the round number configuration register is designed according to the maximum number of rounds of the SM4 encryption algorithm. For example, using a 6-bit register can support up to 64 rounds of encryption operations. The round number configuration register can be configured through software or hardware methods, and a suitable configuration interface is selected according to actual needs. Round function arithmetic module: The round function arithmetic unit contains multiple round function arithmetic modules, and each module is used to perform one round of encryption operation of the SM4 encryption algorithm. The input of each round function arithmetic module is a data packet (128 bits) and the corresponding round key, and the output is the encrypted data packet (128 bits). The round function arithmetic module internally contains two sub-modules: a non-linear transformation sub-module and a linear transformation sub-module. The non-linear transformation sub-module performs S-box (Substitution Box) replacement operations on the input data packet to achieve byte-level non-linear transformation. The S-box can be implemented using a lookup table (LUT) or combinational logic circuits. The linear transformation sub-module performs linear transformation on the result after S-box replacement, including circular left shift and exclusive OR operations. The linear transformation can be implemented using shift registers and exclusive OR gate circuits. The output of the round function arithmetic module is XORed with the corresponding round key to obtain the encryption result of the current round.

[0032] Round Function Selector: The input end of the round function selector is connected to the outputs of each round function operation module, and the output end is connected to the next stage of the pipeline. The round function selector selects and passes the output of the round function operation module corresponding to the round number to the next stage of the pipeline according to the value of the round number configuration register. The round function selector can be implemented using a multiplexer (MUX) or a data selector, and controls the selection signal according to the round number configuration value. Round Key Generation Module: The round key generation module generates the round key for each round through the key expansion algorithm of the SM4 encryption algorithm based on the input initial key. The key expansion algorithm includes the following steps: Divide the initial key into 4 sub-keys (each sub-key is 32 bits). Perform circular left shift and S-box substitution operations on each sub-key to generate the expanded sub-key. Perform exclusive OR operation on the expanded sub-key and the fixed parameter to obtain the round key. Repeat to generate the round keys for all rounds. The round key generation module provides the generated round key to the corresponding round function operation module for encryption operations. The round key generation module can use a look-up table (LUT) or combinational logic circuit to implement the key expansion algorithm and generate the corresponding round key according to the initial key and the round number.

[0033] Round Function Controller: The round function controller controls the transfer of data packets between each stage of the pipeline module according to the process of the SM4 encryption algorithm, controls the round function selector to perform round function switching, controls the round key generation module to provide the round key, and performs multi-round encryption operations of the SM4 encryption algorithm. The working process of the round function controller is as follows: Receive the data packet from the previous stage of the pipeline and transfer it to the first round function operation module. According to the value of the round number configuration register, control the round function selector to select and pass the output of the round function operation module corresponding to the round number to the next stage of the pipeline. Control the round key generation module to generate the round key corresponding to the round number and provide it to the corresponding round function operation module. Wait for the completion of the encryption operation of the current round and transfer the encryption result to the next stage of the pipeline. Repeat until all rounds of encryption operations are completed. The round function controller can be implemented using a finite state machine (FSM), and controls the transfer of data packets, round function selection, and round key generation according to the round number configuration and the process of the SM4 encryption algorithm.

[0034] Furthermore, the key adder includes: Key Group Caching Module: The key group caching module is used to cache the key used in the SM4 encryption algorithm and divide the input initial key into 4 sub-keys for caching. The input of the key group caching module is a 128-bit initial key, and the output is 4 32-bit sub-keys. The key group caching module is implemented using registers or distributed RAM (DRAM), divides the initial key into 4 sub-keys according to the positional relationship of the initial key, and stores them in the corresponding registers or DRAM. The storage order of the sub-keys can be adjusted according to the requirements of the SM4 encryption algorithm to facilitate subsequent key expansion operations.

[0035] Key Expansion Module: The key expansion module expands the sub - keys output by the key grouping cache module through a table - lookup method, and randomly permutes the data bits of the expanded sub - keys to generate the round key for the current encryption round. The input of the key expansion module is 4 32 - bit sub - keys, and the output is a 32 - bit round key. The key expansion module uses a lookup table (LUT) to implement the key expansion operation of the SM4 encryption algorithm, and looks up the corresponding expansion result in the table according to the value of the sub - key. The lookup table can be implemented using ROM or combinational logic circuits, and the corresponding lookup table content is generated according to the key expansion rules of the SM4 encryption algorithm. Randomly permuting the data bits of the expanded sub - keys increases the randomness and security of the key. The data - bit permutation operation can be implemented using a pseudo - random number generator (PRNG) or a fixed permutation pattern. The permuted expanded sub - key is used as the round key for the current encryption round and is used for XOR operation with the output result of the round - function arithmetic unit.

[0036] Key XOR Module: The key XOR module performs a bit - by - bit XOR operation on the intermediate result output by the round - function arithmetic unit and the round key generated by the key expansion module to obtain the output result of the current encryption round. The input of the key XOR module is the 32 - bit intermediate result output by the round - function arithmetic unit and the 32 - bit round key generated by the key expansion module, and the output is a 32 - bit XOR result. The key XOR module uses XOR gate circuits to implement the bit - by - bit XOR operation, performing an XOR operation on each bit of the intermediate result and the round key to obtain the XOR result. The XOR result is used as the output result of the current encryption round and is passed to the next round through a pipeline for subsequent encryption operations.

[0037] Key Controller: According to the round number configuration of the SM4 encryption algorithm, in each encryption round, the key controller sequentially triggers the key expansion module to generate the round key, controls the key XOR module to perform the round - key XOR operation, and passes the XOR result to the next round through a pipeline. The key controller is implemented by a finite - state machine (FSM), and judges whether the key expansion and key XOR are completed according to the current round number and the SM4 encryption round number. The working process of the key controller is as follows: Initialize the round - counter to 0, indicating that the current is the first encryption round. Trigger the key expansion module to generate the corresponding round key according to the current sub - key. Control the key XOR module to perform an XOR operation on the intermediate result output by the round - function arithmetic unit and the round key. Pass the XOR result to the next round through a pipeline. Judge whether the current round number has reached the SM4 encryption round - number configuration value. If not, increment the round - counter and return to continue executing the next round. If the SM4 encryption round - number configuration value has been reached, it means that the encryption process is completed, and the final encryption result is output. The key controller controls the execution times of the key expansion and key XOR by judging the relationship between the round - counter and the SM4 encryption round - number configuration value to ensure the completion of the encryption operation for the specified number of rounds.

[0038] Furthermore, a three-way handshake verification mechanism is adopted. By establishing independent handshake channels among the data fetcher, the round function arithmetic unit, and the key adder, and utilizing the handshake password generated by the pseudo-random number generator and the output of the physical non-deterministic source PHYS, the authentication and communication synchronization among the internal modules of the SM4 encryption core are realized, improving the security and reliability of the SM4 algorithm. Acquisition of handshake signals: A pseudo-random number generator is set on the FPGA chip to generate pseudo-random numbers at different times as handshake passwords. The pseudo-random number generator can be implemented by hardware circuits such as linear feedback shift registers (LFSRs), ensuring the randomness and unpredictability of the handshake passwords.

[0039] Independent handshake channels are set among the data fetcher, the round function arithmetic unit, and the key adder. This channel uses the high-speed serial transceivers of the FPGA (such as LVDS, SerDes, etc.) to achieve high-speed and low-latency transmission of handshake signals, and is independent of the normal data channel without affecting the transmission efficiency of SM4 encrypted data. Physical non-deterministic sources PHYS are set at the sending and receiving ends of the independent channel. PHYS can use random signal sources such as environmental noise and circuit noise. After processing such as ADC sampling, filtering, and quantization, the random characteristics of the physical noise are extracted as an additional factor for handshake password verification. Combining the output of PHYS with the pseudo-random number handshake password constructs a handshake password challenge-response verification mechanism based on physical non-determinism, further enhancing the security of the handshake verification.

[0040] Handshake process: After the data fetcher completes fetching the plaintext block, it obtains the handshake password 1 at the current moment from the pseudo-random number generator, packs the password 1 with the fetch completion identification code and the timestamp, generates the first handshake signal, and sends the first handshake signal to the round function arithmetic unit through the handshake interface on the independent channel. After receiving the first handshake signal, the round function arithmetic unit verifies the password 1 in it. The verification method is to compare the received password 1 with the output of the local PHYS of the round function arithmetic unit to determine the validity of the password 1. After the verification passes, the round function arithmetic unit obtains the handshake password 2 at the next moment from the pseudo-random number generator, packs the password 2 with the round function completion identification code, the timestamp, and the first handshake signal, generates the second handshake signal, and sends the second handshake signal to the key adder through the handshake interface.

[0041] After receiving the second handshake signal, the key adder verifies the password 2 therein, and the verification method is similar to that of the round function arithmetic unit. After successful verification, the key adder obtains the handshake password 3 for the next moment from the pseudo-random number generator, packs the password 3 with the encryption completion identification code, time stamp, first handshake signal and second handshake signal, generates the third handshake signal, and returns the third handshake signal to the data fetcher through the handshake interface. After receiving the third handshake signal, the data fetcher verifies the password 3 therein, and the verification method is similar to the previous two times. After successful verification, the data fetcher completes a complete handshake process, the identities of all modules of the SM4 encryption core are verified, and the communication is also synchronized, and the next round of encryption process can be entered. By using the pseudo-random number handshake password and the physical non-deterministic source, through the three-way handshake verification method, the identity verification and communication synchronization mechanism between the internal modules of the SM4 encryption core are established. The randomness and unpredictability of the handshake password, as well as the introduction of the physical non-deterministic source, enhance the security of the handshake verification and prevent threats such as man-in-the-middle attacks and replay attacks. The independent handshake channel ensures the efficient transmission of the handshake signal without affecting the processing efficiency of the SM4 encrypted data. The introduction of the time stamp prevents the delay and replay of the handshake signal.

[0042] Furthermore, by using the time stamp information carried in the handshake signal, the overall operation efficiency of the SM4 encryption core is obtained by calculating the operation time of each module in the SM4 encryption core. By introducing a weight coefficient, the influence degree of different modules on the encryption efficiency is considered, making the efficiency evaluation more accurate and comprehensive. Parsing of the handshake signal: Parse the collected first handshake signal, second handshake signal and third handshake signal, and extract the time stamp information carried therein. The time stamp information can adopt a high-precision time stamp format, such as the nanosecond-level representation of the UNIX time stamp, to ensure the accuracy of time measurement.

[0043] Calculate the operation times of the data fetcher, round function calculator, and key adder respectively according to the time stamp differences of adjacent handshake signals. The specific methods are as follows: The data fetching controller calculates the operation time T1 of the data fetcher according to the start time stamp T1_start of the first handshake signal and the end time stamp T3_end of the third handshake signal. Here, T1 includes all the time for the data fetcher to complete the plaintext block fetching, generate the first handshake signal, and wait for the return of the third handshake signal. The round function controller calculates the operation time T2 of the round function calculator according to the end time stamp T1_end of the first handshake signal and the end time stamp T2_end of the second handshake signal. Here, T2 includes the time for the round function calculator to receive the first handshake signal, complete the round function operation, and generate the second handshake signal. The key controller calculates the operation time T3 of the key adder according to the end time stamp T2_end of the second handshake signal and the start time stamp T3_start of the third handshake signal. Here, T3 includes the time for the key adder to receive the second handshake signal, complete the key addition operation, and generate the third handshake signal.

[0044] Calculation of operation efficiency: According to the operation times T1, T2, T3 of the data fetcher, round function calculator, and key adder, and their weight coefficients W1, W2, W3 that affect the encryption efficiency, calculate the operation efficiency E of the SM4 encryption core through the formula E = 1 / (W1*T1 + W2*T2 + W3*T3). The values of the weight coefficients W1, W2, W3 can be estimated and adjusted according to factors such as the complexity and operation volume of each module in the SM4 encryption core to reflect their influence degrees on the encryption efficiency. For example, the round function calculator usually has a higher complexity and operation volume, and its weight coefficient W2 can take a larger value; while the operations of the data fetcher and the key adder are relatively simple, and their weight coefficients W1 and W3 can take smaller values. The operation efficiency E calculated through the formula comprehensively considers the operation times and influence weights of each module in the SM4 encryption core and obtains a comprehensive efficiency evaluation result. The larger the value of E, the higher the operation efficiency of the SM4 encryption core and the faster the encryption speed.

[0045] Furthermore, random polynomial operations and logical operations are introduced into each module of the SM4 encryption core. The operation process of each part is coordinated by the hybrid operation control module. When the encrypted data flows inside SM4, it undergoes random encoding, obfuscation, etc. The data flow is complex and changeable, making it difficult for attackers to capture and analyze, thereby improving the security of the SM4 algorithm. When the data fetching controller reads the plaintext data: The polynomial coefficient generation module uses an LFSR to generate random values of variables x, y, and z in the polynomial f(x, y, z) = (ax + b)y + cz. The Bracewell sequence B(n) generates the polynomial coefficients a, b, and c. The random seed of the LFSR and the starting value n of B(n) are provided by the seed generator, ensuring the randomness of the polynomial coefficients. The plaintext data is used as the input variable of the polynomial f(x, y, z), and the plaintext is randomly encoded through polynomial operations. The encoding result is passed as the output of the data fetching controller to the round function arithmetic unit, masking the statistical characteristics of the plaintext. When the round function arithmetic unit performs non-linear transformation: The polynomial operation module expands the operation process of f(x, y, z) through the Horner algorithm. The specific steps are as follows: Initialize the result register res to 0; sequentially read the coefficients c, b, and a of the polynomial z, y, and x terms from B(n); calculate the x term tmp = ax + R, where R is a random number generated by the LFSR; calculate the y term tmp = tmpy + b; calculate the z term res = tmpz + S, where S is another random number generated by the LFSR. The polynomial operation result is used as the intermediate result of the S-box substitution, and then it is obfuscated through logical operations such as exclusive OR, AND, and OR with the output of the S-box through the logical operation circuit to obtain the obfuscated non-linear transformation result. During linear transformation, a randomization factor is introduced by performing logical operations such as exclusive OR between the linear transformation result and the polynomial operation result. The polynomial operation and the logical operation are alternated, increasing the complexity of the round function operation.

[0046] During the process of the key adder generating round keys and performing the round key addition operation, random polynomial operations and logical operations are introduced to randomly encode and obfuscate the round keys and the results of the round function, disrupt the regularity of the key addition operation, and increase the randomness of the key addition result, thereby enhancing the security of the SM4 algorithm. Generation process of round keys: The polynomial coefficient generation module uses an LFSR to generate random values of the variable k in the polynomial f(k) = ak^2 + bk + c, and the Bracewell sequence B(n) generates the polynomial coefficients a, b, and c. The random seed of the LFSR and the starting value n of B(n) are provided by the seed generator, ensuring the randomness of the polynomial coefficients. Each sub-key in the key group is used as the input variable of the polynomial f(k), and the sub-key is randomly encoded through the polynomial operation module. The randomly encoded sub-key then generates the round key through the key expansion algorithm, and the expanded round key is XOR-mixed with the random encoding result to obtain the randomized round key. By introducing random polynomial operations in the round key generation process, the randomness of the round key is increased, making it difficult for attackers to infer the key by analyzing the statistical characteristics of the key.

[0047] Process of the round key addition operation: The result of the round function operation and the randomized round key first perform logical operations such as XOR, AND, and OR through the logical operation circuit to introduce non-linear obfuscation. The logical operation result is then XOR-mixed with the random polynomial operation result to obtain the mixed key addition result. Among them, the random polynomial operation result is generated by the polynomial operation module. Using a method similar to the round key generation process, an LFSR and a Bracewell sequence are used to generate random polynomial coefficients to randomly encode the result of the round function. The key controller controls the timing of the mixed operation, enabling the random polynomial operation, logical operation, and round key addition operation to be executed alternately, disrupting the regularity of the key addition operation and increasing the uncertainty of the key addition result, making it difficult for attackers to infer the key by analyzing the timing regularity of the key addition operation.

[0048] Process of iterative round encryption: After completing one round of SM4 encryption, the encryption result of this round is XOR-mixed with the result of random polynomial operation as the input for the next round of encryption. By introducing random polynomial operation between rounds, the encryption result of the previous round is randomly encoded, making the input of each round random, increasing the diffusion of the SM4 algorithm and making the avalanche effect more obvious. After the last round of encryption is completed, the final result is encoded again through random polynomial operation to obtain the confused SM4 ciphertext output. This step increases the randomness of the ciphertext, making the statistical characteristics of the ciphertext more uniform and difficult to be exploited by attackers. During the polynomial operation process, the random numbers R and S generated by LFSR are inserted to participate in the calculation of the intermediate results tmp and res, introducing randomization factors to dynamically disguise the polynomial operation results; and by changing the clock cycle of the polynomial operation irregularly, the power consumption characteristics of the polynomial operation are adjusted to increase the difficulty of power analysis.

[0049] After the polynomial operation is completed, the results tmp and res are input into the logic operation circuit, and the operation results are further confused through operations such as XOR, AND, and OR; and the logic operation results are fed back to the input of the polynomial operation, repeating the process of polynomial operation and logic operation to form a circular data path, and multiple iterations of the mixed operation improve the data complexity. Control signals for polynomial operation and logic operation are provided, collaborating with the data fetch controller, round function controller, and key controller to ensure that the mixed operation is synchronized with the SM4 encryption process, so that when the encrypted data is transmitted inside SM4, it all undergoes random polynomial mixing operation processing, and the data flow is complex and changeable, making it difficult for attackers to capture.

[0050] Compared with the prior art, the advantages of this application are as follows:

[0051] By configuring multiple SM4 encryption cores with a pipeline architecture on the FPGA chip, the efficiency and throughput of SM4 encryption are improved by using multi-core parallel processing. The pipeline architecture divides the encryption operation into three stages: data fetch, round function operation, and key addition, realizing fine-grained parallelism of the encryption process, giving full play to the parallel computing advantage of the FPGA chip, and improving the encryption speed.

[0052] Collect the handshake signals between each stage inside the SM4 encryption core, accurately reflect the operation status of each stage through the handshake signals, and calculate the real-time operation efficiency of the SM4 encryption core. Dynamically adjust the FPGA resource allocation between SM4 encryption cores according to the operation efficiency using the priority scheduling algorithm, realizing load balancing and resource optimization during the encryption process, and improving the resource utilization rate of the FPGA chip.

[0053] Utilize the physical-noise-based random number extraction circuit integrated on the FPGA chip to generate high-quality random numbers as the encryption keys for the SM4 encryption core. Introducing a physical noise source enhances the randomness of the keys, improves the unpredictability of the keys, and effectively resists attacks such as key guessing and key analysis.

[0054] Adopt a calculation method that combines polynomial operations and logical operations inside the SM4 encryption core. Through non-linear polynomial operations, encode and confuse the plaintext data and intermediate results, mask the statistical characteristics of the data, and increase the difficulty of power analysis attacks. Mixing polynomial operations and logical operations further improves the operation complexity and enhances the ability of the SM4 encryption algorithm to resist side-channel attacks.

[0055] Randomly insert waiting periods during the operation process of the SM4 encryption core. By randomizing the operation timing, introduce time jitter, disrupt the correlation between the encryption operation and the energy consumption, increase the difficulty of side-channel attacks such as power analysis and electromagnetic analysis, and improve the strength of the SM4 encryption against side-channel attacks. Brief Description of the Drawings

[0056] This application will be further described in the manner of exemplary embodiments, and these exemplary embodiments will be described in detail through the drawings. These embodiments are not restrictive. In these embodiments, the same numbers represent the same structures, where:

[0057] Figure 1 is an exemplary flowchart of an SM4 encryption method based on FPGA according to some embodiments of this application;

[0058] Figure 2 is an exemplary flowchart of configuring the encryption core according to some embodiments of this application;

[0059] Figure 3 is an exemplary flowchart of collecting handshake signals according to some embodiments of this application. Detailed Embodiments

[0060] The methods and systems provided in the embodiments of this application will be described in detail below with reference to the drawings.

[0061] Figure 1An exemplary flowchart of an SM4 encryption method based on FPGA according to some embodiments of the present application includes: obtaining plaintext data to be encrypted and inputting the plaintext data into an FPGA chip; configuring multiple SM4 encryption cores on the FPGA chip and using the SM4 encryption cores to perform parallel encryption operations on the plaintext data to obtain ciphertext data; wherein each SM4 encryption core adopts a pipeline architecture, and the pipeline architecture includes a data fetcher, a round function calculator, and a key adder, which respectively execute the three stages of data fetching, round function calculation, and key addition in the encryption operation; using a physical-noise-based random number extraction circuit integrated on the FPGA chip to generate an N-bit random number, and using the generated N-bit random number as the encryption key for each SM4 encryption core for SM4 encryption operations; inside each SM4 encryption core, collecting handshake signals between the data fetcher, the round function calculator, and the key adder, and the handshake signals are used to reflect the operation status of the data fetcher, the round function calculator, and the key adder; calculating the operation efficiency of each SM4 encryption core according to the collected handshake signals; dynamically adjusting the FPGA resource allocation between the SM4 encryption cores according to the calculated operation efficiency by using a priority scheduling algorithm; during the process of each SM4 encryption core performing the encryption operation, calculating in a manner that mixes polynomial operations and logical operations to improve the power analysis attack resistance; and randomly inserting wait cycles in the encryption operation to improve the side-channel attack resistance of SM4 encryption by randomizing the operation timing.

[0062] Obtain plaintext data to be encrypted and input the plaintext data into the FPGA chip: Receive the plaintext data to be encrypted sent from external devices (such as a host computer, a sensor, etc.) through external interfaces of the FPGA chip, such as GPIO, UART, SPI, I2C, etc. Transmit the received plaintext data to the data cache module inside the FPGA chip through a data bus, such as FIFO, RAM, etc. The bit width and depth of the data cache module are designed according to the block length and throughput requirements of the SM4 encryption algorithm. The data cache module is connected to the data input interface of the SM4 encryption core, and data synchronization transmission is achieved through handshake signals. When the amount of plaintext data stored in the data cache module reaches the input data length of the SM4 encryption core, the data fetcher of the SM4 encryption core starts to read the plaintext data from the data cache module.

[0063] Figure 2Exemplary flowchart for configuring encryption cores according to some embodiments of the present application. Configure multiple SM4 encryption cores on the FPGA chip: Using Verilog HDL or VHDL hardware description language, write the RTL-level code of the SM4 encryption IP core according to the specifications of the SM4 encryption algorithm. The code includes three modules: a data fetcher, a round function calculator, and a key adder, and they are connected in a pipelined manner. Use logic synthesis tools such as Synopsys Design Compiler, Cadence RTL Compiler, or Mentor Graphics Precision to perform logic synthesis on the RTL-level code of the SM4 encryption IP core to generate the gate-level netlist file of the SM4 encryption IP core. During the synthesis process, according to the characteristics of the FPGA device, constrain and optimize the timing, area, and power consumption of the SM4 encryption IP core. Use placement and routing tools such as Xilinx Vivado, Intel Quartus, or Lattice Diamond to perform placement and routing on the gate-level netlist of the SM4 encryption IP core, map the logic units in the SM4 encryption IP core to the physical resources (such as LUTs, FFs, etc.) of the FPGA device, and complete the wiring between the physical resources. During the placement and routing process, balance the clock path delays of each physical resource of the SM4 encryption IP core through clock constraints and clock tree synthesis to ensure clock synchronization.

[0064] In the top-level module of the SM4 encryption IP core, initialize the registers and storage units in the SM4 encryption IP core to a preset state through an asynchronous reset signal. Integrate the SM4 encryption IP core into the top-level design of the FPGA chip by port mapping or encapsulating it into a standard interface such as AXI4. According to the throughput requirements of SM4 encryption, instantiate multiple SM4 encryption IP cores in the top-level design to form a parallel SM4 encryption core array. Utilize the on-chip interconnection resources of the FPGA chip, such as the AXI4 internal bus, to connect the SM4 encryption core array with other functional modules (such as key management, data caching, external interfaces, etc.) to build a complete SM4 encryption system. Finally, synthesize, place, and route the FPGA top-level design to generate the FPGA bitstream file and download it to the FPGA chip to complete the configuration and deployment of the SM4 encryption system.

[0065] Using a physical-noise-based random number extraction circuit integrated on an FPGA chip to generate N-bit random numbers, and using the random numbers as the encryption keys for each SM4 encryption core for SM4 encryption operations: Integrate a physical-noise-based random number extraction circuit on the FPGA chip, such as a Ring Oscillator or a Chaotic Mapping Circuit, etc. The physical noise source can come from the thermal noise, shot noise, or 1 / f noise inside the FPGA chip, etc. The random number extraction circuit uses digital post-processing to amplify, filter, and quantize the analog noise signal generated by the physical noise source to obtain a uniformly distributed random bit sequence. Use standard test suites such as NIST SP800-22 to detect the randomness of the random bit sequence to ensure the quality of the random numbers. According to the key length requirement of the SM4 encryption algorithm, group and parallelize the random bit sequence to obtain N-bit (such as 128-bit) random numbers. Structures such as shift registers or FIFOs can be used to cache and read the random bit sequence. Distribute the generated N-bit random numbers to each SM4 encryption core on the FPGA chip through a bus or a dedicated interface as its encryption key. The SM4 encryption core encrypts the plaintext data according to the received random key. To improve the security of the random numbers, a dynamic update method can be adopted to periodically (such as every certain number of encryption operations) generate new random keys from the random number extraction circuit again to update the encryption key of the SM4 encryption core.

[0066] Inside each SM4 encryption core, the data fetcher includes: a packet storage area: Implement the packet storage area using the Block RAM (BRAM) resources on the FPGA chip. BRAM is a high-speed dual-port memory built into the FPGA, with independent read and write ports and a large storage capacity. Design the bit width and depth of the BRAM according to the packet length (128 bits) of the SM4 encryption algorithm and the throughput requirement of the encryption core. For example, a BRAM with a bit width of 128 and a depth of n (n is the number of plaintext data packets to be cached) can be used. Write the plaintext data packets to be encrypted into the BRAM in sequence, and each packet occupies a storage unit of the BRAM. The write port of the packet storage area is connected to the data cache module to receive the plaintext data packets from the data cache module; the read port is connected to the data fetch controller to read the plaintext data packets to be encrypted.

[0067] Circular Buffer: Use the on-chip storage resources of the FPGA (such as Distributed RAM or Shift Register LUT) to implement the linked list nodes of the circular buffer. Each linked list node contains two fields: the starting address (pointing to the starting address of the corresponding plaintext data packet in the packet storage area) and the packet length (recording the length of this plaintext data packet, fixed at 128 bits). The linked list head pointer is implemented using a register (such as a Flip-Flop), and the bit width is the same as the address bit width of the storage unit of the linked list node. Initially, the linked list head pointer points to the first linked list node. The linked list operation interface includes two operations: insertion and deletion, which are implemented using a finite state machine (FSM). Insertion operation: When the data cache module writes a new plaintext data packet to the packet storage area, the insertion operation of the linked list operation interface is triggered. The insertion operation takes an idle node from the pool of idle linked list nodes, writes the starting address and packet length of the new plaintext data packet to this node, and inserts this node at the end of the linked list. At the same time, update the linked list head pointer to point to the newly inserted node. Deletion operation: When the fetch controller has read a plaintext data packet, the deletion operation of the linked list operation interface is triggered. The deletion operation removes the currently read linked list node from the linked list and releases this node back to the pool of idle nodes. At the same time, update the linked list head pointer to point to the next linked list node to be read.

[0068] Fetch Controller: The control logic of the fetch controller is implemented using a finite state machine (FSM). The FSM includes states such as idle, read, and wait. The fetch controller obtains the currently pending linked list node through the linked list head pointer. The starting address and packet length of the plaintext data packet are extracted from this node. The fetch controller uses the starting address as the read address to read the corresponding plaintext data packet from the BRAM in the packet storage area. The read data packet is sent to the round function arithmetic unit of the SM4 encryption core for encryption processing. The fetch controller also controls the read pointer of the circular buffer. The read pointer points to the currently being read linked list node. After each data packet is read, the read pointer is incremented to point to the next pending linked list node. When the read pointer reaches the end of the linked list, the fetch controller determines whether the circular buffer is empty. If it is not empty, the read pointer returns to the head of the linked list to continue reading data packets; if it is empty, the fetch controller enters the wait state, waiting for a new plaintext data packet to be written. The fetch controller also controls the write pointer of the circular buffer. The write pointer points to the current position where a new node can be written. When the data cache module writes a new plaintext data packet, the fetch controller triggers an insert operation on the linked list operation interface, inserts the new node into the position pointed to by the write pointer, and updates the write pointer. The fetch controller determines the state of the circular buffer (empty, full, non-empty non-full) based on the position relationship between the read pointer and the write pointer. When the circular buffer is empty, the fetch controller sends a data request signal to the data cache module to request a new plaintext data packet; when the circular buffer is full, the fetch controller pauses sending data requests and waits for the round function arithmetic unit to read the data packet and release buffer space.

[0069] Workflow of the data fetcher: Initially, the linked list is empty, the circular buffer is empty, and the read pointer and write pointer point to the head of the linked list. When the data cache module writes a new plaintext data packet, the data fetch controller triggers the insertion operation of the linked list operation interface, inserts a new node at the end of the linked list, and updates the write pointer. The data fetch controller reads the first node in the linked list through the head pointer of the linked list to obtain the starting address and packet length of the plaintext data packet to be read. The data fetch controller reads the corresponding plaintext data packet from the BRAM in the packet storage area according to the starting address and sends it to the round function arithmetic unit for encryption processing. The data fetch controller triggers the deletion operation of the linked list operation interface, removes the read node from the linked list, and updates the read pointer. Repeat until the linked list is empty, that is, all cached plaintext data packets have been read and encrypted. When the linked list is empty, the data fetch controller enters the waiting state, waiting for a new plaintext data packet to be written. When the data cache module writes a new packet, repeat. The round function arithmetic unit includes: Round number configuration register: The round number configuration register is implemented using the on-chip register resources (such as Flip-Flop) of the FPGA. The bit width of the round number configuration register is determined by the number of rounds of the SM4 encryption algorithm. For the standard SM4 encryption algorithm, the number of rounds is 32, so the bit width of the round number configuration register is 6 bits (2^6 = 64 > 32). The value of the round number configuration register is written by the control module during the initialization of the SM4 encryption core and stores the round number configuration of the current encryption operation. The output of the round number configuration register is connected to the round function selector to control the selection of the output of the round function arithmetic module corresponding to the corresponding number of rounds.

[0070] Round function arithmetic module: According to the specifications of the SM4 encryption algorithm, each round function arithmetic module consists of a non-linear transformation and a linear transformation. The non-linear transformation uses a lookup table (LUT) to implement the S-box substitution operation. The S-box is a non-linear permutation table with 8-bit input and 8-bit output. The S-box lookup table is implemented using the distributed storage resources (such as Distributed RAM or Shift Register LUT) of the FPGA. The linear transformation is implemented using exclusive OR (XOR) operations and circular shift operations. The linear transformation circuit is implemented using the logic resources (such as LUT) and shift registers of the FPGA. The input of each round function arithmetic module is a 128-bit data packet and a 128-bit round key, and the output is a 128-bit encrypted result. The round function arithmetic modules are connected in a pipeline manner, and each module corresponds to one round of encryption operation of the SM4 encryption algorithm. The depth of the pipeline is designed according to the number of rounds of the SM4 encryption algorithm and the target clock frequency.

[0071] Round Function Selector: The round function selector is implemented using the multiplexer in the FPGA. The input end of the round function selector is connected to the outputs of each round function operation module, and the input data width is 128 bits multiplied by the number of round function operation modules. The selection signal of the round function selector comes from the round number configuration register, and the bit width of the selection signal is log2(number of round function operation modules). The output end of the round function selector is connected to the next-level module of the pipeline, and the output data width is 128 bits. According to the value of the round number configuration register, the round function selector gates the output of the round function operation module corresponding to the current round number to the next level of the pipeline. Round Key Generation Module: The round key generation module is implemented using the logic resources of the FPGA according to the key expansion algorithm of the SM4 encryption algorithm. The input of the round key generation module is the 128-bit initial key, and the output is 32 128-bit round keys. The round key generation module generates the round key for each round based on the initial key through exclusive OR operations, S-box substitutions, and circular shift operations. The generated round keys are provided to the corresponding round function operation modules through dedicated data paths. Each round function operation module uses the round key corresponding to the current round number for encryption operations.

[0072] Round Function Controller: The control logic of the round function controller is implemented using a finite state machine (FSM). The round function controller controls the transfer of data packets between the modules at all levels of the pipeline according to the process of the SM4 encryption algorithm. Data synchronization between modules is achieved through handshake signals (such as data valid signal and data ready signal). The round function controller generates the selection signal of the round function selector, and according to the current encryption round number, gates the output of the corresponding round function operation module to the next level of the pipeline. The round function controller controls the operation of the round key generation module, triggers the round key generation at the start of encryption, and provides the generated round keys to the corresponding round function operation modules. The round function controller is also responsible for controlling the progress of data packets in the pipeline, records the current encryption round number through a counter, and compares it with the value of the round number configuration register to determine whether the encryption is completed.

[0073] Working process of the round function arithmetic unit: At the beginning of encryption, the round function controller resets all round function arithmetic modules and the round key generation module, and initializes the value of the round number configuration register to 0. The round function controller triggers the round key generation module to generate 32 round keys according to the input initial key, and distributes the round keys to the corresponding round function arithmetic modules. The data fetch controller reads the first plaintext data packet from the packet storage area and sends it to the first-stage round function arithmetic module of the pipeline. The round function controller controls the data packet to be passed through the pipeline stage by stage. Each time it passes through a round function arithmetic module, a round of encryption operation is performed on the data packet. The round function controller generates a selection signal for the round function selector according to the current encryption round number, and selects and passes the output of the round function arithmetic module corresponding to the round number to the next stage of the pipeline. Repeat until the data packet passes through all round function arithmetic modules and completes 32 rounds of encryption operations. The last stage of the pipeline outputs the encrypted ciphertext data packet, which is stored in the ciphertext buffer. Repeat until all plaintext data packets are encrypted.

[0074] The key adder includes: Key packet caching module: The key packet caching module is implemented using the on-chip storage resources of the FPGA (such as Block RAM or Distributed RAM). According to the key length of the SM4 encryption algorithm (128 bits), the input initial key is divided into 4 sub-keys of 32 bits. 4 registers or storage units of 32 bits are used to cache the 4 sub-keys respectively. The output of each register or storage unit is connected to the input of the key expansion module. The write port of the key packet caching module is connected to the key input interface for receiving the initial key data; the read port is connected to the key expansion module for providing the sub-key data. Key expansion module: The key expansion module is implemented using the logic resources of the FPGA (such as LUT and registers) according to the key expansion algorithm of the SM4 encryption algorithm. The input of the key expansion module is 4 sub-keys of 32 bits, and the output is 1 extended sub-key of 32 bits. A lookup table (LUT) is used to implement the non-linear transformation (such as S-box substitution) in the key expansion process. According to the specification of the SM4 key expansion algorithm, the S-box substitution results are calculated and stored in advance to speed up the expansion process. The linear transformation in the key expansion process is implemented using exclusive OR (XOR) operations and cyclic shift operations. The data bits of the extended sub-key are randomly permuted to increase the randomness and security of the key. A linear feedback shift register (LFSR) or other pseudo-random number generator can be used to generate the random permutation sequence, and the data bit permutation is implemented through a crossbar switch or a multiplexer. The output of the key expansion module is the round key used in the current encryption round, which is connected to the input of the key XOR module.

[0075] Key XOR Module: The key XOR module is implemented using the logic resources (such as LUT) of the FPGA. The inputs of the key XOR module include the intermediate result (128 bits) output by the round function arithmetic unit and the round key (32 bits) generated by the key expansion module. The intermediate result output by the round function arithmetic unit is divided into 4 sub - groups by 32 bits and performs bit - by - bit XOR operations with the round key respectively. The result of the XOR operation is used as the output result of the current encryption round and is passed to the round function arithmetic unit of the next round through the pipeline. The output of the key XOR module is connected to the input of the next - level pipeline module for data transfer. Key Controller: The control logic of the key controller is implemented using a finite - state machine (FSM). The key controller generates control signals according to the round configuration of the SM4 encryption algorithm to trigger the operations of the key expansion module and the key XOR module. In each encryption round, the key controller first triggers the key expansion module to generate the round key according to the sub - key of the current round. Then, the key controller triggers the key XOR module to perform an XOR operation on the intermediate result output by the round function arithmetic unit and the round key. The key controller records the current encryption round through a counter and compares it with the SM4 encryption round configuration to determine whether the encryption is completed. If the current round has not reached the configured SM4 encryption rounds, the key controller enters the next round and continues to trigger the key expansion and key XOR operations. If the current round has reached the configured SM4 encryption rounds, the key controller outputs an encryption completion signal, indicating the end of the entire encryption process.

[0076] Workflow of the key adder: At the start of encryption, the key controller resets the key expansion module and the key XOR module and initializes the round counter to 0. The key input interface writes the initial key into the key grouping cache module, divides it into 4 sub - keys and caches them. The key controller triggers the key expansion module to generate the round key for the first round according to the first sub - key. The round function arithmetic unit performs the first - round encryption operation on the plaintext data group and outputs the intermediate result. The key controller triggers the key XOR module to perform an XOR operation on the intermediate result output by the round function arithmetic unit and the round key of the first round to obtain the output result of the first round. The key controller passes the output result of the first round to the round function arithmetic unit of the next round through the pipeline. The key controller determines whether the current round has reached the configured SM4 encryption rounds. If not, it enters the next round and repeats. If the current round has reached the configured SM4 encryption rounds, the key controller outputs an encryption completion signal, indicating the end of the entire encryption process.

[0077] Calculate the operation efficiency of the SM4 encryption core operation, and dynamically schedule resources according to the operation efficiency, including: Implementation of the pseudo-random number generator: Implement the pseudo-random number generator on the FPGA chip using a linear feedback shift register (LFSR). The LFSR consists of multiple cascaded D flip-flops and exclusive-OR gates. By selecting an appropriate feedback polynomial, a pseudo-random number sequence with a longer period can be generated. According to the required pseudo-random number bit width, select the appropriate number of LFSR stages. For example, using a 32-stage LFSR can generate 32-bit pseudo-random numbers. Instantiate the LFSR in the FPGA using logic resources (such as LUTs, FFs, etc.), and control its operation through a clock. Each clock cycle, the LFSR generates a new pseudo-random number according to the feedback polynomial. Use the output of the LFSR as the handshake password for the data fetcher, round function calculator, and key adder.

[0078] Figure 3 Is an exemplary flowchart for collecting handshake signals shown in some embodiments of the present application. Setting of the independent handshake channel: Implement an independent handshake channel on the FPGA chip using high-speed serial transceivers (such as LVDS, SERDES, etc.). Set a pair of high-speed serial transceivers between the data fetcher, round function calculator, and key adder respectively for sending and receiving handshake signals. The high-speed serial transceiver at the sending end converts the parallel handshake data into serial data and sends it to the receiving end through a differential signal line. The high-speed serial transceiver at the receiving end converts the received serial data into parallel data and restores the original handshake signal. The handshake channel is independent of the data channel to ensure the high-speed transmission of handshake signals without affecting the bandwidth and latency of the data channel. Setting of the physical non-deterministic source (PHYS): Set physical non-deterministic sources (PHYS) at the sending and receiving ends of the handshake channel for generating random noise. A ring oscillator can be used to implement the PHYS. The ring oscillator consists of an odd number of cascaded inverters, generating an unstable oscillation signal. By sampling the output of the ring oscillator, random noise can be obtained. A quantum random number generator can also be used to implement the PHYS. The quantum random number generator uses quantum noise sources (such as jet noise, photon noise, etc.) to generate true random numbers. Combine the output of the PHYS with the handshake password generated by the pseudo-random number generator to form the final handshake password. This can enhance the randomness and unpredictability of the handshake password.

[0079] Generation and transmission of the first handshake signal: After the fetcher completes fetching the plaintext block, it obtains the handshake password 1 at the current moment from the pseudo-random number generator. The password 1 is packed with the fetch completion identification code and the time stamp to generate the first handshake signal. The identification code is used to identify the fetch completion event, and the time stamp records the time of fetch completion. The first handshake signal is sent to the round function arithmetic unit through the high-speed serial transceiver on the handshake channel. Generation and transmission of the second handshake signal: After receiving the first handshake signal, the round function arithmetic unit verifies the password 1. After successful verification, it obtains the handshake password 2 at the next moment from the pseudo-random number generator. The password 2 is packed with the round function operation completion identification code, the time stamp, and the first handshake signal to generate the second handshake signal. The identification code is used to identify the round function operation completion event, and the time stamp records the time of operation completion. The second handshake signal is sent to the key adder through the high-speed serial transceiver on the handshake channel. Generation and transmission of the third handshake signal: After receiving the second handshake signal, the key adder verifies the password 2. After successful verification, it obtains the handshake password 3 at the next moment from the pseudo-random number generator. The password 3 is packed with the encryption completion identification code, the time stamp, the first handshake signal, and the second handshake signal to generate the third handshake signal. The identification code is used to identify the encryption completion event, and the time stamp records the time of encryption completion. The third handshake signal is returned to the fetcher through the high-speed serial transceiver on the handshake channel.

[0080] Verification of handshake signals: In the handshake interface, the received handshake password is compared with the output of the PHYS to determine the validity of the handshake password. If the received handshake password is consistent with the output of the PHYS, it indicates that the handshake password is valid and the handshake verification passes. If the received handshake password is inconsistent with the output of the PHYS, it indicates that the handshake password is invalid and the handshake verification fails. The handshake interface feeds back the handshake failure information to the sending end and requests to resend the handshake signal. Completion of the handshake process: After the fetcher receives the third handshake signal, it verifies the password 3. After the verification passes, it indicates that a complete handshake process is completed. The fetcher enters the next round of encryption process and starts a new round of handshake process. Calculation of the operation time of the calculation module: Parse the first handshake signal, the second handshake signal, and the third handshake signal collected, and extract the timestamp information therein. According to the timestamp difference between adjacent handshake signals, calculate the operation times of the fetcher, the round function calculator, and the key adder respectively. The fetch controller calculates the operation time T1 = T3_end - T1_start of the fetcher according to the start timestamp T1_start of the first handshake signal and the end timestamp T3_end of the third handshake signal. The round function controller calculates the operation time T2 = T2_end - T1_end of the round function calculator according to the end timestamp T1_end of the first handshake signal and the end timestamp T2_end of the second handshake signal. The key controller calculates the operation time T3 = T3_start - T2_end of the key adder according to the end timestamp T2_end of the second handshake signal and the start timestamp T3_start of the third handshake signal.

[0081] Calculation of the operation efficiency of the SM4 encryption core: According to the collected handshake signals, extract the timestamp information, and calculate the operation time T1 of the data fetcher, the operation time T2 of the round function operator, and the operation time T3 of the key adder. For the data fetcher, round function operator, and key adder in the SM4 encryption core, set the corresponding weight coefficients W1, W2, and W3. The weight coefficients represent the influence degree of each module on the encryption efficiency. Through experimental tests and statistical analysis, determine the appropriate values of the weight coefficients. Specifically, the following method can be adopted: Design a set of test vectors, including plaintext data of different lengths and characteristics, and perform encryption processing on the SM4 encryption core. During the encryption process, record the operation times T1, T2, and T3 of each module, as well as the encryption time T of the entire encryption core. Through multiple tests, collect a certain amount of sample data. Use mathematical optimization methods such as the least squares method to fit the weight coefficients W1, W2, and W3 to minimize the error between the weighted operation time and the overall encryption time. The values of the weight coefficients should satisfy the constraint conditions, such as the sum of the weight coefficients is 1, and the weight coefficients are non-negative, etc. According to the calculated operation times T1, T2, and T3 and the weight coefficients W1, W2, and W3, calculate the operation efficiency E of the SM4 encryption core through the formula E = 1 / (W1 * T1 + W2 * T2 + W3 * T3). The calculation of the operation efficiency E can be implemented using a floating-point arithmetic unit in the FPGA, or the calculation can be simplified through methods such as look-up table optimization.

[0082] Dynamic resource scheduling optimization: According to the calculated operation efficiency E of each SM4 encryption core, use the priority scheduling algorithm to dynamically adjust the FPGA resource allocation. Design a resource management module responsible for monitoring the operation efficiency of each SM4 encryption core and allocating resources according to the priority scheduling algorithm. The resource management module maintains a priority queue, arranging the SM4 encryption cores with higher operation efficiency in the front of the queue and those with lower operation efficiency at the back. When there are available logic resources, storage resources, DSP resources, etc. in the FPGA, the resource management module allocates resources to the SM4 encryption cores in the order of the priority queue. For the SM4 encryption cores with higher operation efficiency, allocate more resources, such as increasing the parallelism, raising the clock frequency, etc., to improve their processing performance. For the SM4 encryption cores with lower operation efficiency, appropriately reduce the resource allocation, such as reducing the parallelism, lowering the clock frequency, etc., and give priority to allocating resources to the cores with higher efficiency. The resource management module dynamically adjusts the resource allocation according to the actual resource usage situation and the requirements of the encryption tasks. When the encryption tasks change or resource usage bottlenecks occur, promptly adjust the priority queue and the resource allocation strategy.

[0083] The selection of the priority scheduling algorithm can be determined according to actual requirements. Common priority scheduling algorithms include: Weight-based priority scheduling: Different weights are assigned according to the operation efficiency of the SM4 encryption core. The higher the weight, the higher the priority. Deadline-based priority scheduling: The SM4 encryption cores are sorted by priority according to the deadlines of the encryption tasks. The more urgent the deadline, the higher the priority. Fairness-based priority scheduling: Consider the resource usage of each SM4 encryption core to ensure the fairness of resource allocation and avoid individual cores occupying resources for a long time. Through dynamic resource scheduling optimization, balance the processing performance of each SM4 encryption core and improve the overall encryption throughput and efficiency.

[0084] Hardware implementation: Design and implement a resource management module in the FPGA, including a priority queue, a resource allocation controller, etc. The resource management module is connected to each SM4 encryption core through a bus or a dedicated interface to achieve dynamic resource allocation and control. The design of the SM4 encryption core needs to consider the scalability and reconfigurability of resources to support dynamic resource scheduling. In the SM4 encryption core, design a resource control interface to receive the control signals from the resource management module and dynamically adjust the internal resource usage, such as parallelism, clock frequency, etc. Implement the design of the resource management module and the SM4 encryption core using a hardware description language (such as Verilog, VHDL), and perform synthesis, placement and routing to generate the FPGA bitstream file. Software control: Integrate a soft-core processor (such as ARM, MicroBlaze, etc.) in the FPGA to run the software program for resource scheduling and control. The software program calculates the priority by reading the operation efficiency of the SM4 encryption core and passes the priority information to the resource management module. The software program dynamically adjusts the priority scheduling algorithm and resource allocation strategy according to the requirements of the encryption tasks and the state of the system. The software program communicates with the resource management module and the SM4 encryption core through configuration registers or control interfaces to achieve the control of dynamic resource scheduling.

[0085] When the SM4 encryption core performs encryption operations, the plaintext data, intermediate results are mixed with random polynomial operations. The polynomial operations are expanded using the Horner algorithm and combined with logical operations to enhance the resistance to side-channel attacks. Random polynomial encoding of plaintext data: Input of plaintext data: The fetch controller reads the plaintext data packets from the external interface or data cache. The length of each packet is 128 bits. The read plaintext data packets are divided into several sub-packets, and the length of each sub-packet can be set according to the needs of polynomial operations, such as 32 bits or 64 bits. Each sub-packet is used as an input variable for polynomial operations and random polynomial encoding is performed in sequence. Generation of random variables: A linear feedback shift register (LFSR) is used to generate the random values of variables x, y, z in the polynomial operation expression f(x, y, z) = (ax + b)y + cz. According to the security requirements of polynomial operations, appropriate LFSR feedback polynomials and initial states are selected. Common feedback polynomials are: x^16 + x^14 + x^13 + x^11 + 1 (16-bit LFSR); x^32 + x^22 + x^2 + x^1 + 1 (32-bit LFSR); x^64 + x^63 + x^61 + x^60 + 1 (64-bit LFSR); The initial state of the LFSR is set to the key or a random seed to ensure that the generated random sequence has sufficient randomness and unpredictability. Through the shift and feedback operations of the LFSR, a pseudo-random sequence is generated as the values of variables x, y, z. According to the needs of the polynomial operation expression, an appropriate number of bits are selected to intercept the random sequence generated by the LFSR to obtain the required random variable values.

[0086] Generation of random coefficients: The coefficients a, b, and c of each term of f(x, y, z) are generated through the Bracewell sequence B(n). The Bracewell sequence is a pseudo-random sequence with good randomness and uniform distribution characteristics. Its generation method is as follows: Select an initial value B(0) as the starting value of the sequence. Generate subsequent sequence values through the recurrence formula B(n) = (B(n - 1) * p) mod q. Among them, p and q are two relatively prime positive integers, usually p = 31 and q = 127 are selected. According to the needs of the polynomial operation expression, select an appropriate number of digits to intercept the Bracewell sequence to obtain the required random coefficients a, b, and c. Multiply the generated random coefficients by the corresponding terms in the polynomial operation expression to obtain the complete polynomial expression. Polynomial operation and encoding: Substitute the plaintext data sub-group as the input variable x of the polynomial operation, the random variables y and z, and the random coefficients a, b, and c into the polynomial expression f(x, y, z) = (ax + b)y + cz. Execute polynomial operations through the polynomial operation unit to calculate the encoded result. The polynomial operation unit can be implemented using hardware resources such as multipliers, adders, and shift registers to improve the operation efficiency. In order to reduce the consumption of hardware resources, the Horner algorithm can be used to expand the polynomial operation and convert the polynomial operation into a series of multiplication and addition operations. For each plaintext data sub-group, execute the corresponding polynomial operation to obtain the encoded result.

[0087] Output of the Encoding Result: Concatenate the results obtained by randomly polynomial encoding each plaintext data sub-group to form a complete encoded plaintext data group. Use the encoded plaintext data group as the output of the data fetching controller and pass it to the round function arithmetic unit for subsequent encryption operations. Through random polynomial encoding, the statistical characteristics of the plaintext data are masked, increasing the randomness and unpredictability of the data and enhancing the ability to resist side-channel attacks. Expansion of the Horner algorithm in the round function arithmetic unit: Expansion of the polynomial: Expand the random polynomial f(x, y, z) = (ax + b)y + cz into the form of the Horner algorithm: f(x, y, z) = (ay + c)x + by + cz. By arranging the polynomial in descending order of the powers of the variable x, the expansion of the Horner algorithm is obtained. The expanded polynomial form can reduce the number of multiplications and improve the computational efficiency. Calculation of the polynomial: In the round function arithmetic unit, calculate the value of the polynomial through iterative loops. First, calculate the value of (ay + c), where a is the polynomial coefficient, y is the random variable, and c is the constant term. Then, multiply the result of (ay + c) by the variable x to get the value of (ay + c)x. Next, add the result of (ay + c)x to the next term by in the polynomial to get the value of (ay + c)x + by. Finally, add the result of (ay + c)x + by to the constant term cz to get the final result of the polynomial f(x, y, z).

[0088] Mixing of Intermediate Results: During each iteration of calculating the polynomial, mix the intermediate results with other operations of the round function to increase the computational complexity and randomness. The intermediate results can be mixed with the S-box substitution to perform a non-linear transformation on the intermediate results. Or the intermediate results can be mixed with a linear transformation to perform a linear confusion on the intermediate results. By mixing the polynomial calculation with other operations of the round function, the computational complexity is increased and the ability to resist side-channel attacks is enhanced. Hardware Implementation: Implement the polynomial arithmetic unit using DSP resources or logic resources in the FPGA, supporting the expansion and calculation of the Horner algorithm. Control the process of polynomial calculation through a finite state machine (FSM) to coordinate the execution of each step. Integrate the polynomial arithmetic unit with other modules such as the S-box and linear transformation to form a complete round function arithmetic unit. Through pipeline design, improve the parallelism and throughput of polynomial calculation and reduce the calculation delay.

[0089] Random Encoding of Keys in the Key Adder: Grouping and Encoding of Keys: The 128-bit key is grouped into several sub-keys, and the length of each sub-key can be set according to the needs of polynomial operations, such as 32 bits or 64 bits. For each sub-key, random variable values and random coefficients are generated and substituted into the polynomial expression f(x, y, z) for calculation to obtain the encoded result. The generation methods of random variable values and random coefficients are similar to those of the random polynomial encoding of plaintext data, and modules such as LFSR and Bracewell sequence generators can be used to generate them. XOR Operation on the Encoding Result: The result obtained by randomly encoding each sub-key is XORed with the corresponding round key generated by the key expansion algorithm. The XOR operation can mix the encoded key with the expanded key, increasing the randomness and unpredictability of the key. The result of the XOR operation is used as the randomized round key for subsequent key addition operations.

[0090] Mixed XOR Operation in the Key Adder: Logical Operation between the Result of the Round Function and the Round Key: When the key adder performs the round key addition operation, first, the result of the round function operation is logically operated with the round key. Different logical operations can be selected, such as AND, OR, XOR, etc., to increase the diversity and randomness of the calculation. By using different logical operations, the result of the key addition operation has higher unpredictability and complexity. XOR between the Result of the Logical Operation and the Result of the Random Polynomial Operation: The result obtained by logically operating the result of the round function with the round key is XORed with the result of the random polynomial operation. The result of the random polynomial operation is obtained by randomly encoding the result of the round function operation. By the XOR operation, the result of the logical operation is mixed with the result of the random polynomial operation to obtain the final key addition result.

[0091] Random Insertion of Waiting Period: Design of Pseudo-Random Number Generator: Use a Linear Feedback Shift Register (LFSR) or other pseudo-random number generator to generate a sequence of random numbers. According to the security requirements of SM4 encryption, select an appropriate LFSR feedback polynomial and initial state to ensure that the generated random numbers have good statistical properties. The output of the pseudo-random number generator can be used to control whether to insert a waiting period and the length of the waiting time. Control of Random Insertion of Waiting Period: Before each round of SM4 encryption operation, decide whether to insert a waiting period according to the output of the pseudo-random number generator. If the value of the pseudo-random number meets the preset condition (such as being greater than a certain threshold), then insert a waiting period before the current round of operation. The length of the waiting period can also be determined according to the value of the pseudo-random number, and different waiting time granularities can be selected, such as 1 clock cycle, 2 clock cycles, etc. Random Polynomial Encoding of the Final Result: Perform random polynomial encoding on the final encryption result: After completing all rounds of SM4 encryption, perform random polynomial encoding on the final 128-bit encryption result. Similar to the encoding process of the plaintext data, divide the final encryption result into several sub-packets, and the length of each sub-packet can be set according to the needs of polynomial operations. For each sub-packet, generate random variable values and random coefficients, substitute them into the polynomial expression f(x, y, z) for calculation, and obtain the encoded result. Concatenate the encoded results to obtain the complete obfuscated ciphertext data.

Claims

1. An SM4 encryption method based on FPGA, comprising: Obtain the plaintext data to be encrypted and input the plaintext data into the FPGA chip; Multiple SM4 encryption cores are configured on the FPGA chip, and the SM4 encryption cores are used to perform parallel encryption operations on plaintext data to obtain ciphertext data; each SM4 encryption core adopts a pipeline architecture, which includes a counter, a round function operator and a key adder, which respectively perform the three stages of the encryption operation: counter, round function operation and key addition; Using the physical noise-based random number extraction circuit integrated on the FPGA chip to generate an N-bit random number, the generated N-bit random number is used as the encryption key of each SM4 encryption core for SM4 encryption operations; In each SM4 encryption core, the handshake signals between the counter, round function operator and key adder are collected. The handshake signals are used to reflect the operation status of the counter, round function operator and key adder. According to the collected handshake signals, the operation efficiency of each SM4 encryption core is calculated; According to the calculated operation efficiency, the priority scheduling algorithm is used to dynamically adjust the FPGA resource allocation between each SM4 encryption core; In the process of each SM4 encryption core performing encryption operations, polynomial operations are mixed with logical operations to improve the power analysis attack capability; and wait cycles are randomly inserted in the encryption operation to improve the SM4 encryption's ability to resist side channel attacks by randomizing the operation timing; Calculate the SM4 encryption core operation efficiency based on the collected handshake signals, including: Parse the collected first handshake signal, second handshake signal and third handshake signal, extract timestamp information, and calculate the operation time of the counter, round function operator and key adder respectively according to the timestamp difference of adjacent handshake signals; The data acquisition controller calculates the operation time T1=T3_end-T1_start of the data acquisition device according to the start time stamp T1_start of the first handshake signal and the end time stamp T3_end of the third handshake signal; The round function controller calculates the operation time T2=T2_end-T1_end of the round function operator according to the end timestamp T1_end of the first handshake signal and the end timestamp T2_end of the second handshake signal; The key controller calculates the operation time T3 of the key adder according to the end timestamp T2_end of the second handshake signal and the start timestamp T3_start of the third handshake signal, which is T3=T3_start-T2_end; The computational efficiency E of the SM4 encryption core is calculated using the following formula: E=1 / (W1*T1+W2*T2+W3*T3), Among them, W1, W2, and W3 are the weight coefficients of the counter, round function operator, and key adder in the SM4 encryption core that affect the encryption efficiency.

2. The FPGA-based SM4 encryption method according to claim 1, characterized in that: Configure multiple SM4 encryption cores on the FPGA chip, including: Use Verilog HDL or VHDL hardware description language to write multiple SM4 encryption algorithm IP core codes and obtain the RTL level description of the SM4 encryption IP core; Input the RTL level description of the SM4 encrypted IP core into the logic synthesis tool, perform logic synthesis on the SM4 encrypted IP core code, convert the RTL level description into a logic gate level netlist, and generate a logic gate level netlist file of the SM4 encrypted IP core; The generated logic gate-level netlist file is used as input, and the layout and routing synthesis tool is used to perform FPGA layout and routing on the SM4 encrypted IP core, map the logic units in the SM4 encrypted IP core to the physical resources of the FPGA device, and complete the layout and routing between the physical resources; wherein the physical resources include the lookup table LUT and the trigger FF; Before FPGA layout and routing, balance the clock tree on the FPGA chip to make the clocks of various physical resources of the SM4 encryption IP core consistent; and put the SM4 encryption IP core into a preset initial state through asynchronous reset; After completing the FPGA layout and routing, the SM4 encryption IP core is connected to the physical resources of the FPGA chip through port mapping to form a group of SM4 encryption cores for parallel SM4 encryption operations.

3. The FPGA-based SM4 encryption method according to claim 2, characterized in that: The logic synthesis tool includes at least one of Synopsys Design Compiler, Cadence RTL Compiler and MentorGraphics Precision; The placement and routing synthesis tool includes at least one of Xilinx Vivado, Intel Quartus, and Lattice Diamond.

4. The FPGA-based SM4 encryption method according to claim 3, characterized in that: In pipeline architecture: The data acquisition device adopts an architecture that supports variable-length packet acquisition. According to the packet length of the SM4 encryption algorithm, the input plaintext data to be encrypted is divided into groups, and the plaintext data is divided into multiple data groups of equal length, and the divided data groups are output to the next level of the pipeline in sequence; The round function operator adopts a programmable architecture that supports dynamic configuration of the number of encryption rounds. The number of SM4 encryption rounds is set through a configurable register, and the number of round function operations is controlled according to the configured number of rounds. In each round of round function operation, nonlinear transformation and linear transformation are performed on the input data group, and the round function operation result is output to the next stage of the pipeline; The key adder adopts an architecture that supports key grouping and group key prediction. It divides the input encryption key into multiple key groups, expands the key groups by table lookup, predicts the next key group based on the current key group, and performs an XOR operation on the round function operation result and the predicted key group to generate the encryption result of the current round.

5. The FPGA-based SM4 encryption method according to claim 4, characterized in that: Data acquisition device, including: A packet storage area, used to cache plaintext data packets to be retrieved; The circular buffer uses a linked list structure to manage plaintext data packets. The linked list structure includes multiple linked list nodes, a linked list head pointer and a linked list operation interface. Each linked list node is used to record the starting address and packet length of a plaintext data packet. The linked list head pointer points to the first node in the linked list. The linked list operation interface is used to manage insertion and deletion of the linked list. The data acquisition controller traverses the linked list through the linked list operation interface according to the packet length of the SM4 encryption algorithm, and reads the plaintext data packets recorded in the linked list in sequence. Each time a data packet is read, the corresponding linked list node is deleted through the linked list operation interface to read the data of the variable-length packet; the data acquisition controller manages the caching and reading of the plaintext data packets in the circular buffer by controlling the write pointer and read pointer of the circular buffer.

6. The FPGA-based SM4 encryption method according to claim 5, characterized in that: Round function operator, including: The round number configuration register is used to store the round number configuration value of the SM4 encryption algorithm; Multiple round function operation modules, each round function operation module is used to perform a round of encryption operation of the SM4 encryption algorithm, the encryption operation includes nonlinear transformation and linear transformation, the nonlinear transformation performs an S-box replacement operation on the input data group, and the linear transformation performs a linear transformation on the result after the S-box replacement; The round function selector has an input end connected to the output of each round function operation module and an output end connected to the next stage of the pipeline. According to the round number configuration value, the output of the round function operation module corresponding to the round number is selected to the next stage of the pipeline; The round key generation module generates the round key of each round through the key expansion algorithm according to the input initial key, and provides the generated round key to the corresponding round function operation module; The round function controller controls the transmission of data packets between pipeline modules at various levels according to the process of the SM4 encryption algorithm, controls the round function selector to switch the round function, controls the round key generation module to provide round keys, and performs multi-round encryption operations of the SM4 encryption algorithm.

7. The FPGA-based SM4 encryption method according to claim 6, characterized in that: Key adder, including: The key grouping cache module caches the key used by the SM4 encryption algorithm and divides the input initial key into 4 subkeys and caches them; The key expansion module expands the subkey output by the key grouping cache module by table lookup; and randomly permutes the data bits of the expanded subkey to generate a round key for the current encryption round; The key XOR module performs a bitwise XOR operation on the intermediate result output by the round function operator and the round key generated by the key expansion module to obtain the output result of the current encryption round; The key controller, according to the round number configuration of the SM4 encryption algorithm, triggers the key expansion module to generate round keys in each encryption round, controls the key XOR module to execute round key XOR, and passes the XOR result to the next round through the pipeline; the key controller uses a finite state machine to determine whether the key expansion and key XOR are completed according to the current round and the number of SM4 encryption rounds. If not, it enters the next round to continue execution until the configured number of SM4 encryption rounds is completed.

8. The FPGA-based SM4 encryption method according to any one of claims 4 to 7, characterized in that: Collect handshake signals, including: A pseudo-random number generator is set on the FPGA chip to generate different pseudo-random numbers as handshake passwords; An independent channel is set between the counter, the round function operator and the key adder. The independent channel uses the high-speed serial transceiver of the FPGA and is independent of the data channel. A physical non-deterministic source PHYS is set at the transmitting and receiving ends of the independent channel. By using the PHYS output as the verification factor of the handshake password, a handshake password challenge response verification mechanism based on physical non-determinism is constructed. After the data acquisition device completes the plaintext group data acquisition, it obtains the handshake password 1 at the current moment from the pseudo-random number generator, packages the password 1 with the data acquisition completion identification code and the timestamp, generates a first handshake signal, and sends the first handshake signal to the round function operator through the handshake interface on the independent channel; The round function operator verifies the password 1 in the first handshake signal. After the verification is passed, the handshake password 2 at the next moment is obtained from the pseudo-random number generator, and the password 2 is packaged with the round function operation completion identification code, the timestamp and the first handshake signal to generate a second handshake signal, and the second handshake signal is sent to the key adder through the handshake interface; The key adder verifies the password 2 in the second handshake signal. After the verification is passed, it obtains the handshake password 3 at the next moment from the pseudo-random number generator, packages the password 3 with the encryption completion identification code, the timestamp, the first handshake signal and the second handshake signal, generates a third handshake signal, and returns the third handshake signal to the data acquirer through the handshake interface; The data reader verifies the password 3 in the third handshake signal. After the verification is passed, a handshake process is completed and the next round of encryption process begins. Among them, the handshake interface determines the validity of the handshake password by comparing the received handshake password with the PHYS output.

9. The FPGA-based SM4 encryption method according to claim 1, characterized in that: In the process of performing encryption operations in each SM4 encryption core, polynomial operations and logical operations are mixed to perform calculations, including: When the data acquisition controller reads the plaintext data group, the plaintext data is used as the input variable of the polynomial operation, and the linear feedback shift register LFSR is used to generate the random values ​​of the variables x, y, and z in the polynomial operation expression f(x, y, z) = (ax+b)y+cz, and the coefficients a, b, and c of each term of f(x, y, z) are generated through the Bracewell sequence B(n); the encoded result is used as the output of the data acquisition controller and passed to the round function operator to cover up the statistical characteristics of the plaintext data; When the round function operator performs nonlinear transformation, the operation process of the polynomial f(x, y, z) is expanded by the Horner algorithm; When the key adder generates the round key, it first uses the polynomial f(x, y, z) to randomly encode each subkey in the key group, and then expands the key, mixes the expanded result with the random encoding result, and obtains the randomized round key; When the key adder performs the round key addition operation, the round function operation result and the round key are first subjected to a logical operation, and then XORed with the random polynomial operation result to obtain a mixed key addition result; After completing a round of SM4 encryption, the encryption result is mixed with the result of the random polynomial f(x, y, z) operation and XORed as the input for the next round of encryption; after the last round of encryption is completed, the final result is encoded again through the random polynomial f(x, y, z) operation to obtain the obfuscated ciphertext output.

Citation Information

Patent Citations

  • FPGA optimization implementation method and system for SM4 cryptographic algorithm and application

    CN113078996A

  • Encryption and decryption method and circuit based on asynchronous circuit

    CN117240430A