Communication encryption acceleration device and method, Internet of Things equipment and system

By designing a communication encryption acceleration device including communication module, asynchronous FIFO, multi-core scheduler and SM4 encryption core, the encryption performance bottleneck and insufficient adaptability of hardware accelerator in IoT devices is solved, and efficient multi-device parallel encryption and decryption processing is achieved.

CN120165866AInactive Publication Date: 2025-06-17ZHEJIANG UNIV +1

Patent Information

Application Number
CN202510647089.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional software-based encryption methods show obvious performance bottlenecks when processing large amounts of data, especially on IoT devices with limited resources, and are difficult to meet the demand for efficient real-time data encryption for IoT devices. At the same time, existing hardware accelerators have the disadvantage of insufficient adaptability to multi-device communication.

Method used

A communication encryption acceleration device is designed, including a communication module, an asynchronous FIFO, a multi-core scheduler, multiple SM4 encryption cores and cache modules. Through asynchronous FIFO, the multi-core scheduler schedules the data stream to be processed to multiple SM4 encryption cores to realize multi-core parallel encryption and decryption processing.

Benefits of technology

It realizes dynamic allocation of computing resources, reasonably schedules encryption and decryption tasks between multiple cores, cope with high concurrency requirements in the IoT device environment, improves the throughput and efficiency of encryption processing, and meets the needs of parallel encryption and decryption of multiple devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165866A_ABST
    Figure CN120165866A_ABST
Patent Text Reader

Abstract

The invention relates to a communication encryption acceleration device, a communication encryption acceleration method, Internet of Things equipment and an Internet of Things system. A communication module in the communication encryption acceleration device is used for receiving a plurality of to-be-processed data streams sent by Internet of Things equipment; the asynchronous FIFO is used for carrying out cross-clock domain synchronous processing and decoupling on transmission of each to-be-processed data stream between a pre-stage module and a post-stage module of the asynchronous FIFO; the multi-core scheduler is used for scheduling each to-be-processed data stream transmitted by the asynchronous FIFO and distributing each to-be-processed data stream to a target SM4 encryption core in the SM4 encryption cores; performing encryption and decryption processing on the corresponding to-be-processed data streams through the target SM4 encryption core to obtain target data, and caching the target data to a cache module; and the communication module is used for acquiring the target data from the cache module and transmitting the target data to the respective corresponding Internet of Things equipment. By adopting the device, the defect of insufficient adaptability of a hardware accelerator to multi-device communication can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication security technologies, and particularly to a communication encryption acceleration device, method, Internet of Things device, and system. Background Art

[0002] With the rapid development of Internet of Things technologies, more and more devices are interconnected through networks, forming a huge data interaction ecosystem. The applications of Internet of Things devices cover a wide range, including smart homes, smart healthcare, vehicle networking, industrial Internet of Things, and other fields. These devices communicate with servers or other devices through various network protocols (such as Wi-Fi, Ethernet, Bluetooth, etc.) and transmit a large amount of sensitive data. However, with the growth in the number of connected devices, data security issues have become increasingly severe. Especially during data transmission, there are risks such as information leakage and tampering. In this context, having an efficient and reliable encryption mechanism has become the key to ensuring data security. Currently, Internet of Things devices usually use symmetric encryption algorithms for data encryption transmission. Among them, SM4 is a block symmetric encryption algorithm designed by the Cryptography Administration and is widely used in security encryption scenarios for various wireless communications. The security of the SM4 algorithm has been widely verified and is especially suitable for the national cryptography standard in China.

[0003] However, traditional software-based encryption methods exhibit obvious performance bottlenecks when processing a large amount of data, especially on resource-constrained Internet of Things devices, and it is difficult to meet the requirements of Internet of Things devices for efficient real-time data encryption. To improve the execution efficiency of encryption algorithms, hardware accelerators are used in traditional technologies to implement the encryption process. However, the hardware accelerators in traditional technologies have the defect of insufficient adaptability to multi-device communication. Summary of the Invention

[0004] Based on this, it is necessary to provide a communication encryption acceleration device, a communication encryption acceleration method, an Internet of Things device, and a communication encryption acceleration system that can solve the defect of insufficient adaptability to multi-device communication for the above technical problems.

[0005] In a first aspect, this application provides a communication encryption acceleration device. The communication encryption acceleration device includes a communication module, an asynchronous FIFO, a multi-core scheduler, a plurality of SM4 encryption cores, and a cache module, where:

[0006] The communication module is configured to receive a plurality of data streams to be processed sent by an Internet of Things device;

[0007] The asynchronous FIFO is configured to perform cross-clock domain synchronization processing and decoupling on the transmission of each data stream to be processed between a pre-stage module and a post-stage module of the asynchronous FIFO;

[0008] The multi-core scheduler is used to schedule each data stream to be processed in the asynchronous FIFO transmission, and distribute each data stream to be processed to a target SM4 encryption core among multiple SM4 encryption cores;

[0009] The target SM4 encryption core is used to encrypt or decrypt the corresponding data stream to be processed, obtain the target data, and cache the target data in the cache module;

[0010] The communication module is further used to obtain the target data from the cache module and transmit the target data to the corresponding Internet of Things device.

[0011] In one embodiment, the asynchronous FIFO is further used to transmit each data stream to be processed between the front-stage module and the rear-stage module of the asynchronous FIFO. If the receiving rate of the front-stage module does not match the processing rate of the rear-stage module, the data stream to be processed is cached.

[0012] In one embodiment, the multi-core scheduler is further used to determine the device ID of each data stream to be processed, determine the target SM4 encryption core that matches the device ID from multiple SM4 encryption cores, and distribute the data stream to be processed to the target SM4 encryption core.

[0013] In one embodiment, if there are multiple candidate SM4 encryption cores that match the device ID among multiple SM4 encryption cores, the multi-core scheduler is further used to determine the idle duration of each candidate SM4 encryption core, and determine the candidate SM4 encryption core with the longest idle duration as the target SM4 encryption core.

[0014] In one embodiment, the SM4 algorithm architecture of each SM4 encryption core includes a control module, a key expansion module, and an encryption / decryption module. Among them, the control module is used to control the key expansion module to stop working in the case of no new key input;

[0015] After the key expansion module completes all-round key generation, the encryption / decryption module starts to work, controls the encryption / decryption module to use the round keys in positive order or reverse order according to the mode selection, and outputs an encryption or decryption result according to a preset clock cycle.

[0016] In one embodiment, the cache module is a synchronous FIFO or an asynchronous FIFO.

[0017] In one embodiment, the communication module supports multiple communication protocols and is further configured to transmit to the respective Internet of Things devices in a communication protocol that matches the target data, and the matching communication protocol is determined from the multiple communication protocols according to the communication requirements of the Internet of Things devices.

[0018] In a second aspect, the present application further provides a communication encryption acceleration method, including:

[0019] Receiving a plurality of data streams to be processed sent by the Internet of Things devices;

[0020] For each of the data streams to be processed, performing cross-clock domain synchronization processing and decoupling on the transmission between the front-stage module and the rear-stage module of the asynchronous FIFO;

[0021] Scheduling each of the data streams to be processed transmitted by the asynchronous FIFO, and distributing each of the data streams to a target SM4 encryption core among a plurality of the SM4 encryption cores;

[0022] Performing encryption or decryption processing on the respective data streams to be processed to obtain target data and caching the target data in the cache module;

[0023] Obtaining the target data from the cache module and transmitting the target data to the respective Internet of Things devices.

[0024] In a third aspect, the present application further provides an Internet of Things device, which includes a memory, a processor, and the communication encryption acceleration device according to any one of the above, the memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.

[0025] In a fourth aspect, the present application further provides a communication encryption acceleration system, which includes an Internet of Things device and the communication encryption acceleration device according to any one of the above, and the communication encryption acceleration device performs encryption and decryption processing on the data streams to be processed sent by the Internet of Things device.

[0026] The above-mentioned communication encryption acceleration device, communication encryption acceleration method, Internet of Things device, and communication encryption acceleration system are based on a hardware architecture of multi-core parallel encryption processing. In the case of multiple data streams to be processed sent by the Internet of Things device, cross-clock domain synchronization processing and decoupling are performed through an asynchronous FIFO, and through a multi-core scheduler, each data stream to be processed transmitted by the asynchronous FIFO is scheduled, and each data stream to be processed is distributed to the target SM4 encryption core among multiple SM4 encryption cores, so that the target SM4 encryption core performs encryption and decryption processing on the data stream to be processed. That is, this method can dynamically allocate computing resources, reasonably schedule encryption and decryption tasks between multiple cores, to cope with the high concurrency requirements in the Internet of Things device environment, and realizes the scheduling of data streams of different devices and the distribution of encryption and decryption tasks in the case of multiple cores. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required to be used in the description of the embodiments of the present application or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a structural block diagram of a communication encryption acceleration device in an embodiment;

[0029] Figure 2 It is a structural block diagram of an Ethernet module in an embodiment;

[0030] Figure 3 It is a schematic diagram of the interface of an SM4 core in an embodiment;

[0031] Figure 4 It is a schematic diagram of a multi-core scheduler in an embodiment;

[0032] Figure 5 It is a schematic diagram of the complete encryption process of the SM4 algorithm in an embodiment;

[0033] Figure 6 It is a schematic diagram of the hardware architecture of the SM4 cryptographic algorithm in an embodiment;

[0034] Figure 7 It is a schematic diagram of a composite domain inversion flowchart in an embodiment;

[0035] Figure 8 It is an architecture diagram of a communication encryption acceleration device in an embodiment;

[0036] Figure 9 It is a schematic diagram of a communication encryption acceleration method in an embodiment;

[0037] Figure 10 It is a structural block diagram of a communication encryption acceleration system in an embodiment. Detailed implementation manners

[0038] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0039] With the growth of the number of Internet-connected devices, the problem of data security has become increasingly severe. Especially during data transmission, there are risks such as information leakage and tampering. In this context, having an efficient and reliable encryption mechanism has become the key to ensuring data security. Currently, Internet of Things (IoT) devices usually use symmetric encryption algorithms for data encryption transmission.

[0040] Traditional software-based encryption methods exhibit obvious performance bottlenecks when processing a large amount of data, especially on resource-constrained IoT devices, and it is difficult to meet the requirements of IoT devices for efficient real-time data encryption. In order to improve the execution efficiency of encryption algorithms, the current technological development trend is to use hardware accelerators to implement the encryption process. Hardware accelerators perform encryption operations through dedicated circuits. Compared with software encryption, they have advantages such as parallel processing, low latency, and high throughput, and can significantly improve encryption efficiency. Especially in the application of IoT devices, hardware accelerators can effectively alleviate the limitations of limited device resources and provide efficient encryption services without increasing the device power consumption.

[0041] Related hardware encryption technologies mainly include pipeline acceleration: using pipeline technology to decompose the encryption process into multiple steps for parallel execution, reducing the encryption latency of each data block. Multi-core parallel computing: running multiple encryption cores simultaneously to improve the parallel processing ability of encryption and meet the requirements of high data throughput. Security enhancement technologies: To defend against side-channel attacks (such as power consumption attacks and timing analysis attacks), many designs introduce technologies such as random masking and composite field inversion to further improve the security of the encryption process.

[0042] Existing hardware encryption accelerators still have some problems, such as insufficient adaptability to multi-device communication, inability to flexibly allocate hardware resources, and the security during the encryption process needs to be further enhanced. Therefore, there is an urgent need for a hardware encryption acceleration solution that can support parallel encryption of multiple IoT devices and has good scalability and high security.

[0043] Moreover, the research in related technologies has focused on the design of the SM4 encryption core. Whether it is variable pipeline design or S-box optimization, these designs are aimed at improving the performance of a single encryption core. In the actual communication scenarios of Internet of Things (IoT) devices, the design of a single encryption core often fails to meet the requirements of multi-device parallel processing. In IoT scenarios, there are usually a large number of devices communicating simultaneously, and each device may require parallel encryption and decryption processing. Therefore, a mechanism for supporting multi-core parallel computing and task scheduling is needed.

[0044] In an exemplary embodiment, as Figure 1 shown, a communication encryption acceleration device is provided. The communication encryption acceleration device includes a communication module, an asynchronous FIFO, a multi-core scheduler, multiple SM4 encryption cores, and a cache module, where:

[0045] The communication module is used to receive multiple data streams to be processed sent by IoT devices.

[0046] It can be understood that the communication module has the functions of receiving and sending data. It can receive plaintext data from IoT devices and send the encrypted ciphertext data back to IoT devices or other communication terminals. The communication module can be responsible for implementing the transceiver functions of various communication protocols, such as Ethernet, Wi-Fi, etc.

[0047] Optionally, the communication module includes a communication receiving module and a communication sending module. Among them, the communication receiving module is responsible for receiving plaintext data from IoT devices. It can be compatible with the communication interfaces of various IoT devices, such as Ethernet, Wi-Fi, etc. The design of the communication receiving module ensures that it can adapt to the common wireless and wired communication protocols in IoT devices and seamlessly receive plaintext data from different devices.

[0048] The communication sending module is responsible for sending the encrypted ciphertext data back to IoT devices or other communication terminals, supports multiple communication protocols, and can select an appropriate communication method (such as Ethernet, Wi-Fi, etc.) to transmit the ciphertext data according to the needs of different IoT devices.

[0049] Exemplarily, the communication module receives plaintext data from IoT devices and stores it in the asynchronous FIFO, takes out the encrypted ciphertext from the synchronous FIFO, and sends it to the IoT device. The data stream to be processed can be plaintext data or ciphertext data. Further, the communication module can simultaneously receive multiple data streams to be processed from different IoT devices.

[0050] In an exemplary embodiment, taking the communication module as an Ethernet module as an example, as Figure 2As shown in the figure, the Ethernet module includes multiple sub-modules, and the sub-modules include User Datagram Protocol (UDP), Phase-Locked Loop (PLL), Gigabit Media Independent Interface to Reduced Gigabit Media Independent Interface (GMII_to_RGMII), Address Resolution Protocol (ARP), Internet Control Message Protocol (ICMP), and Ethernet Controller (ETH_CTRL).

[0051] Among them, the UDP module: unpacks the data from the Ethernet frame into UDP data, or encapsulates the application layer data into UDP packets; PLL: provides the required clock signals for the entire Ethernet transceiver module to ensure the timing synchronization between each sub-module; GMII_to_RGMII module: realizes the interface conversion between the Ethernet MAC layer and the PHY layer, enabling Gigabit Ethernet to communicate with fewer pins; ARP module: in Ethernet communication, ensures that the IP address can be correctly resolved into the MAC address to achieve communication between devices; ICMP (Internet Control Message Protocol module): supports the basic communication diagnosis function between network devices, used to test network connectivity and obtain network status information. ETH_CTRL: is responsible for managing the sending and receiving of network data.

[0052] An asynchronous FIFO is used to perform cross-clock domain synchronization processing and decoupling for the transmission between the front-stage module and the rear-stage module of each data stream to be processed.

[0053] It can be understood that in the scenario of multi-IoT device communication encryption, there is a problem of clock domain mismatch, so an asynchronous FIFO is needed to solve this problem. Among them, the front-stage module of the asynchronous FIFO (First-In-First-Out) can be a communication module, and the rear-stage module can be a multi-core scheduler and / or a target SM4 encryption core. Cross-clock domain synchronization processing can be understood as solving the problem of clock domain mismatch between the front-stage module and the rear-stage module, and ensuring the stability and reliability of data transmission when transferring data between the front-stage module and the rear-stage module. Decoupling can be understood as decoupling the data stream, enabling the front-stage module and the rear-stage module (for example, the receiving and processing modules) to work independently and avoid interfering with each other.

[0054] A multi-core scheduler is used to schedule each data stream to be processed in the asynchronous FIFO transmission and distribute each data stream to be processed to a target SM4 encryption core among multiple SM4 encryption cores.

[0055] Among them, the multi-core scheduler, that is, the multi-core scheduling module, is used to efficiently allocate SM4 cores to process the received tasks. The scheduling mechanism of the multi-core scheduling module includes determining the target SM4 encryption core that matches the device ID corresponding to the data stream to be processed from a data table including the relationship between the device IDs of Internet of Things devices and the corresponding SM4 encryption cores. For example, according to the device ID in the received Ethernet data, the task is assigned to the SM4 core to which the corresponding device belongs, ensuring that the tasks of each device are processed by the SM4 core associated with it.

[0056] It should be noted that each SM4 encryption core is responsible for encrypting and decrypting data using the SM4 algorithm. Through the multi-core parallel architecture, multiple SM4 encryption cores can work simultaneously, greatly improving the throughput of the encryption process and meeting the parallel encryption and decryption requirements of Internet of Things devices.

[0057] The target SM4 encryption core is used to encrypt or decrypt the corresponding data stream to be processed, obtain the target data, and cache the target data in the cache module.

[0058] Among them, the cache module is used to cache the ciphertext data encrypted by the SM4 encryption core to ensure that the data can be transmitted to the subsequent communication sending module in an orderly manner. The synchronous FIFO plays a buffering role to prevent congestion during data transmission. The cache module is a synchronous FIFO or an asynchronous FIFO.

[0059] The SM4 core is an efficient hardware encryption and decryption module implemented based on the commercial cryptography algorithm SM4. Through a flexible interface design, this module supports the loading of user keys, the enabling of key expansion, and status feedback, and can switch between encryption and decryption modes. The module receives 128-bit plaintext or ciphertext data for encryption and decryption operations and outputs the processing results.

[0060] In an exemplary embodiment, as Figure 3 shown, a schematic diagram of the interface of the SM4 core is provided, including an input interface and an output interface. The input interface includes a clock signal clk, a reset signal, a low-level reset rst_n, an enabling signal start for the SM4 core to work, an encryption or decryption selection signal mode, an enabling signal key_exp for SM4 key expansion, a user key valid signal key_valid, a user key key, an enabling signal enc_dec for the encryption and decryption module, a plaintext data valid signal data valid, and plaintext data data. Among the input interfaces, except for the user key key and the plaintext data data with a bit width of 128 bits, the remaining interface bit widths are 1 bit.

[0061] The output interface includes a key expansion completion signal key_done, an encryption / decryption result ready signal done, and an encryption / decryption result (ciphertext) result. The bit width of the encryption / decryption result (ciphertext) result is 128 bits, and the bit widths of the key expansion completion signal key_done and the encryption / decryption result ready signal done are 1 bit.

[0062] Based on the above, the functions of the SM4 core include: key management (supporting key input, expansion, and validity detection), data processing (supporting plaintext or ciphertext input and validity detection), operation mode selection (encryption or decryption), and result status feedback (including a key expansion completion signal and an encryption / decryption completion signal). This design features high performance, low latency, and easy integration, and is suitable for secure communication and data encryption scenarios. After the SM4 core finishes encryption / decryption, the result will be uniformly stored in a cache module (e.g., a synchronous FIFO), and then sent to the original device by the above-mentioned communication module to complete the entire working process.

[0063] The communication module is also used to obtain target data from the cache module and transmit the target data to their respective corresponding Internet of Things devices.

[0064] The above communication encryption acceleration device, based on a hardware architecture of multi-core parallel encryption processing, in the case of multiple data streams to be processed sent by Internet of Things devices, performs cross-clock domain synchronization processing and decoupling through an asynchronous FIFO, and through a multi-core scheduler, schedules each data stream to be processed transmitted by the asynchronous FIFO, and distributes each data stream to be processed to a target SM4 encryption core among multiple SM4 encryption cores, so that the target SM4 encryption core performs encryption / decryption processing on the data stream to be processed. That is, this method can dynamically allocate computing resources, reasonably schedule encryption / decryption tasks among multiple cores, to cope with the high concurrency requirements in the Internet of Things device environment, and realizes scheduling and encryption / decryption task allocation for data streams of different devices in the case of multiple cores.

[0065] In an exemplary embodiment, the communication module supports multiple communication protocols and is also used to transmit to their respective corresponding Internet of Things devices in a communication protocol matching the target data, and the matching communication protocol is determined from multiple communication protocols according to the communication requirements of the Internet of Things devices. It can be realized that the ciphertext data can be transmitted out by selecting a suitable communication method according to the requirements of different Internet of Things devices.

[0066] In an exemplary embodiment, the asynchronous FIFO is further configured to transfer each data stream to be processed between the front - end module and the back - end module of the asynchronous FIFO. If the receiving rate of the front - end module does not match the processing rate of the back - end module, the data to be processed is cached. In this way, the asynchronous FIFO can provide buffering between data reception and encryption processing, preventing data loss or transmission delay. It can be seen from this that the asynchronous FIFO can cache the data received from the communication receiving module, ensuring the stability and smoothness of data transmission. The asynchronous FIFO can provide buffering between data reception and encryption processing, preventing data loss or transmission delay.

[0067] In order to dynamically allocate encryption and decryption tasks according to data traffic and the load conditions of each encryption core, and optimize resource utilization. The scheduling strategy of the multi - core scheduler further includes the least recently used strategy, that is, in the case where a device corresponds to multiple SM4 cores, the task is assigned to the SM4 core that has not been used for the longest time.

[0068] In an exemplary embodiment, the multi - core scheduler is further configured to determine the device ID of each data stream to be processed, determine the target SM4 encryption core that matches the device ID from multiple SM4 encryption cores, and distribute the data stream to be processed to the target SM4 encryption core. Based on this strategy, resource utilization can be optimized and the workload between cores can be balanced. Furthermore, efficient task allocation in a multi - core environment can be achieved and processing delay can be reduced.

[0069] As Figure 4 shown, it is a schematic diagram of a multi - core scheduler in an exemplary embodiment. Among them, the red and yellow SM4 encryption cores respectively represent corresponding to a device ID, and the three green SM4 encryption cores correspond to the same device ID. The multi - core scheduler reads Ethernet data from the asynchronous FIFO to obtain the device ID in the data stream to be processed. According to the device ID, it determines the target SM4 encryption core that matches the device ID from multiple SM4 encryption cores, and distributes the data stream to be processed to the target SM4 encryption core. If there are multiple candidate SM4 encryption cores for this device ID, it determines the idle duration of each candidate SM4 encryption core, and determines the candidate SM4 encryption core with the longest idle duration as the target SM4 encryption core.

[0070] It should be noted that the determination of the above - mentioned two scheduling strategies of the multi - core scheduler is based on considerations of security issues and the encryption, decryption, and key expansion of SM4. The key will not change frequently, saving the time of key expansion, and thus improving efficiency.

[0071] It is understandable that block cipher algorithms divide the plaintext data into blocks of a fixed length (usually 64 bits or 128 bits), and then encrypt these blocks one by one. The encryption process for each block is either carried out independently or associated with the previous block in some modes. The core goal of block cipher algorithms is to improve the data encryption efficiency through block processing, and at the same time enhance security by utilizing the fixed-length feature of the blocks in cryptographic design. The block length of SM4 is 128 bits, and its encryption process is based on 32 rounds of encryption operations, achieving high-strength confusion and diffusion through steps such as the non-linear transformation of the S-box and linear diffusion.

[0072] As a kind of symmetric encryption algorithm, SM4 uses a fixed key to encrypt the plaintext data in blocks, and uses the same key to restore the original data during decryption to ensure the confidentiality and integrity of the data. The complete encryption process of the SM4 algorithm is as Figure 5 shown. First, 128-bit plaintext and 128-bit encryption key are input. The key is expanded through the key expansion algorithm to generate 32 128-bit round keys. The round keys and the plaintext are used as the input of the round function. After 32 rounds of iteration, a 128-bit ciphertext is output after an inverse order transformation. The decryption process of the SM4 algorithm is basically the same as the encryption process, except for the order of using the round keys. During decryption, the round keys start to be used from rk31 and rk0 is used last.

[0073] Furthermore, the hardware of the SM4 cryptographic algorithm includes two architectures: the loop architecture and the pipeline architecture. As Figure 6 shown, for the implementation of the loop architecture of the SM4 algorithm, when the input data is valid, the key expansion module starts to work to generate the round keys of the first round. After one clock cycle, the generation of the first-round round keys is completed, and the encryption / decryption module starts to work, and then starts to perform loop iteration. The key expansion continuously generates round keys, and the encryption / decryption module uses the round keys for encryption or decryption. The loop architecture can ensure that an encryption operation is performed every 32 clock cycles. This implementation method can ensure the optimal area, but there are two problems. First, this method is only applicable to the SM4 encryption process because the SM4 decryption process uses the round keys in reverse order, and there is a sequential dependence between the round keys, so all keys must be generated before decryption can be performed. Second, because the key does not change every time during the SM4 encryption / decryption process, the corresponding round keys also remain unchanged for a period of time. However, the loop architecture only saves the round key of a certain round, and the key expansion module needs to keep working, so more power is consumed.

[0074] For the pipelined architecture implementation of the SM4 algorithm, after the key expansion module receives a new key, it will work continuously for 32 clock cycles to generate all the round keys and then input them into the encryption / decryption module. Secondly, sub-modules can be instantiated in the encryption / decryption module. Each sub-module can implement one round of encryption / decryption operation. According to different requirements, 2, 4, 8, 16 or 32 levels of full pipelines can be formed. In the full pipeline mode, an encryption / decryption result can be output every clock cycle. However, as the number of pipeline levels increases, the area overhead will also gradually increase.

[0075] Based on the problems existing in the above two architectures, the key expansion module in the pipelined architecture is combined with the encryption / decryption module in the loop architecture.

[0076] In an exemplary embodiment, the SM4 algorithm architecture of each SM4 encryption core includes a control module, a key expansion module, and an encryption / decryption module. Among them, the control module is used to control the key expansion module to stop working in the case of no new key input; the encryption / decryption module starts working after the key expansion module completes the generation of all round keys, controls the encryption / decryption module to use the round keys in the forward or reverse order according to the mode selection, and outputs an encryption or decryption result according to the preset clock cycle. Among them, the preset clock can be 32 clock cycles.

[0077] Adopting this architecture, in the case of no new key input, the key expansion module can stop working, and the gated clock technology is used to reduce power consumption; the encryption / decryption module starts working after the key expansion module completes the generation of all round keys, uses the round keys in the forward or reverse order according to the mode selection, and outputs an encryption or decryption result every 32 clock cycles, which can avoid the problems that the key expansion module needs to work all the time, consumes more power, and the area overhead gradually increases as the number of pipeline levels increases.

[0078] Furthermore, the S-box is the core component for implementing the non-linear transformation in the SM4 algorithm. There are also some problems with the implementation based on the lookup table. First of all, table lookup requires fixed storage resources, which may bring additional overhead in resource-constrained embedded devices. Secondly, the table lookup method is prone to exposing the access pattern and may be subject to side-channel attacks. An attacker can infer the key information by monitoring the access behavior. In addition, on some hardware platforms, the storage access latency of the table lookup operation may be uneven, affecting the performance stability. Therefore, the present invention uses a mathematical operation method to implement the S-box to reduce the dependence on storage, improve the anti-attack performance, and enhance the adaptability in the intelligent vehicle environment.

[0079] In an exemplary embodiment, the S-box of the SM4 encryption core is defined in GF(2^8), and the corresponding irreducible polynomial is x^8 + x^7 + x^6 + x^5 + x^4 + x^2 + 1, and can thus be expressed as:

[0080] S(x) = A(Ax + c)^(-1) + c (1)

[0081] Among them, matrix A and matrix c can be respectively expressed as:

[0082]

[0083] The multiplication and addition operations in formula (1) can be implemented through simple exclusive-or operations. The key lies in the problem of finding the inverse operation in the finite field. Let , and the inverse operation in the finite field is expressed as finding the multiplicative inverse a^(-1) of a, such that a × a^(-1) ≡ 1 mod p. Among them, a needs to be a non-zero element, that is, a ≠ 0.

[0084] It can be understood that the problem of finding the inverse operation in the finite field is often implemented using the extended Euclidean algorithm. Based on the Euclidean algorithm to find the greatest common divisor of two integers, and at the same time find the particular solution that satisfies the linear congruence equation through reverse calculation. However, in hardware implementation, there are many iterative steps, which will lead to a decrease in the substitution efficiency of the S-box and affect the overall operation speed of the algorithm. Therefore, the method of composite field inversion is used to reduce the complexity of directly finding the inverse on GF(2^8) and improve the operation efficiency. Among them, the composite field is a special representation form of the finite field. It can be used to represent an element in GF(2^8) with two elements in GF((2^4)^2). After transformation, the inverse operation on GF(2^8) can be transformed into an operation on GF(2^4). Similarly, it can be simplified step by step until it is reduced to an operation on GF(2), significantly reducing the complexity of the inverse element.

[0085] Therefore, in an exemplary embodiment, the inverse operation in GF(2^8) is transformed into the composite field GF(((2^2)^2)^2). The flow chart of the composite field inversion is as Figure 7 shown. Specifically, after the input data of the S-box (such as 8-bit binary data) undergoes affine transformation and isomorphic mapping transformation, the inverse operation in GF(2^8) is sequentially transformed into the composite field GF(((2^2)^2)^2), and the inverse element calculation is performed to output the result. The inverse isomorphic mapping and affine transformation are performed on the output result to output the inverse element. Among them, the isomorphic mapping and inverse isomorphic mapping matrices used are:

[0086]

[0087] Based on the SM4 block cipher algorithm given in the standard GB / T32907-2016, an operation example of encrypting a group of plaintext is given.

[0088] Input plaintext: ;

[0089] Input key: ;

[0090] Ciphertext output: 。

[0091] It should be noted that the implementation of encryption and decryption of the SM4 encryption core can be achieved by combining the key expansion module in the pipeline architecture with the encryption and decryption modules in the loop architecture, implementing the S-box through mathematical operations, and using the method of composite field inversion to replace the specific implementation part in the related technology. On this basis, the implementation principle of encryption and decryption of the SM4 encryption core can be achieved by existing methods and will not be elaborated here.

[0092] The encryption and decryption processes of SM4 have been incorporated into the standard GB / T32907-2016. Its algorithm process, key expansion, round function, and encryption / decryption operations are clearly defined, and standard encryption and decryption examples are provided to ensure the correctness and consistency of different implementation methods. Therefore, the implementation process of encrypting and decrypting the data stream to be processed based on SM4 in this application will not be elaborated.

[0093] It is also worth noting that the above examples involve optimizing the hardware implementation method of SM4 to improve computing efficiency and security, rather than the definition of the algorithm itself. Therefore, the specific implementation of the SM4 principle will not be elaborated here.

[0094] In an exemplary embodiment, as Figure 8 shown, an architecture diagram of a communication encryption acceleration device applied to multiple Internet of Things devices is provided. The communication encryption acceleration device includes a communication receiving module, an asynchronous FIFO, a multi-core scheduler, multiple SM4 encryption cores, a synchronous FIFO, and a communication sending module, where:

[0095] The communication receiving module seamlessly receives plaintext data from different devices based on wireless and wired communication protocols, that is, receives the data streams to be processed sent by multiple Internet of Things devices. The asynchronous FIFO performs cross-clock domain synchronization processing and decoupling on the transmission of each data stream to be processed between the front-stage module and the rear-stage module of the asynchronous FIFO; and during the transmission between the front-stage module and the rear-stage module of the asynchronous FIFO, if the receiving rate of the front-stage module and the processing rate of the rear-stage module do not match, the data to be processed is cached.

[0096] The multi-core scheduler schedules each data stream to be processed transmitted by the asynchronous FIFO, distributes each data stream to be processed to the target SM4 encryption core among multiple SM4 encryption cores; the target SM4 encryption core encrypts or decrypts the corresponding data stream to be processed to obtain the target data and caches the target data in the synchronous FIFO; the communication sending module obtains the target data from the synchronous FIFO and transmits the target data to the corresponding Internet of Things device.

[0097] In this embodiment, by designing a multi-core scheduler and a parallel encryption and decryption architecture for multiple SM4 cores, the limitation of only optimizing the performance of a single encryption core in the prior art is solved. On this basis, not only the performance improvement of the SM4 encryption core is concerned, but also a multi-core parallel computing and task scheduling mechanism is innovatively proposed, which can support the parallel encryption and decryption processing of multiple devices in the Internet of Things scenario. The multi-core scheduler effectively distributes the data stream to multiple SM4 encryption cores through device ID matching and the "least recently used" (LRU) strategy, optimizes the resource utilization rate, improves the overall encryption processing throughput, and meets the requirement of simultaneous encryption and decryption of multiple devices in the communication of Internet of Things devices. This design significantly reduces the latency and improves the parallel processing ability of the system, providing an efficient hardware acceleration solution for Internet of Things communication encryption.

[0098] Based on the same inventive concept, an embodiment of the present application also provides a communication encryption acceleration method for implementing the communication encryption acceleration device involved above. The solution for solving the problem provided by this device is similar to the solution recorded in the above device. Therefore, the specific limitations in one or more embodiments of the communication encryption acceleration method provided below can refer to the limitations on the communication encryption acceleration device in the above text, and will not be repeated here.

[0099] In an exemplary embodiment, as Figure 9 shown, a communication encryption acceleration method is provided, which is applied to the communication encryption acceleration device in any of the above embodiments, and includes the following steps:

[0100] Step 902, receive multiple data streams to be processed sent by Internet of Things devices.

[0101] Step 904, for each data stream to be processed, perform cross-clock domain synchronization processing and decoupling on the transmission between the front-stage module and the rear-stage module of the asynchronous FIFO.

[0102] Step 906, schedule each data stream to be processed transmitted by the asynchronous FIFO, and distribute each data stream to be processed to the target SM4 encryption core among multiple SM4 encryption cores.

[0103] Step 908, perform encryption or decryption processing on the corresponding data stream to be processed, obtain the target data, and cache the target data in the cache module.

[0104] Step 910, obtain the target data from the cache module, and transmit the target data to the corresponding Internet of Things device.

[0105] In the case of multiple data streams to be processed sent by an Internet of Things device, the above-mentioned communication encryption acceleration method performs cross-clock domain synchronization processing and decoupling through an asynchronous FIFO, and through a multi-core scheduler, schedules each data stream to be processed transmitted by the asynchronous FIFO, and distributes each data stream to be processed to a target SM4 encryption core among multiple SM4 encryption cores, so that the target SM4 encryption core performs encryption and decryption processing on the data stream to be processed. That is, this method can dynamically allocate computing resources, reasonably schedule encryption and decryption tasks among multiple cores to cope with the high concurrency requirements in the Internet of Things device environment, and realizes the scheduling of data streams of different devices and the distribution of encryption and decryption tasks in the case of multiple cores.

[0106] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0107] Based on the same inventive concept, an embodiment of the present application also provides an Internet of Things device for implementing the above-mentioned communication encryption acceleration device method. The solution for solving the problem provided by this Internet of Things device is similar to the solution described in the above device, and will not be repeated here.

[0108] In an exemplary embodiment, an Internet of Things device is provided. The Internet of Things device includes a processor, a memory, an input / output interface, a communication interface, a display unit, an input device, and a communication encryption acceleration device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements a communication encryption acceleration method. The display unit of the Internet of Things device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the Internet of Things device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0109] In an exemplary embodiment, as Figure 10 shown, a communication encryption acceleration system is provided. The system includes an Internet of Things device and the communication encryption acceleration device described in any one of the above. The communication encryption acceleration device performs encryption and decryption processing on the data stream to be processed sent by the Internet of Things device.

[0110] Each module in the above communication encryption acceleration device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0111] In an embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0112] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0113] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps in the above method embodiments.

[0114] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0115] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., without limitation.

[0116] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.

[0117] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A communication encryption acceleration device, characterized in that: The communication encryption acceleration device includes a communication module, an asynchronous FIFO, a multi-core scheduler, a plurality of SM4 encryption cores and a cache module, wherein: The communication module is used to receive multiple data streams to be processed sent by the Internet of Things device; The asynchronous FIFO is used to perform cross-clock domain synchronization processing and decoupling of the transmission between the front-stage module and the back-stage module of each of the to-be-processed data streams; The multi-core scheduler is used to schedule each of the to-be-processed data streams transmitted by the asynchronous FIFO, and distribute each of the to-be-processed data streams to a target SM4 encryption core among the plurality of SM4 encryption cores; The target SM4 encryption core is used to encrypt or decrypt the corresponding data stream to be processed, obtain target data and cache the target data to the cache module; The communication module is further used to obtain the target data from the cache module and transmit the target data to the corresponding Internet of Things devices.

2. The communication encryption acceleration device according to claim 1, characterized in that: The asynchronous FIFO is also used to transmit each of the data streams to be processed between the front-stage module and the rear-stage module of the asynchronous FIFO. If the receiving rate of the front-stage module does not match the processing rate of the rear-stage module, the data to be processed is cached.

3. The communication encryption acceleration device according to claim 1, characterized in that: The multi-core scheduler is also used to determine the device ID of each of the data streams to be processed, determine a target SM4 encryption core that matches the device ID from the multiple SM4 encryption cores, and distribute the data streams to be processed to the target SM4 encryption core.

4. The communication encryption acceleration device according to claim 3, characterized in that: The multi-core scheduler is also used to determine the idle time of each candidate SM4 encryption core if there are multiple candidate SM4 encryption cores matching the device ID from the multiple SM4 encryption cores, and determine the candidate SM4 encryption core with the longest idle time as the target SM4 encryption core.

5. The communication encryption acceleration device according to any one of claims 1 to 4, characterized in that: The SM4 algorithm architecture of each SM4 encryption core includes a control module, a key expansion module and an encryption and decryption module, wherein the control module is used to control the key expansion module to stop working when no new key is input; The encryption and decryption module starts working after the key expansion module completes the generation of all round keys, controls the encryption and decryption module to select the use of round keys in forward or reverse order according to the mode, and outputs an encryption or decryption result according to a preset clock cycle.

6. The communication encryption acceleration device according to claim 5, characterized in that: The cache module is a synchronous FIFO or an asynchronous FIFO.

7. The communication encryption acceleration device according to claim 6, characterized in that: The communication module supports multiple communication protocols, and is also used to transmit the target data to the corresponding IoT devices using a communication protocol that matches the target data. The matching communication protocol is determined from the multiple communication protocols based on the communication requirements of the IoT devices.

8. A communication encryption acceleration method, characterized in that: Applied to the communication encryption acceleration device according to any one of claims 1 to 7, the method comprising: Receive multiple data streams to be processed sent by IoT devices; For each of the data streams to be processed, transmission between the front-stage module and the back-stage module of the asynchronous FIFO is synchronously processed and decoupled across clock domains; Scheduling each of the to-be-processed data streams transmitted by the asynchronous FIFO, and distributing each of the to-be-processed data streams to a target SM4 encryption core among the plurality of SM4 encryption cores; Encrypting or decrypting the corresponding data streams to be processed to obtain target data and caching the target data in the cache module; The target data is obtained from the cache module, and the target data is transmitted to the corresponding IoT devices.

9. An Internet of Things device, characterized in that: The Internet of Things device includes a memory, a processor, and a communication encryption acceleration device as described in any one of claims 1 to 7, the memory stores a computer program, and the processor implements the steps of the method described in claim 8 when executing the computer program.

10. A communication encryption acceleration system, characterized in that: The system includes an Internet of Things device and a communication encryption acceleration device as described in any one of claims 1 to 7, and the communication encryption acceleration device performs encryption and decryption processing on the data stream to be processed sent by the Internet of Things device.

Citation Information

Patent Citations

  • Method and device for realizing high-speed SM4 password module based on FPGA (Field Programmable Gate Array)

    CN116488794A

  • Configurable secret key SM4 encryption and decryption system based on FPGA

    CN116506106A

  • SM4 encryption method based on FPGA

    CN119201832A

Cited By

  • SM4 hardware acceleration and anti-side channel protection method and system for domestic FPGA

    CN120358028A

  • High-efficiency compliance tax refund method and system for cross-border e-commerce

    CN120634759A