Data encryption method and system based on wafer-level chip heterogeneous platform

By using the multi-level intermediate representation framework (MLIR) to adapt and optimize symmetric encryption algorithms on wafer-level chips, the problems of performance bottlenecks and low resource utilization in the existing technology are solved, and efficient encryption processing and collaborative computing across hardware platforms are achieved.

CN120068128AActive Publication Date: 2025-05-30INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510564415.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

The performance bottlenecks implemented by software in the prior art, insufficient vectorization and instruction-level parallelism, redundancy defects in key expansion, resource competition and memory inefficiency, and insufficient heterogeneous hardware adaptation, resulting in low encryption throughput, large memory overhead, and difficulty in achieving efficient collaborative computing across CPUs, GPUs, and FPGAs.

Method used

The symmetric encryption algorithm is adapted to the heterogeneous hardware platform of wafer-level chips through the multi-level intermediate representation framework (MLIR), realizing vectorization processing and parallelization of the symmetric encryption algorithm, and combining memory management optimization strategies to dynamically allocate encryption tasks to heterogeneous computing units.

Benefits of technology

It significantly improves encryption throughput, reduces memory overhead, and realizes efficient collaborative computing across CPUs, GPUs, and FPGAs on wafer-level chips, improving algorithm performance and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068128A_ABST
    Figure CN120068128A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the technical field of wafer-level chip heterogeneous computing of a cryptographic algorithm, and provides a data encryption method and system based on a wafer-level chip heterogeneous platform, and the method comprises the steps: obtaining a plurality of groups of original data to be encrypted and an original key; generating a plurality of data encryption subtasks based on the plurality of groups of original data, and performing key expansion on the original key to generate an expansion key; based on a task scheduling mechanism of a multi-stage intermediate representation framework, scheduling the plurality of data encryption sub-tasks to corresponding hardware rear ends in parallel, and performing data encryption processing on the plurality of groups of original data according to the expansion key and the corresponding data encryption sub-tasks to obtain a plurality of encrypted data; and merging the multiple pieces of encrypted data to obtain target encrypted data. Therefore, the encryption throughput is improved, the memory overhead is reduced, and efficient cooperative computing across a CPU, a GPU and an FPGA on a wafer-level chip is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wafer-level chip heterogeneous computing for cryptographic algorithms, and particularly to a data encryption method and system based on a wafer-level chip heterogeneous platform. Background Art

[0002] The current mainstream Advanced Encryption Standard (AES) algorithm mainly relies on software implementation on general-purpose CPUs, and its efficiency is limited in the following aspects: 1. Performance bottlenecks in software implementation and insufficient vectorization and instruction-level parallelism: Traditional software implementations usually adopt byte-by-byte processing logic and do not fully utilize the SIMD (Single Instruction, Multiple Data) instruction set, resulting in limited performance. Since its implementation depends on general-purpose instruction sets (such as x86 or ARM), it is difficult to fully utilize the parallel computing capabilities of modern hardware. For example, although the AES implementation in OpenSSL supports Intel AES-NI instruction acceleration, on platforms without dedicated instruction sets (such as low-end embedded devices), the performance drops significantly, and the parallelization efficiency on multi-core CPUs is limited. Although some SIMD instructions (such as SSE, AVX) are used, the vectorization potential of the entire AES process is not fully covered. For example, row shift and column mixing still rely on byte-by-byte operations and cannot be fully converted into vector instructions, resulting in limited instruction throughput.

[0003] 2. Redundancy defects in key expansion and resource competition and memory inefficiency: In multiple encryption and decryption tasks, traditional implementations may have problems of redundant memory access and multi-thread competition in the key expansion stage. For example, when encrypting multiple data blocks with the same key, the round keys need to be regenerated each time, and the existing results are not reused. This repeated operation significantly increases the computational overhead in batch encryption scenarios, especially in the AES-256 mode with a longer key. In addition, traditional implementations lack optimized memory alignment and sharing strategies, resulting in memory resource competition among multi-threads or hardware units.

[0004] 3. Insufficient adaptation to heterogeneous hardware: Traditional solutions cannot dynamically adapt to multiple hardware backends. For example, GPU implementations usually require writing dedicated code (such as CUDA or OpenCL), with high development costs and difficulty in collaborating with CPU tasks; FPGA implementations are usually fixed pipeline designs and are difficult to flexibly adapt to different encryption modes. The high parallel potential of wafer-level chips has not been fully exploited, and traditional code cannot be automatically mapped to its distributed computing units. Summary of the Invention

[0005] The present invention provides a data encryption method and system based on a wafer-level chip heterogeneous platform, which are used to solve the performance bottleneck in the prior art realized by software, as well as the deficiencies in vectorization and instruction-level parallelism, the redundancy defect in key expansion, the resource competition and memory inefficiency, and the deficiency in heterogeneous hardware adaptation, and to achieve an improvement in encryption throughput, a reduction in memory overhead, and an efficient collaborative computing across CPUs, GPUs, and FPGAs on the wafer-level chip.

[0006] The present invention provides a data encryption method based on a wafer-level chip heterogeneous platform, including: Obtaining multiple groups of original data to be encrypted and an original key; Generating multiple data encryption subtasks based on the multiple groups of original data, and performing key expansion on the original key to generate an expanded key; Based on the task scheduling mechanism of the multi-level intermediate representation framework, parallelly scheduling the multiple data encryption subtasks to corresponding hardware backends, and through the hardware backends, performing data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backends are heterogeneous hardware computing units of the heterogeneous hardware platform of the wafer-level chip; Merging the multiple encrypted data to obtain target encrypted data.

[0007] In a possible implementation manner, the method further includes: Adapting a symmetric encryption algorithm to the heterogeneous hardware platform of the wafer-level chip by using the multi-level intermediate representation framework, where the heterogeneous hardware platform includes multiple heterogeneous hardware computing units; Converting the data processing operations of the symmetric encryption algorithm into vector instructions through the vectorized dialect of the multi-level intermediate representation framework; Performing memory management on the symmetric encryption algorithm based on a memory management optimization strategy, where the memory management optimization strategy includes a data flattening strategy, a shared memory strategy for the expanded key and the T-table, a data alignment optimization strategy, and a dynamic memory allocation strategy.

[0008] In a possible implementation manner, the method further includes: Performing vectorization processing on the byte substitution operation, row shift operation, and round key addition operation of the symmetric encryption algorithm through the vectorized dialect of the multi-level intermediate representation framework to convert them into vector instructions.

[0009] In a possible implementation manner, the method further includes: Flattening high-dimensional data in the symmetric encryption algorithm into one-dimensional data; Storing the expanded key and the T-table in the static shared memory of the wafer-level chip for access by the multiple heterogeneous hardware computing units; Pre-combine the S-box of the symmetric encryption algorithm with every four I8 table entries of the T table into one I32 value.

[0010] In a possible implementation, the method further includes: Generate corresponding target task partitioning code based on the task scheduling mechanism of the multi-level intermediate representation framework and the attributes of each hardware backend; Parallelly schedule the multiple data encryption subtasks to the corresponding hardware backends based on the target task partitioning code.

[0011] In a possible implementation, the method further includes: Based on the data length requirement of the symmetric encryption algorithm, perform data block partitioning on each group of original data, and perform data padding on the last data block to obtain multiple target data blocks; Through the hardware backend, perform data encryption processing on the multiple target data blocks of each group of original data according to the extended key and the symmetric encryption algorithm; Merge the target data blocks of each group of original data after data encryption processing to obtain multiple encrypted data.

[0012] The present invention also provides a data encryption system based on a wafer-level chip heterogeneous platform, including the following modules: A front-end module, configured to obtain multiple groups of original data to be encrypted and an original key; generate multiple data encryption subtasks based on the multiple groups of original data, and perform key expansion on the original key to generate an extended key; A parallel module, configured to parallelly schedule the multiple data encryption subtasks to the corresponding hardware backends based on the task scheduling mechanism of the multi-level intermediate representation framework; through the hardware backend, perform data encryption processing on the multiple groups of original data according to the extended key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backend is a heterogeneous hardware computing unit of a heterogeneous hardware platform of a wafer-level chip; merge the multiple encrypted data to obtain target encrypted data.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the data encryption method based on the wafer-level chip heterogeneous platform as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data encryption method based on the wafer-level chip heterogeneous platform as described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the data encryption method based on a wafer-level chip heterogeneous platform as described in any one of the above.

[0016] The data encryption method and system based on a wafer-level chip heterogeneous platform provided by the present invention obtain multiple groups of original data to be encrypted and an original key; generate multiple data encryption subtasks based on the multiple groups of original data, and expand the original key to generate an expanded key; based on the task scheduling mechanism of a multi-level intermediate representation framework, parallelly schedule the multiple data encryption subtasks to corresponding hardware backends, and through the hardware backends, perform data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, wherein the hardware backends are heterogeneous hardware computing units of a heterogeneous hardware platform of a wafer-level chip; and merge the multiple encrypted data to obtain target encrypted data. Compared with the performance bottleneck of software implementation, the deficiencies of vectorization and instruction-level parallelism, the redundancy defect of key expansion, the resource competition and memory inefficiency, and the deficiency of heterogeneous hardware adaptation in the prior art, in this solution, through a multi-level intermediate representation framework, high-efficiency multi-hardware backend parallelization of the symmetric encryption algorithm on a wafer-level chip is achieved, and at the same time, combined with vectorization optimization and memory management strategies, the algorithm performance and resource utilization rate are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0018] Figure 1 is a flowchart of the data encryption method based on a wafer-level chip heterogeneous platform provided by the present invention.

[0019] Figure 2 is a data flow diagram of the data encryption method based on a wafer-level chip heterogeneous platform provided by the present invention.

[0020] Figure 3 is a structural diagram of the data encryption system based on a wafer-level chip heterogeneous platform provided by the present invention.

[0021] Figure 4 is a flowchart of the parallel module of the data encryption system based on a wafer-level chip heterogeneous platform provided by the present invention.

[0022] Figure 5 is a flowchart of key expansion provided by the present invention.

[0023] Figure 6 It is the schematic diagram of T-table generation provided by the present invention.

[0024] Figure 7 It is the schematic diagram of parallel scheduling provided by the present invention.

[0025] Figure 8 It is the schematic diagram of encryption optimization by look-up table method provided by the present invention.

[0026] Figure 9 It is the schematic diagram of decryption optimization by look-up table method provided by the present invention.

[0027] Figure 10 It is the schematic diagram of column mixing of look-up table provided by the present invention.

[0028] Figure 11 It is the schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] For the convenience of understanding the embodiments of the present invention, the following will further explain and illustrate with specific embodiments in conjunction with the accompanying drawings. The embodiments do not constitute a limitation to the embodiments of the present invention.

[0031] Figure 1 It is one of the schematic diagrams of the data encryption method based on the wafer-level chip heterogeneous platform provided by the present invention. As Figure 1 shown, the method includes the following: S11. Obtain multiple groups of original data to be encrypted and the original key.

[0032] In the embodiments of the present invention, the symmetric encryption algorithm (Advanced Encryption Standard, AES) is adapted to the wafer-level chip through a multi-level intermediate representation framework (MLIR). The AES algorithm is a symmetric encryption algorithm of the advanced encryption standard. By utilizing the hardware dialect support and vectorization optimization capabilities of the MLIR framework, efficient code for the distributed computing units of the wafer-level chip is dynamically generated. Specifically, it includes: implementing the full-process vectorized instruction conversion of the core operations of AES (such as byte substitution, row shift, and column mixing) based on MLIR; optimizing the access to the key and lookup table data through the shared memory strategy to reduce resource competition among multiple cores; and dynamically allocating encryption tasks to the heterogeneous computing units of the wafer-level chip in combination with the MLIR task scheduling mechanism. The embodiments of the present invention can significantly improve the encryption throughput, reduce the memory overhead, and achieve efficient collaborative computing across CPUs, GPUs, and FPGAs on the wafer-level chip.

[0033] Specifically, in combination with Figure 2 As shown in the data flow diagram, at the initial stage of the encryption process, the system first needs to obtain the data to be encrypted and the original key for encryption from the user side to the front-end module. This data can be binary data in any form, such as the content of a text file, a data packet transmitted over the network, or other information that needs to be protected. The size of each group of data can be different, but for the convenience of subsequent processing, the data is usually divided into fixed-size blocks (for example, 16 bytes, which meets the requirements of the AES algorithm).

[0034] Furthermore, the system receives the original key, which is the initial key provided by the user for encryption. The AES algorithm supports multiple key lengths, such as 128 bits, 192 bits, or 256 bits. The length of the original key determines the security of the encryption and the complexity of subsequent key expansion.

[0035] For example, when the system needs to encrypt the content of a group of text files, the user provides a 128-bit original key. The system reads the content of the text file as binary data and divides it into multiple 16-byte data blocks.

[0036] S12. Generate multiple data encryption subtasks based on the multiple groups of original data, and perform key expansion on the original key to generate an expanded key.

[0037] After the system receives multiple groups of original data, to improve processing efficiency, the multiple groups of original data are divided into multiple subtasks. Each subtask contains one or more data blocks, and these data are allocated to different hardware backends for processing through parallel scheduling. The specific division method depends on the parallel processing ability of the system and hardware resources. For example, if the system supports multi-threading or multiple hardware backends, the original data can be divided into multiple subtasks for parallel processing to improve encryption efficiency.

[0038] Furthermore, the original key needs to go through a key expansion operation to generate the expanded keys required for multiple rounds of encryption. As Figure 5 shown in the key expansion flowchart, key expansion is the core process of the AES algorithm. The original 16-byte key is expanded into a 176-byte key after multiple rounds of parallel exclusive OR, word rotation, byte substitution, and round coefficient exclusive OR operations. Whether it is encryption or decryption, the implementation of key expansion is exactly the same. The specific implementation of the embodiments of the present invention is as follows: In the initialization stage of key expansion, the runtime overhead is reduced by pre-computing the round constants (Rcon) and loading the vectorized constants. Rcon is an important parameter in the AES key expansion. The code predefines it as a constant array, avoiding repeated calculations in each round of key expansion. In addition, a mask of vector<4xI1> is created through vector::CreateMaskOp to control the range of subsequent vector operations.

[0039] Only one parallel exclusive OR operation is involved in the ordinary rounds of key expansion, that is, the exclusive OR operation is performed on each byte in two vectors simultaneously through the arith::XOrIOp operation. For the special rounds of key expansion, three additional operations are required: First is the word rotation. The system implements the word rotation step through the vectorized byte rearrangement vector::ShuffleOp operation, that is, 4 bytes are circularly shifted left by 1 byte. The circular left shift of the bytes in the vector is achieved by means of a rearrangement index vector in the form of inv4 = {1, 2, 3, 0}. This rearrangement operation is completed by a single instruction, avoiding multiple shift and splicing operations in the traditional implementation and significantly improving the efficiency.

[0040] Then byte substitution is performed (the byte substitution here is the same as the byte substitution step in the core algorithm of AES encryption and decryption, so it is briefly summarized here), that is, data is collected non-continuously from the S-box using vectorized look-up tables. Specifically, the indexes of 4 bytes are first expanded to 32 bits, and the replacement results of 4 bytes are collected from the S-box at once through vector::GatherOp, avoiding the per-byte look-up table operation in the traditional implementation. In addition, the S-box has been pre-loaded into the static memory before the key expansion starts, which further reduces the memory access latency during the look-up table operation.

[0041] Finally, it is the round coefficient XOR. The Rcon value is XORed with the key bytes through the vectorized XOR operation arith::XOrIOp. The Rcon value is directly obtained through a precomputed constant array. The parallel XOR operation makes full use of multiple execution units of the backend hardware, not only reducing the number of instructions but also improving the instruction-level parallelism.

[0042] Vectorized optimization is also required for data reading and writing. In key expansion, 4 consecutive bytes need to be loaded from memory into a vector register (vectorI8x4) through the vectorized load operation vector::LoadOp. After the key calculation is completed, the system uses the vector::StoreOp operation to complete the vectorized storage of the expanded key and stores it in the static memory area for shared access by multiple parallel module threads. Vectorized reading and writing significantly improve the efficiency of key expansion.

[0043] The theoretical performance improvement points of AES key expansion are as follows: Precomputing Rcon and vectorized constant loading reduce the runtime overhead; Vectorized loading and byte rearrangement optimize memory access and word loop operations; Vectorized table lookup and preloading the S-box accelerate the byte substitution operation; Vectorized XOR operation and branchless design improve the instruction-level parallelism; Vectorized reading and writing reduce the number of memory accesses.

[0044] S13. A task scheduling mechanism based on a multi-level intermediate representation framework parallelly schedules the multiple data encryption subtasks to the corresponding hardware backend. Through the hardware backend, according to the expanded key and the corresponding data encryption subtasks, data encryption processing is performed on the multiple groups of original data to obtain multiple encrypted data.

[0045] S14. Merge the multiple encrypted data to obtain the target encrypted data.

[0046] In the embodiments of the present invention, the hardware backend is a heterogeneous hardware computing unit of a heterogeneous hardware platform of a wafer-level chip. Based on the task scheduling mechanism of the multi-level intermediate representation framework and the attributes of each hardware backend, corresponding target task partitioning codes are generated; Based on the target task partitioning codes, multiple data encryption subtasks are parallelly scheduled to the corresponding hardware backend.

[0047] Further, based on the data length requirements of the symmetric encryption algorithm, each group of original data is divided into data blocks, and the last data block is filled with data to obtain multiple target data blocks; Through the hardware backend, data encryption processing is performed on the multiple target data blocks of each group of original data according to the expanded key and the symmetric encryption algorithm; The target data blocks of each group of original data after data encryption processing are merged to obtain multiple encrypted data.

[0048] It should be noted that the front-end module also needs to pre-allocate static memory spaces for the expansion key, S-box, inverse S-box, and T-table, so that the parallel module can share memory data. After scheduling, each group of original data (a group of original data corresponding to a data encryption sub-task) will be divided into multiple 16-byte data blocks, and then the data blocks will be used as the input of the parallel module in units of data blocks. After the parallel module finishes the operation, all the data blocks will be merged in sequence to form the final result, and padding operations will also be performed on the data blocks at the end of each group of original data; the output of this module is the result after encryption and decryption of multiple groups of original data.

[0049] Such as Figure 6 The schematic diagram of T-table generation provided by the present invention. T-table generation is to pre-calculate polynomial multiplication in the finite field GF(2^8), and then load the result into static memory. It is necessary to perform modular multiplication operations on all I8 values (0x00 - 0xFF) with the column mixing matrix coefficients to obtain the operation results of 256 input values, and then store them in the corresponding positions of the T-table, so as to optimize the matrix multiplication of column mixing into fast bit operations of table lookup and exclusive OR. If the system performs AES encryption, there are only 4 groups of index tables in the T-table related to 0x01 / 0x01 / 0x02 / 0x03; if it is AES decryption, there are only 4 groups of index tables in the T-table related to 0x09 / 0x0B / 0x0D / 0x0E, which ensures the reusability of the T-table. Similar to key expansion, the generation of the T-table is only executed once in multiple parallel AES tasks, avoiding redundant T-table calculations in the parallel module.

[0050] Specific implementation of modular multiplication operation: The algorithm checks the lowest bit of the second multiplier bit by bit to determine whether to perform an exclusive OR operation on the intermediate result with the temporary value temp; at the same time, it determines whether to perform finite field reduction with the irreducible polynomial 0x1b by checking the highest bit of temp to ensure that the result is still within the range of GF(2^8). The system uses arithmetic operations of MLIR such as arith::ShLIOp, arith::ShRUIOp, arith::AndIOp, and arith::XOrIOp to implement efficient calculation logic, significantly improving the performance of the column mixing step.

[0051] The calculation result of the T-table is loaded into continuous static memory (4 In the 256 - byte area, the system further improves performance through the alignment optimization of the T - table and memory - sharing optimization. The storage of the T - table, like that of the S - box, is compressed from a two - dimensional structure to a continuous one - dimensional array. For the storage of the entries in the T - table, the system combines 4 consecutive I8 - byte values into one I32 value to align memory access. In this way, 4 bytes can be loaded at once using vector::LoadOp instead of loading them one by one, thus reducing memory - access latency. Secondly, multiple groups of parallel tasks access the shared T - table memory, avoiding the overhead of multiple groups of tasks repeatedly loading the T - table, thereby increasing the cache hit rate when multiple tasks are executed in parallel. In addition, a mask is created using vector::CreateMaskOp to control the scope of vector operations and avoid unnecessary memory access.

[0052] Such as Figure 7 The parallel - scheduling schematic diagram provided by the present invention. In the specific implementation of multi - group task partitioning, the system, based on the parallel - loop operation hyper::ForOp of the heterogeneous - scheduling dialect and the hardware - information library, dynamically distributes multiple groups of AES tasks to different hardware back - ends and supports parallel computing between devices. First, the system obtains the hardware configuration of the current system (such as the number of devices, computing power) through the hardware - information library and divides the hardware resources into multiple groups, each group containing one or more devices. Then, the system dynamically distributes AES tasks according to the load ratio of each group: calculates the load ratio of each group and determines the number of tasks each group needs to process based on the total number of tasks and the load ratio, so as to maximize the utilization rate of hardware resources. After task distribution, the system uses hyper::ForOp to take each group of tasks as the input of the parallel module and executes the tasks in parallel on heterogeneous devices through the calculation logic described in its sub - regions. Each device independently processes the assigned tasks and at the same time uses reduction operators to perform reduction operations on the calculation results, finally achieving efficient load balancing and parallel computing. This design not only makes full use of the computing power of heterogeneous hardware but also optimizes the overall performance through task partitioning and dynamic scheduling, ensuring that the devices can execute in parallel efficiently after task assignment.

[0053] Before encryption, the original text in each group of tasks needs to be divided into 16-byte data blocks. For the last data block, the system pads it according to the PKCS #7 standard to ensure that the data length meets the block requirements of the AES algorithm (aligned to 16 bytes). For decryption, after all calculations are completed, the padded data is removed according to the PKC #7 standard to restore the original data. The specific implementation is as follows: The system first calculates the remaining number of bytes of the input data and generates a padding value. The size of the padding value is 16 minus the remaining number of bytes, and the padding content is the byte representation of the padding value. The system determines whether to pad the current block through the dynamic masking operation vector::CreateMaskOp. If padding is required, the system loads the original data through the vector::MaskedLoadOp operation and inserts the padding bytes; otherwise, it directly loads the original data. Vectorized masking operations and conditional branch optimizations are used in the padding to avoid multiple memory accesses and branch prediction failures in traditional implementations, improving the padding efficiency. In addition, the system broadcasts the padding value to the entire vector through vector::SplatOp to ensure the unity and efficiency of the padding operation.

[0054] The core idea of the AES algorithm is to confuse and diffuse the data through multiple rounds of encryption operations (including byte substitution, row shift, column mixing, and round key addition) to ensure the encryption strength. AES supports key lengths of 128 bits, 192 bits, and 256 bits, and each round of operation processes a 16-byte data block. The encryption process includes an initial round key addition, multiple rounds of encryption (10 / 12 / 14 rounds, depending on the key length), and a final round of encryption. The decryption process is the inverse operation of encryption, restoring the original data through reverse steps.

[0055] Specifically, the round key addition is the step of exclusive-oring the round key with the data block byte by byte. In traditional implementations, it usually adopts the method of traversing and exclusive-oring byte by byte, with low efficiency. This system uses the vectorized exclusive-or operation of MLIR to implement an efficient vectorized exclusive-or operation. The system first loads the 16-byte key fragment of the current round from the expanded key through vector::LoadOp. Since the total length of the expanded key of AES-128 is 176 bytes, the system locates the key fragment of each round through index calculation. The loaded key fragment and the data block are exclusive-ored byte by byte through arith::XOrIOp to generate the encrypted intermediate result. This vectorized exclusive-or operation not only reduces the number of instructions but also fully utilizes the parallel computing units of the hardware, significantly improving the computing efficiency. In addition, the system combines multiple rounds of key addition operations into a single vectorized operation through loop unrolling technology, reducing the overhead of loop control.

[0056] Byte substitution is a non-linear transformation step in the AES encryption algorithm. Each byte is replaced by querying the S-box to enhance the confusion and security of the algorithm. The S-box is queried during encryption, and the inverse S-box is queried during decryption. In traditional implementations, the efficiency of looking up the table byte by byte is low. This system uses the vector::GatherOp of MLIR to implement vectorized table lookup, processing 16-byte input data at once, significantly reducing the number of table lookups and memory access latency. The specific implementation is as follows: First, the system extends the 16-byte original text to a 32-bit index through arith::ExtUIOp to ensure alignment with the S-box memory address. Then, vector::GatherOp is used to collect the 16-byte replacement results from the pre-loaded S-box at once, and the mask (maskI1x16) controls the position of valid data. Finally, the replaced data is written to the output buffer through vector::StoreOp to complete efficient storage. This vectorized table lookup method greatly improves the execution efficiency of byte substitution.

[0057] Row shift enhances the diffusion of data and the security of the algorithm by cyclically shifting each row of the state matrix by different offsets. In traditional implementations, the efficiency of changing byte by byte is low. This system directly completes the row shift operations for all rows by rearranging the 16-byte vector through vector::ShuffleOp and the predefined index vector inv16. For example, during encryption, inv16 = {0, 5, 10, 15, 4, 9, 14, 3, 8, 13, 2, 7, 12, 1, 6, 11}; during decryption, inv = {0, 13, 10, 7, 4, 1, 14, 11, 8, 5, 2, 15, 12, 9, 6, 3}, and efficient rearrangement is achieved through a single vector instruction. This optimization avoids row-by-row shifting and splicing operations, fully utilizes hardware vectorization support, and significantly improves the efficiency. At the same time, combined with mask operations (vector::CreateMaskOp) to ensure the accuracy and consistency of data processing.

[0058] Table lookup column mixing mixes each column of the state matrix through finite field multiplication to enhance the diffusion of data. However, its traditional implementation relies on complex finite field multiplication operations and has low efficiency. The embodiment of the present invention uses the table lookup method to replace the two steps of byte substitution and column mixing in the main round, as Figure 8It is a schematic diagram of the optimized look-up table method encryption provided by the present invention, that is, by querying the pre-computed T table to replace the finite field multiplication operation with a higher complexity. To meet the correctness of the look-up table method, it is necessary to swap the order of byte substitution and row shift in the main round of the AES algorithm, so that the two steps of byte substitution and column mixing are executed adjacent to each other, facilitating the direct combination of these two steps. The correctness of this operation lies in that the order of byte substitution and row shift in the main round does not affect the result of the AES algorithm. For AES encryption, it is necessary to advance the row shift in the main round, that is, perform look-up table column mixing after the row shift; for AES decryption, it is necessary to place the inverse row shift in the main round at the back, that is, perform the inverse row shift after the inverse look-up table column mixing, as Figure 9 It is a schematic diagram of the optimized look-up table method decryption provided by the present invention.

[0059] Figure 10 It is a schematic diagram of the look-up table column mixing principle provided by the present invention, as Figure 10 shown, the look-up table column mixing needs to intercept the original 16-byte text into 4 segments of size vector<4xi8> using the vector::ExtractStridedSliceOp operation. For each segment, a vector::GatherOp operation similar to byte substitution is used to complete the vectorized look-up table operation. For each segment, 4 different vectorized look-up table operations are required, corresponding to the 4 finite field matrix multiplications in the column mixing respectively. After the look-up table is completed, the 16 vector segments obtained are respectively subjected to the exclusive OR reduction operation of the segment through the vector::ReductionOp to obtain the result of vector<16xI8>. Finally, it is written back to the continuous memory through vector.store.

[0060] The data encryption method based on the wafer-level chip heterogeneous platform provided by the present invention includes obtaining multiple groups of original data and an original key to be encrypted; generating multiple data encryption subtasks based on the multiple groups of original data, and expanding the original key to generate an expanded key; based on the task scheduling mechanism of the multi-level intermediate representation framework, parallelly scheduling the multiple data encryption subtasks to the corresponding hardware backends, and through the hardware backends, performing data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backends are heterogeneous hardware computing units of the heterogeneous hardware platform of the wafer-level chip; merging the multiple encrypted data to obtain the target encrypted data. Compared with the performance bottleneck of software implementation in the prior art, as well as the deficiencies in vectorization and instruction-level parallelism, the redundancy defect of key expansion, resource competition and memory inefficiency, and the deficiency in heterogeneous hardware adaptation, by this method, the symmetric encryption algorithm is efficiently parallelized on the wafer-level chip with multiple hardware backends through the multi-level intermediate representation framework, and at the same time, combined with vectorization optimization and memory management strategies, the algorithm performance and resource utilization rate are significantly improved.

[0061] The data encryption system based on the wafer-level chip heterogeneous platform provided by the present invention will be described below. The data encryption system based on the wafer-level chip heterogeneous platform described below can be correspondingly referred to the data encryption method based on the wafer-level chip heterogeneous platform described above.

[0062] Figure 3 It is a schematic structural diagram of the data encryption system based on the wafer-level chip heterogeneous platform provided by the present invention, specifically including: A front-end module 301, configured to obtain multiple groups of original data and an original key to be encrypted; generate multiple data encryption subtasks based on the multiple groups of original data, and expand the original key to generate an expanded key; A parallel module 302, configured to parallelly schedule the multiple data encryption subtasks to corresponding hardware back-ends based on the task scheduling mechanism of the multi-level intermediate representation framework; through the hardware back-end, perform data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware back-end is a heterogeneous hardware computing unit of the heterogeneous hardware platform of the wafer-level chip; merge the multiple encrypted data to obtain target encrypted data.

[0063] Specifically, the front-end module 301 includes: a parallel scheduling unit, configured to divide multiple groups of original text data into multiple subtasks and allocate them to a specified hardware back-end; a key expansion unit, configured to expand the original key to generate an expanded key; a memory management unit, configured to pre-allocate a static memory space and optimize the management of the memory.

[0064] The parallel module 302 includes: a data block processing unit, configured to divide the original text data into multiple 16-byte data blocks and perform a padding operation on the last data block; an encryption / decryption execution unit, configured to perform AES encryption / decryption operations, including byte substitution, row shift, column mixing, and round key addition.

[0065] Multiple heterogeneous computing units can be deployed on the wafer chip, which has high-throughput parallel computing capabilities and is very suitable for the parallel grouping of multiple groups of AES encryption / decryption operations. By borrowing the cross-architecture support provided by MLIR, only by calling the library of the MLIR framework to descend to the unified intermediate MLIR IR can the compilation process of multiple hardware back-ends be completed, without the need to write independent operator kernels for multiple parallel back-ends or transplant the operator library on a large scale. The wafer chip usually contains a powerful vector instruction set, and the vectorized arithmetic provided by MLIR can make full use of the performance of the vector hardware to directly optimize the key steps of encryption.

[0066] Using the vectorization dialects of MLIR, such as the vector dialect, to convert the core operations of the AES algorithm (such as byte substitution, row shift, and column mixing) into efficient vector instructions, thereby achieving full-process vectorization. For example, the batch processing of the S-box lookup table operation is implemented through vector::GatherOp, and the single-instruction rearrangement of row shift is completed through vector::ShuffleOp, significantly reducing the number of instructions and memory access times. The automatic vectorization ability of MLIR further reduces the manual optimization cost while ensuring instruction-level parallelism across hardware platforms.

[0067] Memory management unit, used to pre-allocate static memory space and optimize the management of memory, including: Data flattening: The traditional implementation of the AES algorithm is based on byte-by-byte calculation of two-dimensional data. This system flattens all high-dimensional data in the traditional AES implementation to one-dimensional, making it easier to load into vector registers and supporting batch processing of SIMD instructions.

[0068] Shared memory for expanded key and T-table: Key expansion and T-table generation are completed in the front-end module, and the expanded key and the generated T-table are stored in the static shared memory of the wafer-level chip for direct access by all parallel computing units, without the need for repeated calculations in the parallel module.

[0069] Data alignment optimization: For the two index tables of the S-box and T-table, every four I8 table entries are pre-computed and merged into one I32 value, leveraging the atomic read-write characteristics of 32-bit integers to improve memory access locality.

[0070] Dynamic memory allocation: Combining the memory management dialect of the MLIR framework, vectorially load and store intermediate vectors, dynamically allocate temporary memory as needed, and reduce memory occupancy.

[0071] Based on MLIR's automatic adaptation to multiple hardware backends, the system adapts the AES algorithm to multiple hardware backends (such as CPUs, GPUs, FPGAs, etc.) on the wafer-level chip, enabling task dynamic allocation and parallel computing. The hardware dialect support of MLIR enables the algorithm to generate optimal code according to different hardware characteristics, maximizing the utilization of heterogeneous computing resources. For example, multi-backend task partitioning is achieved through hyper::ForOp, dynamically allocating encryption tasks to different hardware units, avoiding resource idleness, and improving overall throughput.

[0072] Before optimizing the AES algorithm using this system on a wafer-level chip or simulator, pre-configuration is required, including: Build environment configuration: The user needs to specify the build directory and the compilation toolchain directory in the Linux environment to ensure that the files and tools (g++, make, llvm, etc.) required during the build process can be correctly found and used; Configure AES algorithm parameters: The user needs to specify the plaintext / ciphertext data, plaintext / ciphertext length, plaintext / ciphertext group number, key data, key length, key type, and output address in the front-end module.

[0073] Compile: The user compiles this MLIR project with the help of the Makefile file, and the intermediate files and target files will be generated in the directory specified by the user.

[0074] Execute the AES encryption / decryption execution program: Execute the compiled executable program, and the terminal will output the AES encryption / decryption results of each group of data.

[0075] The embodiment of the present invention also conducts experiments on the wafer-level chip system simulation platform. The test algorithm is AES-128, the encryption / decryption mode is ECB mode, and the padding mechanism is PKCS#7. The same key is used for multiple groups of AES tasks in the experiment.

[0076] The configuration of the wafer-level chip system simulation platform is as follows: CPU (Ariel): Main frequency: 2660 MHz, single-core, maximum number of requests per cycle: 3, network bandwidth: 96 GB / s.

[0077] GPU (NVIDIA GTX 480): Video memory: 1024 KB stack, 8 MB heap, bandwidth: 96 GB / s, latency: 300 ps.

[0078] Memory: Main frequency: 800 MHz, bandwidth: 96 GB / s, latency: 1000 ps, capacity: 4 GB.

[0079] Network: Latency: 25 ps, bandwidth: 96 GB / s, Flit size: 8 B.

[0080] Compiler: Built based on LLVM's MLIR, LLVM version: 17.0.6.

[0081] Experiment 1: Performance comparison Experiment description: Under the same configuration of the wafer-level chip system simulation platform, test the execution time of different algorithms to complete 100 groups of AES encryption / decryption tasks.

[0082] Objective: Verify the performance advantage of the AES algorithm optimization method based on the wafer-level chip heterogeneous platform compared with the traditional method.

[0083] The experimental results are shown in Table 1: Table 1

[0084] Conclusion: The AES optimization method based on wafer-level chips is significantly superior to traditional implementations in throughput. Its high parallel computing ability and the optimization ability of MLIR (such as vectorization and memory optimization) are the keys to performance improvement.

[0085] Experiment 2: Analysis of the Optimization Efficiency of Vectorization Experiment description: Under the same configuration of the wafer-level chip system simulation platform, measure the execution time of different implementation methods to complete 100 groups of AES encryption and decryption tasks and the length of the generated intermediate llvm code.

[0086] Objective: Verify the acceleration effect of MLIR vectorization on key steps.

[0087] The experimental results are shown in Table 2: Table 2

[0088] Conclusion: Vectorization optimization is crucial for improving the performance of the AES algorithm, especially for table lookup operations and XOR operations. After removing the vectorization operation, the execution time increases significantly, and the length of the LLVM intermediate code also increases accordingly.

[0089] Experiment 3: Verification of Memory Optimization Effect Experiment description: Under the same configuration of the wafer-level chip system simulation platform, measure the memory usage of different implementation methods to complete 100 groups of AES encryption and decryption tasks.

[0090] Objective: Verify the impact of each memory optimization strategy on performance.

[0091] The experimental results are shown in Table 3: Table 3

[0092] Conclusion: Memory optimization strategies significantly reduce memory occupancy, especially the memory sharing strategy for the expanded key and T table. Without this optimization, the memory occupancy increases significantly.

[0093] Experiment 4: Heterogeneous Hardware Cooperative Scheduling Test Experiment description: Under different configurations of the wafer-level chip system simulation platform, measure the hardware resource utilization rate when the CPU and GPU task loads are 1:0, 1:1, and 0:1 respectively to complete 100 groups of AES encryption and decryption tasks.

[0094] Objective: Verify whether the system can achieve cross-hardware task allocation.

[0095] The experimental results are shown in Table 4: Table 4

[0096] Conclusion: The experimental results show that the system can dynamically allocate tasks according to the load ratio, and the parallel scheduling mechanism of the system performs excellently in cross-hardware parallelism and resource utilization optimization.

[0097] This system adapts the AES algorithm to the wafer-level chip through the MLIR framework. Utilizing the hardware dialect support and vectorization optimization capabilities of MLIR, it dynamically generates efficient code for the distributed computing units of the wafer-level chip. Specifically, it includes: implementing the full-process vectorized instruction conversion of the AES core operations (such as byte substitution, row shift, column mixing) based on MLIR; optimizing the access to keys and lookup table data through the shared memory strategy to reduce resource competition among multiple cores; combining the MLIR task scheduling mechanism to dynamically allocate encryption tasks to the heterogeneous computing units of the wafer-level chip, significantly improving the encryption throughput, reducing the memory overhead, and achieving efficient collaborative computing across CPUs, GPUs, and FPGAs on the wafer-level chip.

[0098] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communications interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communications interface 820, and the memory 830 complete the communication with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute the data encryption method based on the heterogeneous platform of the wafer-level chip. The method includes: obtaining multiple groups of original data to be encrypted and the original key; generating multiple data encryption subtasks based on the multiple groups of original data, and expanding the original key to generate an expanded key; based on the task scheduling mechanism of the multi-level intermediate representation framework, parallelly scheduling the multiple data encryption subtasks to the corresponding hardware backends, and through the hardware backends, performing data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backends are the heterogeneous hardware computing units of the heterogeneous hardware platform of the wafer-level chip; merging the multiple encrypted data to obtain the target encrypted data.

[0099] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0100] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the data encryption method based on the wafer-level chip heterogeneous platform provided by the above-mentioned various methods. The method includes: obtaining multiple groups of original data to be encrypted and an original key; generating multiple data encryption subtasks based on the multiple groups of original data, and performing key expansion on the original key to generate an expanded key; based on the task scheduling mechanism of the multi-level intermediate representation framework, parallelly scheduling the multiple data encryption subtasks to corresponding hardware backends, and through the hardware backends, performing data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backend is a heterogeneous hardware computing unit of the heterogeneous hardware platform of the wafer-level chip; merging the multiple encrypted data to obtain target encrypted data.

[0101] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the data encryption method based on the wafer-level chip heterogeneous platform provided by the above-mentioned various methods. The method includes: obtaining multiple groups of original data to be encrypted and an original key; generating multiple data encryption subtasks based on the multiple groups of original data, and performing key expansion on the original key to generate an expanded key; based on the task scheduling mechanism of the multi-level intermediate representation framework, parallelly scheduling the multiple data encryption subtasks to corresponding hardware backends, and through the hardware backends, performing data encryption processing on the multiple groups of original data according to the expanded key and the corresponding data encryption subtasks to obtain multiple encrypted data, where the hardware backend is a heterogeneous hardware computing unit of the heterogeneous hardware platform of the wafer-level chip; merging the multiple encrypted data to obtain target encrypted data.

[0102] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data encryption method based on a wafer-level chip heterogeneous platform, characterized in that: include: Obtain multiple groups of original data to be encrypted and original keys; Generating a plurality of data encryption subtasks based on the plurality of groups of original data, and performing key expansion on the original key to generate an expanded key; Based on the task scheduling mechanism of the multi-level intermediate representation framework, the multiple data encryption subtasks are scheduled in parallel to the corresponding hardware backend, and the multiple groups of original data are encrypted by the hardware backend according to the extended key and the corresponding data encryption subtasks to obtain multiple encrypted data, wherein the hardware backend is a heterogeneous hardware computing unit of a heterogeneous hardware platform of a wafer-level chip; The multiple encrypted data are merged to obtain target encrypted data.

2. The method according to claim 1, characterized in that Before obtaining the multiple groups of original data to be encrypted and the original key, the method includes: Adapting a symmetric encryption algorithm to a heterogeneous hardware platform of a wafer-level chip using a multi-level intermediate representation framework, wherein the heterogeneous hardware platform includes a plurality of heterogeneous hardware computing units; Converting data processing operations of the symmetric encryption algorithm into vector instructions through the vectorized dialect of the multi-level intermediate representation framework; The symmetric encryption algorithm is memory managed based on a memory management optimization strategy, wherein the memory management optimization strategy includes a data flattening strategy, a shared memory strategy for an extended key and a T table, a data alignment optimization strategy, and a dynamic memory allocation strategy.

3. The method according to claim 2, characterized in that The step of converting the data processing operations of the symmetric encryption algorithm into vector instructions through the vectorized dialect of the multi-level intermediate representation framework comprises: The byte substitution operation, row shift operation and round key addition operation of the symmetric encryption algorithm are vectorized and converted into vector instructions through the vectorization dialect of the multi-level intermediate representation framework.

4. The method according to claim 2, characterized in that: The memory management of the symmetric encryption algorithm based on the memory management optimization strategy includes: Flattening the high-dimensional data in the symmetric encryption algorithm into one-dimensional data; Storing the extended key and the T table in the static shared memory of the wafer-level chip for access by the multiple heterogeneous hardware computing units; The S box of the symmetric encryption algorithm and every four I8 entries of the T table are pre-calculated and merged into one I32 value.

5. The method according to claim 1, characterized in that The task scheduling mechanism based on the multi-level intermediate representation framework schedules the multiple data encryption subtasks in parallel to the corresponding hardware backend, including: Based on the task scheduling mechanism of the multi-level intermediate representation framework and the properties of each of the hardware backends, generating corresponding target task partitioning codes; Based on the target task division code, the multiple data encryption subtasks are scheduled in parallel to the corresponding hardware backend.

6. The method according to claim 5, characterized in that The hardware backend performs data encryption processing on the multiple groups of original data according to the extended key and the corresponding data encryption subtask to obtain multiple encrypted data, including: Based on the data length requirement of the symmetric encryption algorithm, each group of original data is divided into data blocks, and the last data block is filled with data to obtain multiple target data blocks; Through the hardware backend, data encryption processing is performed on the multiple target data blocks of each group of original data according to the extended key and the symmetric encryption algorithm; The target data blocks of each group of original data after data encryption processing are merged to obtain a plurality of encrypted data.

7. A data encryption system based on a wafer-level chip heterogeneous platform, characterized in that: The system comprises: The front-end module is used to obtain multiple groups of original data to be encrypted and the original key; generate multiple data encryption subtasks based on the multiple groups of original data, and perform key expansion on the original key to generate an expanded key; A parallel module is used to schedule the multiple data encryption subtasks in parallel to the corresponding hardware backend based on the task scheduling mechanism of the multi-level intermediate representation framework; through the hardware backend, the multiple groups of original data are encrypted according to the extended key and the corresponding data encryption subtasks to obtain multiple encrypted data, wherein the hardware backend is a heterogeneous hardware computing unit of a heterogeneous hardware platform of a wafer-level chip; the multiple encrypted data are merged to obtain target encrypted data.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the data encryption method based on the wafer-level chip heterogeneous platform is implemented as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data encryption method based on a wafer-level chip heterogeneous platform as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data encryption method based on a wafer-level chip heterogeneous platform as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Skinny algorithm optimization realization method and system, terminal and storage medium

    CN109995506A

  • Block cipher parallelization method based on embedded platform

    CN111953476A

  • Data encryption method and device, electronic equipment and storage medium

    CN116961958A

  • System-on-chip compiler test method and device, electronic equipment and storage medium

    CN117851270A

  • Encryption method based on RISC-V architecture

    CN117978367A