GPU-accelerated multi-data multi-threaded SHA-256 computation implementation method

CN116088939BActive Publication Date: 2026-09-25北京天数智芯半导体科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211238644.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-11
Publication Date
2026-09-25
Estimated Expiration
2042-10-11

AI Technical Summary

Technical Problem

薛子豪虽然使用CUDA平台将SHA-256算法从CPU上移植到GPU上运行,并进行了一些优化,但其对算法的并行化设计能力不足,无法对同时计算多个数据时的情况进行优化,且没有对CPU与GPU之间的任务划分与通信方法进行优化

Benefits of technology

可以采用多线程的方法处理多个数据的SHA-256摘要的计算任务,并且根据算法的过程设计了并行方案;同时,本发明利用GPU多级存储的结构,将一些常用数据存放于快速存取结构当中,加快了算法的运算速度;此外,本发明设计了CPU的同步调度方案,利用CPU-GPU异构系统不同硬件设备各自的优势实现了合理计算任务分配,降低了计算时间开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116088939B_ABST
    Figure CN116088939B_ABST
Patent Text Reader

Abstract

The application discloses a GPU-accelerated multi-data multi-thread SHA-256 calculation implementation method, comprising the following steps: constructing a parallel model of SHA-256 algorithm message expansion; constructing a parallel model of SHA-256 algorithm cyclic iteration; optimizing a data storage model and a data flow model; and constructing a complete CPU-GPU heterogeneous structure task and data flow. The application adopts a multi-thread method to process SHA-256 digest calculation tasks of multiple data, utilizes a multi-level storage structure of GPU to store some commonly used data in a fast access structure, accelerates the operation speed of the algorithm, utilizes the respective advantages of different hardware devices of a CPU-GPU heterogeneous system to realize reasonable calculation task distribution, and reduces the calculation time cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of encryption technology, specifically relating to a GPU-accelerated multi-data multi-threaded SHA-256 calculation implementation method. Background Technology

[0002] Secure hash algorithms, also known as hash algorithms (SHA), can calculate a fixed-length string corresponding to a digital message; this string is called a message digest. The SHA algorithm is a type of hash algorithm. For hash functions, mapping two different keys to the same value is called a collision, and no hash algorithm can avoid collisions. For messages with different content, because the SHA algorithm performs complex iterative calculations based on the message itself, the probability of generating the same digest is extremely low. This perfectly meets the security requirements of blockchain. Since each transaction in a blockchain is unique, no two pieces of data will be identical, so the digest generated by the SHA algorithm can be used to determine whether the message content has been tampered with. In addition to effectively protecting the security of block data, using the SHA algorithm in blockchain also results in a fixed-length digest, making the block header format neat and easy to parse.

[0003] In blockchain, SHA-256 is currently used to calculate the hash value of the previous block and the hash value of the current block in the block header. Developed by the U.S. National Security Agency and released by the National Institute of Standards and Technology (NIST) in 2001, SHA-256 remains a reliable cryptographic algorithm. The Proof-of-Work (PoW) consensus algorithm is the earliest and most secure public blockchain consensus algorithm. In PoW, each node in the system competes with each other based on its computing power to solve a complex but easily verifiable SHA-256 mathematical problem.

[0004] GPUs (Graphics Processing Units) originated from users' demands for high-quality graphics. By freeing the CPU from graphics rendering, they improved image quality and freed up CPU computing power. GPU cores are much lighter than CPU cores. They constrain the computation of large amounts of data of the same type that can be parallelized and are independent of each other, greatly reducing the chip resource consumption of control circuits and caches. This allows for a far greater number of cores than CPUs with the same chip resources and power, achieving high-performance parallel computing capabilities far exceeding those of traditional CPUs.

[0005] Currently, the main methods for executing the SHA-256 algorithm using CPU-GPU heterogeneous architectures or other hardware platforms (such as FPGAs and ASICs) are as follows:

[0006] CPU-GPU Acceleration Solutions: Cui Yan analyzed and improved the SHA-1 secure hash algorithm, and finally implemented acceleration of the SHA-1 algorithm based on the CUDA platform. Xue Zihao analyzed the evolution of NVIDIA's GPU architecture, accelerated SHA-256 and the Keccak algorithm in the new SHA-3 series of secure hash algorithms based on the CUDA platform, and optimized the data stream processing. Ge Can implemented GPU acceleration of the SHA-2 algorithm on AMD GPUs based on the OpenGL platform, and designed, implemented and optimized GPU acceleration of the SHA-2 password recovery algorithm.

[0007] Specialized circuits: Koziel et al. proposed a hardware architecture to accelerate homology cryptography, one of the post-quantum cryptography candidates. Dai et al. proposed an FFT-based exponential hardware architecture to accelerate the RSA algorithm. Kerckhof et al. summarized the streamlined SHA-3 algorithm implementations based on FPGA platforms by the five finalists in the SHA-3 competition. While specific hardware implementations can offer performance improvements compared to general-purpose processors, they also have some drawbacks. For example, specific hardware implementations are not flexible enough, are difficult to be compatible with different hardware, and are difficult to upgrade and maintain. Furthermore, hardware implementations like ASICs typically require expensive manufacturing costs and specialized development and design skills, leading to longer development cycles.

[0008] Currently, solutions for accelerating SHA-256 using CPU-GPU heterogeneous architectures still have certain shortcomings and limitations. Ge Can's method of porting the SHA-256 algorithm to the GPU is only suitable for small-scale data and lacks practical value. Although Xue Zihao used the CUDA platform to port the SHA-256 algorithm from the CPU to the GPU and made some optimizations, his parallelization design capabilities are insufficient, failing to optimize for simultaneous computation of multiple datasets, and he did not optimize the task partitioning and communication methods between the CPU and GPU. Current optimization schemes have limitations for large-scale data processing. Summary of the Invention

[0009] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing a GPU-accelerated multi-data multi-threaded SHA-256 calculation method. By analyzing the parallel characteristics of SHA-256, a GPU-based multi-threaded method is constructed, which improves the computational efficiency of the algorithm. Compared with similar GPU performance modeling models, it not only ensures modeling speed but also improves modeling accuracy, which has positive significance for the exploration of GPU architecture.

[0010] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: A GPU-accelerated multi-data, multi-threaded SHA-256 computation implementation method includes the following steps: Step 1: Construct a parallel model for message extension of the SHA-256 algorithm; Step 2: Construct a parallel model for iterative SHA-256 algorithm; Step 3: Optimize the data storage model and data flow model; Step 4: Based on steps 1-3, construct a complete CPU-GPU heterogeneous architecture task and data flow.

[0011] To optimize the above technical solution, the specific measures also include: The parallel model for message extension of the SHA-256 algorithm constructed in step 1 above uses a GPU to process the entire data block. The address offset method distinguishes data blocks, and different threads are allocated to different data blocks. The W[j] array, where the value of j is in the range of 16-63, is used to implement the parallel computation mode of SHA-256 algorithm message extension.

[0012] The parallel model of the SHA-256 algorithm iterative loop constructed in step 2 above uses 8 threads to process the iterative calculation tasks of a, b, ..., h simultaneously: a, b, ..., h first obtain the value of the corresponding position in the intermediate variable array H, and then update H according to the calculation result. This loop is repeated 64 times to complete the iterative calculation part of a data block.

[0013] The optimized data storage model in step 3 above stores a, b, ..., h in shared memory. The data in shared memory is copied back to global memory after the calculation is completed. 64 constants K[i] (i = 0, 1, 2 ... 63) are stored in constant memory and used when calculating a, b, ..., h in the iterative update of the H array.

[0014] The data flow model optimized in step 3 above uses two different kernel functions to calculate the message extension part and the subsequent iterative update of the H array, respectively. In this process, the W array is stored in video memory after the first kernel function completes its calculation. The second kernel function uses the W array calculated by the first kernel function as its input, thus reusing intermediate results. Furthermore, when multiple data items are being computed to SHA-256, data transmission and message extension computation are performed asynchronously.

[0015] The complete CPU-GPU heterogeneous architecture task and data flow constructed in step 4 above specifically include: Step 4.1: The CPU initiates an instruction to transfer the constant array K in memory to the constant memory portion in video memory; Step 4.2: The CPU reads data from the disk, determines the data length, allocates memory and video memory, and creates a CUDA stream; Step 4.3: Start running the SHA-256 algorithm and record the start time of the run; Step 4.4: After the data transfer is complete, the GPU begins calculating the message extension. Step 4.5: After the CPU finishes processing all the data padding, it waits for the GPU to finish calculating the message extension part of all the data. Then, the CPU synchronously calls the GPU to calculate the loop iteration part of all the data and store it. Step 4.6: After the GPU completes its calculations, the CPU executes the next step of the code, initiates an asynchronous transfer request, and copies the result array from the video memory to the main memory. Once the data copy is complete, the algorithm finishes execution and the timer stops. Step 4.7: Release memory and video memory, destroy the previously created stream, and calculate the execution time of the SHA-256 algorithm.

[0016] Step 4.2 above specifically involves: the CPU reading data from the disk, determining the data length, allocating memory to store the data based on the data length, allocating the required video memory based on the data length, including the data portion, the W array portion, and the result array portion, and allocating the corresponding result array portion in memory to provide space to store the results calculated by the GPU; finally, creating the corresponding CUDA stream for each piece of data to be processed, and simultaneously determining the total number of data items to be processed.

[0017] The SHA-256 algorithm described in step 4.3 above includes: the CPU first preprocesses the data and performs padding; Then, using the stream that was created for the data in the previous step, the processed data is asynchronously transmitted to the GPU, and the GPU begins asynchronous computation; after the CPU calls the GPU to perform asynchronous operations, it begins to process the padding and filling of the next data.

[0018] The specific process of step 4.4 above is as follows: the GPU directly reads data from the data part of the video memory, calculates it, and stores it in the W array part of the video memory; as the CPU continuously processes the padding of multiple data and completes the data transmission, there will be multiple streams waiting for the GPU to process. When the hardware is not fully loaded, the multiple streams of multiple data are calculated by the GPU at the same time.

[0019] In step 4.5 above, the GPU uses the thread ID to find the address of the corresponding W array and the result array in the video memory. Then each thread starts executing the same code, accessing constant memory, reading the K array, and using shared memory to store some intermediate variables.

[0020] The present invention has the following beneficial effects: The invention employs a multi-threaded approach to process the SHA-256 digest computation task of multiple data sets, and a parallel scheme is designed based on the algorithm's process. Furthermore, the invention utilizes the GPU's multi-level storage structure to store frequently used data in a fast access structure, accelerating the algorithm's computation speed. In addition, the invention designs a CPU synchronous scheduling scheme, leveraging the respective advantages of different hardware devices in the CPU-GPU heterogeneous system to achieve reasonable allocation of computational tasks and reduce computational time overhead. Attached Figure Description

[0021] Figure 1 Here is a pseudocode diagram of the SHA-256 algorithm; Figure 2 Extend the parallelization graph for messages; Figure 3 Parallelization of the loop iteration graph; Figure 4 Flowchart for GPU execution. Detailed Implementation

[0022] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0023] like Figure 1 The image shown is a pseudocode diagram of the SHA-256 algorithm. Figure 1 It fully demonstrates the serial execution process of the SHA-256 algorithm, describing the entire process of dividing a piece of data (such as transaction information or blocks on the blockchain, or web pages of web services) into blocks and calculating a 256-bit digest.

[0024] First, after padding, a piece of data is split into data blocks of 512 bits in length. Each data block is then divided into 16 smaller units of 32 bits in length. In the pseudocode, M represents the padded data and N represents the number of data blocks.

[0025] Secondly, the algorithm enters the message expansion phase, where each data block uses an intermediate parameter group W, which has a total of 64 parameters.

[0026] Then, the algorithm enters the iterative calculation stage. This part first uses the intermediate variable array H obtained from the previous data block to initialize a, b, ..., h, and then enters the 64-round iterative process, writing the result into the intermediate variable array H.

[0027] Finally, the computation of each data block, except for the first data block, must depend on the previous data block, and the intermediate variable array H of the last data block is the digest value.

[0028] A GPU-accelerated multi-data, multi-threaded SHA-256 computation implementation method includes the following steps: Step 1: Construct a parallel model for SHA-256 algorithm message expansion: Analyze the case of performing addition or multiplication calculations on a whole block of data. Use the GPU to distinguish data blocks by address offset and allocate different threads to different data blocks. Based on the data dependency of the W array, the data block adopts a dual-thread parallel computing mode for the W[j] array where the value of j ranges from 16 to 63.

[0029] like Figure 2 The diagram shows the message expansion parallelization. This section assigns different data blocks to different view threads for processing.

[0030] Since there is no data dependency between data blocks when performing message expansion, threads that perform computation tasks on different data blocks can be carried out in parallel; however, for the intermediate parameter group W, since there is a data association between W[j] and W[j-2], each data block only uses two threads to perform message expansion in parallel.

[0031] Step 2: Construct a parallel model for iterative SHA-256 algorithm; For the loop iteration part, consider using 8 threads to calculate the corresponding values ​​of a, b, ..., h. The subsequent process of updating H and assigning it to a, b, ..., h can also be run on the GPU, reducing data transfer between the CPU and GPU.

[0032] like Figure 3 The diagram illustrates the parallelization of the iterative process. This part utilizes eight threads to simultaneously process the iterative computation tasks of a, b, ..., h. Since the operations performed in the iterative computation are similar, a single instruction multiple data (SIMD) approach can be used for parallel processing. a, b, ..., h first obtain the values ​​at the corresponding positions in the intermediate variable array H, and then update the stored content according to the computation results. This process is repeated 64 times to complete the iterative computation of one data block.

[0033] Step 3: Optimize the data storage model and data flow model; Optimize the data storage model, specifically: Based on the characteristics of shared memory, the data that needs to be copied to shared memory is that which is frequently accessed within a thread block and whose size does not exceed its limit (usually 48KB). Therefore, a, b, ..., h need to be placed in shared memory.

[0034] Data in shared memory must be copied back to global memory after calculation is completed; otherwise, the data cannot be transferred back to main memory.

[0035] Because the constant memory cannot modify data and can only store constants, it is only used when calculating a, b, ..., h in the iterative update of the H array. There are 64 constants K[i] (i = 0, 1, 2 ... 63) that can be placed in the constant memory.

[0036] The data flow model is optimized as follows: In this algorithm, two different kernel functions are needed to calculate the message expansion part and the subsequent iterative update of the H array, respectively. The second kernel function requires the W array calculated by the first kernel function as input. Therefore, an intermediate result reuse method can be adopted: after the first kernel function completes its calculation, the W array is retained in GPU memory for subsequent calculations. Furthermore, when multiple data points require SHA-256 calculation, a long-stream segmentation approach can be adopted, using a series of asynchronous operations to calculate multiple data points. Since each data point requires the CPU to perform padding operations before transferring the data to GPU memory for computation, the data transmission and message expansion calculation can be made asynchronous. For the GPU, the message expansion calculation process of one data point and the data transmission process of the next data point occur simultaneously, masking some data communication latency.

[0037] like Figure 4 The diagram shown is a flowchart of GPU operation. Figure 4 This demonstrates how data is processed when multiple data sets are computed to SHA-256 digests simultaneously.

[0038] In a multi-data, single-stream, multi-threaded parallel architecture, the code running on the GPU mainly consists of two kernel functions: kernel1, which calculates the message extension, and kernel2, which iteratively updates the H array. The number of threads required by kernel1 depends on the data size. If the data size is n times 512 bits, then 2n threads are used to calculate the message extension, with each thread calculating 32 numbers in the W array. The number of threads required by kernel2 depends on the amount of data. If m data points need to be calculated, then 8m threads are used, with each thread assigning a number from a, b, c...h to a position in the H array.

[0039] Step 4: Based on steps 1-3, construct a complete CPU-GPU heterogeneous task and data flow. The specific process is as follows: Step 4.1: The CPU initiates an instruction to transfer the constant array K in memory to the constant memory portion in video memory; Step 4.2: Prepare for the calculation by allocating the required memory and video memory.

[0040] The CPU reads data from the disk, determines the data length, allocates an appropriate amount of memory to store the data, and then allocates the required video memory based on the data length. The video memory to be allocated includes the data part, the W array part, and the result array part. In addition, the corresponding result array part needs to be allocated in memory to provide space to store the results calculated by the GPU.

[0041] Finally, create the corresponding CUDA stream for each piece of data to be processed, and determine the total number of data items to be processed. Step 4.3: This step begins the execution of the SHA-256 algorithm and starts recording the start time of the execution.

[0042] The CPU first preprocesses the data and then fills in the missing parts.

[0043] Then, using the stream that was created for the data in the previous step, the processed data is asynchronously transmitted to the GPU, and the GPU begins asynchronous computation.

[0044] After the CPU calls the GPU to perform an asynchronous operation, it begins processing the padding and filling of the next data. Step 4.4: After the data transfer is complete, the GPU begins calculating the message extension. The specific process is as follows: Data is read directly from the data portion of the video memory, and after calculation, it is stored in the W array portion of the video memory.

[0045] After the CPU finishes processing and padding multiple data and completing the data transmission, multiple streams will be waiting for the GPU to process. When the hardware is not at full load, multiple streams of multiple data can be calculated simultaneously. Step 4.5: After the CPU finishes processing all the data padding, it waits for the GPU to finish calculating the message extension part of all the data. Then, the CPU synchronously calls the GPU to calculate the loop iteration part of all the data. Since the W array result is already stored in the video memory, no more data transfer is performed. The GPU uses the thread ID to find the address of the corresponding W array and the result array in the video memory. Then each thread starts to execute the same code, accessing constant memory, reading the K array in it, and using shared memory to store some intermediate variables. Step 4.6: After the GPU completes its calculations, the CPU executes the next step of the code, initiates an asynchronous transfer request, and copies the result array from the video memory to the main memory. Once the data copy is complete, the algorithm finishes execution and the timer stops. Step 4.7: Finally, release the memory and video memory, destroy the previously created stream, and calculate the execution time of the SHA-256 algorithm.

[0046] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A GPU-accelerated multi-data, multi-threaded SHA-256 computation implementation method, characterized in that, Includes the following steps: Step 1: Construct a parallel model for SHA-256 algorithm message extension; The parallel model for SHA-256 algorithm message extension constructed in Step 1 uses the GPU to distinguish data blocks by address offset method, allocates different threads to different data blocks, and uses a dual-thread parallel computing mode to realize the parallelism of SHA-256 algorithm message extension for the W[j] array with j values ​​ranging from 16 to 63. Step 2: Construct a parallel model for iterative SHA-256 algorithm; The parallel model for iterative SHA-256 algorithm constructed in Step 2 uses 8 threads to process the iterative calculation tasks of a, b, ..., h simultaneously: a, b, ..., h first obtain the value of the corresponding position in the intermediate variable array H, and then update H according to the calculation result. This process is repeated 64 times to complete the iterative calculation part of a data block. Step 3: Optimize the data storage model and data flow model; The optimized data storage model stores a, b, ..., h in shared memory. The data in shared memory is copied back to global memory after the calculation is completed. 64 constants K[i] (i = 0, 1, 2 ... 63) are stored in constant memory and used when calculating a, b, ..., h in the iterative update of the H array. The optimized dataflow model uses two different kernel functions to compute the message extension part and the subsequent iterative update of the H array, respectively. In this process, the W array is stored in video memory after the first kernel function completes its calculation. The second kernel function uses the W array calculated by the first kernel function as its input, thus reusing intermediate results. Furthermore, when multiple data are being computed to SHA-256, data transmission and message extension computation are performed asynchronously. Step 4: Construct a complete CPU-GPU heterogeneous architecture task and data flow based on steps 1-3; the complete CPU-GPU heterogeneous architecture task and data flow constructed in step 4 specifically includes: Step 4.1: The CPU initiates an instruction to transfer the constant array K in memory to the constant memory portion in video memory; Step 4.2: The CPU reads data from the disk, determines the data length, allocates memory and video memory, and creates a CUDA stream; Step 4.3: Start running the SHA-256 algorithm and record the start time of the run; Step 4.4: After the data transfer is complete, the GPU begins calculating the message extension. Step 4.5: After the CPU finishes processing all the data padding, it waits for the GPU to finish calculating the message extension part of all the data. Then, the CPU synchronously calls the GPU to calculate the loop iteration part of all the data and store it. Step 4.6: After the GPU completes its calculations, the CPU executes the next step of the code, initiates an asynchronous transfer request, and copies the result array from the video memory to the main memory. Once the data copy is complete, the algorithm finishes execution and the timer stops. Step 4.7: Release memory and video memory, destroy the previously created stream, and calculate the execution time of the SHA-256 algorithm.

2. The GPU-accelerated multi-data multi-threaded SHA-256 calculation implementation method according to claim 1, characterized in that, Step 4.2 specifically involves the CPU reading data from the disk, determining the data length, allocating memory to store the data based on the data length, allocating the required video memory based on the data length, including the data portion, the W array portion, and the result array portion, and allocating the corresponding result array portion in memory to provide space to store the results calculated by the GPU; finally, creating the corresponding CUDA stream for each piece of data to be processed, and simultaneously determining the total number of data items to be processed.

3. The GPU-accelerated multi-data multi-threaded SHA-256 calculation implementation method according to claim 1, characterized in that, Step 4.3 describes the SHA-256 algorithm, which includes: the CPU first preprocesses the data and performs padding; then, using the stream that was created for the data in the previous step, the processed data is asynchronously transmitted to the GPU, and the GPU begins asynchronous computation; after the CPU calls the GPU to perform asynchronous operations, it begins to process the padding and filling part of the next data.

4. The GPU-accelerated multi-data multi-threaded SHA-256 calculation implementation method according to claim 1, characterized in that, The specific process of step 4.4 is as follows: the GPU directly reads data from the data part of the video memory, calculates it, and stores it in the W array part of the video memory; as the CPU continuously processes the padding of multiple data and completes the data transmission, there will be multiple streams waiting for the GPU to process. When the hardware is not fully loaded, the multiple streams of multiple data are calculated by the GPU at the same time.

5. The GPU-accelerated multi-data multi-threaded SHA-256 calculation implementation method according to claim 1, characterized in that, In step 4.5, the GPU uses the thread ID to find the address of the corresponding W array and the result array in the video memory. Then each thread starts executing the same code, accessing constant memory, reading the K array, and using shared memory to store some intermediate variables.

Citation Information

Patent Citations

  • Parallel and collaborative optimization method for big data processing based on CPU multithreading and GPU multi-granularity

    CN106991011A

  • fluid machinery simulation program heterogeneous acceleration method based on a GPU

    CN109522127A