Parallel Computing Method for Large-Scale Elliptic Curve Multi-Scalar Multiplication Accelerated by GPU

By splitting the large-scale elliptic curve multi-scalar multiplication task and storing it in parallel in the sparse matrix, using the GPU for parallel sparse matrix operations, the problem of difficulty in accelerating this type of operation in the GPU high-parallelism environment in the prior art is solved, and efficient parallel accelerated computing is achieved.

CN115481364BActive Publication Date: 2025-06-10ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211138724.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-06-10
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

Existing parallel computing methods are difficult to effectively accelerate large-scale elliptic curve multi-scalar multiplication in GPU high parallelism environments, resulting in low computational efficiency.

Method used

By splitting the original operation tasks into sub-operation tasks and storing these sub-operation tasks in a sparse matrix in parallel, the GPU performs parallel sparse matrix operations, including matrix transposition and matrix vector multiplication, and finally reduces the sub-operation task results to obtain the original task results.

Benefits of technology

This method significantly improves the calculation speed of large-scale elliptic curve multiplication, has lower time complexity, and can achieve parallel acceleration ratio close to the number of threads provided by the GPU, saving computing overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115481364B_ABST
    Figure CN115481364B_ABST
Patent Text Reader

Abstract

The present invention discloses a parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration, comprising the following steps: splitting the original operation task of large-scale elliptic curve multi-scalar multiplication to obtain multiple sub-operation tasks; storing the curve points of the sub-operation tasks in parallel; performing a series of parallel sparse matrix operations; performing weighted summation on the curve point elements in the sparse matrix; and reducing the results of the sub-operation tasks to obtain the final result of large-scale elliptic curve multi-scalar multiplication. This computing method splits the original large-scale elliptic curve multi-scalar multiplication task into sub-tasks, and simplifies various complex operations in the computing process into parallel sparse matrix operations, providing acceleration for large-scale elliptic curve multi-scalar multiplication. The time complexity of this computing method is less than that of the existing methods, and theoretically can achieve parallel acceleration close to the number of threads provided by the GPU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of parallel computing, and particularly relates to a parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration. Background Art

[0002] Large-scale elliptic curve multi-scalar multiplication refers to the multiplication operation of points on an elliptic curve (curve points) with a large number of scalars. This multiplication operation can be represented by the formula where n represents the scale of the operation, k i represents a scalar, P i represents a curve point, k i P i represents the multiplication of the scalar k i and the curve point P i , and Q represents the result of the operation. Large-scale elliptic curve multi-scalar multiplication has been widely used in the field of cryptography, especially in the field of zero-knowledge proof, where it has become one of the most important operators.

[0003] However, completing large-scale elliptic curve multi-scalar multiplication is a very time-consuming operation. For example, in the real-world scenario of zero-knowledge proof, the scale n of this multiplication operation is very large and can reach millions. Even worse, using the traditional operation method to perform a multiplication operation between a scalar and a curve point on an elliptic curve takes far more time than completing a multiplication between two scalars. This results in the time taken to complete a large-scale elliptic curve multi-scalar multiplication operation being unacceptable to users and the computing efficiency being low. Currently, parallel computing for this multiplication operation is one of the core technologies to improve its operation efficiency.

[0004] GPU (Graphic Processing Unit) is a computing platform that can provide thousands of cores for parallel computing. Currently, multiple fields have used GPU to perform parallel acceleration on basic algorithms, such as parallel acceleration of matrix operations in the field of deep learning, acceleration of graphics collision algorithms in the field of graphics, and parallel acceleration of signal processing in the field of communication, etc. However, there is currently no parallel computing method for large-scale elliptic curve multi-scalar multiplication suitable for the GPU environment.

[0005] There are already various parallel computing methods for large-scale elliptic curve multi-scalar multiplication, such as the parallel computing method based on the Chang-Lou algorithm and the parallel computing method based on the Bos-Coster algorithm. However, these parallel computing methods only perform well in low-parallelism environments and are difficult to apply to the high-parallelism environment provided by GPUs. In high-parallelism environments, these parallel computing methods show that their parallel acceleration ratios are lower than the parallelism that GPUs can provide, that is, these parallel computing methods are not applicable on GPUs.

[0006] Therefore, there is an urgent need for a parallel computing method for large-scale elliptic curve multi-scalar multiplication that is suitable for the high-parallelism environment of GPUs, which is a key technology to accelerate this scalar multiplication operation in order to achieve parallel accelerated computing of large-scale elliptic curve multi-scalar multiplication on GPUs. Summary of the Invention

[0007] In view of the above, the object of the present invention is to provide a parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration to improve the operation speed and save computing overhead.

[0008] To achieve the above object of the invention, an embodiment provides a parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration, including the following steps:

[0009] (1) Split the original operation task: Split the original operation task of large-scale elliptic curve multi-scalar multiplication into multiple sub-operation tasks according to the scalars;

[0010] (2) Parallelly store the curve points of the sub-operation tasks: Parallelly store all the curve points of the multiple sub-operation tasks into a sparse matrix according to the sub-scalars corresponding to the curve points;

[0011] (3) Parallel sparse matrix operations: Perform matrix transpose and matrix-vector multiplication operations on the sparse matrix storing the curve points to obtain a curve point vector;

[0012] (4) Parallelly calculate the results of the sub-operation tasks: Perform weighted addition on the curve points in the curve point vector to obtain the results of the sub-operation tasks;

[0013] (5) Calculate the result of the original operation task: Reduce all the results of the sub-operation tasks to obtain the result of the original task.

[0014] Preferably, in step (1), the original operation task is: where n represents the scale of the operation, k i represents the i-th scalar, P i represents the i-th curve point, k i P i represents the scalar k iMultiplication with the curve point P i , and Q represents the result of the operation;

[0015] The scalar k corresponding to the λ bit i is segmented into λ / s sub-scalars m of s bits ij , and multiple sub-scalars m ij satisfy the formula

[0016] In this way, the original calculation task is segmented into λ / s sub-operation tasks where G j is the operation result of the sub-operation task, and j represents the index of the sub-operation task.

[0017] Preferably, in step (2), a sparse matrix of size t×2 s –1 is created, where t represents the total number of GPU threads that the GPU can provide, and s represents the number of bits of the sub-scalar;

[0018] Each GPU thread is evenly assigned n / t curve point tasks. According to the value of the sub-scalar m corresponding to the curve point ij , the curve points are placed at the m ij column positions of each row of the sparse matrix, realizing the storage of all curve points of the sub-operation task in the sparse matrix.

[0019] Preferably, it further includes: according to the value of the sub-scalar m corresponding to the curve point ij , the index of the curve point in the original curve point vector is placed at the m ij column positions of each row of the sparse matrix, realizing the storage of all curve points of the sub-operation task in the sparse matrix, where the curve points in the original operation task form the original curve point vector.

[0020] Preferably, in step (3), first, a parallel matrix transpose operation is performed on the sparse matrix storing the curve points;

[0021] Then, the sparse matrix after the matrix transpose operation and the vector of length n composed of unit scalars are multiplied in parallel once by a vector and a sparse matrix. The result of this multiplication is a curve point vector of length 2 s –1, denoted as [B 1 , B 2 , …, B 2 s -1 , where B 1 , B 2 , …, B 2 s -1 are all curve points.

[0022] Preferably, when performing matrix-vector multiplication operations, different numbers of GPU threads are dynamically scheduled to process different matrix rows to overcome the load imbalance among threads, including:

[0023] First, sort the rows of the sparse matrix according to the row length, and divide all rows into different groups according to the row length; then, schedule different numbers of threads for each group according to the proportion of non-zero matrix entries in the group, so that the number of non-zero entries worked on by each thread is similar, where the row length is the number of non-zero matrix entries in each row.

[0024] Preferably, in step (4), first, evenly distribute the calculation tasks of (2 s -1) / t curve points to each GPU thread, and for the i-th curve point B assigned i find its weighted value with respect to the subscript, that is, iB i ;

[0025] Then, reduce all the weighted values to calculate the result of the sub-operation task

[0026] Preferably, in step (5), reduce all the results of the sub-operation tasks to the original task result where G j is the result of the j-th sub-operation task.

[0027] Compared with the prior art, the beneficial effects of the present invention at least include:

[0028] By splitting the original operation task into sub-operation tasks and simplifying various complex operations in the calculation process into parallel sparse matrix operations, acceleration is provided for large-scale elliptic curve multi-scalar multiplication. The time complexity of this method is less than that of the prior method, and in theory, it can achieve a parallel acceleration ratio close to the number of threads provided by the GPU, improving the operation speed and saving the calculation overhead. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0030] Figure 1 is a flowchart of the parallel computing method provided by the embodiment;

[0031] Figure 2 is a schematic diagram of the curve points for parallel storage of sub-operation tasks provided by the embodiment;

[0032] Figure 3 It is a schematic flow chart of obtaining the results of sub - operation tasks provided by the embodiment. Specific embodiments

[0033] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0034] To save the computational overhead of large - scale elliptic curve multi - scalar multiplication and accelerate the saving of the computational cost of large - scale elliptic curve multi - scalar multiplication, the embodiment provides a parallel computing method for large - scale elliptic curve multi - scalar multiplication based on GPU acceleration. By reducing the computational process of multi - scalar multiplication into various sparse matrix operations to accelerate the operation process. And use GPU acceleration for parallel sparse matrix operations to provide an acceleration scheme for the reduced sparse matrix operations.

[0035] As Figure 1 and as Figure 3 shown, the parallel computing method for large - scale elliptic curve multi - scalar multiplication provided by the embodiment includes the following steps:

[0036] Step 1, splitting the original operation task: splitting the original operation task of large - scale elliptic curve multi - scalar multiplication into multiple sub - operation tasks according to the scalars.

[0037] In the embodiment, the original operation task is: where n represents the scale of the operation, k i represents the i - th scalar, P i represents the i - th curve point, k i P i represents the scalar multiplication of scalar k i and curve point P i , and Q represents the result of the operation. This large - scale elliptic curve multi - scalar multiplication algorithm is mainly used in the high - parallel environment provided by the GPU. The system implementation is very suitable for the job environment of the GPU, adopting a low - storage execution method, which is convenient for changing the implementation details according to the actual production environment and has strong flexibility.

[0038] In the embodiment, the original operation task of large - scale elliptic curve multi - scalar multiplication is applied in the field of cryptography, that is, in public - key encryption, the encryption process is realized by calculating large - scale elliptic curve multi - scalar multiplication.

[0039] In the original operation task, for the scalar k corresponding to λ bits i is split into λ / s sub - scalars m of s bits ij , and multiple sub - scalars m ij satisfy the formula Among them, s can be arbitrarily selected.

[0040] In this way, the original computing task is split into λ / s sub-computing tasks Among them, G j is the operation result of the sub-computing task, and j represents the index of the sub-computing task.

[0041] Step 2, parallelly store the curve points of the sub-computing tasks: All the curve points of multiple sub-computing tasks are parallelly stored into a sparse matrix according to the sub-scalars corresponding to the curve points.

[0042] In the embodiment, an ELL-format sparse matrix is created on the GPU global memory. The number of rows of the sparse matrix is the total number of threads t that the GPU can provide, and the number of columns of the sparse matrix is 2 s -1, where s can be freely selected, and the sparse matrix in the GPU global memory can be accessed by each GPU thread.

[0043] Each GPU thread is evenly assigned n / t curve point tasks. The curve point task is to place the curve points at the m ij column positions of each row of the sparse matrix according to the value of the sub-scalar m corresponding to the curve points, so as to realize the storage of all the curve points of the sub-computing tasks in the sparse matrix. ij

[0044] In the embodiment, each thread should store the curve points in the sparse matrix in the Figure 2 way. In one implementation manner, what is actually stored in the sparse matrix is not the curve points themselves, but also the indices in the original curve point vector. That is, according to the value of the sub-scalar m ij corresponding to the curve points, the indices of the curve points in the original curve point vector are placed at the m ij column positions of each row of the sparse matrix. This helps to save the device storage cost. Each curve point usually has hundreds of bits, while the size of the curve point index is the logarithm of the vector scale and only has dozens of bits.

[0045] When the sparse matrix stores the curve point indices, when the GPU obtains the curve points, the system will obtain the corresponding curve points from the host memory to the GPU memory according to the curve point indices stored in the sparse matrix. The simple method of moving the curve points from the host memory to the device memory requires a lot of time in data transmission, but its latency overhead can be almost eliminated by the overlap of CPU-GPU data transmission and device computing based on the multi-stream technology.

[0046] Step 3, parallel sparse matrix operations: Perform matrix transpose and matrix-vector multiplication operations on the sparse matrix storing the curve points to obtain a curve point vector.

[0047] ​In an embodiment, for matrix device operations, a parallel matrix transpose operation is performed on a sparse matrix storing curve points. There are already many mature solutions for the parallel sparse matrix transpose method based on GPU. In the present invention, a parallel sparse matrix transpose solution based on GPU provided by the cuSP library is deployed. This solution meets the industry's top standards in both time complexity and space complexity.

[0048] In an embodiment, for matrix-vector multiplication operations, the sparse matrix after the matrix transpose operation is multiplied in parallel with a vector of length n composed of unit scalars once. The result of this multiplication is a curve point vector of length 2 s –1, denoted as [B 1 , B 2 , …, B 2 s -1 , where B 1 , B 2 , …, B 2 s -1 are all curve points.

[0049] In an embodiment, different numbers of GPU threads are dynamically scheduled to process different matrix rows to overcome the load imbalance between threads, including:

[0050] First, the rows of the sparse matrix are sorted according to the row length, and all rows are divided into different groups according to the row length; then, different numbers of threads are scheduled for each group according to the proportion of non-zero matrix entries in each group, so that the number of non-zero entries worked on by each thread is similar, where the row length is the number of non-zero matrix entries in each row.

[0051] For example, assuming that there are a total of 1 million non-zero entries in the sparse matrix and 10,000 GPU threads are provided for matrix-vector multiplication operations, so that each GPU thread processes 100 non-zero entries. When grouping according to the row length, assuming that there are 50 rows with a row length of 20 and they are assigned to 1 group, then 10 GPU threads need to be called to perform matrix-vector multiplication operations for this group.

[0052] Step 4, parallelly calculate the result of the sub-operation task: The curve points in the curve point vector are weighted and added to obtain the result of the sub-operation task.

[0053] In an embodiment, for each sub-operation task, first, each GPU thread is evenly assigned (2 s -1) / t curve point calculation tasks. For the i-th curve point B i assigned, its weighted value for the subscript is calculated, that is, iB i ; then, all weighted values are reduced to calculate the result of the sub-operation task

[0054] Step 5, calculate the result of the original operation task: reduce all the results of the sub-operation tasks to obtain the result of the original task.

[0055] In the embodiment, after Steps 2 - 4, the results Gj of each sub-operation task are obtained. 1 , Gj 2 , …, Gj λ / s . On this basis, reduce all the results of the sub-operation tasks to obtain the final result of large-scale elliptic curve multi-scalar multiplication, that is, the result of the original task. where Gj j is the result of the j-th sub-operation task.

[0056] In the embodiment, the parallel computing method for large-scale elliptic curve multi-scalar multiplication provided in the above embodiment can also be applied to multiple GPUs. In this case, the parallel computing method for large-scale elliptic curve multi-scalar multiplication decomposes the original operation task into multiple sub-operation tasks. Here, these sub-operation tasks are evenly distributed to each GPU. Specifically, each GPU allocates its global memory for a sparse matrix and completes its corresponding sub-operation tasks in the manner of Steps 2 - 4 above. The requirement for GPU-GPU data transmission here is very small because after each sub-operation task is completed, finally only the results of the sub-operation tasks need to be reduced to the final result. Note that the result of each sub-operation task is a curve point, which has only hundreds of bits. Therefore, compared with the single-GPU implementation, the multi-GPU implementation of this solution does not introduce a large amount of additional overhead.

[0057] The above specific implementation manners have described in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modification, supplement, equivalent replacement, etc. made within the scope of the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration, characterized in that, it includes the following steps: (1) Split the original operation task: Split the original operation task of large-scale elliptic curve multi-scalar multiplication into multiple sub-operation tasks according to the scalars, including: The original operation task is: where n represents the scale of the operation, k i represents the i-th scalar, P i represents the i-th curve point, k i P i represents the scalar k i and the scalar multiplication of the curve point P i , Q represents the result of the operation; the scalar k corresponding to λ bits i is split into λ / s sub-scalars m of s bits ij , and multiple sub-scalars m ij satisfy the formula In this way, the original calculation task is split into λ / s sub-operation tasks where G j is the operation result of the sub-operation task, and j represents the index of the sub-operation task; (2) Curve points of parallel storage sub - operation tasks: All curve points of multiple sub - operation tasks are stored in a sparse matrix in parallel according to the sub - scalars corresponding to the curve points, creating a sparse matrix with a size of t×2 s –1, where t represents the total number of GPU threads that the GPU can provide, and s represents the number of bits of the sub - scalar; Each GPU thread is evenly assigned n / t curve - point tasks. According to the value of the sub - scalar m ij corresponding to the curve point, the curve point is placed at the m ij -th column position of each row of the sparse matrix, realizing the storage of all curve points of the sub - operation tasks in the sparse matrix; (3) Parallel Sparse Matrix Operations: Perform matrix transpose and matrix-vector multiplication operations on the sparse matrix storing curve points to obtain a curve point vector, including: First, perform a parallel matrix transpose operation on the sparse matrix storing curve points; then, multiply the sparse matrix after the matrix transpose operation with a vector of length n composed of unit scalars in parallel. The result of this multiplication is a curve point vector of length 2 s – 1, denoted as where are all curve points; when performing the matrix-vector multiplication operation, overcome the load imbalance between threads by dynamically scheduling different numbers of GPU threads to process different matrix rows, including: First, sort the rows of the sparse matrix according to the row length and divide all rows into different groups according to the row length; then, schedule different numbers of threads for each group according to the proportion of non-zero matrix entries in the group, so that the number of non-zero entries worked on by each thread is similar, where the row length is the number of non-zero matrix entries in each row; (4) Parallelly calculate the results of sub-operation tasks: perform weighted addition on the curve points in the curve point vector to obtain the results of sub-operation tasks; (5) Calculate the results of the original operation task: reduce all the results of sub-operation tasks to obtain the results of the original task.

2. The parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration according to claim 1, characterized in that, it further includes: According to the value of the sub-scalar m corresponding to the curve point ij Place the index of the curve point in the original curve point vector at the m ij column position of each row of the sparse matrix, so as to store all the curve points of the sub-operation task in the sparse matrix. Among them, the curve points in the original operation task form the original curve point vector.

3. The parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration according to claim 1, characterized in that, In step (4), first, evenly distribute the calculation tasks of \((2 s - 1) / t\) curve points to each GPU thread. For the \(i\)-th curve point \(B\) assigned i , calculate its weighted value with respect to the subscript, that is, \(iB i ;\) Then, reduce all the weighted values and calculate the results of the sub-operation tasks 4. The parallel computing method for large-scale elliptic curve multi-scalar multiplication based on GPU acceleration according to claim 3, characterized in that, In step (5), all the sub - operation task results are reduced to the original task result where G j is the result of the j - th sub - operation task.

Citation Information

Patent Citations

  • Method and system for parallel processing of elliptic curve scalar multiplication

    CN102446088A

  • Accelerator for sparse-dense matrix multiplication

    CN110321525A