Multi-scalar multiplication MSM calculation method and device, program product and storage medium

By dynamically allocating the number of GPU threads based on the number of points to be calculated in each bucket, the problems of low MSM computing efficiency and low GPU resource utilization are solved, and more efficient GPU resource utilization and computing performance are achieved.

CN120074823APending Publication Date: 2025-05-30ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510202718.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, multi-scalar multiplication (MSM) has low computational efficiency on graphics processors (GPUs), resulting in low GPU resource utilization, especially when the number of points to be calculated is uneven.

Method used

The number of allocated GPU threads is determined dynamically based on the number of points to be calculated in each bucket. If there are fewer points to be calculated in the bucket, a single thread is allocated for calculation; if the number of points to be calculated exceeds the preset threshold, multiple threads are allocated for calculation.

Benefits of technology

It improves the resource utilization rate of GPU, avoids waste of thread resources and synchronization overhead, and adapts to the uneven number of points to be calculated in different buckets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074823A_ABST
    Figure CN120074823A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a multi-scalar multiplication MSM calculation method and device, a program product and a storage medium. The method is applied to a graphics processing unit (GPU) for executing MSM calculation. The method comprises the following steps: acquiring a bucket dividing result determined for a current MSM to be calculated; wherein the bucket dividing result represents a bucket to which each point to be calculated in the MSM to be calculated belongs; determining the number of to-be-calculated points in each bucket based on the bucket dividing result; calculating each of the buckets; wherein the buckets, the number of which is greater than or equal to the preset threshold value, are calculated by adopting a plurality of GPU threads, and the buckets, the number of which is smaller than the preset threshold value, are calculated by adopting a single GPU thread.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification belong to the technical field of cryptography, and particularly relate to a calculation method, device, program product, and storage medium for multi-scalar multiplication (MSM). Background Art

[0002] Elliptic Curve Cryptography (ECC) is a public-key cryptography method based on the mathematical theory of elliptic curves. ECC is an alternative to traditional public-key cryptography (such as RSA) and has advantages such as short keys, small computational amount, and small storage space occupation. Under the same security strength, the key length of elliptic curve cryptography is much shorter than that of most other public-key cryptographies.

[0003] The emergence of elliptic curve cryptography mainly aims to solve the following problems:

[0004] · Key length and security strength: With the development of computer technology, the key length required by traditional public-key cryptography has been increasing to ensure sufficient security strength. However, the increase in key length has also led to an increase in computational amount and storage space. Elliptic curve cryptography can achieve a security strength comparable to that of traditional public-key cryptography with a shorter key length, thus improving efficiency.

[0005] · Limitations of mobile devices: Mobile devices such as smartphones and tablets have relatively limited computing power and storage space. The application of traditional public-key cryptography on these devices is restricted. Due to its high efficiency, elliptic curve cryptography is particularly suitable for use on mobile devices.

[0006] · Threat of quantum computing: The development of quantum computers poses a potential threat to traditional public-key cryptography. In theory, quantum computers can efficiently solve the mathematical problems relied on by cryptographies such as RSA using the Shor algorithm. Currently, elliptic curve cryptography is considered a promising solution to resist attacks from quantum computers.

[0007] Elliptic curve cryptography uses the algebraic properties of points on elliptic curves defined over finite fields to construct a security system. An elliptic curve can be described by the following equation:

[0008] y 2 =x 3 +ax + b, where 4a 3 + 27b 2 ≠ 0 (1)

[0009] For example, Figure 1 lists several different shapes of elliptic curves, where b = 1 and a decreases from 2 to -3.

[0010] From Figure 1It can be seen that the characteristic of an elliptic curve is that it is symmetric about the x-axis. In this way, for a point T on one side of the x-axis on the elliptic curve, there exists a point -T that is symmetric to it about the x-axis.

[0011] The following introduces some basics of group theory.

[0012] A group in algebra refers to a set of elements and an operation defined on these elements. For example, all integers form a group, and the operation includes addition. Here, the set is usually denoted by G (group). For a set to become a group, it generally needs to satisfy the following properties:

[0013] 1. Closure: If both a and b belong to G, then a + b also belongs to G;

[0014] 2. Associativity: For any a, b, c belonging to G, (a + b) + c = a + (b + c);

[0015] 3. There exists an identity element O such that for any a ∈ G, a + O = O + a = a. For example, in real numbers, the identity element of addition is 0. Simply put, in a binary operation, the identity element can operate with any element without changing the value of that element. Taking real numbers as an example, the identity element of multiplication is 1, and the identity element of addition is 0;

[0016] 4. Each element has an inverse element (also called an inverse or negative element, or simply an inverse), that is, for any element a, there must exist b such that a + b = O (O is the identity element).

[0017] A set that satisfies the above four properties is called a group.

[0018] There are also some special groups, such as the Abelian Group, which in addition to satisfying the basic properties of a group, also satisfies the commutativity law, that is: a + b = b + a. Therefore, the Abelian Group is also called a commutative group.

[0019] Based on these properties, it can be determined that the set of integers Z is an Abelian group, while the set of natural numbers N is not a group because the set of natural numbers N does not satisfy the fourth property of having an inverse element.

[0020] The above is the basic knowledge of group theory. Although it seems simple, there are many complex characteristics in group theory, which are omitted here for the time being. It should be noted that the elements in a group can not only be numbers, but also other types of elements, such as points in analytic geometry (represented in coordinate form), etc.

[0021] With the basic knowledge of groups, it is possible to further define a group on an elliptic curve in a similar way.

[0022] As mentioned above, the elements of a group can be of any type. Here, the group elements on an elliptic curve are set as the points on the elliptic curve. The identity element is set as the point at infinity, denoted as O. The inverse of any point P on the elliptic curve is the point symmetric to P about the x-axis.

[0023] The addition of an elliptic curve group is also different from the addition of integers. The addition of an elliptic curve group can be defined as: given three non-zero collinear points P, Q, R, their sum is P + Q + R = O. As Figure 2 shown, the geometric meaning of this addition is: draw a straight line L through points P and Q, which intersects the elliptic curve at a third point, and the point symmetric to this point about the X-axis is the required point R.

[0024] Next, handle some exceptional cases:

[0025] 1. O + O = O, that is, for any P, P + O = P; O is regarded as the zero point;

[0026] 2. The negative element of P = (x, y) is the point -P = (x, -y) symmetric about the x-axis rather than about the origin. P + (-P) = O, which can be regarded as the line connecting point P and point -P intersecting the elliptic curve at the point at infinity;

[0027] 3. When calculating 2 times of point P (P ≠ O), draw the tangent line of this point on the elliptic curve, and then take the point symmetric to the intersection point R of the tangent line and the elliptic curve about the X-axis, that is, 2P = P + P = -R. Its geometric meaning is as Figure 3 shown, point -R is point 2P. Furthermore, from the calculation of 2 times value, it can be recursively extended to the calculation of 2 n times value, and it can also be recursively obtained the calculation of any integer multiple value.

[0028] The geometric interpretation is convenient for understanding the meaning of point addition on an elliptic curve, while the algebraic interpretation is more conducive to operation. The points on the curve have coordinates in the two-dimensional plane, that is, the coordinates of the x and y axes. Then, the problem of finding the third intersection point of the curve by drawing a straight line through two points P(x p , y p ) and Q(x q , y q ) that are not negative elements of each other can be described algebraically, that is, to find the solution of the simultaneous equations of the following two equations (2) and (3):

[0029] y 2 = x 3 + ax + b (2)

[0030] y - y p = k(x - x p ) (3)

[0031]

[0032] where k is the slope.

[0033] Substituting equation (3) into equation (2) and using the method of aligning the exponents, it is easy to obtain the symmetric point of the third intersection point about the x-axis, that is, the point R(x r , y r ) representing the sum of point P and point Q is:

[0034] x r = k 2 - x p - x q (4)

[0035] y r = -y p + k(x p - x r ) (5)

[0036] If P = Q, then adding them together is to calculate the multiple. Multiple results can be obtained by repeatedly adding a point to itself. For example, P + P = 2P = R, and the algebraic description is:

[0037]

[0038] So far, the basic operations of elliptic curve group elements in the real number field have been introduced. The addition operation in the real number field cannot meet the needs of actual security because real numbers are continuous, and the inverse operation can be used to solve the problem once the result is known. Therefore, to prevent reversibility, the next step is to discretize the coordinates of the elliptic curve points.

[0039] Here, the operation modulo a prime number p is introduced first. First, the modulo operation represents the remainder obtained when a is divided by p, which can be expressed as:

[0040] a mod p.

[0041] In cryptography, elliptic curves over a finite field are used, that is, the coefficients and variable values of the elliptic equation are all within a finite range. The range of values can be limited to the finite field Z p by using the modulo prime number p. Specifically, introducing the modulo operation into elliptic curve arithmetic means that the variables and coefficients take values from the set [0, p - 1] instead of from real numbers, and it is a finite field (i.e., a finite field). The equation over this field is transformed as follows:

[0042] y 2 mod p = (x 3 + ax + b) mod p, where (4a 3 + 27b 2 ) mod p ≠ 0 (8)

[0043] The above formula (8) is also often expressed as:

[0044] y 2 = (x 3 + ax + b) mod p。

[0045] All integer solutions (x, y) that satisfy the above equation, where x, y ∈ Zp, and the point at infinity O, are denoted by the mathematical symbol E p (a, b), which is a finite discrete (non - continuous) set of points.

[0046] If there exists a point g in this finite field such that every element in the group G can be written in the form of some integer power of the point g (here the power is a generalized concept in the group, which describes the number of repetitions of the group operation; in the additive group, it refers to the number of addition operations), that is:

[0047] G = {g 0 = O, g 1 = g, g 2 = 2g,..., g n-1 = (n - 1)·g},

[0048] where n is the order of the group G (the order refers to the number of elements in the group), then this group is also called a cyclic group, denoted as E p (a, b), and this point g is called a generator. Due to the effect of taking the modulus, the points in this set are distributed in the quadrant from (0, 0) to (p - 1, p - 1). Even if the value of k changes, the obtained results also cycle within this quadrant. In fact, the set E p (a, b) and the addition operation modulo p form a cyclic Abelian group, and this cyclic Abelian group is a subset in the quadrant from (0, 0) to (p - 1, p - 1). Therefore, this cyclic Abelian group is also called a cyclic subgroup.

[0049] Using the prime modulus p ensures consistent and reliable operations. All non - zero elements have inverses, forming a good algebraic structure (finite field). Therefore, it is suitable for elliptic curve cryptography. If a non - prime modulus is used, there will be problems of inconsistent operations and non - existent inverses, resulting in a complex algebraic structure (finite ring). Therefore, it is not suitable for elliptic curve cryptography. In short, elliptic curve cryptography selects the prime modulus p rather than a non - prime number to ensure the reliability and security of mathematical operations. In addition, the order n of the selected cyclic subgroup is usually also chosen to be a prime number to obtain better security and efficiency.

[0050] E p (a, b) has an addition rule similar to that in the aforementioned real number field, that is, similar to the aforementioned formulas (2) - (5). The difference is that there is an additional modulo operation. Similarly, by adding the modulo operation in formulas (6) and (7), the multiplication result obtained by repeatedly adding multiple points to themselves in the discrete - domain cyclic group can be obtained.

[0051] Briefly speaking, for a point P on an elliptic curve, the point Q which is k times of P satisfies the relation:

[0052] Q = k·P, where Q, P ∈ E p (a, b), k < p (9)

[0053] For given k and P, it is easy to calculate Q; but conversely, given Q and P, it is quite difficult to calculate k. This is the intractable problem of discrete logarithm on elliptic curves. Constructing a mathematical problem to ensure the security of encryption is the main idea of encryption algorithms in cryptography. Therefore, Q can be used as the public key and made public; k can be used as the private key and kept secret. It is fast and simple to obtain the public key from the private key, but it is very difficult to crack the private key from the public key.

[0054] In the above addition operation of the group, scalar multiplication is actually defined, that is, adding the same point multiple times. For example, adding k points p can be expressed as:

[0055] k·P = P + P + … + P (10)

[0056] The simplest way to calculate scalar multiplication is to add points one by one. It should be noted that in fact, the points on the elliptic curve have coordinates, and the calculation of the above formula (10) involves coordinate operations similar to those in formulas (6) and (7), including a large number of modular multiplications and modular additions. Therefore, the actual calculation is more complex.

[0057] Although in some optimized algorithms, the number of addition operations can be reduced. For example, 7P can be decomposed into 4P + 2P + P. 2P requires one addition operation, and 4P = 2P + 2P requires one more addition operation based on 2P. In this way, a total of 4 addition operations are required, which is 2 less than the original 6 addition operations for 7P = P + P + P + P + P + P + P (P + P = 2P, 2P + P = 3P, 3P + P = 4P, 4P + P = 5P, 5P + P = 6P, 6P + P = 7P). However, from the perspective of underlying coordinate operations, there are still a large number of modular multiplications and modular additions.

[0058] Generally speaking, addition on the elliptic curve requires a large number of modular multiplications and modular additions. Calculating one addition may require dozens of modular multiplications and modular additions; multiplication on the elliptic curve (i.e., the above scalar multiplication) requires an even larger number of modular multiplications and modular additions. Calculating one multiplication may require hundreds or more modular multiplications and modular additions.

[0059] The above scalar multiplication is a basic operation in elliptic curve cryptography and is used in processes such as key generation, encryption, and signature. Multi-Scalar Multiplication (MSM) is a generalization of scalar multiplication that involves multiple points and multiple scalars. Specifically, given n points P 1 , P 2 , …, P n and n integers k 1 , k 2 , …, k n , the goal of multi-scalar multiplication is to compute:

[0060] Q = k 1 P 1 + k 2 P 2 +... + k n P n (11)

[0061] It can be seen that multi-scalar multiplication is the sum of the results of n scalar multiplications.

[0062] Multi-scalar multiplication is applied in many elliptic curve-based cryptographic protocols, especially in those scenarios that require simultaneous processing of multiple points and scalars. For example:

[0063] · In some digital signature schemes, the generation and verification of signatures require computing linear combinations of multiple points.

[0064] · In some key exchange protocols, the computation of the shared key requires scalar multiplication and summation of multiple public key points with corresponding private key points.

[0065] · In some Zero-Knowledge Proof (ZKP) protocols, the generation and verification of proofs require computing linear combinations of multiple commitment values.

[0066] Taking zero-knowledge proof as an example, illustrate the role of the MSM operation in zero-knowledge proof. First, the protocol principle of ZKP is as Figure 4 shown.

[0067] There are a prover and a verifier. The prover uses a proving key, a private input (also known as a witness), and a public input to create a proof through a program. The prover sends this proof Proof to the verifier. The verifier can use a verification key and a public input, and verify the proof Proof through a verifier, outputting true or false.

[0068] During the above process, the prover does not disclose the private input that it does not want to disclose. The verifier can verify through the proof that the prover knows the private input, but the verifier cannot know what the private input is.

[0069] This property makes zero-knowledge proofs very useful in many scenarios, especially when the prover needs to prove that he has a certain secret but does not want to disclose any information about this secret. For example:

[0070] · In authentication, the prover can prove that he knows a secret associated with the identity (such as a private key) without disclosing this secret.

[0071] · In privacy-preserving transactions, the prover can prove that the transaction is valid (such as proving that he has sufficient balance) without disclosing the details of the transaction (such as the specific amount).

[0072] · In privacy-preserving computations, the prover can prove that he has correctly executed a certain computation (such as proving that he has correctly evaluated a certain function) without disclosing the input of the computation.

[0073] The premise of the above ZKP is the existence of a pair of proving key and verification key. Usually, the prover key and verification key can be generated in a centralized manner, and this generation process is called the trusted setup process. The trusted setup process can be implemented by a trusted third party other than the prover and the verifier to generate the proving key and the verification key, and send the proving key to the prover and the verification key to the verifier respectively. The trusted setup process implemented by the trusted third party can be divided into two categories: general trusted setup and dedicated trusted setup process. The difference lies in:

[0074] For general trusted setup, the generation of the proving key and the verification key does not necessarily need to be based on a specific computational problem. This setup process is a one-time operation, so it can save the overhead of repeated setup for each computational problem.

[0075] A dedicated trusted setup is a setup based on a specific computational problem. The generated proof key and verification key are generated for a specific computational problem. This means that these keys can only be used to prove and verify this specific problem and cannot be directly applied to other problems. On the other hand, these keys can only be used to prove and verify this specific problem and cannot be directly applied to other problems. In this way, even if others obtain these keys, they can only forge proofs for this specific problem and cannot affect other problems. At the same time, the efficiency of the dedicated trusted setup is low, and the setup process needs to be implemented separately for each computational problem.

[0076] Taking the general trusted setup as an example, such as Figure 4 As shown, for example, by selecting an algorithm (such as one of Pinocchio, Groth16, Sonic, etc.) in the ZK-Snark (Zero-Knowledge Succinct Non-interractive Argument of Knowledge) protocol, a trusted third party can generate a proof key and a verification key through the general trusted setup. In addition, some public parameters, reference strings, constraint system templates, proofs of randomness, security proofs, etc. are also involved (step ①). Furthermore, the trusted third party can send the generated proof key and some public parameters, reference strings, constraint system templates, proofs of randomness, security proofs, etc. to the prover (step ②), and send the verification key to the verifier (step ②). The prover can create a proof based on the public input and private input of the computational problem and in combination with the proof key (step ③), and send this proof to the verifier (step ④). The verifier verifies the truth or falsehood of the proof based on the verification key, the proof, and the public input through a verification program (also called a verifier), that is, verifies whether the prover really knows the solution (step ⑤).

[0077] In the process of the prover generating a proof, generally speaking, it can include three parts of work. The first part is a process called arithmetization, which realizes the conversion of the computational problem into a vector of a set of linear constraints; the second part is to convert the vector of these linear constraints into a "polynomial proof"; the third part is to convert the polynomial proof into a "non-interactive proof", which is the final proof, and this process is also called "proof compilation".

[0078] Taking ZK-Snark as an example, the above process includes:

[0079] Step 1: Arithmetization. Arithmetization is the process of converting an original computational problem into a set of algebraic constraints. The core tasks of this process include:

[0080] · Represent the computational problem as an arithmetic circuit or a set of polynomial constraints. This typically involves defining the inputs, outputs, and intermediate variables of the circuit, as well as describing the arithmetic relationships between these variables.

[0081] · Convert the constraints of the circuit into a special linear algebraic structure, usually called R1CS (Rank-1 Constraint System). This involves representing each constraint as a matrix and combining all matrices into a large R1CS instance.

[0082] · Convert the R1CS instance into a set of linear equations or polynomials. This is typically achieved using techniques such as Lagrange interpolation.

[0083] · The goal of arithmetization is to transform the original problem into an algebraic form that is more suitable for cryptographic processing. This lays the foundation for subsequent polynomial proof generation and proof compilation.

[0084] Step 2: Polynomial Proof Generation. In this step, the prover converts the vector of linear constraints obtained from arithmetization into a "polynomial proof". The core tasks of this process include:

[0085] · Randomize the linear constraints using so-called "blinding factors". This is to hide the specific values of the original constraints, thus achieving zero-knowledge.

[0086] · Interpolate the randomized constraints into polynomials.

[0087] · Operate on these polynomials to generate different parts of the proof, such as commitments, proofs, responses, etc. This involves using cryptographic primitives such as homomorphic commitments, verifiable random functions, etc. The operations on polynomials are typically implemented using efficient algorithms such as the Fast Fourier Transform (FFT) or the Number-Theoretic Transform (NTT).

[0088] · The goal of polynomial proof generation is to transform the discrete linear constraints into a continuous polynomial form, which is more suitable for subsequent cryptographic processing. This process also introduces randomization and blinding, which are the keys to achieving zero-knowledge.

[0089] Step 3: Proof Compilation. In the final step, the prover converts the polynomial proof into a "non-interactive proof", which is the final zk-SNARK proof. The core tasks in this process include:

[0090] · Using zk-SNARK protocols such as Pinocchio, Groth16, etc., to convert the polynomial proof into a non-interactive and verifiable computational integrity proof. This involves using advanced cryptographic primitives such as pairing-friendly elliptic curves and pairing operations.

[0091] · Conducting various optimizations, such as using multi-scalar multiplication (MSM) to accelerate calculations and using recursive techniques to compress the proof size, etc. This is to improve the efficiency of proof generation and verification.

[0092] · Packing the optimized proof together with the public input of the original problem to form the final zk-SNARK proof.

[0093] It should also be noted that in the above process, contents such as public parameters, reference strings, constraint system templates, proofs of randomness, proofs of security, etc. (collectively referred to as artifacts together with the proof key and verification key) are also very important. The trusted third party needs to distribute appropriate artifacts to the prover and the verifier. As mentioned above, the third party sends the proof key and some artifacts (such as public parameters, reference strings, constraint system templates, proofs of randomness, and proofs of security) to the prover, while the verification key is sent to the verifier. The following are the functions of these artifacts, namely public parameters, reference strings, constraint system templates, proofs of randomness, and proofs of security:

[0094] · Public parameters: These parameters define the cryptographic primitives on which zk-SNARK is based, such as elliptic curves, pairing-friendly curves, etc. The verifier needs these parameters to execute the verification algorithm, especially when checking certain cryptographic conditions (such as pairing equalities).

[0095] · Reference string: The reference string is a shared input between the prover and the verifier, used to coordinate their interaction. The verifier needs to use the same reference string as the prover to verify the proof.

[0096] · Constraint system template: This template defines the general structure of the computational problems that zk-SNARK can handle. The verifier may need this template to check whether the structure of the public input is compatible with the proof.

[0097] · Proofs of randomness and proofs of security: These proofs provide evidence of the integrity of the setup process and the security of zk-SNARK. The verifier may need these proofs to gain trust in the entire system.

[0098] The dedicated trusted setup is generally similar to the above process. However, as mentioned before, the dedicated trusted setup process needs to be designed in combination with computational problems. Therefore, in the dedicated trusted setup process of a third party, computational problems will be combined to generate a proof key and a verification key. This also means that the generated proof key and verification key will be tightly bound to a specific circuit structure. And during this process, the computational problems can be preprocessed. For example, arithmetization can be set in the preprocessing process because arithmetization can be performed without involving private inputs. Thus, during the process of the prover generating a proof, the second and third steps mentioned above can be executed, and the second and third steps involve private inputs.

[0099] During the proof compilation process of the above third step, it can be said that MSM is everywhere. From generating commitments for polynomial proofs, to checking the correctness of these commitments, to combining these commitments to form the final non-interactive proof, every step involves a large number of MSM operations.

[0100] In fact, the efficiency of MSM is a key factor in the practical usability of zk-SNARK. Since the proof generation and verification of zk-SNARK involve so many MSM operations, the performance of MSM directly affects the speed and scalability of the entire system. This is why many implementations of zk-SNARK invest a lot of effort in optimizing MSM, using various techniques such as precomputation and batch processing to accelerate the calculation of MSM.

[0101] In the process of zk-SNARK proof generation and verification, multi-scalar multiplication (MSM) is a computationally intensive core operation. The performance of MSM has an important impact on the overall efficiency of the zk-SNARK system. To accelerate the calculation of MSM, the industry usually adopts dedicated hardware accelerators, such as acceleration schemes based on FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), or GPU (Graphics Processing Unit). Compared with using a general-purpose CPU, specially designed hardware accelerators can achieve significant advantages in terms of performance and energy efficiency, thus greatly reducing the generation time and energy consumption of zero-knowledge proofs. Therefore, designing efficient MSM hardware accelerators has become a research hotspot in the current field of zk-SNARK hardware optimization.

[0102] Among them, among various dedicated hardware accelerators, GPUs have relatively low costs, are easy to iterate, and are also the most widely used. Currently, there are many GPU accelerations for MSM, including many papers and acceleration libraries. However, if these works are applied to the actual ZKP proof generation process, good acceleration effects are often not achieved. The reason is that the Pippenger algorithm is usually used in existing works to implement the calculation of MSM, and the Pippenger algorithm will distribute the points to be calculated on the elliptic curve to several buckets.

[0103] For example, Figure 5 , which is a schematic diagram of implementing MSM calculation using the Pippenger algorithm according to an exemplary embodiment of this specification; wherein k i The i-th scalar among the multiple scalars representing the MSM, P i represents the i-th point to be calculated among multiple points to be calculated on the elliptic curve of MSM. Q is the final calculation result. 1 Indicates the calculation result of the first scalar and the first point to be calculated, Q 2 Represents the calculation result of the second scalar and the second point to be calculated, and so on. The calculation process of the Pippenger algorithm can include:

[0104] Split the scalar in MSM into windows. Figure 5 As shown in the figure, each scalar k i The binary representation of Figure 5 4 bits in the example) is divided into three windows (i.e. Figure 5 G 0 , G 1 and G 2 ).

[0105] Each window corresponds to a set of buckets. The number of buckets is usually determined based on the number of window bits. The algorithm assigns points with the same value in the same window to the same bucket. Subsequently, only one point addition is required for the points in each bucket, rather than multiple point multiplications.

[0106] like Figure 5 G shown in 0 The 15 barrels under (B 0 To B 15 ). The Pippenger algorithm will also allocate each point to be calculated into buckets according to the strategy. The specific allocation method also depends on the input MSM. Figure 5 In the example shown, G 0 B 1 , not assigned to the point to be calculated; and B 5 There are 4 points to be calculated, B 14 There is one point to be calculated. Each calculation point under each bucket needs to be calculated point by point; for each window, the calculation results of all buckets are accumulated to get G j ,like Figure 5 G shown in 0 , G 1 and G 2 Finally, G j The final result Q is obtained by accumulation. The accumulation of points in a bucket is completely independent and can be processed in parallel. The calculation of different windows can also be parallel.

[0107] The related technology implements MSM calculation based on GPU. Therefore, the main idea is to set a fixed number of GPU threads for each Bucket. Since this method does not require the intervention of the CPU, fixed allocation does not require dynamic decision-making logic, which can reduce the CPU-GPU interaction and synchronization overhead. Moreover, early GPU programming models (such as CUDA) have insufficient support for dynamic parallelism, making it difficult to adjust the number of threads in real time within the GPU kernel.

[0108] Specifically, in the traditional solution, most of the time, each Bucket is assigned to a GPU thread for processing. However, in actual scenarios, the number of points to be calculated assigned to each Bucket is often uneven. The vast majority of Buckets have relatively few points, and the assigned threads can quickly complete the processing. There are a small number of Buckets with an abnormally large number of calculation points, which will cause other threads to complete their tasks early while only one or two threads are still processing, resulting in very low GPU utilization.

[0109] In some other related technologies, instead of assigning one thread to each Bucket, a fixed number N of threads is assigned, where N is a multiple of 32 (because the minimum scheduling unit warp of the GPU consists of 32 threads). In this way, if a Bucket is very large, the computing efficiency of N threads must be far better than that of one thread. At the same time, for Buckets with a small number of points, the thread resources released after the Bucket is processed will be allocated by the GPU to other Buckets. However, this solution also has disadvantages. First, it introduces the overhead of GPU scheduling. Multiple threads need to perform a parallel reduction operation when processing a Bucket. Compared with a single thread, parallel reduction will introduce synchronization overhead and cause waste of some thread resources. The larger N is and the fewer points to be calculated in the bucket, the more obvious the introduced cost will be. Summary of the Invention

[0110] To overcome the problems existing in the related technology, this specification provides a calculation method, device, program product, and storage medium for multi-scalar multiplication MSM.

[0111] According to the first aspect of the embodiments of this specification, a calculation method for multi-scalar multiplication MSM is provided. The method is applied to a graphics processing unit GPU that executes the MSM calculation. The method includes:

[0112] Obtain the bucketing result determined for the current MSM to be calculated; wherein, the bucketing result represents the bucket to which each point to be calculated in the MSM to be calculated belongs;

[0113] Based on the bucketing result, determine the number of points to be calculated in each bucket;

[0114] Calculate for each of the said buckets; among them, the buckets with the quantity greater than or equal to the preset threshold are calculated using multiple GPU threads, and the buckets with the quantity less than the preset threshold are calculated using a single GPU thread.

[0115] According to the second aspect of the embodiments of the present specification, a calculation method for a multi-scalar multiplication (MSM) is provided. The method is applied to a central processing unit (CPU), and the method includes:

[0116] Read the bucket data copied by a graphics processing unit (GPU) to the memory space; the GPU is used to execute the steps of the method in the first aspect; the bucket data refers to the bucket data of the n buckets with the highest number of points to be calculated determined from the bucket result after the GPU determines the bucket result based on the current MSM to be calculated.

[0117] After determining whether the number of points to be calculated in each of the n buckets is greater than or equal to the preset threshold, determine the thread allocation result of allocating multiple GPU threads to the buckets with the quantity greater than or equal to the preset threshold and allocating a single GPU thread to the buckets with the quantity less than the preset threshold, and transmit the thread allocation result to the GPU.

[0118] According to the third aspect of the embodiments of the present specification, a graphics processing unit is provided. The graphics processing unit is used to execute the steps of the method in the first aspect.

[0119] According to the fourth aspect of the embodiments of the present specification, a computer device is provided. The computer device includes a graphics processing unit and a central processing unit. The graphics processing unit is used to execute the steps of the method in the first aspect, and the central processing unit is used to execute the steps of the method in the second aspect.

[0120] According to the fifth aspect of the embodiments of the present specification, a computer program product is provided, including a computer program. When the computer program is executed by a graphics processing unit, it implements the steps of the method in the first aspect; and / or when the computer program is executed by a central processing unit, it implements the steps of the method in the second aspect.

[0121] According to the sixth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a graphics processing unit, it implements the steps of the method in the first aspect; and / or when the computer program is executed by a central processing unit, it implements the steps of the method in the second aspect.

[0122] In the calculation method of the multi-scalar multiplication MSM provided in the above embodiments, the number of threads to be allocated can be dynamically determined according to the number of points to be calculated in the bucket at runtime. If the number of points to be calculated in the bucket is small, one thread is allocated for calculation to avoid synchronization overhead and waste of thread resources. If the number of points to be calculated in the bucket exceeds a certain limit, several threads are allocated according to the bucket size for processing. Therefore, the method of this embodiment can solve the adverse effect of low GPU resource utilization caused by a fixed number of threads. For the uneven situation of the points to be calculated in each bucket, this embodiment makes full use of GPU resources while reducing other additional overheads. BRIEF DESCRIPTION OF THE DRAWINGS

[0123] To more clearly illustrate the technical solutions of the embodiments of this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0124] Figure 1 is a schematic diagram of an elliptic curve;

[0125] Figure 2 is a schematic diagram of defining addition on an elliptic curve;

[0126] Figure 3 is a schematic diagram of defining point doubling on an elliptic curve;

[0127] Figure 4 is a flowchart of the zero-knowledge proof principle in an embodiment;

[0128] Figure 5 is a schematic diagram of implementing MSM calculation by the Pippenger algorithm shown in this specification according to an exemplary embodiment;

[0129] Figure 6 is a flowchart of a calculation method of a multi-scalar multiplication MSM shown in this specification according to an exemplary embodiment;

[0130] Figure 7 is a flowchart of another calculation method of a multi-scalar multiplication MSM shown in this specification according to an exemplary embodiment;

[0131] Figure 8 is a schematic diagram of the structure of a computer device shown in this specification according to an exemplary embodiment;

[0132] Figure 9A is a structural diagram of a calculation device of a multi-scalar multiplication MSM shown in this specification according to an exemplary embodiment;

[0133] Figure 9B This is a structural diagram of another computing device for multi-scalar multiplication (MSM) shown in this specification according to an exemplary embodiment. Detailed implementation manners

[0134] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0135] Based on the foregoing background, the embodiments of this specification provide a calculation method for multi-scalar multiplication (MSM), which can dynamically determine how many threads to allocate at runtime according to the number of points to be calculated in a bucket. If the number of points to be calculated in the bucket is small, one thread is allocated for calculation to avoid synchronization overhead and waste of thread resources. If the number of points to be calculated in the bucket exceeds a certain limit, several threads are allocated according to the bucket size for processing. Therefore, the method in this embodiment can solve the situation where the points to be calculated in each bucket are uneven, and make full use of GPU resources while reducing other additional overheads.

[0136] As Figure 6 shown, this is a flowchart of a calculation method for multi-scalar multiplication (MSM) shown in this specification according to an exemplary embodiment. The method can be applied to a graphics processing unit (GPU) that executes the MSM calculation. The method may include the following steps:

[0137] In step 602, obtain the bucketing result determined for the current MSM to be calculated. Wherein, the bucketing result represents the bucket to which each point to be calculated in the MSM to be calculated belongs.

[0138] In step 604, determine the number of points to be calculated in each bucket based on the bucketing result.

[0139] In step 606, calculate each of the buckets.

[0140] Among them, the buckets with the number greater than or equal to a preset threshold are calculated using multiple GPU threads, and the buckets with the number less than the preset threshold are calculated using a single GPU thread.

[0141] As an example, the GPU of this embodiment can be set on a computer device (host). The computer device may also include a central processing unit CPU, etc. When needed by the host, the host can use the GPU of this embodiment to perform MSM calculations. For example, the GPU of this embodiment can be used in various tasks that require MSM calculations, such as zero-knowledge proof tasks.

[0142] Figure 6 The illustrated embodiment is described with the GPU as the execution entity. In the GPU that performs MSM calculations in this embodiment, the calculation of MSM can be implemented based on the Pippenger algorithm or other algorithms. For example, the host can obtain each scalar and each point to be calculated (i.e., points on the elliptic curve) in the MSM to be calculated, and the CPU can transmit each scalar and each point to be calculated in the MSM calculation to the GPU. The MSM calculation program running on the GPU can perform MSM calculations on the input scalars and points to be calculated.

[0143] As an example, the Pippenger algorithm can be used to implement MSM calculations in the MSM calculation program running on the GPU, as described above Figure 5 As shown, the Pippenger algorithm can perform a bucketing operation to obtain a bucketing result, and the bucketing result represents the bucket to which each point to be calculated in the MSM to be calculated belongs.

[0144] As an example, assume that the points to be calculated involved in a certain MSM calculation include: a total of 10 points to be calculated from P0 to P9; taking the division of two windows (W0 and W1) as an example, there are 16 Buckets (B0 to B15) under each window; the bucketing result obtained by the Pippenger algorithm executed in the GPU can be:

[0145] "(P0, W0, B3), (P0, W1, B5), ……(P6, W0, B3), ……".

[0146] Among them, each parenthesis contains three pieces of information: the information of the point to be calculated, the window information, and the bucket information. Therefore, the bucketing result represents the bucket to which each point to be calculated in the MSM to be calculated belongs. Based on this, in step 604, the number of points to be calculated in each bucket can be determined based on the bucketing result. In practical applications, various methods can be used to determine the number of points to be calculated in each bucket from the bucketing result; for example, the above bucketing result can be traversed, and the points to be calculated in each bucket can be counted, etc. Other efficient methods can also be designed according to needs. This embodiment does not limit this.

[0147] Further, after obtaining the number of points to be calculated in each bucket, this embodiment can reasonably control the number of threads based on the number of points to be calculated in each bucket. For example, in step 206, this embodiment designs a preset threshold to divide large buckets and small buckets. For large buckets with a large number, a sufficient number of multiple GPU threads can be allocated for parallel processing. For small buckets with a number less than the preset threshold, only a single thread is still allocated for processing. Therefore, the solution of this embodiment avoids the situation where one thread is busy and other threads are idle, effectively improving the utilization rate of the GPU, preventing resource waste, and also making the utilization rate of the GPU as close to the maximum value as possible.

[0148] In order to efficiently obtain the number of points to be calculated in each bucket, this embodiment of the specification also provides an efficient acquisition method. As an example, each piece of bucketing data in the bucketing result may include: the point to be calculated, the window index of the window where the bucket to which the point to be calculated belongs, and the bucket index of the bucket to which the point to be calculated belongs. For example, still taking the foregoing embodiment as an example, for the bucketing result:

[0149] “(P0,W0,B3),(P0,W1,B5),……(P6,W0,B3),……”, “(P0,W0,B3)” can be understood as a piece of bucketing data. Among them, the bucketing result is keyed by the point to be calculated.

[0150] In order to efficiently obtain the number of points to be calculated in each bucket in the GPU, determining the number of points to be calculated in each bucket based on the bucketing result may include:

[0151] After splicing the window index and the bucket index in each piece of bucketing data to obtain a splicing index, obtain the splicing index sequence of each piece of bucketing data;

[0152] Sort each splicing index in the splicing index sequence to obtain a sorted splicing index sequence;

[0153] Use a run-length encoding function to count the number of occurrences of each splicing index in the sorted splicing index sequence, and the number of occurrences represents the number of points to be calculated in each bucket.

[0154] In this embodiment, a run-length encoding function is used to determine the number of points to be calculated in each bucket. Run-length encoding is a lossless data compression algorithm, commonly used in image compression or text compression. It records consecutive repeated data as (value, number of repetitions). For example, for the input sequence [A,A,A,B,B,C], it can be encoded by the run-length function as: [A,3),(B,2),

[0155] (C,1)].

[0156] Based on this, in this embodiment, an input sequence that can be encoded by the run - length function needs to be constructed. Since it is necessary to determine the number of points to be calculated in each bucket, a unique identifier can be designed for each bucket, that is, the concatenated index obtained by concatenating the window index and the bucket index is used as the unique identifier of the bucket.

[0157] For example, assume the following bucketing results:

[0158] (P0,W0,B3),(P0,W1,B5),

[0159] (P1,W0,B0),(P1,W1,B8),

[0160] (P2,W0,B3),(P2,W1,B2), ...

[0162] (P9,W0,B15),(P9,W1,B7);

[0163] The concatenated index of each piece of bucketing data is:

[0164] W0B3,W1B5,W0B0,W1B8,W0B3,W1B2,...,W0B15,W1B7.

[0165] Sort the concatenated indexes, and the sorted concatenated index sequence can be:

[0166] W0B0(P1),W0B3(P0),W0B3(P2),...,W1B2(P2),W1B5(P0),W1B7(P9),W1B8(P1).

[0167] In the above example, each concatenated index in the concatenated index sequence carries information about each point to be calculated; for example, information such as the subscript of the point to be calculated in the input scalar array. In practical applications, it is also possible not to carry the points to be calculated. In the concatenated index sequence, since the concatenated indexes are sorted, if there are multiple points to be calculated in a bucket, this concatenated index will appear repeatedly; for example, "W0B3(P0),W0B3(P2)" above.

[0168] Furthermore, this concatenated index sequence can be input into the run - length encoding function, and the encoding result can be:

[0169] "(W0B0,1),(W0B3,2),(W1B2,1),...,"

[0170] Among them, the encoding result contains two pieces of information (concatenation index, number of occurrences); among them, the number of occurrences is the number of times the concatenation index appears in the sequence. For example, "(W0B0,1)" means: the concatenation index W0B0, and the number of times it appears in the sequence is 1.

[0171] Therefore, through the run-length encoding function, the number of occurrences of each concatenation index can be counted, and the number of occurrences characterizes the number of points to be calculated in each bucket.

[0172] In the above embodiments, for the convenience of illustration, strings are used to show the information of the points to be calculated, the window index, and the bucket index. In practical applications, other methods can be used to represent these three pieces of information, such as binary numbers, etc. This embodiment does not limit this.

[0173] For example, the information of the points to be calculated can be the labels of all the points to be calculated in the MSM to be calculated; for example, all the points to be calculated in the MSM to be calculated can be numbered in a continuous and ascending order, for example, starting from 0, so that the serial number of each point to be calculated can be obtained.

[0174] Similarly, for the window and the bucket, they are numbered in a continuous and ascending order, for example, starting from 0. Therefore, the serial number of the index can be used as the window index, and the serial number of the bucket can be used as the bucket index.

[0175] Optionally, in order to reduce the development difficulty and make efficient use of the GPU, some of the above steps can be implemented using the programming tool library of the GPU.

[0176] For example, the sorting of each concatenation index in the concatenation index sequence may include: sorting each concatenation index in the concatenation index sequence through the sorting function provided by the programming tool library of the GPU. Among them, the sorting function can be sorting functions such as sort_pairs.

[0177] For example, the run-length encoding function can also be a function provided by the programming tool library of the GPU, for example, it can be the run_length_encode function.

[0178] Based on this, when implementing this embodiment, the functions provided by the programming tool library of the GPU can be called, and no additional code needs to be written, which can reduce the development difficulty.

[0179] After obtaining the number of points to be calculated in each bucket determined based on the bucketing results, buckets with a quantity greater than or equal to a preset threshold can be calculated using multiple threads, and buckets with a quantity less than the preset threshold can be calculated using a single thread. In practical applications, for some high-version GPUs, the processing flow here can be implemented within the GPU, that is, the logical judgment of the quantity is performed within the GPU, further determining the buckets with a quantity greater than or equal to the preset threshold, and determining the number of threads required for these buckets and scheduling the corresponding number of threads; and determining the buckets with a quantity less than the preset threshold, and scheduling a single thread for these buckets. In some other examples, for some low-version GPUs, it is impossible to determine whether the number of points to be calculated in a bucket is greater than or equal to the preset threshold, nor can it allocate the number of threads for the bucket by itself. The CPU needs to initiate specific tasks and control the task scale to it. Therefore, it can be implemented by the CPU, and then based on the control instructions of the CPU, the GPU calculates the buckets with a quantity greater than or equal to the preset threshold using multiple threads, and calculates the buckets with a quantity less than the preset threshold using a single thread.

[0180] Therefore, in order to be compatible with different versions of GPUs, as an example, the computer device where the GPU is located includes a central processing unit CPU; the method may further include:

[0181] Obtain the bucket data of the n buckets with the highest quantity; where n is a positive integer; for the other buckets except the n buckets with the highest quantity, it is determined to use a single GPU thread for calculation;

[0182] Copy the bucket data of the n buckets with the highest quantity into the memory space of the CPU, and obtain the thread allocation result of the CPU for the n buckets with the highest quantity;

[0183] Wherein, after the CPU reads the bucket data of the n buckets from the memory space, it determines whether the number of points to be calculated in each of the n buckets is greater than or equal to the preset threshold, and then determines the thread allocation result of allocating multiple GPU threads to the buckets with a quantity greater than or equal to the preset threshold and allocating a single GPU thread to the buckets with a quantity less than the preset threshold, and transmits the thread allocation result to the GPU.

[0184] In this embodiment, the CPU can determine the buckets with a quantity greater than or equal to a preset threshold and allocate appropriate threads. The relevant information representing the quantity of points to be calculated for each bucket is stored in the video memory of the GPU. Therefore, it is necessary to copy this relevant information into the memory space of the CPU so that the CPU can read this data and execute the subsequent processing flow. Copying data from the GPU video memory to the CPU memory space incurs a certain overhead, and the larger the data volume, the greater the overhead. In the MSM calculation scenario, many buckets are usually divided. To reduce the overhead, this embodiment designs an idea of not copying all bucket data, but copying the bucket data of the n buckets with the highest quantity among all buckets. In this way, the impact brought by data copying can be controlled.

[0185] For the other buckets except the n buckets with the highest quantity, a single thread is defaultly used for calculation. After the CPU reads the bucket data of the n buckets from the memory space, it can determine whether the quantity of points to be calculated for each of the n buckets is greater than or equal to the preset threshold, allocate multiple threads of the GPU for the buckets with a quantity greater than or equal to the preset threshold for calculation, and allocate a single thread of the GPU for the buckets with a quantity less than the preset threshold for calculation.

[0186] In practical applications, the value of n can be custom-set according to needs. For example, an empirical value can be set based on the actual MSM calculation scenario, and as much as possible, the set n can make the quantity of points to be calculated for the n buckets with the highest quantity greater than the preset threshold.

[0187] In practical applications, there can be various implementation methods for obtaining the bucket data of the n buckets with the highest quantity. For example, methods such as traversing the quantity of each bucket. In some other examples, to improve the processing efficiency, the obtaining of the bucket data of the n buckets with the highest quantity in this embodiment may include:

[0188] Using the reverse sorting function provided by the programming tool library of the GPU, sorting the respective splicing indexes in reverse order according to the occurrence times, and obtaining the bucket data of the top n buckets with the highest quantity based on the reverse sorting result.

[0189] As an example, the reverse sorting function provided by the programming tool library of the GPU can be the sort_pair_descending function, and this function can be called with the data to be sorted in reverse order; for example, taking the data with the structure of "(splicing index, occurrence times)" as an example, such as "(W0B0,1),(W0B3,2),(W1B2,1),...," as an example, sorting the respective splicing indexes in reverse order according to the occurrence times, and the obtained reverse sorting result can be:

[0190] (W1B4,6)...,(W1B2,1)

[0191] It can be seen that the number of points to be calculated in the splicing index "W1B4" is 6, ranking first.

[0192] Based on this, this embodiment can call the reverse sorting function provided by the GPU programming tool library, without the need to write additional code, which can reduce the difficulty of development. In addition, through reverse sorting, each splicing index can be sorted in reverse order, so that the bucket with a high number of points to be calculated will be sorted first, so the top n buckets with the highest number can be quickly determined and the bucket data can be obtained.

[0193] Optionally, the bucket data of the n buckets with the highest number copied to the CPU memory space may include relevant information of the buckets, such as the splicing index of the bucket, the number of points to be calculated in the bucket, etc. As an example, it may be the above-mentioned starting position, splicing index and number of occurrences.

[0194] Optionally, after the CPU reads these data, the number of threads allocated to the buckets whose number is greater than or equal to the preset threshold can be flexibly configured according to actual needs, and this embodiment does not limit this. For example, an upper limit value for the number of points to be calculated processed by each thread can be set, and then the number of threads allocated to the bucket can be determined based on the number of points to be calculated in the bucket and the upper limit value. Among them, the upper limit value can be 10 or 8, that is, each thread processes a maximum of 10 points to be calculated or 8 points to be calculated; of course, other values ​​can be set as needed in actual applications, and this embodiment does not limit this.

[0195] As an example, the calculation formula for the number of threads allocated to a bucket can be: Among them, the symbol Indicates rounding up, T indicates the number of points to be calculated in the bucket, and m indicates the upper limit, which can be a preset constant and can be the same as or different from the preset threshold. For example, if the number of points to be calculated in the bucket is 91 and the upper limit is 10, the number of threads to be allocated can be calculated to be 10 according to the formula.

[0196] Optionally, after allocating threads to each bucket, the CPU can transmit the thread allocation result to the GPU. The thread allocation result can include each splicing index and the corresponding number of threads. For example, the CPU can initiate a GPU computing task to the GPU, and the computing task can include the splicing index of the bucket and the corresponding number of threads (i.e., the thread allocation result), so that the GPU can perform computing based on the computing task.

[0197] like Figure 7As shown, this specification shows a calculation method for a multi-scalar multiplication (MSM) according to an exemplary embodiment. The method is applied to a central processing unit (CPU), and the MSM is implemented based on the Pippenger algorithm running on a graphics processing unit (GPU). The method may include the following steps:

[0198] In step 702, read the bucket data copied by the GPU to the memory space.

[0199] Wherein, the GPU is used to execute the steps of the method described in the foregoing embodiment; the bucket data refers to the bucket data of the n buckets with the highest number of points to be calculated determined from the bucket result after the GPU determines the bucket result based on the current MSM to be calculated.

[0200] In step 704, after determining whether the number of points to be calculated in each of the n buckets is greater than or equal to a preset threshold, determine a thread allocation result in which buckets with a number greater than or equal to the preset threshold are allocated multiple GPU threads, and buckets with a number less than the preset threshold are allocated a single GPU thread, and transmit the thread allocation result to the GPU.

[0201] This embodiment describes, from the perspective of the CPU, the process of the CPU judging whether the number of points to be calculated in the n buckets with the highest number is greater than or equal to a preset threshold based on the bucket data copied by the GPU to the memory space, and allocating multiple threads to large buckets and a single thread to small buckets. This embodiment can specifically refer to the description of the foregoing embodiment and will not be elaborated here.

[0202] As can be seen from the above embodiments, in the solution of this embodiment, buckets with a large number of points to be calculated will be allocated sufficient GPU threads for processing, avoiding the situation where one thread is busy and other threads are idle. At the same time, the number of points to be calculated in the vast majority of buckets is small. Therefore, each small bucket is still allocated a single thread for processing, rather than using multiple GPU threads regardless of the bucket size as in the existing solutions. By reasonably controlling the number of threads, the solution of this embodiment can keep the GPU utilization rate at the maximum value all the time. Among the introduced additional overheads, the relatively high one is the overhead of copying data from the GPU to the CPU, but the time of this overhead is usually also low and can be ignored.

[0203] This specification embodiment also provides a graphics processing unit, which is used to execute the steps of the method described in the first aspect.

[0204] This specification embodiment also provides a computer device, such as Figure 8As shown in the figure, it is a schematic structural diagram of a computer device according to this embodiment. The computer device includes a graphics processor and a central processor. The graphics processor is used to execute the steps of the method in the foregoing embodiment, and the central processor is used to execute the steps of the method in the foregoing embodiment. In practical applications, the computer device may further include other hardware such as a network interface, a memory, and a non-volatile memory, which will not be elaborated in this embodiment.

[0205] An embodiment of this specification also provides a computer program product, including a computer program. When the computer program is executed by the graphics processor, it implements the steps of the method in the foregoing embodiment; and / or, when the computer program is executed by the central processor, it implements the steps of the method described in the foregoing embodiment.

[0206] An embodiment of this specification also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the graphics processor, it implements the steps of the method described in the foregoing embodiment; and / or, when the computer program is executed by the central processor, it implements the steps of the method described in the foregoing embodiment.

[0207] Corresponding to the foregoing embodiment of the calculation method of the multi-scalar multiplication MSM, this specification also provides an embodiment of a calculation device for the multi-scalar multiplication MSM.

[0208] The embodiment of the calculation device for the multi-scalar multiplication MSM in this specification can be applied to a computer device. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory and running them by the processor where it is located.

[0209] As Figure 9A shown in the figure, it is a schematic diagram of a calculation device for a multi-scalar multiplication MSM according to an exemplary embodiment of this specification. The device can be applied to a graphics processing unit GPU that executes the MSM calculation. The device includes:

[0210] An acquisition module 901, configured to: acquire a bucketing result determined for the currently to-be-calculated MSM; wherein, the bucketing result represents the bucket to which each to-be-calculated point in the to-be-calculated MSM belongs.

[0211] A determination module 902, configured to: determine the number of to-be-calculated points in each bucket based on the bucketing result.

[0212] A calculation module 903, configured to: calculate each of the buckets; wherein, for the buckets with the number greater than or equal to a preset threshold, multiple GPU threads are used for calculation, and for the buckets with the number less than the preset threshold, a single GPU thread is used for calculation.

[0213] In some examples, each piece of binned data in the binned result includes: a point to be calculated, a window index of the window where the bin to which the point to be calculated belongs, and a bin index of the bin to which the point to be calculated belongs;

[0214] The calculation module 902 is configured to:

[0215] After concatenating the window index and the bin index in each piece of binned data to obtain a concatenated index, obtain a concatenated index sequence of each piece of binned data;

[0216] Sort each concatenated index in the concatenated index sequence to obtain a sorted concatenated index sequence;

[0217] Statistically analyze the number of occurrences of each concatenated index in the sorted concatenated index sequence through a run-length encoding function, where the number of occurrences represents the number of points to be calculated in each bin.

[0218] In some examples, the calculation module 902 is configured to:

[0219] Sort each concatenated index in the concatenated index sequence through a sorting function provided by the programming tool library of the GPU;

[0220] and / or

[0221] The run-length encoding function is a function provided by the programming tool library of the GPU.

[0222] In some examples, the computer device where the GPU is located includes a central processing unit CPU; the calculation module 902 is configured to:

[0223] Obtain the bin data of the n bins with the highest quantity; where n is a positive integer; for the other bins except the n bins with the highest quantity, it is determined to use a single GPU thread for calculation;

[0224] Copy the bin data of the n bins with the highest quantity to the memory space of the CPU, and obtain the thread allocation result of the CPU for the n bins with the highest quantity;

[0225] Among them, after the CPU reads the bin data of the n bins from the memory space, it determines whether the number of points to be calculated in each of the n bins is greater than or equal to the preset threshold, and then determines the thread allocation result of allocating multiple GPU threads to the bins with the quantity greater than or equal to the preset threshold and allocating a single GPU thread to the bins with the quantity less than the preset threshold, and transmits the thread allocation result to the GPU.

[0226] In some examples, the obtaining the bin data of the n bins with the highest quantity includes:

[0227] Using the reverse sorting function provided by the programming tool library of the GPU, sort the respective splicing indices in reverse order according to the occurrence times, and obtain the bucket data of the top n buckets with the highest quantity based on the reverse sorting result.

[0228] In some examples, the GPU is used to execute the zero-knowledge proof task.

[0229] As Figure 9B shown, it is a schematic diagram of a computing device for multi-scalar multiplication (MSM) according to an exemplary embodiment of this specification. The device is applied to a central processing unit (CPU), and the device includes:

[0230] A reading module 912, configured to: read the bucket data copied by the graphics processing unit (GPU) to the memory space; the GPU is used to execute the steps of the method described in the foregoing embodiment; the bucket data refers to the bucket data of the top n buckets with the highest quantity of points to be calculated determined by the GPU from the bucketing result after determining the bucketing result based on the current MSM to be calculated.

[0231] A determining module 914, configured to: after determining whether the quantity of points to be calculated in each of the n buckets is greater than or equal to a preset threshold, determine a thread allocation result of allocating multiple GPU threads to the buckets with the quantity greater than or equal to the preset threshold and allocating a single GPU thread to the buckets with the quantity less than the preset threshold, and transmit the thread allocation result to the GPU.

[0232] In the 1990s, it was obvious to distinguish whether an improvement in a technology was a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method flows into the hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented with hardware entity modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that by simply performing some logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0233] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to logically program the method steps to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same function. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.

[0234] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, this application does not exclude that with the development of future computer technologies, the computers for implementing the functions of the above embodiments can be, for example, personal computers, laptop computers, in-vehicle human-machine interaction devices, cellular phones, camera phones, smart phones, personal digital assistants, media players, navigation devices, email devices, game consoles, tablet computers, wearable devices, or any combination of these devices.

[0235] Although one or more embodiments of this specification provide method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among many orders of step execution and does not represent the only execution order. When the actual device or terminal product is executing, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, product or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. For example, if terms such as first and second are used to denote names, they do not denote any particular order.

[0236] For convenience of description, when describing the above device, it is described by dividing it into various modules according to functions. Of course, when implementing one or more of this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0237] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0238] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction means that implements the function specified in one or more of the processes Figure 1 or processes and / or boxes Figure 1 specified in one or more boxes.

[0239] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the function specified in one or more of the processes Figure 1 or processes and / or boxes Figure 1 specified in one or more boxes.

[0240] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0241] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of a computer-readable medium.

[0242] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage, graphene storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0243] Those skilled in the art should understand that one or more embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, one or more embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0244] One or more embodiments of this specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of this specification can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0245] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments. In the description of this specification, the description of reference terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0246] The above is only the embodiments of one or more embodiments of this specification and is not used to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims.

Claims

1. A method for calculating a multi-scalar multiplication (MSM), the method being applied to a graphics processor (GPU) for performing the MSM calculation, the method comprising: Obtaining a bucketing result determined for the current MSM to be calculated; wherein the bucketing result represents the bucket to which each point to be calculated in the MSM to be calculated belongs; Determine the number of points to be calculated in each bucket based on the bucketing result; Calculation is performed on each of the buckets; wherein, buckets whose number is greater than or equal to a preset threshold are calculated using multiple GPU threads, and buckets whose number is less than the preset threshold are calculated using a single GPU thread.

2. According to the method of claim 1, each bucket data in the bucket result includes: The point to be calculated, the window index of the window where the bucket to which the point to be calculated belongs is located, and the bucket index of the bucket to which the point to be calculated belongs; The determining the number of points to be calculated in each bucket based on the bucketing result includes: After splicing the window index and the bucket index in each bucket data to obtain the splicing index, the splicing index of each bucket data is obtained to obtain a splicing index sequence; Sorting each splicing index in the splicing index sequence to obtain a sorted splicing index sequence; The number of occurrences of each splicing index in the sorted splicing index sequence is counted by a run-length encoding function, and the number of occurrences represents the number of points to be calculated in each bucket.

3. The method according to claim 2, wherein the step of sorting each splicing index in the splicing index sequence comprises: Sorting each splicing index in the splicing index sequence by using a sorting function provided by a programming tool library of the GPU; and / or, The run-length encoding function is a function provided by a programming tool library of the GPU.

4. The method according to claim 1 or 2, wherein the computer device where the GPU is located includes a central processing unit (CPU); the method further comprises: Obtaining bucket data of n buckets with the highest number, wherein n is a positive integer; determining to use a single GPU thread for calculation for buckets other than the n buckets with the highest number; Copying the bucket data of the n buckets with the highest number to the memory space of the CPU, and obtaining the thread allocation result of the CPU for the n buckets with the highest number; Among them, the CPU is used to read the bucket data of the n buckets from the memory space, determine whether the number of points to be calculated in each bucket of the n buckets is greater than or equal to the preset threshold, determine the thread allocation result of allocating multiple GPU threads to the buckets with the number greater than or equal to the preset threshold and allocating a single GPU thread to the buckets with the number less than the preset threshold, and transmit the thread allocation result to the GPU.

5. According to the method of claim 4, the step of obtaining the bucket data of the n buckets with the highest number comprises: By using the reverse sorting function provided by the programming tool library of the GPU, each splicing index is sorted in reverse order according to the number of occurrences, and the bucket data of the first n buckets with the highest quantity are obtained based on the reverse sorting result. The method according to claim 1 , wherein the GPU is used to perform zero-knowledge proof tasks.

7. A method for calculating a multi-scalar multiplication MSM, the method being applied to a central processing unit (CPU), the method comprising: Read the bucket data copied to the memory space by the graphics processor GPU; The GPU is used to execute the steps of the method according to any one of claims 1 to 6; the bucket data refers to the bucket data of n buckets with the highest number of points to be calculated determined from the bucket results after the GPU determines the bucket results based on the current MSM to be calculated; After determining whether the number of points to be calculated in each of the n buckets is greater than or equal to a preset threshold, determine the thread allocation result of allocating multiple GPU threads to the buckets whose number is greater than or equal to the preset threshold and allocating a single GPU thread to the buckets whose number is less than the preset threshold, and transmit the thread allocation result to the GPU.

8. A graphics processor, configured to execute the steps of the method according to any one of claims 1 to 6.

9. A computer device, comprising a graphics processor and a central processing unit, wherein the graphics processor is used to execute the steps of the method according to any one of claims 1 to 6, and the central processing unit is used to execute the steps of the method according to claim 7.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a graphics processor, the computer program implements the steps of the method according to any one of claims 1 to 6; and / or when the computer program is executed by a central processing unit, the computer program implements the steps of the method according to claim 7.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a graphics processor, implements the steps of the method described in any one of claims 1 to 6; and / or, when executed by a central processing unit, implements the steps of the method described in claim 7.

Citation Information

Cited By

  • Multi-scalar multiplication (MSM) computation

    WO2026174829A1