Concurrent computing system for multi scalar multiplication

The concurrent computing system addresses inefficiencies in multi-scalar multiplication by sequencing point additions to subsume higher-order buckets, achieving a significant performance boost in cryptographic computations.

GB2629232BActive Publication Date: 2025-06-11MARK ALEXANDER BLUNDEN +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2023018835
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-11
Estimated Expiration
2043-12-11

AI Technical Summary

Technical Problem

Existing methods for computing multi-scalar multiplications in elliptic curve cryptography are inefficient due to sequential processing in the AGGREGATE stage, which limits concurrent or parallel processing capabilities.

Method used

A concurrent computing system with multiple curve adding units and a controller that sequences point additions to minimize conflicts, allowing for concurrent execution by subsuming higher-order buckets into lower-order buckets, reducing the number of point-scalar multiplications through a process called bucket subsumption.

Benefits of technology

The system significantly reduces the number of required point additions by half, enabling efficient concurrent computation of multi-scalar multiplications, thereby improving performance in cryptographic applications like zk-snarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

The invention is a multi scalar multiplication scheme where polynomial elliptic curve points are to be multiplied by scalar coefficient values. Curve points with the same scalar coefficient value (i.e
Need to check novelty before this filing date? Find Prior Art

Description

The present invention relates to a system for performing multi-scalar multiplications, and thus providing faster processing of a cryptographic primitive in elliptic curve cryptography. 5 BACKGROUND TO THE INVENTION Elliptic curves are of interest in cryptography. Particularly useful is a prime elliptic curve, which is the set of solutions to an equation in the form y2 = x3 +ux + vy, where u and v, as well as points (x,y) that lie on the curve, belong to a finite field Fp defined by a prime p. Fpis the set {0,1,..., p - 1}, with addition and multiplication being modulo 10 p. Typically, for cryptographic applications there exists a subgroup of points on the curve of order some large prime number, q. I.e, the number of points in the subgroup is prime q. Points on such a curve can be “added” with the result being another point on the curve (point addition is defined as the negation of the point resulting from the intersection of C\J 15 the curve and the straight line defined by the points being added, with point doubling 00 defined as a special case when a point is added to itself). Multiplication of a point by a scalar n is the same as adding the point to itself n times. Elliptic curves of this type provide an effective trapdoor function. In short, a public point on a (public) curve can be added to itself a secret number of times to generate a public 20 key. In general, for z = w ■ G, where G is a point on the curve, finding w given only z and G is a “hard” problem - i.e. for an appropriate choice of elliptic curve parameters it is unfeasible to find w in a reasonable amount of time. In other words it is not possible, given two public points on a public elliptic curve, to 25 efficiently discover the secret number of times one point has to be added to itself to arrive at the other point. On the other hand, computing z = w-G with knowledge of w and G is relatively straightforward using known elliptic curve point multiplication algorithms. By careful selection of the particular curve, and the public point (known as the 30 “generator point” or “base point”), any point at all on the curve (or at least, any of a large number of points in a cyclic subgroup) can be arrived at by a scalar multiplication of the generator point. Elliptic curve cryptography is not often used to directly encrypt data (although it is possible to do so). It is more often used in key-agreement protocols, and in pairing-5 based zero-knowledge proofs which form part of many useful security protocols. However, suffice to say that operations (encryptions) can be carried out on values using knowledge of the public key (point z), and those operations can only be feasibly reversed with knowledge of the private key (scalar w by which generator G was multiplied to arrive at z). 10 Elliptic curves have some homomorphic properties which allow computations to be carried out on encrypted values. In particular, encrypted values can be multiplied by a scalar or added to arrive at encrypted result values which in principle can be decrypted using the private key(s) to retrieve the result of the addition or multiplication. These homomorphic properties are useful since, in particular, they allow the evaluation of a 15 polynomial p(x) at a secret value x = s, using only encrypted forms of s. As will be seen, this is critical for the KZG commitment scheme. Commitment schemes are an important cryptographic primitive, and allow a proving party to commit to a value while keeping the value itself secret. It is not possible to change the committed value once the commitment has been created, and the 20 committed value can optionally be revealed at a later time. The whole process can be thought of as equivalent to putting a note (the “committed value”) into a locked box (“the commitment”) and then giving the locked box to a counterparty or putting it in a public place. Later, when the proving party wants to reveal the committed value, they provide the key or combination to open the box. 25 Commitment schemes play a vital role in a number of cryptographic applications, such as zero-knowledge proofs and secure multi-party computation. With polynomial commitment schemes, the committing party commits to a polynomial p(x). In particular, KZG is a pairing-based polynomial commitment scheme in which the commitment to a polynomial p(x) is computed by evaluating p(x) at a secret (in 30 this case, secret in the sense that it is known to neither party) value x = s. The commitment is given as a point on an elliptic curve p(s) • G. Importantly, the homomorphic properties of elliptic curve encryption allow p(s) • G to be computed without knowing s, but being provided with points s ■ G, s2 • G, ..., sn ■ G where n is the degree of polynomial p(x). These points P, = ■ 6 for 1 = 1, ...,n are in essence encrypted forms of s. The initial creation of points Pt requires a trusted setup in which the secret s is used and then forgotten and made unrecoverable, so that s remains secret from all parties. 5 Given the commitment p(s) ■ G it is not possible to deduce p(x). However, in the “reveal” it is possible to prove that the revealed p(x) was the authentic polynomial when p(s) ■ G was calculated. Computing the commitment p(s) ■ G using points Pt= sl • G, without knowing s, involves what is called a multi scalar multiplication (MSM), i.e. it involves computing 10 the sum: n Zat ■ Pi i=0 where aL are the coefficients of polynomial p(x) (scalar values). Pt are points on an elliptic curve, as explained above. The multiplication (dot) operator is defined as CM repeated “addition” of point Pt to itself. Point addition (or more particularly point CO 15 doubling) is defined as explained above. KZG commitments are cryptographic primitives which find application in different cryptographic systems and protocols. In particular, commitments are used in zeroknowledge proofs, particularly zk-snarks, which in turn have many applications. A recent example of a zero-knowledge proof scheme is PLONK, and without going into 20 detail, PLONK requires nine MSMs. It is the computation of MSMs which dominate the computational expense of these schemes. The standard method for computing MSMs is Pippenger’s Algorithm, or the “bucket” algorithm. The bucket algorithm is applied to “reduced” MSMs, which are MSMs with scalar coefficients which are all less than some power of two 2C. In general MSMs can 25 be solved by breaking them down into a number of reduced MSMs, solving the reduced MSMs, and then computing the solution to the original MSM from the solutions to the reduced MSMs. This procedure of breaking down MSMs into reduced MSMs and reconstructing the results is known, and is not described in detail here. This disclosure addresses the problem of efficiently solving reduced MSMs. 30 The bucket algorithm comprises an ACCUMULATE stage and an AGGREGATE stage. In the ACCUMULATE stage, each point Pt is added to a bucket Ba. according to its associated scalar. In other words, all the points which are to be multiplied by the same scalar are put in the same “bucket”. Points are added to the bucket simply by adding the point to the existing point in the bucket. 21 08 24 Once the ACCUMULATE stage has been completed, the point in each bucket needs 5 to be multiplied by its associated scalar, and the results added together. In other words the purpose of the AGGREGATE function is to compute as efficiently as possible: T ' kSk = + 2 ■ S2 + • • • + T ■ Sf fc=i Where Sk is the point in bucket Bk following completion of the ACCUMULATE stage (that is, the sum of all points allocated to bucket Bk, or the sum of all points having 10 associated scalar k). The standard approach to the AGGREGATE stage is based on the observation that T ^kSk = S1 + 2-S2 + - + T-St k=l = Sy + (Sf + + •" + (Sy + Sf-i + •" + Sq) and hence T point multiplications can be replaced by 2T point additions as follows: 15 1. Fori = T-l, ...,1 letSj ^Sj+1+Si 2. Let SF1NAL = ^j=1Sj This method is optimal in terms of the number of required operations. The time taken for AGGREGATE is dependent on the number of buckets. However, it can be seen that the additions are sequential, and are not suited to concurrent processing. This is 20 because each addition uses the result of the previous addition. In FPGA Acceleration of Multi-Scalar Multiplication, Jump Crypto (2022), available at https: / / qjthub.com / 'z:^ msrn4pqa / nickray / zprize-fpga-msm.pdf, an FPGA-based system for computing MSMs is proposed. The authors optimise the ACCUMULATE stage by carrying out the point 25 additions concurrently. A scheduler detects conflicting points and defers those additions to avoid the conflict. Hence parallel processing can be utilised to speed up the ACCUMULATE stage. However, the AGGREGATE stage cannot be optimised in 21 08 24 this way - a scheduler of the type used for the ACCUMULATE stage is no good because every single point addition relies on the result of the previous point addition. In the existing scheme (optimal in terms of the number of additions), there can be no concurrent or parallel processing. 5 Accordingly it is an object of the present invention to provide a concurrent processing system for computing multi scalar multiplications. STATEMENT OF INVENTION According to the present invention, there is provided a concurrent computing system configured to compute a multi scalar multiplication in a cryptographic computation, 10 the computing system having a plurality of curve adding units, each curve adding unit being capable of performing an addition of two points on an elliptic curve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being Hmax, and the computing system further having a memory accessible to the curve adding 15 units, and a controller, the multi scalar multiplication being in the form of: ai-Pi i=0 where Pt is a point on an elliptic curve and is a scalar, the maximum value of being T, 20 and the controller being adapted to carry out the multi scalar multiplication by: defining a memory location Bk associated with each distinct value of ab and storing an identity point 0 in each memory location Bk, Sk being the point stored in memory location Bk\ for each point Pb adding the point Pt to the point Sk at Bk where k = ah and 25 storing the result in Bk, for some d <T, sequencing point additions for all values of k where d <k <T, the point additions being: Bg Sg + Sk (i.e. add Sk to Sg and store the result in Bg) B h <- Sh + Sk(i.e. add Sk to Sh and store the result in Bh) where (g + h) = k, g <d, h< d, the point additions being sequenced such that the minimum distance in the sequence between two point additions acting on the 5 same memory location Bz is greater than some H, and the controller causing concurrent execution of the sequenced point additions on each of H curve adding units where H <Hmax', 21 08 24 It can be seen that the invention works using a variation on the known “bucket” 10 algorithm. The ACCUMULATE phase involves putting each point Pt to bucket Bk, the bucket to which Pt is added depending on the associated scalar at, i.e. k = This is known, and in embodiments may be optimised to make use of the concurrently operable point adding units as described in Jump (2022). However, in the AGGREGATE phase, the system of the invention is able to subsume 15 the points in higher-order buckets into lower-order buckets. In other words, it turns the sum 'k into j=l 20 In general, call the transformation from the first sum above into the second sum “bucket subsumption”. Critically, d <T. In a typical embodiment, d is about half of T. Hence the number of point-scalar multiplications in the sum, and hence the number of point additions required to solve it using known sequential addition methods, is cut in half. To achieve 25 this, of course point additions need to be carried out. Specifically the system has to compute Bg <- Sg + Sk and Bh <- Sh+Sk, for all k where d <k <T. The system of the invention does not therefore reduce the total number of point additions. However, unlike in known systems, these point additions are not sequentially dependent on each other, and therefore they can be sequenced efficiently for concurrent execution. Note that the step of sequencing and concurrently executing additions may be repeated for a new d' <d , preferably d' is about half of d. Hence the number of terms 5 in the sum may be repeatedly reduced (preferably repeatedly halved), a process which is preferably repeated either until there is only one term, or the number of terms is so small that it becomes faster to use known sequential methods to solve the remaining sum rather than having the overheads associated with concurrent operation. As the number of terms reduces, so will the maximum distance at which two point additions 10 acting on the same memory location Bz can be sequenced. This may require H to be set to less than the maximum number of available concurrent curve adding units Hmax, to ensure concurrent operations do not conflict. Hence the number of point additions which can be performed concurrently, and the benefit of the concurrent system compared to known sequential methods, will reduce as the number of terms reduces. 15 At some stage, embodiments may revert to solving the remaining sum Yg^jSj using known sequential methods. In some embodiments, H is always set to Hmax, and once C\l the number of terms in the sum reduces below the minimum for which two point 00 additions acting on the same memory location Bz can be sequenced at a distance of more than Hmax apart, the remaining sum will have to be solved using known v— 20 sequential techniques. CM In principle g and h can be any integers such that (g + h) = k, g <d,h< d. However, in one straightforward and efficient sequencing scheme, g and h are selected as: for even k\ g = h = i j-i , k+1 , k-1 for odd k: q = —, h = — 25 For even k, g and h are the same. One implication of this is that in some embodiments point doubling may be supported as a single (fast) operation in the curve adding units. Hence instead of two additions: Bk^Sk+ Sk 2 2 Bk Sk + Sk 2 2 30 a single addition plus a point doubling may be sequenced: Bk^Sk + 2-Sk 2 2 For odd k, the relevant additions are: Bkn Sk+1 + Sk 2 2 Bk-i Sk-i + Sk 2 2 5 In any case, choosing g and h in this way enables a straightforward but efficient sequence for the point additions. Typically T = 2C - 1 for some c, and d = 2^1. Accordingly, d is about half of T, and the “top half” of the set of buckets gets subsumed into the “bottom half” of the set of buckets. Note that in the first iteration, T will typically be one less than a power of two (because the MSM being solved is a reduced MSM 10 as mentioned above). However, in subsequent iterations, the number of terms will be d = 2C-1, i.e. a power of two exactly. Preferably, to sequence these point additions, they are dealt with in groups, each of the additions associated with odd k and of the form Bk+i <- Sk+i + Sk being grouped in 2 2 a first group, each of the additions associated with odd k and of the form Bk-i «- Sk-i + 2 2 15 Sk being grouped in a second group. Where point doubling operations are available in the curve adding units, the set of additions of the form Bk «- Sk + 2 • Sk form a third group. Where point doubling operations are not available 2 2 there may be two identical groups, a third and fourth group, each of which contains the set of additions of the form Bk Sk + Sk. 2 2 20 Within each group, the additions are ordered by the index of the relevant bucket. It does not matter whether the order is increasing or decreasing (or indeed ordered in any other well-defined way), as long as that is consistent among the groups. The additions are then sequenced with all the additions in a single group being sequenced together. For example, the sequence could be all the additions in the first group, then 25 all the additions in the second group, then all the additions in the third group, then all the additions in the fourth group. But equally it could be the fourth group, then the first group, then the second, then the third group, or any other order. Each group is guaranteed not to contain a conflicting addition (i.e. a group never contains two additions which write to the same bucket). Furthermore, because the additions are 30 ordered within the group, when two groups are sequenced together, there is a minimum distance between any two conflicting additions which is at least as large as the group. All the groups are at least as large as t, and t = - 1 (where T is one less than a power of two; in the subsequent iteration the starting number of terms d is exactly a 5 power of two, the ending number of terms d' is exactly half of d, and therefore all the groups are at least as large as t = y = |). Accordingly, H can be safely set to t. In other embodiments, rather than setting H, all Hmax curve adding units are always used and the controller ensures that concurrent operation only continues for as long as t >Hmax, that is that the groups remain sufficiently large such that the distance between 10 conflicting additions in the sequence will always be at least as large as the number of concurrent operations. When t >Hmax is no longer satisfied, the controller will cease concurrent operation and solve the remaining sum by known sequential methods. Note that the description so far assumes that the concurrently operable curve adding units are to some extent synchronised - i.e. it assumes that (say 96) curve adding units \J 15 will all perform their addition at the same time, finish at the same time, and then CM together perform the next batch of 96 additions. Or at least, that the curve adding units CO perform an addition in predetermined timeslots. The timeslots may be overlapping rather than exactly in parallel, but the point is that the order in which add operations start and finish is entirely predictable, so for example it can be guaranteed that the 97th CM 20 operation will not begin until the 1st operation has completed, the 98th operation will not begin until the 2nd operation has completed, and so on. In some embodiments, the concurrently operable curve adding units may be asynchronous, and some additions may complete and free the relevant unit for another addition more quickly than others. In such a case, a scheduler may ensure that additions are deferred if they are going to 25 conflict with pending additions which are not yet completed. The hardware of the computing system may be similar to that described in Jump (2022). In particular the key components are the plurality of curve adders, the MSM controller and the scheduler. Note that the scheduler is optional in embodiments of this invention, but may be included so that the ACCUMULATE phase can also be operated 30 using concurrent additions as described in Jump, and possibly to further optimise the AGGREGATE phase as well if the curve adders are allowed to operate asynchronously. This may all be implemented in an FPGA. The curve adders preferably support both ADD and DOUBLE (and / or DOUBLE-AND-ADD) operations. Software implementations using multi-core processors are also possible. In some embodiments, multiple devices may be used, operating in parallel. Each device may have a controller, multiple curve adders, and a memory. The devices may be able to communicate with each other to share data between devices, and may be controlled by a master controller (the master controller may be for example a 5 conventional PC). Advantageously, the buckets may be distributed between devices so that each device carries out mainly point additions on its own buckets, with a small number of intermediate results having to be copied between devices. In particular, let RD1 represent the buckets used with device DI, and RD2 represent the 10 buckets used with device DI. As before, St is the point in bucket Bt. As explained, in iterations of the bucket subsumption stage, typically the starting number of buckets is a power of two, i.e. 2c(or possibly the number of buckets is one less than a power of two 2C - 1, but equivalently the bucket B2c can just be set empty, i.e. to the identity element). The ending number of buckets after one bucket 15 subsumption iteration is half that number, i.e. 2C1. CM CO In an iteration of the bucket subsumption stage, the points in the higher order buckets (i.e. the buckets from B2c-i+1 to B2c) are added to the points in the top half of the lower order buckets (i.e. the buckets from B2c-2+1 to B2c-i). c\i Therefore, if buckets are allocated to devices DI and D2 as follows: 20 R^1 = {B2c-i+1, ,B3.2c-2} R^2 = {B2.2c-2+1, ..., B2c} then the points in the buckets of R^1 are added to the points in the buckets of B2<^2+v ...,B3.2c^3, except forS2c i+1 (the point in B2c^+1) which is added to the point S2c^2 (the point in B2c^i). 25 Similarly, the points in the buckets of R^2 are added to the points in buckets D3.2c-3+i, ...,B2c-i, except for S3.2c-2+1 (the point in B3.2c-2+1) which is added to the point S3.2c-3 (the point in B3.2c-3y Accordingly the points in buckets of Df1 are subsumed into the points in the buckets of R^, with all the necessary point additions involving points on device DI only, bar 30 the exceptional cases noted above which will require copying between devices. 21 08 24 Generally, for multiple iterations of the point subsumption, buckets can be allocated to two devices by defining RDj = R^UR^^U ... II R^it for j e {1,2}, for some integer t. The maximum value of t is determined by the requirement that the cardinality of R^t+i is sufficiently large as to ensure that the point additions that subsume the points in the 5 buckets of 'nto Rc^t remain independent. As explained above, in some embodiments the number of concurrent curve adders in use H may be dynamic as set sizes get smaller, or in other embodiments a constant number of curve adders Hmax may determine an absolute maximum value of t. More generally, for more than two devices (any number h of devices where h = 2d 10 andd<c-l): Rc 2^c~^ + (J —l)-2*-c—+1' R2^-1) and as before RD} = R^^R^L^ ... U R^Lt, where R1^ is defined for any integer w <c as: 15 BRIEF DESCRIPTION OF THE DRAWINGS For a better understanding of the present invention, and to show more clearly how it may be carried into effect, reference will now be made by way of example only to the accompanying drawings, in which: Figure 1 is a schematic of a concurrent computing system according to the invention; 20 Figure 2 is a schematic of another embodiment of a concurrent computing system according to the invention, comprising two of the devices of Figure 1 together with a master controller, linked together by a network bus; and Figure 3 illustrates how points are stored in buckets during use of the computing system of Figure 1 or Figure 2. 25 DESCRIPTION OF PREFERRED EMBODIMENTS Referring firstly to Figure 1, a broad schematic of a computing system according to the invention is shown. The system comprises Hmax curve adding units indicated at 12, a controller 14 and a memory 16. Each curve adder is capable of adding a point on a curve to another point. Preferably, each curve adder also supports an efficient point doubling operation, i.e., adding a point to itself. 21 08 24 Each curve adder can obtain its operands from the memory 16 and store the result of 5 its computation (which is a new point) in the memory 16. The controller 14 allocates point additions to the curve adders 12. In one straightforward embodiment the controller 14 will allocate a number H of parallel curve additions to the adders 12, monitor for when the curve adders 12 have finished these calculations and stored their results in memory 16, and then allocate another set of 10 curve additions. However, in other embodiments the curve adders may operate asynchronously. The controller may include a scheduler to avoid conflicting additions from occurring at the same time, or at overlapping times. The computing system of the invention is designed to compute a multi-scalar multiplication of the form: n 15 t=o To do so, the controller is adapted to firstly: define a memory location Bk associated with each distinct value of ah and store an identity point 0 in each memory location Bk, Sk being the point stored in memory location Bk, 20 for each point Ph add the point Pt to the point Sk at Bk where k = and store the result in Bk. This is known as the ACCUMULATE phase and is conventional. This phase may make use of concurrent multiple curve adders using the techniques disclosed in Jump (2022). 25 The next stage is the AGGREGATE stage. The purpose of the AGGREGATE stage is to compute as efficiently as possible the sum: t=i kSk = + 2 ■ S2 + ••• + T ■ ST where T is the number of buckets and is typically 2C - 1 for some c. 21 08 24 The key to doing this concurrently is to sequence point additions so that, subject to H point additions taking place concurrently, there will be no conflicting concurrent additions. This means making sure that the maximum distance between any two 5 conflicting point additions in the sequence is greater than H. The point additions which are sequenced are in the form: Bg <- Sg + Sk (i.e. add Sk to Sg and store the result in Bg) Bh^ Sh + Sk(i.e. add Sk to Sh and store the result in Bh). For some d <T, these two additions are put in the sequence for all values of k where 10 d< k <T. g and h are chosen so that g + h = k, g <d, h< d. This process is called POINT SUBSUMPTION. Its effect is to subsume the points in higher-order buckets into lower-order buckets. In other words, to turn the sum: T kSk k=l 15 into Where d <T. It can be seen that this process can be repeated iteratively to continually reduce the number of buckets, either until there is only one bucket or until the number of buckets is below a threshold after which it is more efficient to solve the sum using 20 traditional methods, than to accept the overheads of running concurrent additions. Typically d is about half of T, both being typically powers of two (or, if T starts off at one less than a power of two, it may be convenient to say that T = 2C and bucket B2c just contains the identity element). A particular embodiment may sequence and carry out concurrent additions as follows: 21 08 24 Input: c, Si,..., £2=.-1, H Output: := c u := 2C - I v := 2' 5 •= 1 t = 2^1 -- I {That is, f = ,u — v + 1, the munber of values i such that t* <i <« } while t >2H do for i = 'it- downfo v do if i is even then S=2 + / ¾ {That is, St is added to the point in bucket end if end for for i = u downto v do if i is even then + Si end if end for for i — u downfo v do if i is odd then, — ^4) / 2) % end if end for for i = a downto 7; do if i is odd then ^{( / -1) / 2} 1) / 2) "r end if end for d := d - 1 « := 2" v = 2fZ 1 4- 1 t ;= 2^-1 {That is, tc= a — v T 1} end while Referring now to Figure 3, it will be seen that, according to the system set out above, the number of buckets is reduced about by half on each iteration (i.e. for each distinct value of t and each iteration of the outer while loop). In particular, the upper-order 5 buckets labelled R^1 and R®2 in Figure 3 are “subsumed” into the “top half” of the lower-order buckets, labelled R^ and Rc-x- In the next iteration, when t has been halved, the buckets labelled R^ and R^ will be “subsumed” into the buckets labelled R^ and R^, and so on until there are not enough buckets for concurrent operation to be worthwhile. In this embodiment, the while loop continues while t >2H, effectively 10 ensuring that the loop will continue for as long as the full concurrent capacity of the system can be used (i.e. all of the curve adders can be used at once) without sequencing a conflicting addition. In other embodiments, iterations could continue for 21 08 24 lower values of t, the controller reducing the number of adders in use as required in the later iterations. As shown by the arrows on Figure 3, buckets in Rf1 are subsumed into buckets in Rc\, and buckets in 2 are subsumed into buckets in R^apart from the exceptional 5 cases noted above. Therefore, if buckets RD1 are allocated to a first device and buckets RD2 are allocated to a second device, the subsumption operations can take place concurrently with only minimal copying of values between devices. Figure 2 shows a suitable architecture with two devices, connected by a network bus. It will be appreciated that the idea can be extended to use more than two devices in parallel. 10 The system according to the invention provides for concurrent computation of the AGGREGATE stage of multi scalar multiplications in elliptic curves. This represents a significant performance improvement compared to the sequential additions carried out on known systems. This has applications for cryptographic primitives, particularly zk-snarks, which rely on these multi scalar multiplications. 15 The embodiments described above are provided by way of example only, and various changes and modifications will be apparent to persons skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

21 08 241. A computing system configured to compute a multi scalar multiplication in a cryptographic computation,the computing system having a plurality of curve adding units, each curve5 adding unit being capable of performing an addition of two points on an ellipticcurve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being Hmax,and the computing system further having a memory accessible to the curve adding units, and a controller,10 the multi scalar multiplication being in the form of:nZat ■ Pii=0where Pt is a point on an elliptic curve and at is a scalar, the maximum value of a, being T,and the controller being adapted to carry out the multi scalar multiplication by:15 defining a memory location Bk associated with each distinct value of ah andstoring an identity point 0 in each memory location Bk, Sk being the point stored in memory location Bk;for each point Ph adding the point P, to the point Sk at Bk where k = ah and storing the result in Bk,20 for some d <T, sequencing point additions for all values of k where d <k <T, the point additions being:Bg Sg + Sk (i.e. add Sk to Sg and store the result in Bg)Bh<- Sh + Sk(je. add Sk to Sh and store the result in Bh)where (g + h) = k, g <d, h< d, the point additions being sequenced such25 that the minimum distance in the sequence between two point additions acting on the same memory location Bz is greater than some H, and the controller causingconcurrent execution of the sequenced point additions on each of H curve adding units where H <Hmax;2. A system as claimed in claim 1, in which the step of sequencing and 5 concurrently executing point additions is repeated for a new d' <d.

3. A system as claimed in claim 1 or claim 2, in which g and h are selected as:for even k: g = h =21 08 244. A system as claimed in claim 3, in which the point additions are sequenced in10 groups, wherein:a first group comprises the additions associated with odd k and of the form Bk+i <- Sk+i +2 2a second group comprises the additions associated with odd k and of the form Bk-i <- Sk-i + Sk',2 215 a third and optionally a fourth group comprises the additions associatedwith even k,in which the additions are ordered within each group, and in which the additions are sequenced with all the additions in a single group being sequenced together.20 5. A system as claimed in any of the preceding claims, further comprising ascheduler for preventing conflicts between sequenced additions executed concurrently.

6. A system as claimed in any of the preceding claims, in which the plurality of curve adding units are provided as a first set of curve adding units on a first 25 device, and a second set of curve adding units on a second device, the firstdevice having a first memory associated with and accessible to the first set of curve adding units and the second device having a second memory associatedwith and accessible to the second set of curve adding units, and the first and second devices being linked together by a communication means.

7. A system as claimed in claim 6, in which memory locations in some subset RD1 of R - {Bo,..., Bt] are in the memory of the first device and memory locations5 in another subset of R, RD2, are in the memory of the second device, and inwhich additions of the form Bg^ Sg + Sk and Bh <- Sh + Sk are sequenced on the first device for Bn g Rd1 and on the second device for Bh g RD2.

8. A system as claimed in claim 7, in which:RWJ = 1^2^^10 for w <c, and RDi = R^U Rf-iU - U Rc-f f°r some integer t and where j g{1,2,..., / 1} for some integer h, wherein h devices are provided, each device having a plurality of curve adding units and a memory accessible by the curve adding units of the device, and the devices being linked together by a communication means, and in which memory locations in are in the15 memory of the jth device and in which additions of the form Bg Sg + Sk andBh Sh + Sk are sequenced on the yth device for Bg g Rdj ."1“ 9. A non-transient computer-readable medium having instructions thereon which,when executed on a suitable machine, cause the machine to carry out a multi scalar multiplication in a cryptographic computation,20 in which the machine includes a plurality of curve adding units, each curveadding unit being capable of performing an addition of two points on an elliptic curve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being Hmax,and the machine further comprises a memory accessible to the curve adding25 units, and a controller,and in which the multi scalar multiplication is of the formZai ' Pi21 08 24where P^ is a point on an elliptic curve and aL is a scalar, the maximum value of at being T,and in which the multi scalar multiplication is carried out by:defining a memory location Bk associated with each distinct value of and 5 storing an identity point 0 in each memory location Bk, Sk being the point storedin memory location Bk,for each point Ph adding the point Pt to the point Sk at Bk where k = ait and storing the result in Bk,for some d <T, sequencing point additions for all values of k where d <k <10 T, the point additions being:Bg <- Sg + Sk (i.e. add Sk to Sg and store the result in Bg)Bh $h + Skfte. add Sk to Sh and store the result in Bh)where (g + h) = k, g <d, h< d, the point additions being sequenced such that the minimum distance in the sequence between two point additions acting15 on the same memory location Bz is greater than some H, and the controllercausing concurrent execution of the sequenced point additions on each of H curve adding units where H <Hmax,computing ■ S,.

10. A computer readable medium as claimed in claim 9, in which the step of 20 sequencing and concurrently executing point additions is repeated for a newd' <d.

11. A computer readable medium as claimed in claim 9 or claim 10, in which g and h are selected as:for even k: g = h =r -I-I >k + 1 , k-125 for odd k: q = —, h = —.W 2 212. A computer readable medium as claimed in claim 11, in which the point additions are sequenced in groups, wherein:a first group comprises the additions associated with odd k and of the form Bk+i <- Sk+i + Sk;2 2a second group comprises the additions associated with odd k and of the form Hk i <- Si^t +2 25 a third and optionally a fourth group comprises the additions associatedwith even k,in which the additions are ordered within each group, and in which the additions are sequenced with all the additions in a single group being sequenced together.1021 08 24

Citation Information

Patent Citations

  • Formalized multi-scalar multiplication analysis and calculation acceleration method

    CN116932991A