Concurrent computing system for multi scalar multiplication

The concurrent computing system addresses the inefficiency of sequential processing in multi-scalar multiplications by utilizing multiple curve adding units and a controller to sequence concurrent point additions, resulting in improved computational efficiency and performance for cryptographic computations.

WO2025125771A1PCT designated stage expired Publication Date: 2025-06-19BLUNDEN MARK +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2024/052200
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-08-22
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing methods for computing multi-scalar multiplications in elliptic curve cryptography are inefficient due to sequential processing in the AGGREGATE stage, which limits concurrent computation and increases computational expense.

Method used

A concurrent computing system is designed to perform multi-scalar multiplications using a plurality of curve adding units that can operate concurrently, with a controller that sequences point additions to avoid conflicts and subsume points in higher-order buckets into lower-order buckets, allowing for efficient concurrent execution.

Benefits of technology

The system significantly reduces the number of point additions required, enabling concurrent processing and thereby improving computational efficiency and performance in cryptographic computations, such as those required for zk-snarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2024052200_19062025_PF_FP_ABST
    Figure GB2024052200_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A concurrent computing system is disclosed for efficiently carrying out multi scalar multiplications on elliptic curve points Solving providing a fast and efficient solution to these multiplications is crucial for the performance of various cryptographic primitives, in particular zk-snarks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CONCURRENT COMPUTING SYSTEM FOR MULTI SCALAR MULTIPLICATIONThe present invention relates to a system for performing multi-scalar multiplications,and thus providing faster processing of a cryptographic primitive in elliptic curve cryptography. BACKGROUND TO THE INVENTIONElliptic curves are of interest in cryptography. Particularly useful is a prime ellipticcurve, which is the set of solutions to an equation in the form ^^ = ^^ + ^^ + ^^, where^ and ^, as well as points (^, ^) that lie on the curve, belong to a finite field ^^ definedby a prime ^. ^^is the set {0, 1, … , ^ − 1}, with addition and multiplication being modulo^. Typically, for cryptographic applications there exists a subgroup of points on thecurve of order some large prime number, ^. I.e, the number of points in the subgroupis prime ^.Points on such a curve can be “added” with the result being another point on the curve(point addition is defined as the negation of the point resulting from the intersection of the curve and the straight line defined by the points being added, with point doublingdefined as a special case when a point is added to itself). Multiplication of a point by ascalar ^ is the same as adding the point to itself ^ times.Elliptic curves of this type provide an effective trapdoor function. In short, a public pointon a (public) curve can be added to itself a secret number of times to generate a publickey.In general, for ^ = ^ ∙ ^, where ^ is a point on the curve, finding ^ given only ^ and ^is a “hard” problem – i.e. for an appropriate choice of elliptic curve parameters it isunfeasible to find ^ in a reasonable amount of time.In other words it is not possible, given two public points on a public elliptic curve, to efficiently discover the secret number of times one point has to be added to itself toarrive at the other point. On the other hand, computing ^ = ^ ∙ ^ with knowledge of ^and ^ is relatively straightforward using known elliptic curve point multiplicationalgorithms. By careful selection of the particular curve, and the public point (known as the “generator point” or “base point”), any point at all on the curve (or at least, any of a large number of points in a cyclic subgroup) can be arrived at by a scalar multiplication of the generator point. Elliptic curve cryptography is not often used to directly encrypt data (although it is possible to do so). It is more often used in key-agreement protocols, and in pairing-based zero-knowledge proofs which form part of many useful security protocols.However, suffice to say that operations (encryptions) can be carried out on values using knowledge of the public key (point ^), and those operations can only be feasiblyreversed with knowledge of the private key (scalar ^ by which generator ^ wasmultiplied to arrive at ^). Elliptic curves have some homomorphic properties which allow computations to be carried out on encrypted values. In particular, encrypted values can be multiplied by ascalar or added to arrive at encrypted result values which in principle can be decryptedusing the private key(s) to retrieve the result of the addition or multiplication. Thesehomomorphic properties are useful since, in particular, they allow the evaluation of apolynomial ^(^) at a secret value ^ = ^, using only encrypted forms of ^. As will beseen, this is critical for the KZG commitment scheme.Commitment schemes are an important cryptographic primitive, and allow a provingparty to commit to a value while keeping the value itself secret. It is not possible tochange the committed value once the commitment has been created, and thecommitted value can optionally be revealed at a later time. The whole process can be thought of as equivalent to putting a note (the “committed value”) into a locked box (“the commitment”) and then giving the locked box to acounterparty or putting it in a public place. Later, when the proving party wants to revealthe committed value, they provide the key or combination to open the box. Commitment schemes play a vital role in a number of cryptographic applications, such as zero-knowledge proofs and secure multi-party computation.With polynomial commitment schemes, the committing party commits to a polynomial^(^). In particular, KZG is a pairing-based polynomial commitment scheme in whichthe commitment to a polynomial ^(^) is computed by evaluating ^(^) at a secret (inthis case, secret in the sense that it is known to neither party) value ^ = ^. Thecommitment is given as a point on an elliptic curve ^(^) ∙ ^. Importantly, thehomomorphic properties of elliptic curve encryption allow ^(^) ∙ ^ to be computedwithout knowing ^, but being provided with points ^ ∙ ^, ^^ ∙ ^, …, ^^ ∙ ^ where ^ is thedegree of polynomial ^(^). These points ^^ = ^^ ∙ ^ for ^ = 1,… , ^ are in essenceencrypted forms of ^. The initial creation of points ^^ requires a trusted setup in whichthe secret ^ is used and then forgotten and made unrecoverable, so that ^ remainssecret from all parties.Given the commitment ^(^) ∙ ^ it is not possible to deduce ^(^). However, in the“reveal” it is possible to prove that the revealed ^(^) was the authentic polynomialwhen ^(^) ∙ ^ was calculated.Computing the commitment ^(^) ∙ ^ using points ^ ^^ = ^ ∙ ^, without knowing ^,involves what is called a multi scalar multiplication (MSM), i.e. it involves computing the sum: ^ where the coefficients of polynomial ^(^) (scalar values). ^^ are points on anelliptic curve, as explained above. The multiplication (dot) operator is defined as repeated “addition” of point ^^to itself. Point addition (or more particularly point doubling) is defined as explained above. KZG commitments are cryptographic primitives which find application in different cryptographic systems and protocols. In particular, commitments are used in zero- knowledge proofs, particularly zk-snarks, which in turn have many applications. A recent example of a zero-knowledge proof scheme is PLONK, and without going into detail, PLONK requires nine MSMs. It is the computation of MSMs which dominate the computational expense of these schemes. The standard method for computing MSMs is Pippenger’s Algorithm, or the “bucket” algorithm. The bucket algorithm is applied to “reduced” MSMs, which are MSMs withscalar coefficients which are all less than some power of two 2^. In general MSMs canbe solved by breaking them down into a number of reduced MSMs, solving the reduced MSMs, and then computing the solution to the original MSM from the solutions to the reduced MSMs. This procedure of breaking down MSMs into reduced MSMs and reconstructing the results is known, and is not described in detail here. This disclosureaddresses the problem of efficiently solving reduced MSMs.The bucket algorithm comprises an ACCUMULATE stage and an AGGREGATE stage. In the ACCUMULATE stage, each point ^^is added to a bucket ^^^according to its associated scalar. In other words, all the points which are to be multiplied by the same scalar are put in the same “bucket”. Points are added to the bucket simply by adding the point to the existing point in the bucket. Once the ACCUMULATE stage has been completed, the point in each bucket needs to be multiplied by its associated scalar, and the results added together. In other wordsthe purpose of the AGGREGATE function is to compute as efficiently as possible:^ Where ^^is the point in bucket ^^following completion of the ACCUMULATE stage (that is, the sum of all points allocated to bucket ^^, or the sum of all points having associated scalar ^).The standard approach to the AGGREGATE stage is based on the observation that and hence ^ point multiplications can be replaced by 2^ point additions as follows:1. For ^ = ^ − 1,… , 1 let ^^ ← ^^^^ + ^^2. Let This method is optimal in terms of the number of required operations. The time taken for AGGREGATE is dependent on the number of buckets. However, it can be seen that the additions are sequential, and are not suited to concurrent processing. This is because each addition uses the result of the previous addition. In FPGA Acceleration of Multi-Scalar Multiplication, Jump Crypto (2022), available at https: / / github.com / z-prize / 2022-entries / blob / main / open-division / prize1-msm / prize1b- msm-fpga / nickray / zprize-fpga-msm.pdf, an FPGA-based system for computing MSMs is proposed. The authors optimise the ACCUMULATE stage by carrying out the point additions concurrently. A scheduler detects conflicting points and defers those additions to avoid the conflict. Hence parallel processing can be utilised to speed up the ACCUMULATE stage. However, the AGGREGATE stage cannot be optimised inthis way – a scheduler of the type used for the ACCUMULATE stage is no goodbecause every single point addition relies on the result of the previous point addition. In the existing scheme (optimal in terms of the number of additions), there can be no concurrent or parallel processing.Accordingly it is an object of the present invention to provide a concurrent processingsystem for computing multi scalar multiplications. STATEMENT OF INVENTION According to the present invention, there is provided a concurrent computing systemconfigured to compute a multi scalar multiplication in a cryptographic computation,the computing system having a plurality of curve adding units, each curve adding unit being capable of performing an addition of two points on an elliptic curve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being ^^^^, and the computing system further having a memory accessible to the curve adding units, and a controller, the multi scalar multiplication being in the form of: ^^^^ ∙ ^^^^^where a point on an elliptic curve and ^^ is a scalar, the maximum value of ^^being ^,and the controller being adapted to carry out the multi scalar multiplication by:defining a memory location ^^associated with each distinct value of ^^, andstoring an identity point ^ in each memory location ^^, ^^ being the point stored inmemory location ^^; for each point ^^, adding the point ^^ to the point ^^ at ^^ where ^ = ^^, andstoring the result in ^^; for some ^ < ^, sequencing point additions for all values of ^ where ^ < ^ ≤^, the point additions being:^^ ← ^^ + ^^ (i.e. add ^^ to ^^ and store the result in ^^)^^ ← ^^ + ^^(i.e. add ^^ to ^^ and store the result in ^^)where (^ + ℎ) = ^, ^ ≤ ^, ℎ ≤ ^, the point additions being sequenced suchthat the minimum distance in the sequence between two point additions acting on thesame memory location ^^ is greater than some ^, and the controller causingconcurrent execution of the sequenced point additions on each of ^ curve adding unitswhere ^ ≤ ^^^^ ;computing ^ ∙ It can be seen that the invention works using a variation on the known “bucket” algorithm. The ACCUMULATE phase involves putting each point ^^to bucket ^^, thebucket to which ^^ is added depending on the associated scalar ^^ , i.e. ^ = ^^. This isknown, and in embodiments may be optimised to make use of the concurrently operable point adding units as described in Jump (2022).However, in the AGGREGATE phase, the system of the invention is able to subsumethe points in higher-order buckets into lower-order buckets. In other words, it turns the sum into ^ ^^^^^^^In general, call the transformation from the first sum above into the second sum “bucket subsumption”.Critically, ^ < ^. In a typical embodiment, ^ is about half of ^. Hence the number ofpoint-scalar multiplications in the sum, and hence the number of point additionsrequired to solve it using known sequential addition methods, is cut in half. To achieve this, of course point additions need to be carried out. Specifically the system has tocompute ^^ ← ^^ + ^^ and ^^ ← ^^ + ^^ , for all ^ where ^ < ^ ≤ ^. The system of theinvention does not therefore reduce the total number of point additions. However, unlike in known systems, these point additions are not sequentially dependent on each other, and therefore they can be sequenced efficiently for concurrent execution. Note that the step of sequencing and concurrently executing additions may berepeated for a new ^′ < ^ , preferably ^′ is about half of ^. Hence the number of termsin the sum may be repeatedly reduced (preferably repeatedly halved), a process which is preferably repeated either until there is only one term, or the number of terms is so small that it becomes faster to use known sequential methods to solve the remaining sum rather than having the overheads associated with concurrent operation. As the number of terms reduces, so will the maximum distance at which two point additionsacting on the same memory location ^^ can be sequenced. This may require ^ to beset to less than the maximum number of available concurrent curve adding units ^^^^, to ensure concurrent operations do not conflict. Hence the number of point additions which can be performed concurrently, and the benefit of the concurrent system compared to known sequential methods, will reduce as the number of terms reduces.At some stage, embodiments may revert to solving the remaining sum ∑^^^^ ^^^ usingknown sequential methods. In some embodiments, ^ is always set to ^^^^, and oncethe number of terms in the sum reduces below the minimum for which two point additions acting on the same memory location ^^can be sequenced at a distance ofmore than ^^^^ apart, the remaining sum will have to be solved using knownsequential techniques.In principle ^ and ℎ can be any integers such that (^ + ℎ) = ^, ^ ≤ ^, ℎ ≤ ^. However,in one straightforward and efficient sequencing scheme, ^ and ℎ are selected as:for even ^: ^ = ℎ = ^^for odd ^: ^ = ^^^ ^^^^ , ℎ = ^For even ^, ^ and ℎ are the same. One implication of this is that in some embodimentspoint doubling may be supported as a single (fast) operation in the curve adding units. Hence instead of two additions: ^^ ← ^^ + ^^^ ^^^ ← ^^ + ^^^ ^a single addition plus a point doubling may be sequenced:^^ ← ^^ + 2 ⋅ ^^^ ^For odd ^, the relevant additions are: ^^^^ ← ^^^^ + ^^^ ^^^^^ ← ^^^^ + ^^^ ^In any case, choosing ^ and ℎ in this way enables a straightforward but efficientsequence for the point additions. Typically ^ = 2^ − 1 for some ^, and ^ = 2^^^.Accordingly, ^ is about half of ^, and the “top half” of the set of buckets gets subsumedinto the “bottom half” of the set of buckets. Note that in the first iteration, ^ will typicallybe one less than a power of two (because the MSM being solved is a reduced MSMas mentioned above). However, in subsequent iterations, the number of terms will be^ = 2^^^, i.e. a power of two exactly.Preferably, to sequence these point additions, they are dealt with in groups, each ofthe additions associated with odd ^ and of the form ^^^^ ← ^^^^ + ^^ being grouped in^ ^a first group, each of the additions associated with odd ^ and of the form ^^^^ ^^^^← ^^+^^ being grouped in a second group. Where point doubling operations are available inthe curve adding units, the set of additions of the form^^ ← ^^ + 2 ⋅ ^^ form a third group. Where point doubling operations are not available^ ^there may be two identical groups, a third and fourth group, each of which contains theset of additions of the form ^^ ← ^^ + ^^ .^ ^Within each group, the additions are ordered by the index of the relevant bucket. Itdoes not matter whether the order is increasing or decreasing (or indeed ordered inany other well-defined way), as long as that is consistent among the groups. The additions are then sequenced with all the additions in a single group being sequenced together. For example, the sequence could be all the additions in the first group, then all the additions in the second group, then all the additions in the third group, then allthe additions in the fourth group. But equally it could be the fourth group, then the firstgroup, then the second, then the third group, or any other order. Each group isguaranteed not to contain a conflicting addition (i.e. a group never contains two additions which write to the same bucket). Furthermore, because the additions are ordered within the group, when two groups are sequenced together, there is aminimum distance between any two conflicting additions which is at least as large asthe group.All the groups are at least as large as ^, and ^ = ^ − 1 (where ^ is one less than apower of two; in the subsequent iteration the starting number of terms ^ is exactly apower of two, the ending number of terms ^′ is exactly half of ^, and therefore all thegroups are at least as large as ^ = ^^ ^^ = ^). Accordingly, ^ can be safely set to ^. Inother embodiments, rather than setting ^, all ^^^^curve adding units are always usedand the controller ensures that concurrent operation only continues for as long as ^ ≥^^^^, that is that the groups remain sufficiently large such that the distance betweenconflicting additions in the sequence will always be at least as large as the number ofconcurrent operations. When ^ ≥ ^^^^ is no longer satisfied, the controller will ceaseconcurrent operation and solve the remaining sum by known sequential methods. Note that the description so far assumes that the concurrently operable curve addingunits are to some extent synchronised – i.e. it assumes that (say 96) curve adding unitswill all perform their addition at the same time, finish at the same time, and thentogether perform the next batch of 96 additions. Or at least, that the curve adding unitsperform an addition in predetermined timeslots. The timeslots may be overlapping rather than exactly in parallel, but the point is that the order in which add operationsstart and finish is entirely predictable, so for example it can be guaranteed that the 97thoperation will not begin until the 1stoperation has completed, the 98thoperation will notbegin until the 2nd operation has completed, and so on. In some embodiments, theconcurrently operable curve adding units may be asynchronous, and some additions may complete and free the relevant unit for another addition more quickly than others. In such a case, a scheduler may ensure that additions are deferred if they are going to conflict with pending additions which are not yet completed. The hardware of the computing system may be similar to that described in Jump (2022). In particular the key components are the plurality of curve adders, the MSM controller and the scheduler. Note that the scheduler is optional in embodiments of this invention, but may be included so that the ACCUMULATE phase can also be operatedusing concurrent additions as described in Jump, and possibly to further optimise theAGGREGATE phase as well if the curve adders are allowed to operateasynchronously. This may all be implemented in an FPGA. The curve adderspreferably support both ADD and DOUBLE (and / or DOUBLE-AND-ADD) operations.Software implementations using multi-core processors are also possible. In some embodiments, multiple devices may be used, operating in parallel. Each device may have a controller, multiple curve adders, and a memory. The devices may be able to communicate with each other to share data between devices, and may becontrolled by a master controller (the master controller may be for example aconventional PC). Advantageously, the buckets may be distributed between devices so that each device carries out mainly point additions on its own buckets, with a small number of intermediate results having to be copied between devices. In particular, let ^^^represent the buckets used with device ^1, and ^^^represent the buckets used with device ^1. As before, ^^is the point in bucket ^^. As explained, in iterations of the bucket subsumption stage, typically the starting number of buckets is a power of two, i.e. 2^(or possibly the number of buckets is oneless than a power of two 2^ − 1, but equivalently the bucket ^^^ can just be set empty,i.e. to the identity element). The ending number of buckets after one bucket subsumption iteration is half that number, i.e.2^^^. In an iteration of the bucket subsumption stage, the points in the higher order buckets (i.e. the buckets from ^^^^^^^to ^^^) are added to the points in the top half of the lower order buckets (i.e. the buckets from ^^^^^^^to ^^^^^).Therefore, if buckets are allocated to devices ^1 and ^2 as follows:^^^^ = {^^^^^^^, … , ^^∙^^^^} … ,then the points in the buckets of ^^^^are added to the points in the buckets of^^^^^^^, … , ^^∙^^^^, except for ^^^^^^^ (the point in ^^^^^^^) which is added to the point^^^^^ (the point in ^^^^^).Similarly, the points in the buckets of ^^^^ are added to the points in buckets^^∙^^^^^^ , … , ^^^^^, except for ^^∙^^^^^^ (the point in ^^∙^^^^^^) which is added to thepoint ^^∙^^^^(the point in ^^∙^^^^). Accordingly the points in buckets of ^^^^are subsumed into the points in the bucketsof ^^^^^^ , with all the necessary point additions involving points on device ^1 only, barthe exceptional cases noted above which will require copying between devices. Generally, for multiple iterations of the point subsumption, buckets can be allocated totwo devices by defining ^^^ = ^^^ ^^ ^^^ ⋃ ^^^^ ⋃…⋃ ^^^^ for ^ ∈ {1,2}, for some integer ^. The maximum value of ^ is determined by the requirement that the cardinality of ^^^^^^^^is sufficiently large as to ensure that the point additions that subsume the points in the buckets of ^^^into ^^^^^^ remain independent. As explained above, inembodiments of concurrent curve adders in use ^ may be dynamic as setsizes get smaller, or in other embodiments a constant number of curve adders ^^^^may determine an absolute maximum value of ^.More generally, for more than two devices (any number ℎ of devices where ℎ = 2^and ^ ≤ ^ − 1):^^ and as before ^ ^^ ^^ where ^^^^ is defined for any integer ^ ≤ ^ ^^: BRIEF DESCRIPTION OF THE DRAWINGS For a better understanding of the present invention, and to show more clearly how it may be carried into effect, reference will now be made by way of example only to the accompanying drawings, in which: Figure 1 is a schematic of a concurrent computing system according to the invention; Figure 2 is a schematic of another embodiment of a concurrent computing system according to the invention, comprising two of the devices of Figure 1 together with a master controller, linked together by a network bus; and Figure 3 illustrates how points are stored in buckets during use of the computing system of Figure 1 or Figure 2. DESCRIPTION OF PREFERRED EMBODIMENTSReferring firstly to Figure 1, a broad schematic of a computing system according to theinvention is shown. The system comprises ^^^^curve adding units indicated at 12, a controller 14 and a memory 16. Each curve adder is capable of adding a point on a curve to another point. Preferably, each curve adder also supports an efficient point doubling operation, i.e., adding a point to itself. Each curve adder can obtain its operands from the memory 16 and store the result of its computation (which is a new point) in the memory 16. The controller 14 allocates point additions to the curve adders 12. In onestraightforward embodiment the controller 14 will allocate a number ^ of parallel curveadditions to the adders 12, monitor for when the curve adders 12 have finished these calculations and stored their results in memory 16, and then allocate another set ofcurve additions. However, in other embodiments the curve adders may operateasynchronously. The controller may include a scheduler to avoid conflicting additions from occurring at the same time, or at overlapping times. The computing system of the invention is designed to compute a multi-scalar multiplication of the form: To do so, the controller is adapted to firstly: define a memory location ^^associated with each distinct value of ^^, and storean identity point ^ in each memory location ^^, ^^ being the point stored in memorylocation ^^; for each point ^^, add the point ^^ to the point ^^ at ^^ where ^ = ^^, and storethe result in ^^. This is known as the ACCUMULATE phase and is conventional. This phase may makeuse of concurrent multiple curve adders using the techniques disclosed in Jump(2022). The next stage is the AGGREGATE stage. The purpose of the AGGREGATE stage is to compute as efficiently as possible the sum: ^^^where ^ is the number of buckets and is typically 2^ − 1 for some ^.The key to doing this concurrently is to sequence point additions so that, subject to ^ point additions taking place concurrently, there will be no conflicting concurrent additions. This means making sure that the maximum distance between any two conflicting point additions in the sequence is greater than ^. The point additions which are sequenced are in the form: ^^ ← ^^ + ^^ (i.e. add ^^ to ^^ and store the result in ^^)^^ ← ^^ + ^^(i.e. add ^^ to ^^ and store the result in ^^).For some ^ < ^, these two additions are put in the sequence for all values of ^ where^ < ^ ≤ ^.^ and ℎ are chosen so that ^ + ℎ = ^, ^ ≤ ^, ℎ ≤ ^.This process is called POINT SUBSUMPTION. Its effect is to subsume the points in higher-order buckets into lower-order buckets. In other words, to turn the sum: into Where ^ < ^. It can be seen that this process can be repeated iteratively to continuallyreduce the number of buckets, either until there is only one bucket or until the number of buckets is below a threshold after which it is more efficient to solve the sum using traditional methods, than to accept the overheads of running concurrent additions.Typically ^ is about half of ^, both being typically powers of two (or, if ^ starts off atone less than a power of two, it may be convenient to say that ^ = 2^ and bucket ^^^just contains the identity element). A particular embodiment may sequence and carry out concurrent additions as follows:

[0002] Referring now to Figure 3, it will be seen that, according to the system set out above, the number of buckets is reduced about by half on each iteration (i.e. for each distinctvalue of ^ and each iteration of the outer while loop). In particular, the upper-orderbuckets labelled ^^^^^^^^are “subsumed” into the “top half” of the lower- order buckets, labelled ^^^ ^^^ . In the next iteration, when ^ has beenabelled ^^^^^^ and ^^ the buckets l ^^^^ will be “subsumed” into the buckets ^^^and ^^^^^^ , and so on until there are not enough buckets for concurrent operation to beworthwhile. In this embodiment, the while loop continues while ^ ≥ 2^, effectivelyensuring that the loop will continue for as long as the full concurrent capacity of the system can be used (i.e. all of the curve adders can be used at once) without sequencing a conflicting addition. In other embodiments, iterations could continue for lower values of ^, the controller reducing the number of adders in use as required in the later iterations.As shown by the arrows on Figure 3, buckets in ^^^^ are subsumed into buckets in^^^^^^ , and buckets in ^^^^ are subsumed into buckets in ^^^^^^ , apart from the exceptionalcases noted above. Therefore, if buckets ^^^ are allocated to a first device and buckets^^^ are allocated to a second device, the subsumption operations can take placeconcurrently with only minimal copying of values between devices. Figure 2 shows a suitable architecture with two devices, connected by a network bus. It will be appreciated that the idea can be extended to use more than two devices in parallel. The system according to the invention provides for concurrent computation of the AGGREGATE stage of multi scalar multiplications in elliptic curves. This represents a significant performance improvement compared to the sequential additions carried out on known systems. This has applications for cryptographic primitives, particularly zk- snarks, which rely on these multi scalar multiplications. The embodiments described above are provided by way of example only, and various changes and modifications will be apparent to persons skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

CLAIMS 1. A computing system configured to compute a multi scalar multiplication in acryptographic computation, the computing system having a plurality of curve adding units, each curve adding unit being capable of performing an addition of two points on an elliptic curve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being ^^^^, and the computing system further having a memory accessible to the curve adding units, and a controller, the multi scalar multiplication being in the form of: ^^^^ ∙ ^^^^^where a point on an elliptic curve and ^^is a scalar, the maximum value of ^^being ^, and the controller being adapted to carry out the multi scalar multiplication by: defining a memory location ^^associated with each distinct value of ^^, andstoring an identity point ^ in each memory location ^^, ^^ being the point stored inmemory location ^^; for each point ^^, adding the point ^^ to the point ^^ at ^^ where ^ = ^^, andstoring the result infor some ^ < ^, sequencing point additions for all values of ^ where ^ < ^ ≤^, the point additions being:^^ ← ^^ + ^^ (i.e. add ^^ to ^^ and store the result in ^^)^^ ← ^^ + ^^(i.e. add ^^ to ^^ and store the result in ^^)where (^ + ℎ) = ^, ^ ≤ ^, ℎ ≤ ^, the point additions being sequenced suchthat the minimum distance in the sequence between two point additions acting on the same memory location ^^is greater than some ^, and the controller causingconcurrent execution of the sequenced point additions on each of ^ curve adding unitswhere ^ ≤ ^^^^ ;computing ∑^^^^ ^ ∙ ^^ .

2. A system as claimed in claim 1, in which the step of sequencing andconcurrently executing point additions is repeated for a new ^′ < ^.

3. A system as claimed in claim 1 or claim 2, in which ^ and ℎ are selected as:for even ^: ^ = ℎ = ^^for odd ^: ^ = ^^^ ^^^^ , ℎ = ^ .

4. A system as claimed in claim 3, in which the point additions are sequenced ingroups, wherein: afirst group comprises the additions associated with odd ^ and of theform ^^^^ ← ^^^^ + ^^;^ ^a second group comprises the additions associated with odd ^ and ofthe forma third and optionally a fourth group comprises the additions associated with even ^, in which the additions are ordered within each group, and in which the additions are sequenced with all the additions in a single group being sequenced together.

5. A system as claimed in any of the preceding claims, further comprising ascheduler for preventing conflicts between sequenced additions executed concurrently.

6. A system as claimed in any of the preceding claims, in which the plurality ofcurve adding units are provided as a first set of curve adding units on a first device, and a second set of curve adding units on a second device, the first device having a first memory associated with and accessible to the first set of curve adding units and the second device having a second memory associatedwith and accessible to the second set of curve adding units, and the first and second devices being linked together by a communication means.

7. A system as claimed in claim 6, in which memory locations in some subset ^^^of ^ = {^^, … , ^^} are in the memory of the first device and memory locationsin another subset of ^, ^^^, are in the memory of the second device, and in which additions of the form ^^ ← ^^ + ^^ and ^^ ← ^^ + ^^ are sequenced onthe first device for ^ ∈ ^^^ ^^^ and on the second device for ^^ ∈ ^ .

8. A system as claimed in claim 7, in which:^^^and where ^ ∈{1, 2, … , ℎ} for some integer ℎ, wherein ℎ devices are provided, each devicehaving a plurality of curve adding units and a memory accessible by the curve adding units of the device, and the devices being linked together by a communication means, and in which memory locations in ^^^are in the memory of the ^th device and in which additions of the form ^^ ← ^^ + ^^ and^ ← th ^^^ ^^ + ^^ are sequenced on the ^ device for ^^ ∈ ^ .

9. A non-transient computer-readable medium having instructions thereon which,when executed on a suitable machine, cause the machine to carry out a multi scalar multiplication in a cryptographic computation,in which the machine includes a plurality of curve adding units, each curve adding unit being capable of performing an addition of two points on an elliptic curve, the curve adding units being operable concurrently, the total number of concurrently operable curve adding units being ^^^^, and the machine further comprises a memory accessible to the curve adding units, and a controller, and in which the multi scalar multiplication is of the form^^^where ^^ is a point on an elliptic curve and ^^ is a scalar, the maximum value of^^ being ^,and in which the multi scalar multiplication is carried out by: defining a memory location ^^associated with each distinct value of ^^, and storing an identity point ^ in each memory location ^^, ^^ being the point storedin memory location ^^; for each point ^^, adding the point ^^ to the point ^^ at ^^ where ^ = ^^, andstoring the result in ^^; for some ^ < ^, sequencing point additions for all values of ^ where ^ < ^ ≤^, the point additions being:^^ ← ^^ + ^^ (i.e. add ^^ to ^^ and store the result in ^^)^^ ← ^^ + ^^(i.e. add ^^ to ^^ and store the result in ^^)where (^ + ℎ) = ^, ^ ≤ ^, ℎ ≤ ^, the point additions being sequenced suchthat the minimum distance in the sequence between two point additions acting on the same memory location ^^is greater than some ^, and the controller causing concurrent execution of the sequenced point additions on each of ^ curve adding units where ^ ≤ ^^^^;computing10. A computer readable medium as claimed in claim 9, in which the step ofsequencing and concurrently executing point additions is repeated for a new^′ < ^.

11. A computer readable medium as claimed in claim 9 or claim 10, in which ^ andℎ are selected as:for even ^: ^ = ℎ = ^^for odd ^: ^ = ^^^ ^^^^ , ℎ = ^ .

12. A computer readable medium as claimed in claim 11, in which the pointadditions are sequenced in groups, wherein:a first group comprises the additions associated with odd ^ and of theform ^^^^ ← ^^^^ + ^^;^ ^a second group comprises the additions associated with odd ^ and ofthe form ^^^^ ← ^^^^ + ^^;^ ^a third and optionally a fourth group comprises the additions associated with even ^, in which the additions are ordered within each group, and in which the additions are sequenced with all the additions in a single group being sequenced together.