Method and apparatus for implementing large-number operation, and addition operation unit, subtraction operation unit and shifter

By adding hidden bits to the destination register to store carry or borrow values ​​and loading them into the carry input bits when needed, the problem of too many addition and carry addition operations in large numbers in Montgomery domain is solved, and the calculation efficiency is improved.

WO2025035472A9PCT designated stage expired Publication Date: 2025-06-26SUNLUNE (SINGAPORE) PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/113626
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-08-17
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When performing large-number operations in the Montgomery domain, a large number of addition and carry addition operations are required in the prior art, resulting in waste of hardware resources and reduced computing efficiency.

Method used

Reduce unnecessary ADC or SBB operations by adding a hidden bit to the destination register, storing the carry value of the ADD operation or the borrow value of the SUB operation, and loading it into the carry input bit of the addition or subtraction component by the instruction getMGBtoStat srcReg when needed.

Benefits of technology

Reduces the overall total number of computation instructions, improves computational efficiency, and avoids the collision of carry bits in the status register.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023113626_26062025_PF_FP_ABST
    Figure CN2023113626_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application is a method for implementing a large-number operation. The method comprises: in a first iterative operation, storing a carry value of an ADD operation in the most significant bit of a destination register, wherein the destination register is an (L+1)-bit register, which is used for storing the calculation result of the ADD operation; and in iterative operations other than the first iterative operation, corresponding to the ADD operation or an ADC operation in the previous iterative operation, taking out a carry value of the previous iteration, using the ADC operation to perform addition processing on two pieces of L-bit data having the taken-out carry value, and storing a carry value of the ADC operation in the most significant bit of the destination register.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for realizing large number operation, addition and subtraction operator and shifter Technical Field

[0001] The present application relates to, but is not limited to, large number arithmetic technology, and in particular to a method, device, addition and subtraction operators, and shifters for implementing large number arithmetic. Background Art

[0002] Zero-knowledge proofs (ZKPs) were proposed by S. Goldwasser, S. Micali, and C. Rackoff in the early 1980s. As a highly secure encryption technology, ZKPs hold broad application prospects in future information transmission. ZKPs involve a large number of finite field calculations, namely, addition, subtraction, and multiplication operations bounded by a very large prime number. To simplify the computational process and eliminate the most complex division and remainder calculations, related technologies typically convert values ​​from a finite field to a Montgomery field, where the corresponding addition, subtraction, multiplication, and squaring operations are performed. Finally, the results are transferred from the Montgomery field back to the corresponding finite field.

[0003] In practical applications, the modulus corresponding to the finite field, i.e., the prime number, is very large (more than several hundred bits). The corresponding Montgomery multiplication on a general-purpose 32-bit or 64-bit processor will truncate the multiplier and multiplicand into several 32-bit or 64-bit parts before performing decomposition calculations.

[0004] In the related art Montgomery multiplication, to add two large numbers, a large number of addition (ADD) and carry-add (ADC) operations are required. This significantly increases hardware resources, increases the total number of computational instructions, and reduces computational efficiency. In the related art squaring of large numbers, a large number of 2-multiplications (multiplying two identical unsigned integers by 2) are involved. This requires a large number of left shifts and bitwise OR operations, which increases the total number of computational instructions and reduces computational efficiency.

[0005] SUMMARY OF THE INVENTION

[0006] The present application provides a method for implementing large number operations, and an addition and subtraction operator and a shifter device, which can solve any of the above technical problems.

[0007] The present invention provides a method for implementing large number operations, including:

[0008] In the first iteration of the Montgomery field multiplication operation, the carry value of the ADD operation is stored in the highest bit of the destination register, wherein the ADD operation is used to add two L-bit data without adding the carry bit; the destination register is a (L+1)-bit register used to store the calculation result of the ADD operation;

[0009] In the remaining iterative operations except the first iterative operation, corresponding to the ADD operation or ADC operation in the previous iterative operation, the carry value of the previous iteration is taken out, and the ADC operation is used to implement the addition of the carry values ​​taken out of the two L-bit data bands, and the carry value of the ADC operation is stored in the highest bit of the destination register.

[0010] In an exemplary embodiment, extracting the carry value of the previous iteration includes:

[0011] The instruction getMGBtoStat srcReg is used to copy the information in the (L+1)th bit hidden in the source register srcReg to the status register Stat;

[0012] Wherein, the source register srcReg is the destination register for storing the calculation result after the addition calculation in the corresponding previous iterative operation;

[0013] The status register Stat is a 1-bit register used to store a carry addition flag.

[0014] In an exemplary embodiment, the Montgomery multiplication includes: implementing Montgomery multiplication of two M-bit data using an L-bit processor; wherein M is greater than L; and M varies according to the size of the finite field data selected by the algorithm.

[0015] The present application also provides a method for implementing large number operations, including:

[0016] In the first iteration, the borrowed value of the SUB operation is stored in the highest bit of the destination register, where the SUB operation is used to implement the subtraction of two L-bit data without borrowing. The destination register is a (L+1)-bit register used to store the calculation result of the SUB operation.

[0017] In the remaining iterative operations except the first iterative operation, corresponding to the SUB operation or SBB operation in the previous iterative operation, the borrow value of the previous iteration is taken out, and the SBB operation is used to implement the subtraction of the borrow values ​​taken out of the two L-bit data bands, and the borrow value of the SBB operation is stored in the highest bit of the destination register.

[0018] In an exemplary embodiment, extracting the borrow value of the previous iteration includes:

[0019] The instruction getMGBtoStat srcReg is used to copy the information in the (L+1)th bit hidden in the source register srcReg to the status register Stat;

[0020] Wherein, the source register srcReg is the destination register for storing the calculation result after the subtraction calculation in the corresponding previous iterative operation;

[0021] The status register Stat is a 1-bit register used to store a flag bit of a borrow subtraction.

[0022] The embodiment of the present application further provides a device for implementing large number operations, comprising: an operation unit and a destination register; wherein,

[0023] an operation unit for implementing an ADD / ADC operation on two L-bit data; in a first iterative operation, storing a carry value of the ADD operation in the highest bit of a destination register; in the remaining iterative operations except the first iterative operation, corresponding to the ADD operation or ADC operation in the previous iterative operation, extracting the carry value of the previous iteration, performing an addition operation with the extracted carry value by using an ADC operation, and storing the carry value of the ADC operation in the highest bit of the destination register;

[0024] The destination register is a (L+1)-bit register used to store the calculation result of the ADD / ADC operation, and the carry value is stored in the (L+1)th bit;

[0025] or,

[0026] an arithmetic unit for implementing a SUB / SBB operation on two L-bit data; in a first iterative operation, storing a borrow value from the SUB operation in the highest bit of a destination register; in the remaining iterative operations except the first iterative operation, corresponding to the SUB operation or SBB operation in the previous iterative operation, extracting the borrow value from the previous iteration, performing a subtraction with the extracted carry value using an SBB operation, and storing the carry value from the SBB operation in the highest bit of the destination register;

[0027] The destination register is a (L+1)-bit register used to store the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L+1)th bit.

[0028] In one exemplary embodiment, the device further includes: a status register;

[0029] The operation unit extracts the carry value / borrow value of the previous iteration by using the instruction getMGBtoStat srcReg to copy the information in the (L+1)th bit hidden in the source register srcReg to the status register. The source register srcReg is the destination register.

[0030] The embodiment of the present application further provides an adder, comprising: an operation unit, a destination register, and a status register; wherein,

[0031] an operation unit for implementing an ADD / ADC operation on two L-bit data; in a first iterative operation, storing a carry value of the ADD operation in the highest bit of a destination register; in remaining iterative operations except the first iterative operation, corresponding to an ADD operation or an ADC operation in a previous iterative operation, taking out the carry value of the previous iteration and storing it in a status register, performing an addition operation with the carry value in the status register using an ADC operation, and storing the carry value of the ADC operation in the highest bit of the destination register;

[0032] The destination register is a (L+1)-bit register used to store the calculation result of the ADD / ADC operation, and the carry value is stored in the (L+1)th bit;

[0033] The status register is used to store carry values ​​corresponding to the ADD operation or ADC operation in the previous iterative operation in the remaining iterative operations except the first iterative operation.

[0034] The embodiment of the present application further provides a subtraction operator, comprising: an operation unit, a destination register, and a status register; wherein,

[0035] an operation unit for implementing a SUB / SBB operation on two L-bit data; in a first iterative operation, storing a borrow value from the SUB operation in the highest bit of a destination register; in remaining iterative operations except the first iterative operation, corresponding to a SUB operation or SBB operation in a previous iterative operation, taking out the borrow value from the previous iteration and storing it in a status register, performing a subtraction operation with the borrow value in the status register using an SBB operation, and storing the borrow value from the SBB operation in the highest bit of the destination register;

[0036] The destination register is a (L+1)-bit register used to store the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L+1)th bit;

[0037] The status register is used to store the borrow value of the SUB operation or SBB operation in the previous iterative operation in the remaining iterative operations except the first iterative operation.

[0038] The embodiment of the present application further provides a device for implementing large number operations, comprising: a plurality of first operation units and a plurality of second operation units; wherein,

[0039] a first operation unit, configured to store a carry value resulting from an ADD operation in a highest bit of a destination register in a first iteration of a Montgomery field multiplication operation, wherein the ADD operation is configured to add two L-bit data; and the destination register is a (L+1)-bit register configured to store a calculation result of the ADD operation;

[0040] a second operation unit configured to, in the remaining iterations of the Montgomery field multiplication operation except the first iteration, extract a carry value from the previous iteration corresponding to an ADD operation or an ADC operation in the previous iteration, perform addition processing with the extracted carry value using an ADC operation, and store the carry value from the ADC operation in the most significant bit of the destination register;

[0041] or,

[0042] a first operation unit, configured to store, in a first iteration operation, a borrow value of a SUB operation in the highest bit of a destination register, wherein the SUB operation is used to implement subtraction of two L-bit data; the destination register is an (L+1)-bit register, configured to store a calculation result of the SUB operation;

[0043] The second operation unit is used to, in the remaining iterative operations except the first iterative operation, correspond to the SUB operation or SBB operation in the previous iterative operation, extract the borrow value of the previous iteration, use the SBB operation to implement subtraction processing with the extracted borrow value, and store the borrow value of the SBB operation in the highest bit of the destination register.

[0044] In an exemplary embodiment, extracting the carry value of the previous iteration in the second operation unit includes: using the instruction getMGBtoStat srcReg to copy the information in the (L+1)th bit hidden in the source register srcReg to the status register Stat;

[0045] Wherein, the source register srcReg is the destination register for storing the calculation result after the addition calculation in the corresponding previous iterative operation;

[0046] The status register Stat is a 1-bit register used to store a carry addition flag.

[0047] The present application further provides a method for implementing large number operations, which includes:

[0048] Perform a multiplication operation on the two numbers involved in the 2-times multiplication operation, shift the high-order register of the multiplication result left by one bit, store the shifted-out information in the highest hidden bit of the high-order register, and retrieve the information from the hidden bit of the high-order register and store it in the lowest bit of the third destination register; the high-order register is a (L+1)-bit register used to store the highest bit of the multiplication result;

[0049] Shift the low-order register of the multiplication result left by one bit, save the information of the shifted-out bit in the hidden bit of the low-order register, and extract the information in the hidden bit of the low-order register and store it in the lowest bit of the second destination register; the low-order register is a (L+1)-bit register, which is used to store the low bit of the multiplication result;

[0050] A bitwise AND operation is performed on the second destination register and the high-order register of the multiplication result, and the result is stored in the second destination register; the information of the low-order register of the multiplication result is moved to the first destination register; the first destination register is used to store the low-order information of the 2x multiplication result, the second destination register is used to store the high-order information of the 2x multiplication result, and the third destination register is used to store the carry value of the 2x multiplication result.

[0051] In an exemplary embodiment, the instruction getMGBtoReg drcReg,srcReg is used to retrieve the information in the hidden bit and store it in the third destination register / the second destination register;

[0052] The instruction getMGBtoReg drcReg, srcReg indicates copying the data stored in the (L+1)th bit of the source register srcReg to the lowest bit of the destination register drcReg.

[0053] An embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the method for implementing large number operations described in any one or any combination of the embodiments of the present application.

[0054] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of the method for implementing large number operations described in any one or any combination of the embodiments of the present application.

[0055] The embodiment of the present application further provides a shifter, comprising: a shifter, a result register of (L+1) bits, and a destination register; wherein,

[0056] The shifter is used to shift the L-bit input information according to the specified shift number, store the shift result in the result register, and copy the information in the (L+1)th bit hidden in the result register to the lowest bit of the destination register.

[0057] In an exemplary embodiment, the instruction getMGBtoReg drcReg,srcReg is used to retrieve the information in the hidden (L+1)th bit and store it in the destination register;

[0058] The instruction getMGBtoReg drcReg,srcReg indicates that the data stored in the (L+1)th bit of the result register serving as the source register srcReg is copied to the lowest bit of the destination register serving as the destination register drcReg.

[0059] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0060] Summary of the Figures

[0061] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.

[0062] FIG1 is a schematic diagram of an implementation process of Montgomery multiplication in an embodiment of the present application;

[0063] FIG2 is a flow chart of a first embodiment of a method for implementing large number operations in an embodiment of the present application;

[0064] FIG3 is a flow chart of a second embodiment of a method for implementing large number operations in an embodiment of the present application;

[0065] FIG4 is a schematic diagram of the composition structure of a first embodiment of a device for implementing large number operations in an embodiment of the present application;

[0066] FIG5 is a schematic diagram of the structure of a second embodiment of a device for implementing large number operations in an embodiment of the present application;

[0067] FIG6 is a flow chart of a third embodiment of a method for implementing large number operations according to an embodiment of the present application;

[0068] FIG7 is a schematic diagram of the composition structure of a shifter for implementing large number operations in an embodiment of the present application.

[0069] Details

[0070] To make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of this application can be combined with each other in any way.

[0071] To facilitate understanding of the present application, the present application will be described more fully below with reference to the accompanying drawings. The accompanying drawings provide embodiments of the present application. However, the present application may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to make the disclosure of the present application more thorough and comprehensive.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0073] FIG1 is a schematic diagram illustrating an implementation process of a Montgomery multiplication in an embodiment of the present application. As shown in FIG1 , a Montgomery multiplication of two Mbit data, such as 256-bit data, is implemented using an L-bit processor, such as a 64-bit processor, where M is greater than L. The embodiment shown in FIG1 is a Montgomery multiplication operation process of two 256-bit data in 64-bit units. The implementation process, as shown in FIG1 , includes a total of four iterations, each with the same structure. FIG1 only illustrates the first iteration. In FIG1 , each minimum rectangle is a 64-bit computation unit. Thus, each 256-bit data can be represented as four 64-bit computation units combined together.

[0074] In Figure 1, the variable result represents the result of the operation and is initialized to all 0s at the beginning; the variable ptr_self represents the multiplicand (256 bits); the variable other represents the multiplier, which is actually 256 bits, but is actually divided into 4 iterations, each time processing 64 bits from low to high. Figure 1 shows the first iteration, so the variable other represents the low 64 bits of the multiplier; the variable po (320 bits) is the product of the multiplicand (256 bits) and a part of the multiplier (64 bits); the variable por (320 bits) is the sum of the variable po (320 bits) and the variable result (256 bits); the low 64 bits of the variable por are taken and multiplied by the constant INV (64 bits) to obtain the variable K (128 bits); the low 64 bits of the variable K are multiplied by the constant MODU (256 bits) to obtain the variable MK (320 bits); the high 256 bits of the variable MK are added to the high 256 bits of the intermediate variable por, and the result of the addition updates the variable result. In this way, the other bits of the multiplier are processed according to the same process as shown in Figure 1, each time 64 bits are processed. In this way, after four iterative operations, the value of the variable resule is two 256-bit data, which is the result of the Montgomery multiplication in 64-bit units.

[0075] In the program implementation of the Montgomery multiplication operation, in order to save temporary memory resources, the variable ptr_self is multiplied by the variable other, and the variable K is multiplied by the constant MODU at the same time. In the implementation process of this method, many calculations similar to: d = a + b + c will appear, where a and d are logically 128 bits, and b and c are 64 bits. The implemented calculation is uint128d = uint128a + uint64b + uint64c.

[0076] In a typical 64-bit computer architecture, there is no 128-bit computing component, so it is implemented through software, such as uint128_t a will be represented as uint128_t a[2]. Later, in an architecture with more instruction optimization, the code for d=a+b+c (the following four operations) can be implemented by combining the addition instruction ADD (that is, the ADD operation represents the addition of two numbers without adding a carry bit) and the addition with carry instruction ADC (that is, the ADD operation represents the addition of two numbers with a carry bit):

[0077] ADD d[0],a[0],b / / Add the lower 64 bits of b and a, assign the result to d[0], and retain the carry;

[0078] ADC d[1],a[1],0 / / Add the carry bit to the high 64 bits of a and assign the result to d[1];

[0079] ADD d[0],d[0],c / / Add the lower 64 bits of c and d, and then assign the result to d[0] while retaining the carry;

[0080] ADC d[1],d[1],0 / / Add the carry bit to the high 64 bits of d, and then assign the result to d[1].

[0081] As can be seen from the code implementing d = a + b + c, the ADC operation only stores the carry information from the previous ADD or ADC operation. In other words, the carry bit input value for the current beat cannot store more information from earlier times. Therefore, an ADC operation is required after each ADD operation to handle the carry. If the carry is 0, the ADC operation is unnecessary. The addition of half the number of zero-addition operations actually makes the entire operation inefficient.

[0082] In conjunction with the embodiment shown in FIG1 , the Montgomery multiplication fragment, as shown in FIG1 , can be expressed as:

[0083] The above Montgomery multiplication fragment can be expressed as follows using pseudo-assembly code, where there are two calculations of the form d = a + b + c, as shown in the bold and underlined font in the following code:

[0084] In the above code, in order to perform a 128-bit addition with a 64-bit number, a carry-add operation must be added between the upper 64 bits and 0, i.e., ADC temp[1], temp[1], 0. However, if the carry is 0, the ADC operation is unnecessary, and the added ADC operation will reduce the efficiency of the entire operation. To reduce unnecessary ADC carry-add operations and the total number of overall computational instructions, the present embodiment proposes a method for implementing large number operations based on the Montgomery field. This method aims to reduce the total number of computational instructions for finite field computations without significantly increasing hardware resources, thereby improving computational efficiency.

[0085] FIG2 is a flow chart of a first embodiment of a method for implementing large number operations in an embodiment of the present application. FIG2 shows a typical loop calculation process. A typical loop content, corresponding to a 256-bit processor corresponding to a 64-bit processor, requires four loops. As shown in FIG2 , based on a Montgomery domain operation scenario, it may include:

[0086] Step 200: In the first iteration of the Montgomery field multiplication operation, the carry value of the ADD operation is stored in the highest bit of the destination register, wherein the ADD operation is used to implement the addition of two L-bit data without adding the carry bit; the destination register is an (L+1)-bit register used to store the calculation result of the ADD operation.

[0087] In an exemplary embodiment, Montgomery multiplication can be Montgomery multiplication of two M-bit data implemented by an L-bit processor; wherein M is greater than L; M will vary depending on the size of the finite field data selected by the algorithm, ranging from tens of bits to thousands of bits. In one embodiment, the L-bit processor can be a 64-bit processor or a 32-bit processor; the M-bit data can be 256-bit data. It should be noted that in the method for implementing addition provided in the embodiment of the present application, L includes but is not limited to 32 bits, 64 bits, 128 bits, 256 bits, etc. M includes but is not limited to 256, and varies from tens of bits to thousands of bits depending on different algorithms and the selected finite field value. In one embodiment, for example, for a 384-bit finite field calculation, if a 64-bit processor is used, then the data involved in the calculation can be divided into 6 equal parts of 64-bit data.

[0088] In the embodiment of the present application, the width of the destination register is expanded by adding a hidden bit to the highest bit to store the carry value of the ADD operation. That is, the width of the destination register is changed to (L+1) bits, but the processing width is still L bits.

[0089] In one embodiment, when performing unsigned addition calculation, when two L-bit data are added, the result (L+1) bits are assigned to the destination register, which is equivalent to storing the calculation result of the two L-bit data together with the 1-bit carry into the (L+1)-bit destination register.

[0090] Still taking the above-mentioned Montgomery multiplication fragment as an example, according to step 200 of the embodiment of the present application, the pseudo assembly code in the first iteration operation can be represented as follows, where each ADD operation in the first iteration operation is shown in bold and underlined font:

[0091] Among them, the processor is a 64-bit processor. Therefore, register temp1 is a (64+1)=65-bit register, whose highest bit is a hidden bit, used to store the carry value of the first ADD operation; register temp2 is a (64+1)=65-bit register, whose highest bit is a hidden bit, used to store the carry value of the second ADD operation; register temp3 is a (64+1)=65-bit register, whose highest bit is a hidden bit, used to store the carry value of the third ADD operation; register temp4 is a (64+1)=65-bit register, whose highest bit is a hidden bit, used to store the carry value of the fourth ADD operation.

[0092] From the first iterative operation after step 200, it can be seen that the carry addition operation of the upper 64 bits and 0 required for the addition of 128 bits and 64 bits, i.e., ADC temp[1], temp[1], 0, is omitted. Instead, the carry value of each ADD operation is directly stored in the expanded destination register.

[0093] Step 201: In the remaining iterative operations of the Montgomery domain multiplication operation except the first iterative operation, corresponding to the ADD operation or ADC operation in the previous iterative operation, the carry value of the previous iteration is extracted, and the ADC operation is used to implement the addition of the extracted carry values ​​of two L-bit data bands, and the carry value of the ADC operation is stored in the highest bit of the destination register.

[0094] In one exemplary embodiment, the carry value from the previous iteration in step 201 can be retrieved using the instruction getMGBtoStat srcReg, which copies the information hidden in the (L+1)th bit of the source register srcReg to the status register Stat. The source register srcReg is the destination register that stores the result of the addition calculation in the previous iteration. The status register Stat can be a 1-bit register used to store a carry flag.

[0095] It should be noted that the instruction getMGBtoStat srcReg extracts the information from the hidden bit and places it into a flag bit in the processor, namely, the 1-bit carry register, for subsequent carry-addition (ADC) or borrow-subtraction (SBB) operations. In other words, when performing an ADC operation, the information in the 1-bit register is input into the 1-bit carry or borrow input of the adder or subtractor. This processing in the embodiment of the present application continues to utilize the 1-bit carry register, and therefore does not significantly alter the traditional computer architecture.

[0096] Still taking the above-mentioned Montgomery multiplication fragment as an example, according to step 201 of the embodiment of the present application, the remaining iterative operations except the first iterative operation can be represented as follows using pseudo assembly code. In the remaining iterative operations except the first iterative operation, each of the carry values ​​of the previous iteration and the ADC operation is shown in bold and underlined font:

[0097] In the above procedure Taking this as an example, the low-bit addition with the carry value generated in the previous iteration is implemented, that is, the carry value stored in the hidden bit of temp1 is used as the carry input of the ADC operation to implement the addition with carry. At the same time, the current carry value is stored in the hidden bit of temp1. At this time, the value in the hidden bit of temp1 will overwrite the carry value stored in the previous iteration and become the carry value generated in this iteration.

[0098] The method for implementing large number operations provided by the embodiment shown in Figure 2 of the present application, based on the Montgomery domain operation scenario, eliminates unnecessary ADC operations. For example, there is no need to add a carry addition operation between the upper 64 bits and 0 in order to perform addition of 128 bits and 64 bits, thereby reducing the total number of overall calculation instructions and improving calculation efficiency.

[0099] The present application also provides a computer-readable storage medium storing computer-executable instructions for executing the method for implementing large number operations described in any one of FIG. 2 .

[0100] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of the method for implementing large number operations as described in any one of Figure 2.

[0101] The present application also provides a method for implementing large number operations, similar to the addition operation shown in FIG2 , as shown in FIG3 , including:

[0102] Step 300: In the first iterative operation, the borrow value of the SUB operation is stored in the highest bit of the destination register, where the SUB operation is used to implement the subtraction of two L-bit data without borrow; the destination register is an (L+1)-bit register, which is used to store the calculation result of the SUB operation.

[0103] In one embodiment, when performing unsigned subtraction, when two L-bit data are subtracted, the (L+1)-bit result is assigned to the destination register. This is equivalent to storing the calculation result of the two L-bit data together with the 1-bit sign bit in the (L+1)-bit destination register.

[0104] Step 301: In the remaining iterative operations except the first iterative operation, corresponding to the SUB operation or SBB operation in the previous iterative operation, the borrow value of the previous iteration is taken out, and the SBB operation is used to implement the subtraction of the borrow values ​​taken out of two L-bit data bands, and the borrow value of the SBB operation is stored in the highest bit of the destination register.

[0105] In one exemplary embodiment, the borrow value from the previous iteration in step 201 can be retrieved using the instruction getMGBtoStat srcReg, which copies the information hidden in the (L+1)th bit of the source register srcReg to the status register Stat. The source register srcReg is the destination register that stores the result of the subtraction calculation in the previous iteration. The status register Stat can be a 1-bit register used to store a flag bit for the borrow subtraction.

[0106] The method for implementing large number operations using the Montgomery domain provided by the embodiment shown in FIG3 of the present application reduces the total number of overall computing instructions, thereby improving computing efficiency.

[0107] The method for implementing large number operations provided by the embodiment of the present application provides a hidden bit to save the carry value of addition or the borrow value of subtraction by expanding the width of the destination register. When the carry value or borrow value is needed, it can be loaded into the carry input bit of the addition unit by simply using the instruction getMGBtoStat srcReg.

[0108] The present application also provides a computer-readable storage medium storing computer-executable instructions for executing the method for implementing large number operations described in any one of FIG. 3 .

[0109] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of the method for implementing large number operations as described in any one of Figure 3.

[0110] In an exemplary embodiment, considering the process of implementing addition in Montgomery field multiplication, the method for implementing large number operations provided by the embodiments of the present application eliminates many ADC reg,reg,0 operations for simply carrying to higher bits. At the same time, since the carry value of the addition is directly assigned to the hidden bit of the destination register, the problem of carry bit conflict in the status register is also effectively avoided.

[0111] This is illustrated by the following two instructions:

[0112] ADD A,B,C / / A=B+C, carry value is stored in status register

[0113] ADD D,E,F / / D=E+F, carry value is stored in status register

[0114] Obviously, the execution of these two ADD instructions will cause a status register carry bit conflict due to the write-after-write dependency of the status register carry.

[0115] However, if the method for implementing large number operations based on the Montgomery field provided in the embodiment of the present application is adopted, 1 most significant bit is added to register A as a hidden bit, and 1 most significant bit is added to register D as a hidden bit, the above two instructions will be transformed into:

[0116] ADD A,B,C / / A=B+C, the carry value is stored in the highest hidden bit of register A

[0117] ADD D,E,F / / D=E+F, the carry value is stored in the highest hidden bit of register D

[0118] The carry values ​​of the two additions are stored in the newly added highest hidden bit of their respective destination registers, and will not affect each other, so there is no correlation conflict problem.

[0119] Consider the following three instructions:

[0120] ADD A,B,C / / A=B+C

[0121] ADC G,A,C / / G=A+C

[0122] ADD D,E,F / / D=E+F

[0123] Although the instructions ADD D, E, and F have nothing to do with the previous two instructions, they cannot be executed in advance due to resource conflicts in the status register carry bit. However, if the method for implementing large number operations based on the Montgomery field provided in the embodiment of the present application is adopted, the most significant bit of register A is increased by 1 as a hidden bit, and the most significant bit of register G is increased by 1 as a hidden bit, the above three instructions will be changed to the following:

[0124] ADD A,B,C / / A=B+C, the carry value is stored in the highest hidden bit of register A

[0125] GetMGBtoStat A / / Get the carry value in the highest hidden bit of register A and store it in the status register Stat

[0126] ADC G,A,C / / G=A+C / / Carry addition with the carry value taken from the highest hidden bit of register A, and store the calculated carry value in the highest hidden bit of register G

[0127] ADD D,E,F / / D=E+F

[0128] After this processing, the instructions ADD D, E, and F can be executed in parallel, improving processor efficiency.

[0129] The present application also provides a device for implementing large number operations. FIG4 is a schematic diagram of the composition structure of the first embodiment of the device for implementing large number operations in the present application. As shown in FIG4 , the device includes: an operation unit and a destination register; wherein,

[0130] an operation unit for implementing an ADD / ADC operation on two L-bit data; in a first iterative operation, storing a carry value of the ADD operation in the highest bit of a destination register; in the remaining iterative operations except the first iterative operation, corresponding to the ADD operation or ADC operation in the previous iterative operation, extracting the carry value of the previous iteration, performing an addition operation with the extracted carry value by using an ADC operation, and storing the carry value of the ADC operation in the highest bit of the destination register;

[0131] The destination register is a (L+1)-bit register used to store the calculation result of the ADD / ADC operation, and the carry value is stored in the (L+1)th bit.

[0132] or,

[0133] an arithmetic unit for implementing a SUB / SBB operation on two L-bit data; in a first iterative operation, storing a borrow value from the SUB operation in the highest bit of a destination register; in the remaining iterative operations except the first iterative operation, corresponding to the SUB operation or SBB operation in the previous iterative operation, extracting the borrow value from the previous iteration, performing a subtraction with the extracted carry value using an SBB operation, and storing the carry value from the SBB operation in the highest bit of the destination register;

[0134] The destination register is a (L+1)-bit register used to store the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L+1)th bit.

[0135] In one exemplary embodiment, as indicated by a dotted arrow in FIG4 , the carry / borrow value from the previous iteration in the arithmetic unit can be retrieved using the instruction getMGBtoStat srcReg , thereby copying the information in the (L+1)th bit hidden in the source register srcReg to the status register Stat . Here, the source register srcReg is the destination register.

[0136] The embodiment of the present application further provides an adder, as shown in FIG4 , comprising: an operation unit, a destination register, and a status register; wherein,

[0137] an operation unit for implementing an ADD / ADC operation on two L-bit data; in a first iterative operation, storing a carry value of the ADD operation in the highest bit of a destination register; in remaining iterative operations except the first iterative operation, corresponding to an ADD operation or an ADC operation in a previous iterative operation, taking out the carry value of the previous iteration and storing it in a status register, performing an addition operation with the carry value in the status register using an ADC operation, and storing the carry value of the ADC operation in the highest bit of the destination register;

[0138] The destination register is a (L+1)-bit register used to store the calculation result of the ADD / ADC operation, and the carry value is stored in the (L+1)th bit;

[0139] The status register is used to store carry values ​​corresponding to the ADD operation or ADC operation in the previous iterative operation in the remaining iterative operations except the first iterative operation.

[0140] It should be noted that the registers used in the calculations herein are all (L+1)-bit registers. However, when used as source registers in the relevant calculations, only the L bit is included in the calculation. In other words, the hidden bit, the (L+1)th bit, is not included in the calculation. The hidden bit is only used to store additional information when the register is used as the destination register, and can be read using the GetMGBtoStat and GetMGBtoReg instructions in the embodiments of this application.

[0141] The embodiment of the present application provides a subtraction operator, as shown in FIG4 , comprising: an operation unit, a destination register, and a status register; wherein,

[0142] an operation unit for implementing a SUB / SBB operation on two L-bit data; in a first iterative operation, storing a borrow value from the SUB operation in the highest bit of a destination register; in remaining iterative operations except the first iterative operation, corresponding to a SUB operation or SBB operation in a previous iterative operation, taking out the borrow value from the previous iteration and storing it in a status register, performing a subtraction operation with the borrow value in the status register using an SBB operation, and storing the borrow value from the SBB operation in the highest bit of the destination register;

[0143] The destination register is a (L+1)-bit register used to store the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L+1)th bit;

[0144] The status register is used to store the borrow value of the SUB operation or SBB operation in the previous iterative operation in the remaining iterative operations except the first iterative operation.

[0145] FIG5 is a schematic diagram of the composition structure of the second embodiment of the device for implementing large number operations in the embodiment of the present application. As shown in FIG5 , the device may include: a plurality of first operation units and a plurality of second operation units; wherein,

[0146] a first operation unit, configured to store a carry value resulting from an ADD operation in the highest bit of a destination register during a first iteration of a Montgomery field multiplication operation, wherein the ADD operation is configured to add two L-bit data (i.e., both source registers are physically (L+1) bits, but their hidden bits do not participate in the normal operation); and the destination register is a (L+1)-bit register configured to store a calculation result of the ADD operation;

[0147] The second operation unit is configured to, in the remaining iteration operations except the first iteration operation of the Montgomery field multiplication operation, corresponding to the ADD operation or the ADC operation in the previous iteration operation, extract the carry value of the previous iteration, implement addition processing with the extracted carry value by using the ADC operation, and store the carry value of the ADC operation in the highest bit of the destination register.

[0148] or,

[0149] a first operation unit, configured to store, in a first iteration operation, a borrow value of a SUB operation in the highest bit of a destination register, wherein the SUB operation is used to implement subtraction of two L-bit data; the destination register is an (L+1)-bit register, configured to store a calculation result of the SUB operation;

[0150] The second operation unit is used to, in the remaining iterative operations except the first iterative operation, correspond to the SUB operation or SBB operation in the previous iterative operation, extract the borrow value of the previous iteration, use the SBB operation to implement subtraction processing with the extracted borrow value, and store the borrow value of the SBB operation in the highest bit of the destination register.

[0151] In an exemplary embodiment, the second operation unit can extract the carry value of the previous iteration by using the instruction getMGBtoStat srcReg to copy the information in the (L+1)th bit hidden in the source register srcReg to the status register Stat.

[0152] In an exemplary embodiment, for square calculation in a finite field, taking the square calculation of a 256-bit number as an example, with 64 bits as the calculation unit length, that is, the processor bit is 64 bits, then the data to be squared will first be divided into four segments with 64 bits as the unit and then the square operation will be performed. For example: a 256-bit number is divided into four segments A, B, C and D, each segment is 64 bits, then: (D|C|B|A)×(D|C|B|A), which can be expressed as: (D|C|B|A)×(D|C|B|A)=A 2 +B 2 +C 2 +D 2 +2AB+2AC+2AD+2BC+2BD+2CD

[0153] The above formula includes a large number of 2-times multiplication operations. Taking AB as an example, the multiplication of AB is a 128-bit value. The program is expressed as follows:

[0154] In the upper half of the program segment above (that is, before performing the 2x multiplication operation, i.e., calculating 2AB), A×B is a 128-bit result, which can be expressed as: (temp[1])(temp[0]), that is, two 64-bit values ​​are concatenated, where temp[1] is the high-order register for storing the result of the multiplication operation of the two numbers involved in the 2x multiplication operation, and temp[0] is the low-order register for storing the result of the multiplication operation of the two numbers involved in the 2x multiplication operation; in the lower half of the program segment above, to calculate 2AB, i.e., to calculate 2 times (temp[1])(temp[0]), since the calculation cannot be guaranteed to be within 128 bits, it may be 129 bits, so the result must be expressed as (result[2])(result[1])(result[0]) concatenated, as shown in the program, which will generate a large number of bit shift operations.

[0155] To this end, an embodiment of the present application further provides a method for implementing large number operations. As shown in FIG6 , based on a finite field scenario, in the 2x multiplication operation in the finite field square operation process, the method includes:

[0156] Step 600: Perform a multiplication operation on the two numbers involved in the 2x multiplication operation, shift the high-order register of the multiplication result left by one bit, save the shifted information in the hidden bit of the highest bit of the high-order register, and take out the information in the hidden bit of the high-order register and store it in the lowest bit of the third destination register; the high-order register is a (L+1)-bit register used to store the high bit of the multiplication result.

[0157] In the embodiment of the present application, the third destination register is an (L+1)-bit register, which is also provided with a hidden bit, but the assignment operation will not change its hidden bit.

[0158] In one embodiment, the information stored in the hidden bit after the 1-bit left shift operation can be stored in the destination register using the instruction getMGBtoReg drcReg,srcReg. The instruction getMGBtoReg drcReg,srcReg indicates that the data stored in the hidden bit, i.e., the (L+1)th bit, in the source register srcReg (the high-order register in this step) is copied to the lowest bit, i.e., the 0th bit, of the destination register drcReg (the third destination register in this step).

[0159] Step 601: Shift the low-order register of the multiplication result left by one bit, save the information of the shifted bit in the hidden bit of the low-order register, and take out the information in the hidden bit of the low-order register and store it in the lowest bit of the second destination register; the low-order register is a (L+1)-bit register, which is used to store the low bit of the multiplication result.

[0160] In the embodiment of the present application, the second destination register is an (L+1)-bit register, which is also provided with a hidden bit, but the assignment operation will not change the hidden bit.

[0161] In one embodiment, the information stored in the hidden bit after the 1-bit left shift operation can be stored in the destination register using the instruction getMGBtoReg drcReg,srcReg. The instruction getMGBtoReg drcReg,srcReg indicates that the data stored in the hidden bit, i.e., the (L+1)th bit, in the source register srcReg (the low-order register in this step) is copied to the lowest bit, i.e., the 0th bit, of the destination register drcReg (the second destination register in this step).

[0162] It should be noted that there is no strict order in which step 600 and step 601 are performed.

[0163] Step 602: Perform a bitwise AND operation on the second destination register and the high-order register of the multiplication result, and store the result in the second destination register; move the information of the low-order register of the multiplication result to the first destination register; the first destination register is used to store the low-order information of the 2x multiplication result, the second destination register is used to store the high-order information of the 2x multiplication result, and the third destination register is used to store the carry value of the 2x multiplication result.

[0164] Taking the above (D|C|B|A)×(D|C|B|A) as an example, according to the method for implementing large number operations provided in the embodiment of the present application, temp[1] is directly shifted left by 1 bit, and the highest bit is saved in the corresponding highest hidden bit. Use getMGBtoReg result[2],temp[1] to take out the hidden highest bit of temp[1] and assign it to the lowest bit of the third destination register result[2]; temp[0] is shifted left by 1 bit, and the highest bit is saved in the corresponding highest hidden bit. Use getMGBtoReg result[1],temp[0] to take out the hidden highest bit of temp[0] and assign it to the lowest bit of the second destination register result[1]; perform a bitwise AND (OR) operation on result[1] and temp[1], and assign them to the second destination register result[1]; and move the information of temp[0] of the multiplication result to result[0]. You can get the carry value of (D|C|B|A)×(D|C|B|A), which is the value in the third destination register, the high-order information of (D|C|B|A)×(D|C|B|A), which is the result[1] value, and the low-order information of (D|C|B|A)×(D|C|B|A), which is the result[0] value. The program is expressed as follows:

[0165] result[2]=(temp[1]<<1); / / The result of result[2] is the hidden bit of the result of temp[1]<<1, which is directly assigned to result[2]

[0166] result[1]=(temp[0]<<1);

[0167] result[1]=result[1]|x; / / x represents the hidden bit of the result of temp[0]<<1 calculation;

[0168] result[0]=temp[0]<<1;

[0169] The assembly can be expressed as follows:

[0170] SLL temp[0],temp[0],1 / / temp[0] is shifted left by 1 bit and stored in the hidden bit of temp[0]

[0171] SLL temp[1],temp[1],1 / / temp[1] is shifted left by 1 bit and stored in the hidden bit of temp[1]

[0172] getMGBtoReg result[2],temp[1] / / The hidden bit of temp[1] is directly assigned to result[2] to get the carry value

[0173] getMGBtoReg result[1],temp[0] / / The hidden bit of temp[0] is directly assigned to result[2]

[0174] OR result[1],result[1],temp[1]; / / Perform bitwise AND operation on temp[1] and result[1] to get 126~63 bits

[0175] MOV result[0],temp[0]; / / get the low bit of 2 times multiplication

[0176] The method for implementing large number operations provided in the embodiments of the present application reduces a certain number of bit shift operations for square calculations, reduces the total number of overall calculation instructions, and thus improves calculation efficiency.

[0177] The present application also provides a computer-readable storage medium storing computer-executable instructions for executing the method for implementing large number operations described in any one of FIG6.

[0178] An embodiment of the present application further provides a computer device, including a memory and a processor, wherein the memory stores the following instructions that can be executed by the processor: for executing the steps of the method for implementing large number operations as described in any one of Figure 6.

[0179] An embodiment of the present application also provides a shifter. Figure 7 is a schematic diagram of the composition structure of the shifter used in large number operations in an embodiment of the present application. As shown in Figure 7, it includes: a shifter, an (L+1)-bit result register, and a destination register; wherein the shifter is used to shift the L-bit input information (such as input a in Figure 7) according to a specified shift number, and store the shift result in the result register; and copy the information in the (L+1)th bit hidden in the result register to the lowest bit of the destination register, that is, the 0th bit.

[0180] Although the embodiments disclosed in this application are as described above, the contents described are merely embodiments adopted to facilitate understanding of this application and are not intended to limit this application. Any person skilled in the art to which this application belongs may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in this application. However, the scope of patent protection of this application shall still be based on the scope defined by the attached claims.

Claims

1. A method for implementing large number operations, characterized in that, Including: In the first iteration operation of Montgomery domain multiplication, store the carry value of the ADD operation in the highest bit of the destination register, where the ADD operation is used to perform an addition of two L-bit data without carry; the destination register is an (L + 1)-bit register for storing the calculation result of the ADD operation; In the remaining iteration operations except the first iteration operation, corresponding to the ADD operation or ADC operation in the previous iteration operation, take out the carry value of the previous iteration, use the ADC operation to perform an addition of two L-bit data with the taken-out carry value, and store the carry value of the ADC operation in the highest bit of the destination register.

2. The method according to claim 1, wherein The taking out of the carry value of the previous iteration includes: Use the instruction getMGBtoStat srcReg to implement copying the information in the hidden (L + 1)th bit in the source register srcReg to the status register Stat; where the source register srcReg is the destination register that stores the calculation result after the addition calculation in the corresponding previous iteration operation; The status register Stat is a 1-bit register for storing the flag bit of carry addition.

3. The method according to claim 1 or 2, wherein, The Montgomery multiplication includes: implementing the Montgomery multiplication of two M-bit data through an L-bit processor; where M is greater than L; M changes according to the size of the finite field data selected by the algorithm.

4. A method for implementing large number operations, characterized in that, Including: In the first iteration operation, store the borrow value of the SUB operation in the highest bit of the destination register, where the SUB operation is used to perform a subtraction of two L-bit data without borrow; the destination register is an (L + 1)-bit register for storing the calculation result of the SUB operation; In the remaining iteration operations except the first iteration operation, corresponding to the SUB operation or SBB operation in the previous iteration operation, take out the borrow value of the previous iteration, use the SBB operation to perform a subtraction of two L-bit data with the taken-out borrow value, and store the borrow value of the SBB operation in the highest bit of the destination register.

5. The method according to claim 4, wherein The taking out of the borrow value of the previous iteration includes: Use the instruction getMGBtoStat srcReg to implement copying the information in the hidden (L + 1)th bit in the source register srcReg to the status register Stat; where the source register srcReg is the destination register that stores the calculation result after the subtraction calculation in the corresponding previous iteration operation; The status register Stat is a 1-bit register for storing the flag bit of borrow subtraction.

6. An apparatus for implementing large number operations, characterized in that Including: An arithmetic unit, a destination register; where An arithmetic unit for performing ADD / ADC operations on two L-bit data; in the first iteration operation, store the carry value of the ADD operation in the highest bit of the destination register; in the remaining iteration operations except the first iteration operation, corresponding to the ADD operation or ADC operation in the previous iteration operation, take out the carry value of the previous iteration, perform an addition process with the taken-out carry value using the ADC operation, and store the carry value of the ADC operation in the highest bit of the destination register; The destination register is an (L + 1)-bit register for storing the calculation result of the ADD / ADC operation, and the carry value is stored in the (L + 1)-th bit; Or, An arithmetic unit for performing SUB / SBB operations on two L-bit data; in the first iteration operation, store the borrow value of the SUB operation in the highest bit of the destination register; in the remaining iteration operations except the first iteration operation, corresponding to the SUB operation or SBB operation in the previous iteration operation, take out the borrow value of the previous iteration, perform a subtraction process with the taken-out carry value using the SBB operation, and store the carry value of the SBB operation in the highest bit of the destination register; The destination register is an (L + 1)-bit register for storing the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L + 1)-th bit.

7. The apparatus according to claim 6, further comprising: A status register; Taking out the carry value / borrow value of the previous iteration in the arithmetic unit includes: using the instruction getMGBtoStat srcReg to copy the information in the hidden (L + 1)-th bit in the source register srcReg to the status register, where the source register srcReg is the destination register.

8. An adder, characterized in that, Including: An arithmetic unit, a destination register, and a status register; where, An arithmetic unit for performing ADD / ADC operations on two L-bit data; in the first iteration operation, store the carry value of the ADD operation in the highest bit of the destination register; in the remaining iteration operations except the first iteration operation, corresponding to the ADD operation or ADC operation in the previous iteration operation, take out the carry value of the previous iteration and store it in the status register, perform an addition process with the carry value in the status register using the ADC operation, and store the carry value of the ADC operation in the highest bit of the destination register; The destination register is an (L + 1)-bit register for storing the calculation result of the ADD / ADC operation, and the carry value is stored in the (L + 1)-th bit; The status register is used to store the carry value corresponding to the ADD operation or ADC operation in the previous iteration operation in the remaining iteration operations except the first iteration operation.

9. A subtraction calculator, characterized in that, Including: An arithmetic unit, a destination register, and a status register; where, An arithmetic unit for implementing SUB / SBB operations on two L-bit data; in the first iteration operation, store the borrow value of the SUB operation in the highest bit of the destination register; in the remaining iteration operations except the first iteration operation, corresponding to the SUB operation or SBB operation in the previous iteration operation, take out the borrow value of the previous iteration and store it in the status register, and use the SBB operation to implement subtraction processing with the borrow value in the status register, and store the borrow value of the SBB operation in the highest bit of the destination register; The destination register is an (L + 1)-bit register for storing the calculation result of the SUB / SBB operation, and the borrow value is stored in the (L + 1)-th bit; The status register is used to store the borrow value corresponding to the SUB operation or SBB operation in the previous iteration operation in the remaining iteration operations except the first iteration operation.

10. An apparatus for implementing large number operations, characterized in that, It includes: Multiple first arithmetic units and multiple second arithmetic units; among them, The first arithmetic unit is used for the first iteration operation in the Montgomery domain multiplication operation, Store the carry value of the ADD operation in the highest bit of the destination register, where the ADD operation is used to implement the addition of two L-bit data; the destination register is an (L + 1)-bit register for storing the calculation result of the ADD operation; The second arithmetic unit is used for the remaining iteration operations except the first iteration operation in the Montgomery domain multiplication operation. Corresponding to the ADD operation or ADC operation in the previous iteration operation, take out the carry value of the previous iteration, use the ADC operation to implement addition processing with the taken-out carry value, and store the carry value of the ADC operation in the highest bit of the destination register; Or, The first arithmetic unit is used for the first iteration operation to store the borrow value of the SUB operation in the highest bit of the destination register, where the SUB operation is used to implement the subtraction of two L-bit data; the destination register is an (L + 1)-bit register for storing the calculation result of the SUB operation; The second arithmetic unit is used for the remaining iteration operations except the first iteration operation. Corresponding to the SUB operation or SBB operation in the previous iteration operation, take out the borrow value of the previous iteration, use the SBB operation to implement subtraction processing with the taken-out borrow value, and store the borrow value of the SBB operation in the highest bit of the destination register.

11. The device according to claim 10, wherein, Taking out the carry value of the previous iteration in the second arithmetic unit includes: using the instruction getMGBtoStat srcReg to implement copying the information in the hidden (L + 1)-th bit in the source register srcReg to the status register Stat; Wherein, the source register srcReg is the destination register that stores the calculation result after the addition calculation in the corresponding previous iteration operation; The status register Stat is a 1-bit register for storing the flag bit of the carry addition.

12. A method for implementing large number operations, characterized in that, In the 2-fold multiplication operation in the finite field squaring operation process, it includes: Perform a multiplication operation on two numbers participating in a double multiplication operation, perform a left shift operation on the high-order register of the multiplication result by one bit, save the shifted-out information in the hidden bit of the highest bit of the high-order register, and take out the information in the hidden bit of the high-order register and store it in the lowest bit of the third destination register; the high-order register is an (L + 1)-bit register for storing the high-order part of the multiplication result. Perform a left shift operation on the low-order register of the multiplication result by one bit, save the information of the shifted-out bit in the hidden bit of the low-order register, and take out the information in the hidden bit of the low-order register and store it in the lowest bit of the second destination register; the low-order register is an (L + 1)-bit register for storing the low-order part of the multiplication result. Perform a bitwise AND operation on the second destination register and the high-order register of the multiplication result, and store the result in the second destination register; move the information of the low-order register of the multiplication result to the first destination register; the first destination register is used to store the low-order information of the double multiplication operation result, the second destination register is used to store the high-order information of the double multiplication operation result, and the third destination register is used to store the carry value of the double multiplication operation result.

13. The method according to claim 12, wherein, Use the instruction getMGBtoReg drcReg,srcReg to take out the information in the hidden bit and store it in the third destination register / the second destination register. The instruction getMGBtoReg drcReg,srcReg means copying the data stored in the (L + 1)th bit in the source register srcReg to the lowest bit of the destination register drcReg.

14. A computer-readable storage medium storing computer-executable instructions for performing the method for implementing large number operations according to any one of claims 1 to 3, and / or claim 4 or 5, and / or claim 12 or 13.

15. A computer device includes a memory and a processor, wherein, Instructions executable by a processor are stored in a memory: steps for performing the method for implementing large number operations according to any one of claims 1 to 3, and / or claim 4 or 5, and / or claim 12 or 13.

16. A shifter, characterized in that, Including: A shifter, an (L + 1)-bit result register, and a destination register; wherein, The shifter is used to perform a shift operation on the L-bit input information according to a specified shift number, store the shift result in the result register, and copy the information in the hidden (L + 1)th bit in the result register to the lowest bit of the destination register.

17. The shifter according to claim 16, wherein, Use the instruction getMGBtoReg drcReg,srcReg to take out the information in the hidden (L + 1)th bit and store it in the destination register. The instruction getMGBtoReg drcReg,srcReg means copying the data stored in the (L + 1)th bit in the result register as the source register srcReg to the lowest bit of the destination register as the destination register drcReg.