Coprocessor and computer equipment

By designing a coprocessor including instruction decoding module and calculation module, the calculation process of the elliptic curve encryption algorithm is optimized, and the problems of high complexity and poor adaptability in hardware implementation are solved, and efficient computing performance improvement is achieved.

CN120263391APending Publication Date: 2025-07-04CHONGQING XINLIANXIN INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510395469.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art has the problem of high design complexity and inability to adapt to different optimization algorithms in the hardware implementation of elliptic curve encryption algorithms, and it is difficult to improve computing performance while reducing redundant computing logic.

Method used

A coprocessor is designed, including an instruction decoding module, a calculation module and a register module. By decomposing the scalar operation instructions calculated by elliptic curve encryption into a domain operation instruction stream, and through the combination of the modular operation scheduling module, a large number operation scheduling module, a multiplication operation scheduling module and a multiplication calculation module, the calculation process is optimized to reduce redundant logic and improve performance.

Benefits of technology

It realizes the reduction of redundant computing logic in elliptic curve encryption calculation, improves computing performance, supports the adaptability of different optimization algorithms, and improves the efficiency and flexibility of hardware implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263391A_ABST
    Figure CN120263391A_ABST
Patent Text Reader

Abstract

A coprocessor and computer equipment are applied to elliptic curve encryption calculation, and the coprocessor comprises an instruction decoding module, a calculation module and a register module, the instruction decoding module is used for decomposing a scalar operation instruction in elliptic curve encryption calculation into a corresponding domain operation instruction stream; the domain operation instruction stream comprises a modular operation instruction and a large number operation instruction; the calculation module comprises a modular operation scheduling module, a large number operation scheduling module, a multiply-add operation scheduling module and a multiply-add calculation module which are connected in sequence; when the domain operation instruction is a modular operation instruction, the modular operation scheduling module completes corresponding modular operation by sequentially calling the large number operation scheduling module, the multiply-add operation scheduling module and the multiply-add operation module according to the modular operation instruction; when the domain operation instruction is a large number operation instruction, the large number operation scheduling module completes corresponding large number operation by calling the multiply-add operation scheduling module and the multiply-add operation module in sequence; and the register module is used for storing data participating in operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of processors, and in particular, to a coprocessor and a computer device. Background Art

[0002] In the field of information security, based on the elliptic curve encryption algorithm, different security protocols formulated by countries and different organizations in the world call different operations, with great flexibility. In addition, in the face of group operations based on elliptic curves, different scholars may adopt different optimization algorithms. If all operations based on the elliptic curve encryption algorithm are implemented in hardware, not only will there be a relatively large design complexity, but also it will not be able to adapt to different optimization algorithms. Therefore, the research on a fast implementation scheme of elliptic curve encryption that can avoid both large hardware complexity and adapt to different optimization algorithms is of great significance to the development of the public key cryptography industry.

[0003] Therefore, there is an urgent need for a coprocessor that can reduce redundant calculation logic and improve calculation performance during the elliptic curve encryption and decryption process. Summary of the Invention

[0004] This application provides a coprocessor that can reduce redundant calculation logic and improve calculation performance during the elliptic curve encryption and decryption process.

[0005] In a first aspect, this application provides a coprocessor applied to elliptic curve encryption calculation. The coprocessor includes an instruction decoding module, a calculation module, and a register module;

[0006] The instruction decoding module is used to decompose scalar operation instructions in elliptic curve encryption calculation into corresponding domain operation instruction streams; the domain operation instruction streams include modulo operation instructions and large number operation instructions;

[0007] The calculation module includes a modulo operation scheduling module, a large number operation scheduling module, a multiply-add operation scheduling module, and a multiply-add calculation module connected in sequence;

[0008] When the domain operation instruction is a modulo operation instruction, the modulo operation scheduling module completes the corresponding modulo operation by sequentially calling the large number operation scheduling module, the multiply-add operation scheduling module, and the multiply-add calculation module;

[0009] When the domain operation instruction is a large number operation instruction, the large number operation scheduling module completes the corresponding large number operation by sequentially calling the multiply-add operation scheduling module and the multiply-add calculation module.

[0010] The register module is used to store data participating in the operation.

[0011] In a possible design, the scalar operation instruction is a macro instruction; the domain operation instruction is a micro instruction;

[0012] The instruction decoding module includes an instruction decoder and a microprogram memory;

[0013] If the instruction is a macro instruction, a micro-instruction stream corresponding to the macro instruction is called from the microprogram memory, and the instruction decoder performs a decoding operation on the micro-instruction stream;

[0014] If the instruction is a micro-instruction, the instruction decoder directly performs a decoding operation on the micro-instruction.

[0015] In a possible design, the modular arithmetic scheduling module includes different modular arithmetic sub-scheduling units, and any modular arithmetic sub-scheduling unit includes a comparator and a plurality of controllers; the comparator is used to perform a comparison operation in the modular arithmetic; the controller is used to call the large number arithmetic scheduling module;

[0016] The large number arithmetic scheduling module includes different large number arithmetic sub-scheduling units; any large number arithmetic sub-scheduling unit includes a plurality of controllers; the controller is used to call the multiply-add operation scheduling module.

[0017] In a possible design, different modular arithmetic sub-scheduling units include modular addition and subtraction sub-scheduling units; different large number arithmetic sub-scheduling units include large number addition and subtraction sub-scheduling units;

[0018] If the modular arithmetic instruction is a modular addition instruction or a modular subtraction instruction, the modular addition and subtraction sub-scheduling unit selects to call the large number addition and subtraction sub-scheduling unit to perform large number addition and subtraction operations or performs a large number comparison operation through its own comparator according to the operation types in different stages of the modular addition or modular subtraction operation by its own controller.

[0019] In a possible design, different modular arithmetic sub-scheduling units include modular inverse sub-scheduling units; different large number arithmetic sub-scheduling units include large number addition and subtraction sub-scheduling units;

[0020] If the modular arithmetic instruction is a modular inverse instruction, the modular inverse sub-scheduling unit selects to call the large number addition and subtraction sub-scheduling unit to perform large number addition and subtraction operations or performs a large number comparison operation through its own comparator according to the operation types in different stages of the modular inverse operation by its own controller.

[0021] In a possible design, different modular arithmetic sub-scheduling units include modular multiplication sub-scheduling units; different large number arithmetic sub-scheduling units include large number addition and subtraction sub-scheduling units and large number multiplication sub-scheduling units;

[0022] If the modulo operation instruction is a modular multiplication operation instruction, the modular multiplication sub-scheduling unit, through its own controller, selects and calls the large number addition and subtraction sub-scheduling unit to perform large number addition and subtraction operations, or calls the large number multiplication sub-scheduling unit to perform large number multiplication operations, or performs large number comparison operations through its own comparator according to the operation types in different stages of the modular multiplication operation process.

[0023] In a possible design, different large number operation sub-scheduling units include a large number addition and subtraction sub-scheduling unit; the large number addition and subtraction sub-scheduling unit includes multiple large number addition and subtraction execution units; each large number addition and subtraction execution unit includes a first selector, a second selector group, and a third selector group;

[0024] When the large number operation instruction is a large number addition operation instruction, the first selector is used to select to input 0 to the first multiply-accumulate calculation unit, and each selector in the second selector group is used to select to directly input the subtrahend in the array participating in the operation to the first multiply-accumulate calculation unit or the corresponding multiply-accumulate calculation unit group;

[0025] When the large number operation instruction is a large number subtraction operation instruction, the first selector is used to select to input 1 to the first multiply-accumulate calculation unit, and each selector in the second selector group is used to select to input the inverted subtrahend in the array participating in the operation to the first multiply-accumulate calculation unit or the corresponding multiply-accumulate calculation unit group;

[0026] Each multiply-accumulate calculation unit group includes a second multiply-accumulate calculation unit and a third multiply-accumulate calculation unit, the carry input of the second multiply-accumulate calculation unit is 0, and the carry input of the third multiply-accumulate calculation unit is 1;

[0027] Each selector in the third selector group is used to select to output the calculation result of the second multiply-accumulate calculation unit if there is no carry in the calculation result of the previous multiply-accumulate calculation unit group; otherwise, select to output the calculation result of the third multiply-accumulate calculation unit;

[0028] Each multiply-accumulate calculation unit is used to perform an addition operation on the array participating in the operation.

[0029] In a possible design, different large number operation sub-scheduling units further include a large number multiplication sub-scheduling unit; the large number multiplication sub-scheduling unit includes multiple large number multiplication execution units;

[0030] The large number multiplication execution unit includes a fourth selector, a fifth selector, and a cache unit;

[0031] When the large number operation instruction is a large number multiplication operation instruction, the fourth selector is used to input the multiplier in the array participating in the operation to the corresponding multiply-accumulate calculation unit according to the clock cycle;

[0032] The fifth selector is used to input the calculation results of each multiply-accumulate calculation unit in the previous clock cycle into the corresponding multiply-accumulate calculation unit;

[0033] Each multiply-accumulate calculation unit is used to first perform a multiplication operation on the array participating in the operation, and then perform an addition operation on the result of the multiplication operation and the corresponding calculation result in the previous clock cycle;

[0034] The cache unit is used to store the calculation results of each multiply-accumulate calculation unit in each clock cycle.

[0035] In a possible design, the multiply-accumulate calculation module includes multiple multiply-accumulate calculation rows, and each multiply-accumulate calculation row further includes three registers. The first register is used to save the operation type of the multiply-accumulate calculation row, the second register is used to save whether the multiply-accumulate calculation row is in an idle state, and the third register is used to save the return channel of the calculation result of the multiply-accumulate calculation row.

[0036] In a possible design, the multiply-accumulate operation scheduling module is specifically used to obtain the values in the second registers corresponding to each multiply-accumulate calculation row, and count the number of multiply-accumulate calculation rows in an idle state;

[0037] If the number of multiply-accumulate calculation rows in an idle state is greater than or equal to 2, two idle multiply-accumulate calculation rows are preferentially called to perform large number addition and subtraction operations; if the number of multiply-accumulate calculation rows in a space state is less than 2, the idle multiply-accumulate calculation rows are preferentially called to perform multiply-accumulate operations.

[0038] In a second aspect, an embodiment of the present application provides a computer device, including the coprocessor according to any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0040] Figure 1 It is a schematic structural diagram of a coprocessor provided by an embodiment of the present application;

[0041] Figure 2 It is a schematic diagram of an instruction decoding module provided by an embodiment of the present application;

[0042] Figure 3 It is a schematic diagram of another instruction decoding module provided by an embodiment of the present application Figure 3 ;

[0043] Figure 4Schematic diagram of a computing module provided by an embodiment of the present application;

[0044] Figure 5 Schematic diagram of another computing module provided by an embodiment of the present application;

[0045] Figure 6 Schematic diagram of a state machine for modular addition and subtraction operations provided by an embodiment of the present application;

[0046] Figure 7 Schematic diagram of a state machine for modular inverse operations provided by an embodiment of the present application;

[0047] Figure 8 Schematic diagram of a state machine for modular multiplication operations provided by an embodiment of the present application;

[0048] Figure 9 Schematic diagram of a large number addition and subtraction execution unit provided by an embodiment of the present application;

[0049] Figure 10 Schematic diagram of a large number addition and subtraction execution unit provided by an embodiment of the present application;

[0050] Figure 11 Schematic diagram of the principle of a large number multiplication execution unit provided by the present application;

[0051] Figure 12 Schematic diagram of a multiply - add computing module and a multiply - add scheduling module provided by an embodiment of the present application. Detailed implementation manners

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0053] In the embodiments of the present application, "a plurality of" means two or more. Words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order.

[0054] The following introduces the security protocols related to the elliptic curve encryption algorithm

[0055] In the cryptographic protocols constructed based on the elliptic curve discrete logarithm problem, common applications include digital signature schemes, public key encryption algorithms, and key exchange schemes.

[0056] (1) The relevant algorithms for digital signatures include SM2 digital signature and signature verification, and digital signature (ECDSA) and signature verification of elliptic curve encryption algorithms. Table 1 exemplarily lists the pseudocodes of these algorithms.

[0057] Table 1

[0058]

[0059]

[0060] The following conclusions can be drawn from Table 1: The operations of data signature and signature verification of SM2 and ECDSA include hash calculation, random number generation, large number modular addition calculation, large number addition calculation, large number subtraction calculation, large number multiplication calculation, large number inverse calculation, large number comparison, and elliptic curve point multiplication calculation.

[0061] (2) The relevant algorithms for password exchange include shared key generation algorithm, ECIES data encryption algorithm, ECIES data decryption algorithm, SM2 data encryption algorithm, and SM2 data decryption algorithm. Table 2 exemplarily shows the pseudocodes of password exchange algorithms.

[0062] Table 2

[0063]

[0064]

[0065] The following conclusions can be drawn from Table 2: The password exchange algorithms of ECIES and SM2 include elliptic curve point multiplication calculation, key calculation of KDF function, large number exclusive OR calculation, MAC calculation, and large number comparison calculation.

[0066] (3) The encryption and decryption protocols of elliptic curve encryption algorithms include SM2 encryption algorithm, SM2 decryption algorithm, ECC encryption algorithm, and ECC decryption algorithm. Table 3 exemplarily shows the pseudocodes of encryption and decryption algorithms.

[0067] Table 3

[0068]

[0069]

[0070] The following conclusions can be drawn from Table 3: The operations of encryption calculation and decryption calculation of ECC and SM2 include elliptic curve point multiplication calculation, key calculation of KDF function, hash calculation, large number exclusive OR calculation, large number multiplication calculation, and large number addition calculation.

[0071] The core operation of the elliptic curve encryption algorithm is the scalar point multiplication operation based on elliptic curves, and the scalar point multiplication operation is derived from the point addition operation and double point operation of elliptic curves. The following describes the derivation process and optimization methods of the scalar point multiplication operation based on elliptic curves over different finite fields.

[0072] The most commonly used equation of an elliptic curve is the plane curve determined by the Weierstrass equation. Related elliptic curve domain parameters are recommended for elliptic curves over prime fields and elliptic curves over binary extension fields, as shown in Table 4.

[0073] Table 4

[0074]

[0075] (1) Scalar point addition operation and scalar double point operation

[0076] There are two commonly used representation forms for the coordinates of an elliptic curve: affine coordinates and Jacobian projective coordinates. Elliptic curves over prime fields and elliptic curves over binary extension fields each have their own affine coordinates and Jacobian projective coordinates. The point addition operation and double point operation based on elliptic curves over prime fields and elliptic curves over binary extension fields have representation forms based on affine coordinates and representation forms based on Jacobian projective coordinates, as shown in Table 5.

[0077] Table 5

[0078]

[0079]

[0080] The computational amounts of the point addition operation and double point operation based on elliptic curves over prime fields and elliptic curves over binary extension fields calculated using affine coordinates and Jacobian projective coordinates are different. The computational amounts of the point addition operation and double point operation under different coordinates are shown in Table 6.

[0081] Table 6

[0082] Point operation type Affine coordinate representation Jacobian coordinate representation Point addition operation in prime field 1I + 2M + 1S 12M + 4S Point doubling operation in prime field 1I + 2M + 2S 4M + 6S Point addition operation in binary extension field 13M + 1S + 7A Point doubling operation in binary extension field 7M + 5S + 4A 5M + 5S + 4A

[0083] Among them, I represents the inverse operation, M represents the multiplication operation, and S represents the squaring operation. In the affine coordinate, the computational complexity of the point addition operation in the prime field is 1 inverse operation, 2 multiplication operations, and 1 squaring operation; in the Jacobian coordinate, the computational complexity of the point addition operation in the prime field is 12 multiplications and 4 squarings. In the affine coordinate, the computational complexity of the point addition operation in the prime field is 1 inverse operation, 2 multiplication operations, and 2 squaring operations; while in the Jacobian coordinate, the computational complexity of the point doubling operation in the prime field is 4 multiplications and 6 squarings. From the above analysis of the computational complexity, the point addition and point doubling operations in the prime field in the Jacobian coordinate use more multiplication and squaring operations, but do not use the more complex inverse operation. Therefore, when implementing the hardware of elliptic curve encryption in this application, before performing relevant calculations using the Jacobian projective coordinate, first map the point (x, y) in the affine coordinate to the point (x, y, 1) in the standard projective coordinate; after completing the relevant calculations based on the Jacobian projective coordinate, then map the Jacobian projective coordinate point (x, y, z) to the point (x / z, y / z) in the affine coordinate. It is more efficient to perform the point addition operation and the point doubling operation in the Jacobian coordinate.

[0084] (2) Field operations of large numbers

[0085] The field operations used in the elliptic curve encryption algorithm include modular addition operation, modular subtraction operation, modular multiplication operation, and modular inverse operation. Table 7 exemplarily shows the pseudocode of the field operations.

[0086] Table 7

[0087]

[0088]

[0089]

[0090] It can be seen that the basic calculations of the modular addition, modular subtraction, and modular multiplication of large numbers are the calculations of large number addition, large number subtraction, large number multiplication, etc. In the modular addition, modular subtraction, and modular multiplication of large numbers, each step of the calculation depends on the calculation result of the previous step, and this dependence causes the calculation process to be unable to be processed in parallel. In the modular inverse operation, except that the operations within each conditional branch can be processed in parallel, other operations also cannot be processed in parallel due to dependence on the calculation result of the previous step.

[0091] The calculations of large number addition and large number subtraction are implemented by a number of basic calculation units connected in series. The calculation of each calculation unit depends on the carry output of the previous calculation unit. This calculation dependence causes a number of basic calculation units for large number addition and large number subtraction to be unable to be processed in parallel. Large number multiplication is also affected by the memory bandwidth and the calculation dependence between calculation units, making multiple basic calculation units unable to be processed in parallel. Based on the characteristic that large number calculations cannot be processed in parallel at the basic calculation unit level, in this application, when performing large number calculations, the multiplication accumulator MADD with a basic granularity is used as the basic calculation unit, taking into account the calculation requirements of large number addition and multiplication.

[0092] Figure 1 The figure is a schematic diagram of a coprocessor provided by an embodiment of this application, which is applied to elliptic curve encryption calculations, such as Figure 1 shown. The coprocessor includes an instruction decoding module 100, a calculation module 200, and a register module 300.

[0093] The instruction decoding module 100 is used to decompose the scalar operation instructions in the elliptic curve encryption calculation into corresponding domain operation instruction streams. Among them, the scalar operations in the elliptic curve encryption calculation include point multiplication operation, point addition operation, and point doubling operation, specifically, the point multiplication operation, point addition operation, and point doubling operation of elliptic curves in a prime field or a binary extension field. The domain operation instruction stream includes modulo operation instructions and large number operation instructions.

[0094] The calculation module 200 includes a modulo operation scheduling module 210, a large number operation scheduling module 220, a multiply-accumulate operation scheduling module 230, and a multiply-accumulate calculation module 240 connected in sequence.

[0095] When the domain operation instruction is a modulo operation instruction, the modulo operation scheduling module 210 completes the corresponding modulo operation by sequentially calling the large number operation scheduling module 220, the multiply-accumulate operation scheduling module 230, and the multiply-accumulate calculation module 240. When the domain operation instruction is a large number operation instruction, the large number operation scheduling module 220 completes the corresponding large number operation by sequentially calling the multiply-accumulate operation scheduling module 230 and the multiply-accumulate calculation module 240.

[0096] The register module 300 is used to store the data participating in the operation.

[0097] The modular arithmetic instructions include modular addition, modular subtraction, modular multiplication, and modular inversion. Since modular addition, modular subtraction, and modular inversion are mainly implemented by calling large number addition and large number subtraction operations, and modular multiplication is mainly implemented by calling large number addition, large number multiplication, large number subtraction, and double-word multiplication operations, while large number addition and subtraction and large number multiplication operations are mainly implemented by calling double-word multiply-add operations. Based on this, when the coprocessor proposed in this application performs elliptic curve encryption calculations, the modular arithmetic scheduling module, large number arithmetic scheduling module, and multiply-add arithmetic scheduling module sequentially perform scheduling control on the lower-level modules, and finally the multiply-add calculation module completes the basic calculation work. The calculation logic is simple, and it can support scalar operations and large number domain operations on elliptic curves based on prime fields and binary extension fields, etc.

[0098] Figure 2 It is a schematic diagram of an instruction decoding module provided by an embodiment of this application. As Figure 2 shown, the instruction decoding module 100 includes an instruction decoder 110 and a microprogram memory 120.

[0099] This application defines scalar operation instructions as macro instructions, and scalar operations include dot multiplication, dot addition, and point doubling operations; it defines field operation instructions as micro instructions, and field operations include modular arithmetic and large number arithmetic. In addition, micro instructions also include some necessary transfer instructions, jump instructions, synchronization instructions, and exit instructions during the calculation process. The microprogram memory 120 stores the micro instruction stream corresponding to each pre-written macro instruction. If the instruction is a macro instruction, the micro instruction stream corresponding to the macro instruction is called from the microprogram memory 120, and the instruction decoder 110 performs a decoding operation on the micro instruction stream; if the instruction is a micro instruction, the instruction decoder 110 directly performs a decoding operation on the micro instruction.

[0100] In addition, referring to Figure 3 , the instruction decoding module 100 also includes an instruction interface 130 and an address generator 140. The instruction interface 130 is used to receive instructions. If the instruction is a macro instruction, the address generator 140 generates the storage address of the micro instruction stream corresponding to the macro instruction according to the macro instruction, then calls the micro instruction stream corresponding to the macro instruction from the corresponding storage address in the microprogram memory 120, and then the instruction decoder 110 performs a decoding operation on the micro instruction stream; if the instruction is a micro instruction, the instruction decoder 110 directly performs a decoding operation on the micro instruction.

[0101] Exemplarily, Table 8 exemplarily shows the instruction list of the macro instructions and micro instructions provided by this application.

[0102] Table 8

[0103]

[0104]

[0105]

[0106]

[0107] If the instruction is a download instruction, according to the data address in the instruction, the specified operand is downloaded in the form of an array to the specified large number register through DMA; if the instruction is a calculation instruction, the corresponding calculation is performed on the specified operand; if the instruction is a conditional branch instruction, the calculation unit calculates the operand, makes a conditional judgment according to the calculation result, and if the condition is satisfied, the instruction pointer jumps to the specified instruction position and reads the instruction at the specified position; if the instruction is a storage instruction, the data in the specified register is stored to the specified memory address through DMA.

[0108] Taking some macro instructions as examples, the micro-instruction streams corresponding to each macro instruction refer to Table 9.

[0109] Table 9

[0110]

[0111]

[0112]

[0113] Taking the prime number elliptic curve point doubling instruction ECP_POINT_DBL rd, rs as an example, according to the order of each micro-instruction in its corresponding micro-instruction stream, the large number modular multiplication instruction BN_MOD_MUL$1, ECC_PX, ECC_PX, the large number left shift instruction BN_SHL$1, $1, 1, the large number addition instruction BN_MOD_ADD$1, $1, ECC_PX, the large number addition instruction BN_MOD_ADD$1, $1, ECC_PA, the large number left shift instruction BN_SHL$2, ECC_PY, 1, the large number modular inverse instruction BN_MOD_INV$2, $2, the large number modular multiplication instruction BN_MOD_MUL$1, $1, $2 / / m, the large number modular multiplication instruction BN_MOD_MUL$2, $1, $1, the large number addition instruction BN_MOD_ADD$3, ECC_PX, ECC_QX, the large number modular subtraction instruction BN_MOD_SUB$2, $2, $3 / / xo, the large number modular subtraction instruction BN_MOD_SUB$3, $2, ECC_PX, the large number modular multiplication instruction BN_MOD_MUL$1, $3, $1, the large number addition instruction BN_MOD_ADD$1, $1, ECC_PY / / yo, and the elliptic curve coprocessor data equality exit instruction ECCP_RET rd, $1, $2 are executed in sequence.

[0114] Since the elliptic curve encryption algorithm includes modular arithmetic, large number arithmetic, and multiplication and addition operations, based on the calculation characteristics of elliptic curves, when the coprocessor proposed in this application performs elliptic curve encryption calculations, the modular arithmetic scheduling module, the large number arithmetic scheduling module, and the multiplication and addition arithmetic scheduling module sequentially control the lower-level modules, and finally the multiplication and addition calculation module completes the basic calculation work.

[0115] Figure 4 It is a schematic diagram of a calculation module provided by an embodiment of this application, as Figure 4 shown, the modular arithmetic scheduling module 210 includes different modular arithmetic sub-scheduling units, and any modular arithmetic sub-scheduling unit includes a comparator and multiple controllers. The comparator is used to perform the comparison operation in modular arithmetic, and the controller of the modular arithmetic sub-scheduling unit is used to call the large number arithmetic scheduling module 220. The large number arithmetic scheduling module 220 includes different large number arithmetic sub-scheduling units, and any large number arithmetic sub-scheduling unit includes multiple controllers. The controller of the large number arithmetic sub-scheduling unit is used to call the multiplication and addition arithmetic scheduling module 230. The multiplication and addition arithmetic scheduling module 230 calls the multiplication and addition calculation module 240 to perform calculations. Among them, the types of modular arithmetic sub-scheduling units can be modular addition and subtraction sub-scheduling units, modular inverse sub-scheduling units, modular multiplication sub-scheduling units; the types of large number arithmetic sub-scheduling units can be large number addition and subtraction sub-scheduling units, large number multiplication sub-scheduling units. The quantity and types of modular arithmetic sub-scheduling units and large number arithmetic sub-scheduling units can be set according to actual needs, and this application does not make specific limitations.

[0116] In a possible implementation manner, Figure 5 It is another schematic diagram of a calculation module provided by an embodiment of this application. The different modular arithmetic sub-scheduling units of the modular arithmetic scheduling module 210 include a modular addition and subtraction sub-scheduling unit 211, and the different large number arithmetic sub-scheduling units of the large number arithmetic scheduling module 220 include a large number addition and subtraction sub-scheduling unit 221. If the modular arithmetic instruction is a modular addition instruction or a modular subtraction instruction, the modular addition and subtraction sub-scheduling unit 211 selects to call the large number addition and subtraction sub-scheduling unit 221 to perform large number addition and subtraction operations or perform large number comparison operations through its own comparator according to the operation types in different stages of the modular addition or modular subtraction operation process.

[0117] Exemplarily, referring to Figure 5, the modulo addition / subtraction sub-scheduling unit includes two modulo addition / subtraction controllers, both of which are used to control the invocation of lower-level operations to support the simultaneous operation of two modulo addition (subtraction) operations. As long as one of the controllers in the modulo addition / subtraction sub-scheduling unit is in an idle state, it can receive a new modulo addition (subtraction) operation instruction; if all the scheduling controllers in the modulo addition / subtraction operation module are not in an idle state, subsequent related instructions have to be paused to wait for an idle scheduling controller. The large number addition / subtraction sub-scheduling unit and the subsequent modulo inverse sub-scheduling unit, modulo multiplication sub-scheduling unit, and large number multiplication sub-scheduling unit also have the same working mechanism. In addition, the number of controllers set in each sub-scheduling unit can be set according to actual requirements.

[0118] The following specifically explains the operation of the modulo addition operation instruction or the modulo subtraction operation instruction. Refer to Figure 6 The state machine of the modulo addition / subtraction operation has four states: MDAS_IDLE stage, MDAS_CALC stage, MDAS_CMP stage, and MDAS_MOD stage. The calculations in the modulo addition / subtraction operation are divided into three categories according to the calculation stage: (1) calculations in the MDAS_CALC stage (2) judgment calculations in the MDAS_CMP stage (3) result adjustment calculations of the MDAS_MOD result. The calculations in these three stages are presented in different types of calculations when performing modulo addition and modulo subtraction operations. When performing the modulo addition operation, the calculation in the MDAS_CALC stage is the addition calculation of large numbers, the calculation in the MDAS_CMP stage is the comparison calculation of the calculation result of the large number addition in the MDAS_CALC stage with the maximum value p of the domain boundary, and the result adjustment calculation of the MDAS_MOD result is the large number subtraction calculation of the calculation result of the large number addition in the MDAS_CALC stage with the maximum value p of the domain boundary. When performing the modulo subtraction operation, the calculation in the MDAS_CALC stage is the subtraction calculation of large numbers, the calculation in the MDAS_CMP stage is the comparison calculation of the calculation result of the large number subtraction in the MDAS_CALC stage with the minimum value 0 of the domain boundary, and the result adjustment calculation of the MDAS_MOD result is the large number addition calculation of the calculation result of the large number subtraction in the MDAS_CALC stage with the maximum value P of the domain boundary. The call functions for different stages of the modulo addition / subtraction operation refer to Table 10.

[0119] Table 10

[0120] Operation type MDAS_CALC MDAS_CMP MDAS_MOD Modular addition operation Large number addition Large number comparison Large number subtraction Modular subtraction operation Large number subtraction Large number comparison Large number addition

[0121] That is to say, the modulo addition operation needs to perform large number addition operations, large number comparison operations, and large number subtraction operations; the modulo subtraction operation needs to perform large number subtraction operations, large number comparison operations, and large number addition operations. Among these operations to be executed, the large number comparison is performed in the comparator of the modulo addition / subtraction sub-scheduling unit 211, while the large number addition and large number subtraction are realized by calling the large number addition / subtraction sub-scheduling unit 221.

[0122] In a possible design, different modulo operation sub-scheduling units include a modular inverse sub-scheduling unit 212, and different large number operation sub-scheduling units include a large number addition and subtraction sub-scheduling unit 221. If the modulo operation instruction is a modular inverse operation instruction, the modular inverse sub-scheduling unit 212, through its own controller, selects and calls the large number addition and subtraction sub-scheduling unit 221 to perform large number addition and subtraction operations, or performs large number comparison operations through its own comparator, according to the operation types in different stages of the modular inverse operation process.

[0123] The operations of the modular inverse operation instruction will be specifically explained below. Refer to Figure 7 , the calculations in the modular inverse operation are executed step by step according to the state of the modular inverse operation scheduling controller. In each clock cycle of the MDINV_LP stage, two large number subtraction operations (u - v, v - u), one large number addition operation, and one large number comparison operation are performed. In the MDINV_CMP stage, large number comparison operations are performed, and in the MDINV_BNS stage, large number subtraction operations are performed. The calling functions for different stages of the modular inverse operation refer to Table 11. Among these operations to be executed, The large number addition operation and the large number subtraction operation are entrusted the large number addition and subtraction sub-scheduling unit 221 to be implemented, The large number comparison operation is carried out the comparator of the modular inverse sub-scheduling unit 212 completed.

[0124] Table 11

[0125] State machine state Calculation type MDINV_LP Two large number subtractions (u - v, v - u), one large number addition, and one large number comparison MDINV_CMP Large number comparison MDINV_BNS Large number subtraction

[0126] In a possible design, different modulo operation sub-scheduling units include a modular multiplication sub-scheduling unit 213, and different large number operation sub-scheduling units include a large number addition and subtraction sub-scheduling unit 221 and a large number multiplication sub-scheduling unit 222. If the modulo operation instruction is a modular multiplication operation instruction, the modular multiplication sub-scheduling unit 213, through its own controller, selects and calls the large number addition and subtraction sub-scheduling unit 221 to perform large number addition and subtraction operations, or calls the large number multiplication sub-scheduling unit 222 to perform large number multiplication operations, or performs large number comparison operations through its own comparator, according to the operation types in different stages of the modular multiplication operation process.

[0127] The operations of the modular multiplication operation instruction will be specifically explained below. Refer to Figure 8, the calculations in the modular multiplication operation are performed step by step according to the calculation stages. The large number multiplication operation is performed in the MDML_BNM0 stage and the MDML_BNM1 stage, the 64-bit multiplication operation is performed in the MDML_DWM stage, the large number addition operation is performed in the MDML_BNA0 stage and the MDML_BNA1 stage, the large number comparison operation is performed in the MDML_BNC stage, and the large number subtraction operation is performed in the MDML_BNS stage. The calling functions for different stages of the modular multiplication operation refer to Table 12. Among these calculations to be performed, the large number multiplication operation is performed by delegating to the large number multiplication sub-scheduling unit 222 to implement of, the large number addition operation and the large number subtraction operation are implemented through delegation large number addition and subtraction sub-scheduling unit 221 to be implemented, and the large number ratio comparison calculation is carried out The comparator of the modular multiplication sub-scheduling unit 213 completes the calculation.

[0128] Table 12

[0129] Scheduling controller state Operations to be scheduled MDML_BNM0 Large number multiplication MDML_DWM Double - word multiplication MDML_BNA1 Large number addition MDML_BNM1 Large number multiplication MDML_BNC Large number comparison MDML_BNA2 Large number addition MDML_BNS Large number subtraction

[0130] In a possible design, different large number operation sub-scheduling units include a large number addition and subtraction sub-scheduling unit 221, and the large number addition and subtraction sub-scheduling unit 221 includes multiple large number addition and subtraction execution units. Figure 9 FIG. is a schematic diagram of a large number addition and subtraction execution unit provided for an embodiment of the present application. Each large number addition and subtraction execution unit includes a first selector 2211, a second selector group 2212, and a third selector group 2213.

[0131] When the large number operation instruction is a large number addition operation instruction, the first selector 2211 is used to select to input 0 to the first multiply-add calculation unit, and each selector in the second selector group 2212 is used to select to directly input the minuend in the array participating in the operation to the first multiply-add calculation unit or the corresponding multiply-add calculation unit group.

[0132] When the large number operation instruction is a large number subtraction operation instruction, the first selector is used to select to input 1 to the first multiply-add calculation unit, and each selector in the second selector group is used to select to input the inverted minuend in the array participating in the operation to the first multiply-add calculation unit or the corresponding multiply-add calculation unit group.

[0133] Each multiply-add calculation unit group includes a second multiply-add calculation unit and a third multiply-add calculation unit. The carry input of the second multiply-add calculation unit is 0, and the carry input of the third multiply-add calculation unit is 1.

[0134] Each selector in the third selector group 2213 is used to select to output the calculation result of the second multiply-add calculation unit if there is no carry in the calculation result of the previous multiply-add calculation unit group; otherwise, select to output the calculation result of the third multiply-add calculation unit.

[0135] Each multiply-accumulate calculation unit is used to perform an addition operation on the arrays participating in the operation.

[0136] The large number addition and subtraction execution unit provided by this application has the following characteristics: (1) The high-order calculation depends on the carry of the low-order calculation; (2) The large number subtraction operation is converted into a large number addition operation by taking the inverse and adding one. Through the conversion of taking the inverse and adding one, a set of large number addition and subtraction execution unit circuits can be used to implement both large number addition operations and large number subtraction operations.

[0137] To improve the performance of large number addition and subtraction operations, the high-order large number operations use the method of selecting carry adders, which greatly shortens the carry chain length of large number addition and subtraction operations. The calculations of the first to eighth groups and the zero group of large numbers are carried out simultaneously, and their results are calculated simultaneously. Only the calculation results of the current group need to be selected according to the carry output generated by the previous group. The calculation results selected from all groups are put together to form the calculation result of the large number addition and subtraction operation. Its critical path is only the carry chain of a set of 64-bit adders and 8 cascaded selectors. These logics can be completed within one or two clock cycles, greatly improving the performance of large number addition and subtraction operations. The large number addition and subtraction execution unit only includes the inverse-selection logic that fuses addition and subtraction operations and the selection logic for grouped calculation results. The calculation part is entrusted to the multiply-accumulate calculation unit. Each modulo addition and subtraction operation's fast calculation calls 17 64-bit multiply-accumulate calculation units. In a possible design, different large number operation sub-scheduling units further include a large number multiplication sub-scheduling unit 222; the large number multiplication sub-scheduling unit 222 includes multiple large number multiplication execution units. Figure 10 It is a schematic diagram of a large number multiplication execution unit provided for an embodiment of this application. The large number multiplication execution unit includes a fourth selector 2221, a fifth selector, and a cache unit 2222. When the large number operation instruction is a large number multiplication operation instruction, the fourth selector 2221 is used to input the multiplicand in the array participating in the operation to the corresponding multiply-accumulate calculation unit according to the clock cycle;

[0138] The fifth selector 2222 is used to input the calculation results of each multiply-accumulate calculation unit in the previous clock cycle to the corresponding multiply-accumulate calculation unit;

[0139] Each multiply-accumulate calculation unit is used to first perform a multiplication operation on the arrays participating in the operation, and then perform an addition operation on the result of the multiplication operation and the corresponding calculation result in the previous clock cycle;

[0140] The cache unit is used to store the calculation results of each multiply-accumulate calculation unit in each clock cycle.

[0141] Figure 11 It is a schematic diagram of the principle of the large number multiplication execution unit provided by this application. The large number multiplication execution unit is based on a 64-bit grouped shift-and-add algorithm, and the grouped product results are accumulated to obtain the partial products of the corresponding groups. Its pseudocode is

[0142] Input: a = (a n-1 , a n-2 , …, a0), b = (b n-1 , b n-2 , …`b0), n1 = n / w, B = 2 w

[0143] Output: c = a * b

[0144] Calculation process

[0145] C = 0;

[0146] For I from 0 to n1 - 1

[0147] U = 0;

[0148] For j from 0 to n1 - 1

[0149] (u || v) = c i+j + a i b j + u;

[0150] Ci + j = v;

[0151] C i+n1 = u;

[0152] End for

[0153] End for

[0154] Return c

[0155] The large - number multiplication execution unit of this application adopts 9 64 - bit multiply - add calculation logics. It completes 9 multiply - add operations per clock cycle, and completes the large - number multiplication operation of 576 bits in 9 clock cycles. The data input of the 9 multiply - add calculations per clock cycle selects different source operand groups according to the counter of the state machine. The large - number multiplication execution unit only contains the call control logic of 9 multiply - add calculations, and the 9 multiply - add calculations are entrusted to the multiply - add calculation unit.

[0156] Figure 12 It is a schematic diagram of a multiply - add calculation module and a multiply - add scheduling module provided by an embodiment of this application. The multiply - add calculation module 240 includes multiple multiply - add calculation rows, and each multiply - add calculation row further includes three registers. The first register is used to save the operation type of this multiply - add calculation row, the second register is used to save whether this multiply - add calculation row is in an idle state, and the third register is used to save the return channel of the calculation result of this multiply - add calculation row.

[0157] The multiplication and addition operation scheduling module 230 is specifically configured to obtain the values in the second registers corresponding to each multiplication and addition calculation row, and count the number of multiplication and addition calculation rows in the idle state; if the number of multiplication and addition calculation rows in the idle state is greater than or equal to 2, two idle multiplication and addition calculation rows are preferentially called to perform large number addition and subtraction operations; if the number of multiplication and addition calculation rows in the space state is less than 2, the idle multiplication and addition calculation rows are preferentially called to perform multiplication and addition operations.

[0158] The number of multiplication and addition calculation rows in the multiplication and addition calculation module and the number of multiplication and addition calculation units in the multiplication and addition calculation rows can be set according to actual needs, and the present application does not make specific limitations. Refer to Figure 12 Taking the multiplication and addition calculation module including 4 multiplication and addition calculation rows, and each calculation row containing 9 64-bit multiplication and addition calculation units as an example, this multiplication and addition calculation module supports 2 large number addition and subtraction operations or 4 groups of multiplication and addition operations to be performed simultaneously. Each multiplication and addition calculation row has three registers respectively used to store the operation type of the calculation in this row, the usage flag (whether it is in the idle state), and the return channel of the calculation result. The multiplication and addition operation scheduling module works according to the following principles: (1) When there are more than 2 idle multiplication and addition calculation rows, the calculation priority of large number addition is the highest; when there are less than 2 idle multiplication and addition calculation rows, the multiplication and addition calculation for large number multiplication has a higher priority; in the multiplication and addition calculation sequence of large number multiplication, the priority follows the principle of first come, first served. (2) If there are no idle multiplication and addition calculation rows available in the current multiplication and addition calculation module, subsequent calculation requests will continue to wait until there are idle multiplication and addition calculation rows available.

[0159] According to the above principles, the multiplication and addition operation scheduling module on the one hand constantly checks the working status of the four multiplication and addition calculation rows and the number of idle multiplication and addition calculation rows, and on the other hand checks the number and type of calculation requests. The hardware scheduling method of the multiplication and addition operation scheduling module is: according to the number of large number additions, different scheduling combinations are selected, and then filled according to the positions and numbers of the idle multiplication and addition calculation rows to complete the scheduling of calculation requests and calculation resources. Refer to Table 13 for resource scheduling when the number of large number additions is different, and Table 14 for calculation scheduling under different effective bit conditions.

[0160] Table 13

[0161] Number of large number addition requests Calculation request combination 0 Mulx0, mulx1, mulx2, mul3 1 Addx0, Mulx0, mulx1 2 Add0, add1

[0162] Mulx0 represents the multiplication and addition calculation request for the first effective multiplication operation, Mulx1 represents the multiplication and addition calculation request for the second effective multiplication operation, Mulx2 represents the multiplication and addition calculation request for the third effective multiplication operation, and Addx0 represents the first effective large number addition calculation request.

[0163] Table 14

[0164]

[0165] As long as there are enough idle multiply - add computing rows to support an operation, the operation can start working. The multiply - add computing rows of 4 threads have significant performance advantages compared with the traditional single - thread computing method. Based on the same technical concept, the embodiment of the present application also provides a computer device, including the coprocessor listed in any of the above - mentioned ways.

[0166] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0167] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these changes and modifications.

Claims

1. A coprocessor, applied to elliptic curve cryptography calculations, characterized in that It includes an instruction decoding module, a computing module, and a register module; The instruction decoding module is used to decompose the scalar operation instructions in elliptic curve encryption calculations into corresponding domain operation instruction streams; the domain operation instruction streams include modulo operation instructions and large number operation instructions; The computing module includes a modulo operation scheduling module, a large number operation scheduling module, a multiply-accumulate operation scheduling module, and a multiply-accumulate computing module connected in sequence; When the domain operation instruction is a modulo operation instruction, the modulo operation scheduling module completes the corresponding modulo operation by sequentially calling the large number operation scheduling module, the multiply-accumulate operation scheduling module, and the multiply-accumulate computing module according to the modulo operation instruction; When the domain operation instruction is a large number operation instruction, the large number operation scheduling module completes the corresponding large number operation by sequentially calling the multiply-accumulate operation scheduling module and the multiply-accumulate computing module; The register module is used to store the data participating in the operation.

2. The coprocessor according to claim 1, wherein The scalar operation instruction is a macro instruction; the domain operation instruction is a micro instruction; The instruction decoding module includes an instruction decoder and a microprogram memory; If the instruction is a macro instruction, the micro instruction stream corresponding to the macro instruction is called from the microprogram memory, and the instruction decoder performs a decoding operation on the micro instruction stream; If the instruction is a micro instruction, the instruction decoder directly performs a decoding operation on the micro instruction.

3. The coprocessor according to claim 1, wherein The modulo operation scheduling module includes different modulo operation sub-scheduling units. Any modulo operation sub-scheduling unit includes a comparator and multiple controllers; the comparator is used to perform comparison operations in the modulo operation; the controller is used to call the large number operation scheduling module; The large number operation scheduling module includes different large number operation sub-scheduling units; any large number operation sub-scheduling unit includes multiple controllers; the controller is used to call the multiply-accumulate operation scheduling module.

4. The coprocessor according to claim 3, wherein different modulo operation sub-scheduling units include modulo addition and subtraction sub-scheduling units; different large number operation sub-scheduling units include large number addition and subtraction sub-scheduling units; If the modulo operation instruction is a modulo addition operation instruction or a modulo subtraction operation instruction, the modulo addition and subtraction sub-scheduling unit selects to call the large number addition and subtraction sub-scheduling unit to perform large number addition and subtraction operations or performs large number comparison operations through its own comparator according to the operation types in different stages of the modulo addition or modulo subtraction operation process by its own controller.

5. The coprocessor according to claim 3, wherein different modulo operation sub-scheduling units include modulo inverse sub-scheduling units; different large number operation sub-scheduling units include large number addition and subtraction sub-scheduling units; If the modulo operation instruction is a modulo inverse operation instruction, the modulo inverse sub-scheduling unit selects to call the large number addition and subtraction sub-scheduling unit to perform large number addition and subtraction operations or performs large number comparison operations through its own comparator according to the operation types in different stages of the modulo inverse operation process by its own controller.

6. The coprocessor according to claim 3, wherein different modulo operation sub-scheduling units include modulo multiplication sub-scheduling units; different large number operation sub-scheduling units include large number addition and subtraction sub-scheduling units and large number multiplication sub-scheduling units; If the modulo operation instruction is a modulo multiplication operation instruction, the modulo multiplier scheduling unit, through its own controller, selects and calls the large number addition and subtraction scheduling unit to perform large number addition and subtraction operations, or calls the large number multiplication scheduling unit to perform large number multiplication operations, or performs large number comparison operations through its own comparator according to the operation types in different stages of the modulo multiplication operation process.

7. The coprocessor according to claim 3, wherein The different large number operation scheduling units include a large number addition and subtraction scheduling unit; the large number addition and subtraction scheduling unit includes a plurality of large number addition and subtraction execution units; each large number addition and subtraction execution unit includes a first selector, a second selector group, and a third selector group; When the large number operation instruction is a large number addition operation instruction, the first selector is used to select to input 0 to the first multiply-accumulate calculation unit, and each selector in the second selector group is used to select to directly input the subtractor in the array participating in the operation to the first multiply-accumulate calculation unit or the corresponding multiply-accumulate calculation unit group; When the large number operation instruction is a large number subtraction operation instruction, the first selector is used to select to input 1 to the first multiply-accumulate calculation unit, and each selector in the second selector group is used to select to input the inverted subtractor in the array participating in the operation to the first multiply-accumulate calculation unit or the corresponding multiply-accumulate calculation unit group; Each multiply-accumulate calculation unit group includes a second multiply-accumulate calculation unit and a third multiply-accumulate calculation unit, the carry input of the second multiply-accumulate calculation unit is 0, and the carry input of the third multiply-accumulate calculation unit is 1; Each selector in the third selector group is used to select to output the calculation result of the second multiply-accumulate calculation unit if there is no carry in the calculation result of the previous multiply-accumulate calculation unit group; otherwise, select to output the calculation result of the third multiply-accumulate calculation unit; Each multiply-accumulate calculation unit is used to perform an addition operation on the array participating in the operation.

8. The coprocessor according to claim 7, wherein The different large number operation scheduling units further include a large number multiplication scheduling unit; the large number multiplication scheduling unit includes a plurality of large number multiplication execution units; The large number multiplication execution unit includes a fourth selector, a fifth selector, and a cache unit; When the large number operation instruction is a large number multiplication operation instruction, the fourth selector is used to input the multiplier in the array participating in the operation to the corresponding multiply-accumulate calculation unit according to the clock cycle; The fifth selector is used to input the calculation results of each multiply-accumulate calculation unit in the previous clock cycle to the corresponding multiply-accumulate calculation unit; Each multiply-accumulate calculation unit is used to first perform a multiplication operation on the array participating in the operation, and then perform an addition operation on the result of the multiplication operation and the corresponding calculation result in the previous clock cycle; The cache unit is used to store the calculation results of each multiply-accumulate calculation unit in each clock cycle.

9. The coprocessor according to claim 1, wherein The multiply-accumulate calculation module includes a plurality of multiply-accumulate calculation rows, and each multiply-accumulate calculation row further includes three registers. The first register is used to save the operation type of the multiply-accumulate calculation row, the second register is used to save whether the multiply-accumulate calculation row is in an idle state, and the third register is used to save the return channel of the calculation result of the multiply-accumulate calculation row.

10. The coprocessor according to claim 9, wherein The multiplication and addition operation scheduling module is specifically configured to obtain the values in the second registers corresponding to each multiplication and addition calculation row, and count the number of multiplication and addition calculation rows in the idle state; When the number of multiplication and addition calculation rows in the idle state is greater than or equal to 2, two idle multiplication and addition calculation rows are preferentially called to perform large number addition and subtraction operations; When the number of multiplication and addition calculation rows in the space state is less than 2, the idle multiplication and addition calculation rows are preferentially called to perform multiplication and addition operations.

11. A computer device, characterized in that, Comprising: The coprocessor according to any one of claims 1 to 10.