High-radix montgomery modular multiplication circuit and parallel computation method therefor

By introducing multiple sets of cyclic control units into the cyclic control module of the modular multiplication circuit, the parallel operation of multiple rounds of cyclic operations is solved, and a high-efficiency circuit design is achieved.

WO2025130779A1PCT designated stage expired Publication Date: 2025-06-26AMICRO SEMICONDUCTOR CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139210
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-19
Filing Date
2024-12-13
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

When existing modem multiplication circuits are implemented in hardware, the circuit performance and circuit area cannot be well balanced, and there is a problem that the circuit area is larger when the circuit performance is improved, or the circuit area is smaller and the circuit performance is poor.

Method used

By introducing multiple sets of cyclic control units into the cyclic control module, the same storage module is called and the parallel operation of multiple rounds of cyclic operations is realized based on the same operation module, thereby reducing the circuit area while maintaining high efficiency of circuit cyclic operations.

Benefits of technology

It realizes the high-efficiency operation of circuit cyclic operations while reducing the circuit area, solving the problem of balance between circuit performance and area.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139210_26062025_PF_FP_ABST
    Figure CN2024139210_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present application are a high-radix Montgomery modular multiplication circuit and a parallel computation method therefor. The circuit comprises: a loop control module, which comprises m groups of loop control units, is used for running loop computation in parallel to complete k+1 rounds of loop computation, and is used for caching intermediate data during a loop computation process, reading from a storage module data required for loop computation, and transmitting the corresponding data to a computation module on the basis of a loop computation step; a core state machine module, which is responsible for controlling the starting and ending of each round of first-stage loop computation among the k+1 rounds of loop computation that are run by the m groups of loop control units in the loop control module; the computation module, which is used for receiving the data transmitted by the loop control module, and performing multiplication computation and addition computation on the basis of the received data; and the storage module, which is used for storing the data that is required by the loop computation, such that the loop control module can read and call the data. In the present application, parallel running by m groups of loop control units reduces the circuit area and also maintains the high-efficiency running of circuit loop computation.
Need to check novelty before this filing date? Find Prior Art

Description

A high-radix Montgomery modular multiplication circuit and its parallel operation method Technical Field

[0001] The present application relates to the field of modular multiplication circuits, and in particular to a high-radix Montgomery modular multiplication circuit and a parallel operation method thereof. Background Art

[0002] The Montgomery modular multiplication algorithm is one of the most widely used methods for implementing modular multiplication algorithms. As the basic unit of asymmetric encryption and decryption algorithms such as RAS and ECC, the Montgomery modular multiplication algorithm's computing speed determines the algorithm's overall computing efficiency. With the development and application of information technology, the security of encryption and decryption algorithms has received increasing attention in order to improve information security. Currently, most encryption and decryption algorithms are implemented in hardware to prevent the theft of encrypted and decrypted data due to software vulnerabilities, which can cause serious adverse effects. Currently, most modular multiplication circuits are designed based on the Montgomery algorithm and its variants. When implemented in hardware, a good balance between circuit performance and circuit area is not achieved. There are problems with increased circuit performance resulting in a larger circuit area, or decreased circuit performance resulting in poorer circuit performance. Summary of the Invention

[0003] The present application provides a high-radix Montgomery modular multiplication circuit, which specifically includes: a loop control module, including m groups of loop control units, which are used to run loop operations in parallel to complete k+1 rounds of loop operations, and are used to cache intermediate data during the loop operations, read data required for the loop operations from the storage module, and transmit the corresponding data to the operation module according to the loop operation steps; a core state machine module, which is responsible for controlling the start and end of each round of first-level loop operations in the k+1 rounds of loop operations run by the m groups of loop control units in the loop control module; an operation module, which is used to receive data transmitted by the loop control module and perform multiplication and addition operations based on the received data; a storage module, which is used to store data required for loop operations for reading and calling by the loop control module; wherein the k+1 rounds of loop operations include k+1 rounds of first-level loop operations, and each round of first-level loop operations includes a pre-start loop operation and k rounds of second-level loop operations; the loop control module runs at most m rounds of loop operations in parallel at the same time; k and m are both positive integers, and k is a positive integer multiple of m.

[0004] The present application also provides a parallel operation method for a high-radix Montgomery modular multiplication circuit, specifically including: when the rth group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the r+1th group of loop control units to be in an idle state to start executing the next round of primary loop operations; wherein, when r is equal to m, when the rth group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the first group of loop control units to start executing the next round of primary loop operations; when the loop control unit starts to execute a round of primary loop operations, it switches from an idle state to a working state, and correspondingly, when the loop control unit ends a round of primary loop operations, it switches from a working state to an idle state.

[0005] The high-radix Montgomery modular multiplication circuit and its parallel operation method described in the present application achieve the technical effect of reducing circuit area while maintaining high-efficiency operation of circuit loop operations by having multiple groups of loop control units in the loop control module call the same storage module and realize parallel operation of multiple rounds of loop operations based on the same operation module. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] FIG1 is a module diagram of a high-radix Montgomery modular multiplication circuit according to an embodiment of the present application.

[0007] FIG2 is a schematic diagram of a module of a computing module according to an embodiment of the present application.

[0008] FIG3 is a module diagram of a calculation module according to another embodiment of the present application.

[0009] FIG4 is a module schematic diagram of a circulation control unit according to an embodiment of the present application. DETAILED DESCRIPTION

[0010] The following will describe the embodiments of the present application in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described below are only used to explain the present application and are not intended to limit the present application.

[0011] Currently, most modular multiplication circuits are designed based on the Montgomery algorithm and its variants. However, when implemented in hardware, they often fail to achieve a good balance between circuit performance and circuit area. This results in a larger circuit area when circuit performance is improved, or poorer circuit performance when circuit area is reduced. To improve the balance between circuit area and circuit performance in high-radix Montgomery modular multiplication circuits, the present application provides a high-radix Montgomery modular multiplication circuit designed to achieve the technical effect of reducing circuit area while maintaining high efficiency in circuit loop operations by enabling multiple groups of loop control units in a loop control module to call the same storage module and implement multiple rounds of loop operations in parallel based on the same operation module.

[0012] Specifically, as shown in FIG1 , the high-radix Montgomery modular multiplication circuit, as shown in FIG1 , specifically includes:

[0013] A loop control module includes m groups of loop control units, which are used to run loop operations in parallel to complete k+1 rounds of loop operations, and are used to cache intermediate data during the loop operations, read data required for the loop operations from the storage module, and transmit the corresponding data to the operation module according to the loop operation steps; wherein the m groups of loop control units in the loop control module can run at most m rounds of loop operations in parallel at the same time; the intermediate data cached during the loop operations refers to data transmitted by the loop control units according to the loop operation steps, and the operation module feeds back to the corresponding loop control units based on the operation results; the intermediate data can be used in subsequent loop operation steps;

[0014] The core state machine module is responsible for controlling the start and end of each round of first-level loop operations in the k+1 rounds of loop operations run by the m groups of loop control units in the loop control module; specifically, in order to enable the k+1 rounds of loop operations to be smoothly and parallelly operated in the m groups of loop control units, the core state machine module controls the node at which each group of loop control units starts to execute a round of loop operations; wherein, the core state machine module controls the start execution node of a new round of loop operations to be, but not limited to, controlling the first group of loop control units to start executing the first round of loop operations after the high-radix Montgomery modular multiplication circuit is powered on, or controlling the next group of loop control units to start executing the next round of loop operations after the specified intermediate data operation is completed during the execution of a round of loop operations by the current group of loop control units; the completion of the specified intermediate data operation can be, but not limited to, the completion of the first round of second-level loop operations during a round of loop operations.

[0015] The calculation module is configured to receive data transmitted by the loop control module and perform multiplication and addition operations based on the received data. The calculation module includes at least a multiplier and an adder to perform these operations, with the number of multipliers and adders configured based on actual computational requirements. The loop control module implements time-sharing multiplexing on the calculation modules, achieving the technical effect of enabling simple calculation modules to perform complex calculations through time-sharing multiplexing, thereby optimizing the algorithmic efficiency of the high-radix Montgomery modular multiplication circuit.

[0016] The storage module is used to store the data required for loop operation for reading and calling by the loop control module; the storage module is called by the m groups of loop control units in the loop control module when performing parallel loop operation, realizing data reuse of storage space, improving storage space utilization, ensuring circuit performance while reducing the required circuit area.

[0017] The high-radix Montgomery modular multiplication circuit described in this application is used to implement the high-radix Montgomery modular multiplication algorithm. The high-radix Montgomery modular multiplication algorithm is an efficient modular multiplication algorithm that can perform modular multiplication operations without using division and is widely used in the field of cryptography. Wherein, the k is equal to the radix of the high-radix Montgomery modular multiplication algorithm, and 2 k A k+1-round loop operation consists of k+1 rounds of primary loop operations. Each round of primary loop operations includes a pre-start loop operation and k rounds of secondary loop operations. Both k and m are positive integers, and k is a positive integer multiple of m.

[0018] As a preferred embodiment of the present application, as shown in FIG1 , the storage module includes three single-port random access static memories, one dual-port random access static memory, and two registers; wherein the three single-port random access static memories include:

[0019] The first single-port random static memory is used to store k+1 groups of multipliers A, each group includes 2 k bit multipliers A, wherein the lower k groups of multipliers A are valid data and the k+1th group of multipliers A is 0; the first single-port random static memory is used for each group of loop control units to read and call a corresponding group of multipliers A when performing the above-mentioned pre-start loop operation and the secondary loop operation; the number of groups of multipliers A called by the loop control unit from the first single-port random static memory is determined based on the number of primary loop operation rounds and the number of secondary loop operation rounds performed by it; preferably, the loop control unit calls the same number of groups of multipliers A from the first single-port random static memory based on the number of execution rounds of its primary loop operation, for example, when the loop control unit performs the first round of primary loop operation, it calls the first group of multipliers A from the first single-port random static memory for the first round of primary loop operation.

[0020] The second single-port random static memory is used to store k+1 groups of multipliers B, each group includes 2 kbit multiplier B, the lower k groups of multipliers B are valid data, and the k+1th group of multipliers B is 0; the second single-port random static memory is used for each group of loop control units to read and call a corresponding group of multipliers B when performing a pre-start loop operation and a secondary loop operation; when the number of groups of multipliers B called by the loop control unit from the second single-port random static memory is used to perform a pre-start loop operation, the first group of multipliers B is called; when the number of groups of multipliers B called by the loop control unit from the second single-port random static memory is used to perform a secondary loop operation, it is determined based on the number of secondary loop operation rounds performed by it; preferably, the loop control unit calls, based on the number of execution rounds of its secondary loop operation, a group of multipliers B whose number is increased by 1 compared with the number of execution rounds of the secondary loop operation from the second single-port random static memory, such as: when the loop control unit performs the first round of secondary loop operation, the second group of multipliers B is called from the second single-port random static memory for the first round of secondary loop operation.

[0021] The third single-port random static memory is used to store k+1 groups of moduli N, each group includes 2 k bit modulus N, the lower k groups of modulus N are valid data, and the k+1th group of modulus N is 0; the third single-port random static memory is used for each group of loop control units to read and call a corresponding group of modulus N when performing the above-mentioned pre-start loop operation and secondary loop operation; when the loop control unit calls the modulus N from the third single-port random static memory for performing the pre-start loop operation, the first group of modulus N is called; when the loop control unit calls the modulus N from the third single-port random static memory for performing the secondary loop operation, the number of calling groups of modulus N is determined based on the number of execution rounds of the secondary loop operation of the loop control unit; preferably, the loop control unit calls the number of groups of modulus N increased by 1 compared with the number of execution rounds of the secondary loop operation from the third single-port random static memory based on the number of execution rounds of its secondary loop operation, such as: when the loop control unit performs the first round of secondary loop operation, the second group of modulus N is called from the third single-port random static memory for the first round of secondary loop operation.

[0022] A dual-port random access static memory is used to store k+1 groups of initial results S and also to store the results of each round of secondary loop operations; wherein the k+1 groups of initial results S are used by the loop control unit when performing the first round of primary loop operations; when the loop control unit performs the second to k+1th rounds of primary loop operations, it is implemented based on the results of each round of secondary loop operations stored in the dual-port random access static memory in the previous round of loop operations.

[0023] The two registers include: a constant q register and a result S register; wherein the constant q register is used to store a set of k-bit constants q, and the calculation formula of the constant q is: q = -N -1 mod2 kThe result S register is used to store the result S(i) of each round of the first-level loop operation. In this embodiment, the memory used to store the result S of each round of the second-level loop operation adopts a dual-port random access static memory to ensure that the memory storing the result S has the parallel execution capability of reading and storing, so that when multiple groups of loop control units are stored in parallel in the loop control module, the initial result S can be called and the second-level loop operation result can be written to the dual-port random access static memory at the same time.

[0024] As a preferred embodiment of the present application, as shown in Figure 2, the operation module includes a first multiplier, a second multiplier, and an adder. Specifically, the first multiplier includes a first input terminal, a second input terminal, and a first output terminal; the second multiplier includes a third input terminal, a fourth input terminal, and a second output terminal; and the adder includes a fifth input terminal, a sixth input terminal, a seventh input terminal, an eighth input terminal, and a third output terminal. The first input terminal, the second input terminal, the third input terminal, the fourth input terminal, the fifth input terminal, and the sixth input terminal serve as input ports of the operation module, connected to the loop control module, and configured to receive data transmitted by the loop control module.

[0025] Specifically, the first output of the first multiplier is connected to the seventh input of the adder, enabling the first multiplier to transmit data to the adder, typically for transmitting data after multiplication by the first multiplier to the adder for addition. The second output of the second multiplier is connected to the eighth input of the adder, enabling the second multiplier to transmit data to the adder, typically for transmitting data after multiplication by the second multiplier to the adder for addition. The third output of the adder serves as an output port of the operation module for outputting the added data. In this embodiment, the operation module is limited to including two multipliers and one adder, so that the two multipliers can perform multiplication operations in parallel.

[0026] Preferably, in some embodiments of the present application, in order to improve the parallel computing capability, a plurality of operation modules including two multipliers and one adder can be configured in the high-radix Montgomery modular multiplication circuit, and m groups of loop operation units are respectively configured to correspond to one operation module according to a preset number of groups, such as: one operation module including two multipliers and one adder is configured for every 4 groups of loop operation units, so as to improve the parallel computing efficiency of multiple groups of loop operations in the loop control module.

[0027] As a preferred embodiment of the present application, as shown in FIG3 , the operation module further includes: a first register, a second register, and a third register; wherein the first register is arranged between the first multiplier and the adder, the first output end is connected to the input end of the first register, and the output end of the first register is connected to the seventh input end of the adder; the second register is arranged between the second multiplier and the adder, the second output end is connected to the input end of the second register, and the output end of the second register is connected to the eighth input end of the adder; the input end of the third register is connected to the third output end of the adder, and the output end of the third register serves as the output port of the operation module. This embodiment, by arranging the first register between the first multiplier and the adder, the second register between the second multiplier and the adder, and the third register between the adder output and the cyclic operation device, interrupts the first multiplier / second multiplier and the adder by the register, effectively improving the operation execution efficiency of the operation module.

[0028] As a preferred embodiment of the present application, as shown in FIG3 , the first register further includes a fourth output terminal, which serves as an output port of the operation module for outputting data obtained through the multiplication operation of the first multiplier. This embodiment, based on the existence of an operation step in the pre-start loop operation of the high-radix Montgomery modular multiplication algorithm that does not require an addition operation, provides a fourth output terminal on the first register, allowing the fourth output terminal to directly output the data obtained through the multiplication operation of the first multiplier outside the operation module, without passing through an adder and directly feeding back to the loop operation device, thereby improving the feedback efficiency of the intermediate operation results and reducing unnecessary occupation of operation resources.

[0029] As a preferred embodiment of the present application, as shown in FIG4 , each group of loop control units includes a loop controller, which is used to respond to the start control signal of the first-level loop operation of the core state machine module, read the data stored in the storage module, and perform the first-level loop operation based on the data stored in the read storage module using the first multiplier, the second multiplier, and the adder in the time-division multiplexing operation module, and is also used to control the start and end of each round of the second-level loop operation. Specifically, when the loop controller responds to the start control signal of the first-level loop operation of the core state machine module, it controls the group of loop control units to perform a corresponding round of the first-level loop operation. The loop controller reads and calls the corresponding data from the storage module based on the execution progress of the first-level loop operation. For example, when executing the pre-start loop operation, the first group of results S(1) in the dual-port random static memory is read from the storage module, the corresponding group of multipliers A(i) in the first single-port random static memory is read according to the execution round number i of the first-level loop operation, and the first group of multipliers B(1) in the second single-port random static memory is read. The loop controller controls the start and end of each round of secondary loop calculations based on the progress of the pre-start loop calculations executed by the loop control unit. When the pre-start loop calculations are completed, the loop controller controls the start of the first round of secondary loop calculations. When the results of a round of secondary loop calculations are obtained, the loop controller controls the end of that round of secondary loop calculations and the start of the next round of secondary loop calculations. When the results of the kth round of secondary loop calculations are obtained, the loop controller controls the end of the kth round of secondary loop calculations. In the loop control unit provided in this embodiment, the loop controller controls the data transmission and reception between the loop control unit and the storage module, the core state machine, and the computing module.

[0030] As a preferred embodiment of the present application, as shown in FIG4 , each group of loop control units further includes: an intermediate data cache module for caching data read from the storage module and intermediate computation data fed back by the computation module during the loop computation process, and for the loop controller to read and call; wherein the intermediate data cache module includes:

[0031] The first calculation result t register is used to cache the first calculation result t fed back by the operation module; wherein the first calculation result t refers to the first calculation result t obtained by the loop operation unit when executing the first calculation step of a round of pre-start loop operation.

[0032] A second calculation result u register is used to cache a second calculation result u fed back by the operation module; wherein the second calculation result u refers to a second calculation result u obtained by the loop operation unit performing a second calculation step of a round of pre-start loop operation;

[0033] The third calculation result d register is used to cache the third calculation result d fed back by the operation module; wherein the third calculation result d refers to the third calculation result d obtained by the loop operation unit when executing the third calculation step of a round of pre-start loop operation; the third calculation result d is used in k rounds of secondary loop operations in the same round of primary loop operation.

[0034] The multiplier A register is used to cache the multiplier A read by the loop control unit from the first single-port random access static memory of the storage module. The multiplier B register is used to cache the multiplier B read by the loop control unit from the second single-port random access static memory of the storage module. The modulus N register is used to cache the modulus N read by the loop control unit from the third single-port random access static memory of the storage module. The intermediate result S register is used to cache the result S read by the loop control unit from the dual-port random access static memory of the storage module. This embodiment provides an intermediate data cache module in each group of loop control units. On the one hand, it achieves caching of the first calculation result t, second calculation result u, and third calculation result d fed back by the operation module to facilitate the call of subsequent loop operation steps. On the other hand, it achieves the call cache of the multiplier A, multiplier B, modulus N, and result S in the storage module, allowing the data stored in the storage module to be read and called by multiple groups of loop control units running in parallel. Based on the caching of data required during the loop operation by the intermediate data cache module, the parallel operation efficiency of the m groups of loop control units is improved.

[0035] Preferably, the first calculation result register t, the second calculation result register u, the third calculation result register d, the multiplier A register, the multiplier B register, the modulus N register and the intermediate result S register included in the intermediate data cache module are all registers that can cache k-bit data; when there is new k-bit data to be read or written, the new data is cached by overwriting the old data.

[0036] As a preferred embodiment of the present application, the high-radix Montgomery modular multiplication circuit further includes:

[0037] The first input end of the first multiplier is respectively connected to the multiplier A register in each group of loop control units to receive the multiplier A cached in the multiplier A register transmitted by the loop control unit; the first input end of the first multiplier is also respectively connected to the first calculation result t register in each group of loop control units to receive the first calculation result t cached in the first calculation result t register transmitted by the loop control unit; the second input end of the first multiplier is respectively connected to the multiplier B register in each group of loop control units to receive the multiplier B cached in the multiplier B register transmitted by the loop control unit.

[0038] The third input end of the second multiplier is respectively connected to the second calculation result u register in each group of loop control units to receive the second calculation result u cached in the second calculation result u register transmitted by the loop control unit; the fourth input end of the second multiplier is respectively connected to the modulus N register in each group of loop control units to receive the modulus N cached in the modulus N register transmitted by the loop control unit.

[0039] The fifth input of the adder is connected to the intermediate result S register in each group of loop control units to receive the result S cached in the intermediate result S register transmitted by the loop control unit; the sixth input of the adder is connected to the third calculation result d register in each group of loop control units to receive the third calculation result d cached in the third calculation result d register transmitted by the loop control unit. This embodiment limits the registers in the loop control unit to which the inputs of the first multiplier, the second multiplier, and the adder are correspondingly connected, so that the first multiplier is only used to perform the multiplication operation of the multiplier A and the multiplier B and the multiplication operation of the first calculation result t and the constant q, and the second multiplier is only used to perform the multiplication operation of the second calculation result u and the modulus N, thereby distinguishing the multiplication operation steps executed by the first multiplier and the second multiplier in the loop operation, and avoiding the situation where the execution of the first multiplier and the second multiplier is confused due to the parallel operation of multiple groups of loop control units.

[0040] As a preferred embodiment of the present application, a parallel operation method for high-radix Montgomery modular multiplication is provided, specifically comprising: when the rth group of loop control units completes the first round of secondary loop operations in the first round of loop operations, the core state machine controls the r+1th group of loop control units to be in an idle state and start executing the next round of primary loop operations; wherein, when r equals m, then when the rth group of loop control units completes the first round of secondary loop operations in the first round of loop operations, the core state machine controls the first group of loop control units to start executing the next round of primary loop operations. Specifically, when the loop control unit starts executing a round of primary loop operations, it transitions from an idle state to a working state, and correspondingly, when the loop control unit completes a round of primary loop operations, it transitions from a working state to an idle state. Because the pre-start loop calculation step on the third calculation result d in a round of loop operations requires the results of the first secondary loop in the previous round of loop operations, and the execution of the secondary loop operations requires the results of the second through kth secondary loops in the previous round of loop operations, this embodiment uses the completion of the first secondary loop operation in a primary loop operation as the trigger for the start of the next primary loop operation. This ensures that when the next loop operation starts, it will have at least the data required for the pre-start loop calculation step on the third calculation result d. This improves parallel computing efficiency while ensuring the smooth execution of a single round of loop operations.

[0041] As a preferred embodiment of the present application, the logic of the core state machine module controlling the loop control unit to start executing the first-level loop operation is based on whether the first round of secondary loop operations in the first-level loop operation is completed; specifically, the core state machine module controls the first group of loop control units to start executing the first round of first-level loop operations. When the first group of loop control units completes the first round of secondary loop operations in the first round of first-level loop operations, the core state machine module controls the second group of loop control units to start executing the second round of first-level loop operations. When the second group of loop control units completes the first round of secondary loop operations in the second round of first-level loop operations, the core state machine module controls the third group of loop control units to start executing the second round of first-level loop operations, and so on. When the m-1th group of loop control units completes the first round of secondary loop operations in the m-1th round of first-level loop operations, the core state machine controls the mth group of loop control units to start executing the mth round of first-level loop operations.

[0042] As a preferred embodiment of the present application, a set of loop control units performs a round of primary loop operations, specifically including: Step 1: performing a pre-start loop operation and obtaining a pre-start loop operation result; Step 2: performing k rounds of secondary loop operations based on the pre-start loop operation result. The pre-start loop operation result refers to the obtained third calculation result d. This embodiment divides a round of primary loop operations into a pre-start loop operation and k rounds of secondary loop operations. The k rounds of secondary loop operations are executed based on the pre-start loop operation result, thereby nesting k rounds of secondary loop operations within each round of primary loop operations.

[0043] As a preferred embodiment of the present application, the method of performing the pre-start cycle operation and obtaining the pre-start cycle operation result in step 1 specifically includes:

[0044] Step 11: The loop control unit reads the multiplier A(i) stored in the first single-port random access static memory, the multiplier B(1) stored in the second single-port random access static memory, and the result S(1) stored in the dual-port random access static memory and transmits the results to the operation module to obtain a first calculation result t of the pre-start loop operation;

[0045] Step 12: The loop control unit reads the constant q stored in the constant q register and transmits the constant q in combination with the first calculation result t of the pre-start loop operation to the operation module to obtain the second calculation result u of the pre-start loop operation;

[0046] Step 13: The loop control unit reads the modulus N(1) stored in the third single-port random access static memory, combines the multiplier A(i) stored in the first single-port random access static memory, the multiplier B(1) stored in the second single-port random access static memory, the result S(1) stored in the dual-port random access static memory, and the second calculation result u of the pre-start loop operation, and transmits them to the operation module to obtain a third calculation result d of the pre-start loop operation.

[0047] Specifically, i in the multiplier A(i) refers to the number of the first-level loop operation round currently executed by the loop control unit; the multiplier A(i) refers to the i-th group of multipliers A pre-stored in the first single-port random static memory; the multiplier B(1) refers to the first group of multipliers B pre-stored in the second single-port random static memory; the result S(1) refers to the result of the first round of the second-level loop operation in the previous round of loop operation. In particular, if the first round of loop operation is currently being executed, the result S(1) is the first group of initial results S pre-stored in the dual-port random static memory.

[0048] As a preferred implementation of the present application, step 11 specifically includes:

[0049] The loop control unit initiates a read operation of a multiplier A(i) to a first single-port random access static memory of a storage module based on the number i of first-level loop rounds executed by the loop control unit, and initiates a read operation of a multiplier B(1) to a second single-port random access static memory of the storage module; wherein the multiplier A(i) refers to the i-th group of multipliers A pre-stored in the first single-port random access static memory; and the multiplier B(1) refers to the first group of multipliers B pre-stored in the second single-port random access static memory.

[0050] The loop control unit caches the read multiplier A(i) into the multiplier A register and caches the read multiplier B(1) into the multiplier B register. At the same time, the loop control unit initiates a result S(1) read operation to the dual-port random access static memory; wherein the result S(1) refers to the result of the first round of secondary loop operation in the previous round of loop operation. In particular, if the first round of loop operation is currently being executed, the result S(1) is the first group of initial results S pre-stored in the dual-port random access static memory.

[0051] The loop control unit caches the read result S(1) into the intermediate result S register, and at the same time, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(1);

[0052] The first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal, and the loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, so that the adder performs an addition operation on the product of the multiplier A(i) and the multiplier B(1) and the result S(1) and transmits the addition operation result as the first calculation result t to the corresponding group of loop control units, and the loop control unit caches the first calculation result t in the first calculation result t register; wherein i is a positive integer less than or equal to k+1. In this embodiment, the calculation step of the first calculation result of the pre-start loop in the high-radix Montgomery modular multiplication algorithm is divided into two steps. The first step is to calculate the product of the multiplier A(i) and the multiplier B(1), and the second step is to calculate the sum of the product of the multiplier A(i) and the multiplier B(1) and the result S(1) based on the product of the multiplier A(i) and the multiplier B(1), thereby obtaining the first calculation result t.

[0053] As a preferred implementation of the present application, step 12 specifically includes: the loop control unit transmits the first calculation result t cached in the first calculation result t register to the first multiplier through the first input end, the loop control unit reads the constant q from the constant q register and transmits it to the first multiplier through the second input end, so that the first multiplier performs a multiplication operation of the first calculation result t and the constant q, and the first multiplier transmits the multiplication result of the first calculation result t and the constant q as the second calculation result u to a corresponding group of loop control units.

[0054] As a preferred implementation of the present application, step 13 specifically includes:

[0055] The loop control unit caches the second calculation result u to the second calculation result u register, and at the same time, the loop control unit initiates a modulus N (1) read operation to the third single-port random static memory of the storage module; the loop control unit caches the read modulus N (1) to the modulus N register; wherein, the modulus N (1) refers to the first group of modulus N pre-stored in the third single-port random static memory.

[0056] The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and the loop control unit transmits the modulus N(1) cached in the modulus N register to the second multiplier through the fourth input terminal, so that the second multiplier performs a multiplication operation of the second calculation result u and the modulus N(1).

[0057] The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, and at the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs an addition operation on the result S(1), the product of the multiplier A(i) and the multiplier B(1), and the product of the modulus N(1) and the second calculation result u, and transmits the addition operation result as a third calculation result d to a corresponding group of loop control units; the loop control unit caches the third calculation result d in the third calculation result d register, and obtains the third calculation result d as a pre-start loop operation result.

[0058] As a preferred implementation of the present application, step 2 of performing k rounds of secondary loop calculations based on the pre-start loop calculation results specifically includes:

[0059] The loop control unit initiates a read operation of the multiplier B(j+1) to the second single-port random access static memory of the storage module based on the number j of execution rounds of the secondary loop operation, and initiates a read operation of the modulus N(j+1) to the third single-port random access static memory of the storage module; wherein j is a positive integer less than or equal to k. The j+1 is determined based on the number j of execution rounds of the secondary loop operation of the loop control unit, the multiplier B(j+1) refers to the j+1th group of multipliers B pre-stored in the second single-port random access static memory, and the modulus N(j+1) refers to the j+1th group of moduli N pre-stored in the third single-port random access static memory.

[0060] The loop control unit caches the read multiplier B(j+1) into the multiplier B register, caches the read modulus N(j+1) into the modulus N register, and initiates a result S(j+1) read operation to the dual-port random access static memory; wherein the result S(j+1) refers to the j+1th group of results S stored in the dual-port random access static memory; specifically, when the loop control unit executes the first round of the first-level loop operation, the result S(j+1) refers to the j+1th group of initial results S pre-stored in the dual-port random access static memory; on the contrary, when the loop control unit does not execute the first round of the first-level loop operation, the result S(j+1) refers to the j+1th group of results S among the k groups of results S stored in the dual-port random access static memory based on the results of the second-level loop operation in the previous round of loop operation.

[0061] The loop control unit caches the read result S(j+1) to the intermediate result S register, and the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input end, and transmits the multiplier B(j+1) cached in the multiplier B register to the first multiplier through the second input end, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(j+1) to obtain the product of the multiplier A(i) and the multiplier B(j+1); the loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input end, and transmits the modulus N(j+1) cached in the modulus N register to the second multiplier through the fourth input end, so that the second multiplier performs a multiplication operation of the second calculation result u and the modulus N(j+1) to obtain the product of the second calculation result u and the modulus N(j+1).

[0062] The loop control unit transmits the result S(j+1) cached in the intermediate result S register to the adder through the fifth input terminal, and the loop control unit transmits the third calculation result d cached in the third calculation result d register to the adder through the sixth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(j+1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(j+1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs an addition operation on the result S(1), the product of the multiplier A(i) and the multiplier B(j+1), the product of the modulus N(j+1) and the second calculation result u, and the pre-start loop operation result third calculation result d, and caches the addition operation result as the result S(j) in the dual-port random static memory for reading in the next round of loop operation, thereby completing one round of secondary loop operation.

[0063] Determine whether the number of secondary loop operation rounds j is equal to k. If so, terminate step 2. If not, increment the number of secondary loop operation rounds j by 1 and proceed to the next secondary loop operation. In this embodiment, the loop controller determines whether the number of secondary loop operation rounds j is equal to k, thereby controlling the termination of the secondary loop operation according to the number of secondary loop operation rounds. If the number of secondary loop operation rounds has not reached k, control the loop control unit to continue executing the secondary loop operation.

[0064] As a preferred embodiment of the present application, step 2 of executing k rounds of secondary loop operations based on the pre-start loop operation result also includes: when the number of secondary loop operation execution rounds j is equal to k, the loop control unit transmits the addition operation result as the current round of primary loop operation result to the result S register of the storage module for cache.

[0065] As a preferred embodiment of the present application, there is an upper limit on the number of primary loop operations that can be completed by each group of loop control units; wherein, the number of primary loop operations that can be completed by the first group of loop control units is equal to 1 + k / m rounds; and the number of primary loop operations that can be completed by the remaining groups of loop control units is equal to k / m rounds; wherein k / m is a positive integer. For example, in some embodiments, the loop control module is configured to have four groups of loop control units, i.e., m = 4, and a radix-32 Montgomery modular multiplication algorithm is implemented based on the four groups of loop control units, i.e., k = 32. In this loop control module, the first group of loop control units is used to implement nine rounds of primary loop operations, and the second to fourth groups of loop control units are used to implement eight rounds of primary loop operations.

[0066] As a preferred embodiment of the present application, in order to prevent the reading, writing, and calling of the storage module and the operation module from interfering with each other when multiple groups of loop control units are running in parallel, the timing of the m groups of loop control units in the loop control module time-sharing the first multiplier, the second multiplier, and the adder in the operation module is limited, the timing of the time-sharing reading of the first single-port random access static memory, the second single-port random access static memory, the third single-port random access static memory, the dual-port random access static memory, and the constant q in the storage module is limited, and the timing of the time-sharing writing of the result S register in the storage module is limited. This ensures that only one loop control module performs a read operation on the first single-port random access static memory / the second single-port random access static memory / the third single-port random access static memory in each clock cycle, and ensures that only one loop control module performs a read operation and a write operation on the dual-port random access static memory in each clock cycle. The specific limiting logic can be, but is not limited to: the timing is divided into 4 time units into 1 group, and in each group of timing, the first time unit allows a read operation to be performed on the dual-port random access static memory; the second time unit allows the call of the multiplier; the third time unit allows the call of the adder; the fourth time unit allows a read operation to be performed on the first single-port random access static memory, the second single-port random access static memory, and the third single-port random access static memory, and a write operation to the dual-port random access static memory. This embodiment controls the loop control unit to perform read calls and transfer operations on each memory and register in a time-sharing manner, so that the storage module can be read and written in a time-sharing manner and the operation module can be multiplexed in a time-sharing manner, thereby improving the multiplexing efficiency of the storage module and the operation module when multiple groups of loop control units operate in parallel.

[0067] Those skilled in the art will appreciate that all or part of the steps in the above-described method embodiments can be implemented by hardware associated with program instructions. These programs can be stored in a computer-readable storage medium (e.g., ROM, RAM, magnetic disk, optical disk, or other medium capable of storing program code). When executed, the program performs the steps of the above-described method embodiments.

[0068] It should be noted that the aforementioned loop control modules, calculation modules, and storage modules, etc., may be, but are not limited to, digital circuit modules described by a designer using the hardware description language Verilog HDL, or digital circuit modules created by a designer by drawing or compiling circuits using software with circuit drawing or compilation capabilities. Furthermore, the functional units in various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A high-radix Montgomery modular multiplication circuit, characterized in that: The high-base Montgomery modular multiplication circuit comprises: A loop control module, including m groups of loop control units, is used to run loop operations in parallel to complete k+1 rounds of loop operations, and is used to cache intermediate data during the loop operations, read data required for the loop operations from the storage module, and transmit the corresponding data to the operation module according to the loop operation steps; The core state machine module is responsible for controlling the start and end of each round of primary loop operation in the k+1 rounds of loop operations run by the m groups of loop control units in the loop control module; An operation module, used for receiving data transmitted by the cycle control module, and performing multiplication and addition operations based on the received data; The storage module is used to store the data required for the loop operation so that the loop control module can read and call it; Among them, the k+1 rounds of cyclic operations include k+1 rounds of primary cyclic operations, each round of primary cyclic operations includes a pre-start cyclic operation and k rounds of secondary cyclic operations; the cyclic control module runs at most m rounds of cyclic operations in parallel at the same time; k and m are both positive integers, and k is a positive integer multiple of m.

2. The high-radix Montgomery modular multiplication circuit according to claim 1, characterized in that: The storage module includes 3 single-port random static memories, 1 dual-port random static memory and 2 registers; The three single-port random static memories include a first single-port random static memory for storing k+1 groups of multipliers A, a second single-port random static memory for storing k+1 groups of multipliers B, and a third single-port random static memory for storing k+1 groups of moduli N; The dual-port random static memory is used to store k+1 groups of initial results S and also to store the results of each round of secondary loop operation; the k+1 groups of initial results S are used for the first round of loop operation; The two registers include a constant q register for storing a group of constants q and a result S register for storing the result of each round of first-level loop operation.

3. The high-radix Montgomery modular multiplication circuit according to claim 1, characterized in that: The operation module includes a first multiplier, a second multiplier and an adder; The first multiplier includes a first input terminal, a second input terminal and a first output terminal; The second multiplier includes a third input terminal, a fourth input terminal and a second output terminal; The adder comprises a fifth input terminal, a sixth input terminal, a seventh input terminal, an eighth input terminal and a third output terminal; The first input terminal, the second input terminal, the third input terminal, the fourth input terminal, the fifth input terminal and the sixth input terminal are used as input ports of the operation module to receive data transmitted by the cycle control module; The first output terminal of the first multiplier is connected to the seventh input terminal of the adder, and is used to transmit the data after the multiplication operation of the first multiplier to the adder for addition operation; The second output terminal of the second multiplier is connected to the eighth input terminal of the adder, and is used to transmit the data after the multiplication operation of the second multiplier to the adder for addition operation; The third output terminal of the adder is used as an output port of the operation module to output the data after the addition operation.

4. The high-radix Montgomery modular multiplication circuit according to claim 3, characterized in that: The operation module also includes: a first register, a second register and a third register; wherein the first register is arranged between the first multiplier and the adder, the first output end is connected to the input end of the first register, and the output end of the first register is connected to the seventh input end of the adder; the second register is arranged between the second multiplier and the adder, the second output end is connected to the input end of the second register, and the output end of the second register is connected to the eighth input end of the adder; the input end of the third register is connected to the third output end of the adder, and the output end of the third register serves as the output port of the operation module.

5. The high-radix Montgomery modular multiplication circuit according to claim 4, characterized in that: The first register further includes a fourth output terminal, which is used as an output port of the operation module to output data obtained through the multiplication operation of the first multiplier.

6. The high-radix Montgomery modular multiplication circuit according to claim 3, characterized in that: Each group of loop control units includes a loop controller, which is used to respond to the start control signal of the first-level loop operation of the core state machine module, read the data stored in the storage module, and perform the first-level loop operation by the first multiplier, the second multiplier and the adder in the time-sharing multiplexing operation module based on the data stored in the read storage module, and is also used to control the start and end of each round of second-level loop operation.

7. The high-radix Montgomery modular multiplication circuit according to claim 6, characterized in that: Each group of loop control units also includes: an intermediate data cache module, which is used to cache the data read from the storage module and the intermediate data fed back by the operation module during the loop operation process, and is read and called by the loop control unit; wherein the intermediate data cache module includes: A first calculation result t register, used for caching a first calculation result t fed back by the operation module; A second calculation result u register, used for caching a second calculation result u fed back by the operation module; A third calculation result d register, used for caching the third calculation result d fed back by the operation module; A multiplier A register, used for caching the multiplier A read by the loop control module from the first single-port random static memory of the storage module; A multiplier B register, used for caching the multiplier B read by the loop control module from the second single-port random static memory of the storage module; A modulus N register, used for caching the modulus N read by the loop control module from the third single-port random static memory of the storage module; The intermediate result S register is used to cache the result S read by the loop control module from the dual-port random static memory of the storage module.

8. The high-radix Montgomery modular multiplication circuit according to claim 7, characterized in that: The high-radix Montgomery modular multiplication circuit further comprises: The first input end of the first multiplier is respectively connected to the multiplier A register in each group of loop control units to receive the multiplier A cached in the multiplier A register transmitted by the loop control module; The first input end of the first multiplier is also connected to the first calculation result t register in each group of loop control units respectively, so as to receive the first calculation result t cached in the first calculation result t register transmitted by the loop control module; The second input end of the first multiplier is respectively connected to the multiplier B register in each group of loop control units to receive the multiplier B cached in the multiplier B register transmitted by the loop control module; The third input terminal of the second multiplier is respectively connected to the second calculation result u register in each group of loop control units to receive the second calculation result u cached in the second calculation result u register transmitted by the loop control module; The fourth input terminal of the second multiplier is respectively connected to the modulus N register in each group of loop control units to receive the modulus N cached in the modulus N register transmitted by the loop control module; The fifth input terminal of the adder is respectively connected to the intermediate result S register in each group of loop control units to receive the result S cached in the intermediate result S register transmitted by the loop control module; The sixth input terminal of the adder is respectively connected to the third calculation result d register in each group of loop control units to receive the third calculation result d cached in the third calculation result d register transmitted by the loop control module.

9. A parallel operation method for a high-radix Montgomery modular multiplication circuit, characterized in that: The parallel operation method of the high-radix Montgomery modular multiplication circuit specifically includes: When the rth group of loop control units completes the first round of secondary loop operations in the primary loop operations, the core state machine controls the r+1th group of loop control units to be in an idle state to start executing the next round of primary loop operations; Among them, when r is equal to m, when the rth group of loop control units completes the first round of secondary loop operations in the first-level loop operations, the core state machine controls the first group of loop control units to start executing the next round of primary loop operations; when the loop control unit starts to execute a round of primary loop operations, it is converted from an idle state to a working state, and accordingly, when the loop control unit ends a round of primary loop operations, it is converted from a working state to an idle state.

10. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 9, characterized in that: A set of loop control units performs a round of primary loop operations, including: Step 1: Execute the pre-start cycle operation and obtain the pre-start cycle operation result; Step 2: Perform k rounds of secondary loop calculations based on the pre-start loop calculation results to obtain primary loop calculation results, thereby ending a round of primary loop calculations.

11. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 10, characterized in that: The method of performing the pre-start cycle operation and obtaining the pre-start cycle operation result described in step 1 specifically includes: Step 11: the loop control unit reads the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, and the result S(1) stored in the dual-port random static memory and transmits them to the operation module to obtain the first calculation result t of the pre-start loop operation; Step 12: The loop control unit reads the constant q stored in the constant q register and transmits the constant q in combination with the first calculation result t of the pre-start loop operation to the operation module to obtain the second calculation result u of the pre-start loop operation; Step 13: The loop control unit reads the modulus N(1) stored in the third single-port random static memory, combines the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, the result S(1) stored in the dual-port random static memory, and the second calculation result u of the pre-start loop operation and transmits it to the operation module to obtain the third calculation result d of the pre-start loop operation.

12. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 11, characterized in that: The step 11 specifically includes: The loop control unit initiates a read operation of a multiplier A(i) to a first single-port random static memory of the storage module based on the number of first-level loop rounds executed by the loop control unit, and initiates a read operation of a multiplier B(1) to a second single-port random static memory of the storage module; The loop control unit caches the read multiplier A(i) into the multiplier A register, and caches the read multiplier B(1) into the multiplier B register. At the same time, the loop control unit initiates a read operation of the result S(1) to the dual-port random static memory. The loop control unit caches the read result S(1) into the intermediate result S register, and at the same time, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(1); The first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder, and the loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, so that the adder performs an addition operation on the product of the multiplier A(i) and the multiplier B(1) and the result S(1) and transmits the addition operation result as a first calculation result t to a corresponding group of loop control units, and the loop control unit caches the first calculation result t of the pre-start loop operation to the first calculation result t register; wherein i is a positive integer less than or equal to k+1.

13. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 12, characterized in that: The step 12 specifically includes: The loop control unit transmits the first calculation result t cached in the first calculation result t register to the first multiplier through the first input terminal, the loop control unit reads the constant q from the constant q register and transmits it to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the first calculation result t and the constant q, and the first multiplier transmits the multiplication operation result of the first calculation result t and the constant q as the second calculation result u of the pre-start loop operation to a corresponding group of loop control units.

14. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 13, characterized in that: The step 13 specifically includes: The loop control unit caches the second calculation result u to the second calculation result u register, and at the same time, the loop control unit initiates a modulus N(1) read operation to the third single-port random static memory of the storage module; The loop control unit caches the read modulus N(1) into the modulus N register; The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and the loop control unit transmits the modulus N(1) cached in the modulus N register to the second multiplier through the fourth input terminal, so that the second multiplier performs a multiplication operation of the second calculation result u and the modulus N(1). The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, and at the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs an addition operation on the result S(1), the product of the multiplier A(i) and the multiplier B(1), and the product of the modulus N(1) and the second calculation result u, and transmits the addition operation result as the third calculation result d to a corresponding group of loop control units; The loop control unit caches the third calculation result d in the third calculation result d register, and obtains the third calculation result d as the pre-start loop operation result.

15. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 14, characterized in that: Step 2 performs k rounds of secondary loop calculation based on the pre-start loop calculation result, specifically including: The loop control unit initiates a multiplier B(j+1) read operation to the second single-port random static memory of the storage module based on the number of execution rounds of the secondary loop operation, and initiates a modulus N(j+1) read operation to the third single-port random static memory of the storage module; The loop control unit caches the read multiplier B(j+1) into the multiplier B register, caches the read modulus N(j+1) into the modulus N register, and initiates a result S(j+1) read operation to the dual-port random static memory; The loop control unit caches the read result S(j+1) into the intermediate result S register, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(j+1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation of the multiplier A(i) and the multiplier B(j+1), the loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and transmits the modulus N(j+1) cached in the modulus N register to the second multiplier through the fourth input terminal, so that the second multiplier performs a multiplication operation of the second calculation result u and the modulus N(j+1); The loop control unit transmits the result S(j+1) cached in the intermediate result S register to the adder through the fifth input terminal, and the loop control unit transmits the third calculation result d cached in the third calculation result d register to the adder through the sixth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(j+1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(j+1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs an addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(j+1), the product of the modulus N(j+1) and the second calculation result u, and the third calculation result d of the pre-start loop operation result, and caches the addition operation result as the result S(j) in the dual-port random static memory for reading in the next round of loop operation, thereby completing a round of secondary loop operation; Determine whether the number of execution rounds of the secondary loop operation is equal to k. If so, end step 2. If not, increase the value of the number of execution rounds of the secondary loop operation by 1 and enter the next round of secondary loop operation. Wherein, j is a positive integer less than or equal to k.

16. The parallel operation method of the high-radix Montgomery modular multiplication circuit according to claim 14, characterized in that: Step 2, performing k rounds of secondary loop operations based on the pre-start loop operation result, also includes: when the number of secondary loop operation execution rounds is equal to k, the loop control unit transmits the addition operation result as the current round of primary loop operation result to the result S register of the storage module for cache.

Citation Information

Patent Citations

  • Optimized Montgomery modular multiplication method, optimized modular square method and optimized modular multiplication hardware

    CN103761068A

  • Chip and batch modular operation method for chip

    CN113031920A

  • System and method for quickly realizing Montgomery modular multiplication by using multiplier

    CN114840174A

  • Montgomery modular multiplication method and device based on 2

    CN115268839A

  • Montgomery modular multiplier based on parallel structure and coding

    CN116893800A