High-base Montgomery modular multiplication circuit and parallel operation method thereof

By designing a high-key Montgomery modular multiplication circuit in the modular multiplication circuit, and using multiple sets of cycle control units to parallel operation, the problem of circuit performance and area imbalance in the prior art is solved, and an efficient and low-area circuit design is achieved.

CN120179209AActive Publication Date: 2025-06-20AMICRO SEMICONDUCTOR CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202311746840.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20
Estimated Expiration
2043-12-19

AI Technical Summary

Technical Problem

When existing modem multiplication circuits are implemented in hardware, the circuit performance and circuit area cannot be well balanced, and there is a problem that the circuit area is larger when the circuit performance is improved, or the circuit area is smaller and the circuit performance is poor.

Method used

A high-key Montgomery module multiplication circuit is designed, and the same storage module is called by multiple sets of cyclic control units in the cyclic control module and the parallel operation of multiple rounds of cyclic operations is realized based on the same operation module, reducing the circuit area while maintaining high efficiency.

Benefits of technology

It realizes the high-efficiency operation of circuit cyclic operations while reducing the circuit area, solving the problem of circuit performance and area balance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179209A_ABST
    Figure CN120179209A_ABST
Patent Text Reader

Abstract

The invention discloses a high-base Montgomery modular multiplication circuit and a parallel operation method thereof, and the circuit comprises a loop control module which comprises m groups of loop control units and is used for carrying out the parallel operation of loop operation to complete k + 1 rounds of loop operation, and is used for caching intermediate data in a loop operation process, reading data required by loop operation from the storage module and transmitting corresponding data to the operation module according to loop operation steps; the core state machine module is used for controlling the start and end of each round of first-stage cyclic operation in k + 1 rounds of cyclic operation operated by m groups of cyclic control units in the cyclic control module; the operation module is used for receiving the data transmitted by the loop control module and carrying out multiplication operation and additive operation based on the received data; and the storage module is used for storing data required by loop operation for the loop control device to read and call. The circuit area is reduced based on parallel operation of the m groups of cycle control units, and high-efficiency operation of cycle operation of the circuit is maintained at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of modular multiplication circuits, and particularly to a high-radix Montgomery modular multiplication circuit and its parallel operation method. Background Art

[0002] The Montgomery modular multiplication algorithm is one of the widely used methods for implementing modular multiplication algorithms. As a basic unit of asymmetric encryption and decryption algorithms such as RAS and ECC, its operation speed determines the overall operation efficiency of the algorithm. With the development and application of information technology, in order to improve information security, the security issues of encryption and decryption algorithms have received increasing attention. Currently, most encryption and decryption algorithms choose to be implemented in hardware to avoid the theft of encrypted and decrypted data caused by software vulnerabilities, resulting in serious adverse effects. Currently, most modular multiplication circuits are designed based on the Montgomery algorithm and its variant algorithms. When implemented in hardware, the circuit performance and circuit area cannot achieve a good balance, and there are problems such as a large circuit area when the circuit performance is improved, or poor circuit performance when the circuit area is small. Summary of the Invention

[0003] This application provides a high-radix Montgomery modular multiplication circuit and its parallel operation method. The specific technical solutions are as follows: A high-radix Montgomery modular multiplication circuit specifically includes: a loop control module, including m groups of loop control units, which are used to perform loop operations in parallel to complete k + 1 rounds of loop operations, and are used to cache intermediate data during the loop operations, read the data required for the loop operations from the storage module, and transfer the corresponding data to the operation module according to the loop operation steps; a core state machine module, which is used to control the start and end of each round of primary loop operation in the k + 1 rounds of loop operations performed by the m groups of loop control units in the loop control module; an operation module, which is used to receive the data transmitted by the loop control module and perform multiplication operations and addition operations based on the received data; a storage module, which is used to store the data required for the loop operations for the loop control device to read and call; where the k + 1 rounds of loop operations include a total of k + 1 rounds of primary loop operations, and each round of primary loop operation includes a pre-start loop operation and k rounds of secondary loop operations; the loop control device can perform at most m rounds of loop operations in parallel at the same time; both k and m are positive integers, and k is a positive integer multiple of m.

[0004] Further, the storage module includes three single-port random static memories, one dual-port random static memory, and two registers. Among them, the three single-port random static memories include a first single-port random static memory for storing k + 1 groups of multiplicand A, a second single-port random static memory for storing k + 1 groups of multiplier B, and a third single-port random static memory for storing k + 1 groups of modulus N. The dual-port random static memory is used to store k + 1 groups of initial results S and also used to store the results of each round of secondary loop operations. The k + 1 groups of initial results S are used for the first round of loop operations. The two registers include a constant q register for storing one group of constant q and a result S register for storing the results of each round of primary loop operations.

[0005] Further, the operation module includes a first multiplier, a second multiplier, and an adder. The first multiplier includes a first input terminal, a second input terminal, and a first output terminal. The second multiplier includes a third input terminal, a fourth input terminal, and a second output terminal. The adder includes a fifth input terminal, a sixth input terminal, a seventh input terminal, an eighth input terminal, and a third output terminal. Among them, the first input terminal, the second input terminal, the third input terminal, the fourth input terminal, the fifth input terminal, and the sixth input terminal serve as the input ports of the operation module for receiving data transmitted by the loop control device. The first output terminal of the first multiplier is connected to the seventh input terminal of the adder for transmitting the data after the multiplication operation of the first multiplier to the adder for addition operation. The second output terminal of the second multiplier is connected to the eighth input terminal of the adder for transmitting the data after the multiplication operation of the second multiplier to the adder for addition operation. The third output terminal of the adder serves as the output port of the operation module for outputting the data after the addition operation.

[0006] Further, the operation module further includes: a first register, a second register, and a third register. Among them, the first register is arranged between the first multiplier and the adder, the first output terminal is connected to the input terminal of the first register, and the output terminal of the first register is connected to the seventh input terminal of the adder. The second register is arranged between the second multiplier and the adder, the second output terminal is connected to the input terminal of the second register, and the output terminal of the second register is connected to the eighth input terminal of the adder. The input terminal of the third register is connected to the third output terminal of the adder, and the output terminal of the third register serves as the output port of the operation module.

[0007] Further, the first register further includes a fourth output terminal, and the fourth output terminal serves as the output port of the operation module for outputting the data after the multiplication operation of the first multiplier.

[0008] Further, each group of loop control units includes a loop controller, which is configured to read data stored in the storage module in response to the start control signal of the first-level loop operation of the core state machine module, perform the first-level loop operation by time-division multiplexing the first multiplier, the second multiplier, and the adder in the operation module based on the read data stored in the storage module, and is further configured to control the start and end of each round of the second-level loop operation.

[0009] Further, each group of loop control units further includes: an intermediate data cache module, which is configured to cache the data read from the storage module and the intermediate operation data fed back by the operation module during the loop operation, and provide the data for the loop control unit to read and call; wherein, the intermediate data cache module includes: a first calculation result t register, which is configured to cache the first calculation result t fed back by the operation module; a second calculation result u register, which is configured to cache the second calculation result u fed back by the operation module; a third calculation result d register, which is configured to cache the third calculation result d fed back by the operation module; a multiplier A register, which is configured to cache the multiplier A read by the loop control module from the first single-port random static memory of the storage module; a multiplier B register, which is configured to cache the multiplier B read by the loop control module from the second single-port random static memory of the storage module; a modulus N register, which is configured to cache the modulus N read by the loop control module from the third single-port random static memory of the storage module; and an intermediate result S register, which is configured to cache the result S read by the loop control module from the dual-port random static memory of the storage module.

[0010] Furthermore, the high-base Montgomery modular multiplication circuit further includes: the first input end of the first multiplier is respectively connected to the multiplier A register in each group of loop control units to receive the multiplier A cached in the multiplier A register transmitted by the loop control module; the first input end of the first multiplier is also respectively connected to the first calculation result t register in each group of loop control units to receive the first calculation result t cached in the first calculation result t register transmitted by the loop control module; the second input end of the first multiplier is respectively connected to the multiplier B register in each group of loop control units to receive the multiplier B cached in the multiplier B register transmitted by the loop control module; the third input end of the second multiplier is respectively connected to the second calculation result u register in each group of loop control units to receive the second calculation result u cached in the second calculation result u register transmitted by the loop control module; the fourth input end of the second multiplier is respectively connected to the modulus N register in each group of loop control units to receive the modulus N cached in the modulus N register transmitted by the loop control module; the fifth input end of the adder is respectively connected to the intermediate result S register in each group of loop control units to receive the result S cached in the intermediate result S register transmitted by the loop control module; the sixth input end of the adder is respectively connected to the third calculation result d register in each group of loop control units to receive the third calculation result d cached in the third calculation result d register transmitted by the loop control module.

[0011] The present application also provides a parallel operation method for a high-base Montgomery modular multiplication circuit, which specifically includes: when the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the (r + 1)-th group of loop control units to start executing the next round of first-level loop operations when they are in the idle state; wherein, when r is equal to m, when the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the first group of loop control units to start executing the next round of first-level loop operations; when a loop control unit starts to execute a round of first-level loop operations, it transitions from the idle state to the working state, and correspondingly, when a loop control unit ends a round of first-level loop operations, it transitions from the working state to the idle state.

[0012] Furthermore, for a group of loop control units to execute a round of first-level loop operations, it specifically includes: Step 1: Execute a pre-start loop operation to obtain the pre-start loop operation result; Step 2: Based on the pre-start loop operation result, execute k rounds of secondary loop operations.

[0013] Further, the method for performing the pre-start loop operation in step 1 to obtain the pre-start loop operation result specifically includes: Step 11: The loop control unit reads the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, and the result S(1) stored in the dual-port random static memory and transmits them to the operation module to obtain the first calculation result t of the pre-start loop operation; Step 12: The loop control unit reads the constant q stored in the constant q register and combines it with the first calculation result t of the pre-start loop operation and transmits them to the operation module to obtain the second calculation result u of the pre-start loop operation; Step 13: The loop control unit reads the modulus N(1) stored in the third single-port random static memory, combines it with the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, the result S(1) stored in the dual-port random static memory, and the second calculation result u of the pre-start loop operation and transmits them to the operation module to obtain the third calculation result d of the pre-start loop operation.

[0014] Further, step 11 specifically includes: The loop control unit initiates a read operation for the multiplier A(i) to the first single-port random static memory of the storage module based on the number of first-level loop rounds it executes, and initiates a read operation for the multiplier B(1) to the second single-port random static memory of the storage module; The loop control unit caches the read multiplier A(i) into the multiplier A register, caches the read multiplier B(1) into the multiplier B register. At the same time, the loop control unit initiates a read operation for the result S(1) to the dual-port random static memory; The loop control unit caches the read result S(1) into the intermediate result S register. At the same time, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs the multiplication operation of the multiplier A(i) and the multiplier B(1); The first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder. The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, so that the adder executes the addition operation of the product of the multiplier A(i) and the multiplier B(1) and the result S(1) and takes the addition operation result as the first calculation result t and transmits it to the corresponding group of loop control units. The loop control unit caches the first calculation result t into the first calculation result t register; where i is a positive integer less than or equal to k + 1.

[0015] Further, step 12 specifically includes: The loop control unit transmits the first calculation result t cached in the first calculation result t register to the first multiplier through the first input terminal. The loop control unit reads the constant q from the constant q register and transmits it to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation on the first calculation result t and the constant q. The first multiplier transmits the multiplication operation result of the first calculation result t and the constant q as the second calculation result u to a corresponding set of loop control units.

[0016] Further, step 13 specifically includes: The loop control unit caches the second calculation result u in the second calculation result u register. At the same time, the loop control unit initiates a modulo N(1) read operation on the third single-port random static memory of the storage module; the loop control unit caches the read modulo N(1) in the modulo N register; the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs a multiplication operation on the multiplier A(i) and the multiplier B(1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and the loop control unit transmits the modulo N(1) cached in the modulo N register to the second multiplier through the fourth input terminal, so that the second multiplier performs a multiplication operation on the second calculation result u and the modulo N(1); the loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulo N(1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs an addition operation on the result S(1), the product of the multiplier A(i) and the multiplier B(1), and the product of the modulo N(1) and the second calculation result u, and transmits the addition operation result as the third calculation result d to a corresponding set of loop control units; the loop control unit caches the third calculation result d in the third calculation result d register and obtains the third calculation result d as the pre-start loop operation result.

[0017] Further, step 2, the k-round secondary loop operation performed based on the pre-start loop operation result specifically includes: The loop control unit initiates a read operation for the multiplier B(j + 1) to the second single-port random static memory of the storage module and a read operation for the modulus N(j + 1) to the third single-port random static memory of the storage module based on the number of rounds j of the secondary loop operation; the loop control unit caches the read multiplier B(j + 1) into the multiplier B register and the read modulus N(j + 1) into the modulus N register, and the loop control unit initiates a read operation for the result S(j + 1) to the dual-port random static memory; the loop control unit caches the read result S(j + 1) into the intermediate result S register, and the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(j + 1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs the multiplication operation of the multiplier A(i) and the multiplier B(j + 1), and the loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and transmits the modulus N(j + 1) cached in the modulus N register to the second multiplier through the fourth input terminal, so that the second multiplier performs the multiplication operation of the second calculation result u and the modulus N(j + 1); the loop control unit transmits the result S(j + 1) cached in the intermediate result S register to the adder through the fifth input terminal, and the loop control unit transmits the third calculation result d cached in the third calculation result d register to the adder through the sixth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(j + 1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(j + 1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs the addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(j + 1), the product of the modulus N(j + 1) and the second calculation result u, and the third calculation result d of the pre-start loop operation, and caches the addition operation result as the result S(j) into the dual-port random static memory for reading in the next round of loop operation, completing one round of secondary loop operation; determine whether the number of rounds j of the secondary loop operation is equal to k. If so, end step 2. If not, increase the value of the number of rounds j of the secondary loop operation by 1 and enter the next round of secondary loop operation; where j is a positive integer less than or equal to k.

[0018] Further, the k rounds of secondary loop operations performed based on the pre-start loop operation result in step 2 further include: when the number of rounds j of the secondary loop operation is equal to k, the loop control unit caches the addition operation result as the result of the current round of primary loop operation into the result S register of the storage module.

[0019] The high-radix Montgomery modular multiplication circuit and its parallel operation method described in this application realize the parallel operation of multiple rounds of cyclic operations by multiple groups of cyclic control units in the cyclic control module calling the same storage module and based on the same operation module, achieving the technical effect of reducing the circuit area while maintaining the high-efficiency operation of the circuit cyclic operation. Description of the Drawings

[0020] Figure 1 It is a schematic diagram of the modules of the high-radix Montgomery modular multiplication circuit according to an embodiment of this application.

[0021] Figure 2 It is a schematic diagram of the modules of the operation module according to an embodiment of this application.

[0022] Figure 3 It is a schematic diagram of the modules of the operation module according to another embodiment of this application.

[0023] Figure 4 It is a schematic diagram of the modules of the cyclic control unit according to an embodiment of this application. Embodiment

[0024] Next, the embodiments of this application will be described in detail with reference to the drawings. It should be understood that the specific embodiments described below are only used to explain this application and are not used to limit this application.

[0025] Currently, most modular multiplication circuits are designed based on the Montgomery algorithm and its variant algorithms. When implemented in hardware, there is no good balance between circuit performance and circuit area. There are problems such as a large circuit area when the circuit performance is improved, or poor circuit performance when the circuit area is small. In order to improve the balance between the circuit area and circuit performance of the high-radix Montgomery modular multiplication circuit, this application provides a high-radix Montgomery modular multiplication circuit, aiming to realize the parallel operation of multiple rounds of cyclic operations by multiple groups of cyclic control units in the cyclic control module calling the same storage module and based on the same operation module, achieving the technical effect of reducing the circuit area while maintaining the high-efficiency operation of the circuit cyclic operation.

[0026] Specifically, as Figure 1 shown, the high-radix Montgomery modular multiplication circuit, as Figure 1 shown, specifically includes: The loop control module includes m groups of loop control units, which are used to run loop operations in parallel to complete k + 1 rounds of loop operations, and are used to cache intermediate data during the loop operation process, read the data required for the loop operation from the storage module, and transfer the corresponding data to the operation module according to the loop operation steps. Among them, at most m rounds of loop operations can be run in parallel by the m groups of loop control units in the loop control module. The caching of the intermediate data during the loop operation process means that the loop control unit transfers the corresponding data to the operation module according to the loop operation steps, and the data fed back by the operation module based on the operation result to the corresponding loop control unit. The intermediate data can be used for subsequent loop operation steps. The core state machine module is used to control the start and end of each round of the first-level loop operation in the k + 1 rounds of loop operations run by the m groups of loop control units in the loop control module. Specifically, in order to enable the k + 1 rounds of loop operations to be smoothly run in parallel in the m groups of loop control units, the core state machine module controls the node at which each group of loop control units starts to execute a round of loop operation. Among them, the node at which the core state machine module controls the start of a new round of loop operation can be, but is not limited to, that after the high-base Montgomery modular multiplication circuit is powered on and completed, it controls the first group of loop control units to start executing the first round of loop operation, or when the specified intermediate data operation is completed during the execution of a round of loop operation by the current group of loop control units, it controls the next group of loop control units to start executing the next round of loop operation, etc. The completion of the specified intermediate data operation can be, but is not limited to, for example, the completion of the first round of the second-level loop operation during a round of loop operation.

[0027] The operation module is used to receive the data transmitted by the loop control module and perform multiplication and addition operations based on the received data. Among them, the operation module includes at least a multiplier and an adder to implement multiplication and addition operations, and the number of multipliers and adders is configured according to the actual operation requirements. The loop control module multiplexes the operation module in a time-sharing manner to achieve the technical effect of enabling a simple operation module to implement a complex operation process based on time-sharing multiplexing means, and optimizing the algorithm operation efficiency of the high-base Montgomery modular multiplication circuit.

[0028] The storage module is used to store the data required for the loop operation for the loop control device to read and call. The storage module is called by the m groups of loop control units in the loop control device to run the loop operation in parallel, realizing data multiplexing of the storage space, improving the utilization rate of the storage space, ensuring the circuit performance while reducing the required circuit area occupied.

[0029] The high-base Montgomery modular multiplication circuit described in this application is used to implement the high-base Montgomery modular multiplication algorithm. The high-base Montgomery modular multiplication algorithm is an efficient modular multiplication algorithm that can perform modular multiplication operations without using division and is widely used in the field of cryptography. Among them, k is equal to the radix of the high-base Montgomery modular multiplication algorithm, realizing 2 kModular multiplication of bits. The k + 1 - round cyclic operation includes k + 1 rounds of primary cyclic operations. Each round of primary cyclic operation includes a pre - start cyclic operation and k rounds of secondary cyclic operations. Both k and m are positive integers, and k is a positive - integer multiple of m.

[0030] As a preferred embodiment of this application, as Figure 1 shown, the storage module includes 3 single - port random static memories, 1 dual - port random static memory, and 2 registers; among them, the 3 single - port random static memories include: The first single - port random static memory is used to store k + 1 groups of multiplicand A, each group including 2 k bit multiplicand A. Among them, the lower k groups of multiplicand A are valid data, and the (k + 1) - th group of multiplicand A is 0; the first single - port random static memory is used for each group of loop control units to read and call the corresponding group of multiplicand A when performing the above - mentioned pre - start cyclic operation and secondary cyclic operation; the number of groups of multiplicand A called by the loop control unit from the first single - port random static memory is determined based on the number of rounds of primary cyclic operation and the number of rounds of secondary cyclic operation it executes; preferably, the loop control unit calls the same number of groups of multiplicand A from the first single - port random static memory based on the number of rounds of its primary cyclic operation. For example, when the loop control unit executes the first round of primary cyclic operation, it calls the first group of multiplicand A from the first single - port random static memory for the first round of primary cyclic operation.

[0031] The second single - port random static memory is used to store k + 1 groups of multiplier B, each group including 2 k bit multiplier B. The lower k groups of multiplier B are valid data, and the (k + 1) - th group of multiplier B is 0; the second single - port random static memory is used for each group of loop control units to read and call the corresponding group of multiplier B when performing the pre - start cyclic operation and secondary cyclic operation; when the number of groups of multiplier B called by the loop control unit from the second single - port random static memory is used to perform the pre - start cyclic operation, it calls the first group of multiplier B; the number of groups of multiplier B called by the loop control unit from the second single - port random static memory when performing the secondary cyclic operation is determined based on the number of rounds of its secondary cyclic operation; preferably, the loop control unit calls the number of groups of multiplier B that is 1 more than the number of rounds of its secondary cyclic operation from the second single - port random static memory based on the number of rounds of its secondary cyclic operation. For example: when the loop control unit executes the first round of secondary cyclic operation, it calls the second group of multiplier B from the second single - port random static memory for the first round of secondary cyclic operation.

[0032] The third single - port random static memory is used to store k + 1 groups of modulus N, each group including 2 kThe bit modulus N, the modulus N of the low-k group is valid data, and the modulus N of the (k + 1)-th group is 0; the third single-port random static memory is used for each group of loop control units to read and call the corresponding group of modulus N when performing the above pre-start loop operation and the secondary loop operation; when the loop control unit calls the modulus N from the third single-port random static memory for performing the pre-start loop operation, it always calls the modulus N of the first group; when the loop control unit calls the modulus N from the third single-port random static memory for performing the secondary loop operation, the number of groups of modulus N called is determined based on the number of execution rounds of the secondary loop operation of the loop control unit; preferably, the loop control unit calls the modulus N of the number of groups increased by 1 compared to the number of execution rounds of the secondary loop operation from the third single-port random static memory based on the number of execution rounds of its secondary loop operation, for example: when the loop control unit executes the first round of the secondary loop operation, it calls the modulus N of the second group from the third single-port random static memory for the first round of the secondary loop operation.

[0033] A dual-port random static memory is used to store k + 1 groups of initial results S and is also used to store the results of each round of secondary loop operation; among them, the k + 1 groups of initial results S are used for the loop control unit to use when performing the first round of the first-level loop operation; when the loop control unit performs the second round to the (k + 1)-th round of the first-level loop operation, it is implemented based on the results of each round of secondary loop operation stored in the dual-port random static memory in the previous round of loop operation.

[0034] The two registers include: a constant q register and a result S register; among them, the constant q register is used to store a group of k-bit constant q, and the calculation formula of the constant q is: q = -N -1 mod2 k . The result S register is used to store the result S(i) of each round of the first-level loop operation. In this embodiment, the memory used to store the result S of each round of the secondary loop operation adopts a dual-port random static memory, ensuring that the memory storing the result S has the parallel execution ability of reading and storing, so that when multiple groups of loop control units are parallel in the loop control module, the call operation of the initial result S and the write operation of the result of the secondary loop operation can be performed on the dual-port random static memory simultaneously.

[0035] As a relatively preferred embodiment of the present application, such as Figure 2As shown, the operation module includes a first multiplier, a second multiplier, and an adder. Specifically, the first multiplier includes a first input terminal, a second input terminal, and a first output terminal; the second multiplier includes a third input terminal, a fourth input terminal, and a second output terminal; the adder includes a fifth input terminal, a sixth input terminal, a seventh input terminal, an eighth input terminal, and a third output terminal. Among them, the first input terminal, the second input terminal, the third input terminal, the fourth input terminal, the fifth input terminal, and the sixth input terminal serve as the input ports of the operation module and are connected to the loop control device for receiving the data transmitted by the loop control device.

[0036] Specifically, the first output terminal of the first multiplier is connected to the seventh input terminal of the adder to realize the data transmission from the first multiplier to the adder, which is usually used to transmit the data after the multiplication operation of the first multiplier to the adder for addition operation; the second output terminal of the second multiplier is connected to the eighth input terminal of the adder to realize the data transmission from the second multiplier to the adder, which is usually used to transmit the data after the multiplication operation of the second multiplier to the adder for addition operation; the third output terminal of the adder serves as the output port of the operation module for outputting the data after the addition operation. In this embodiment, the operation module is defined to include two multipliers and an adder, so that the two multipliers can perform multiplication operations in parallel.

[0037] Preferably, in some embodiments of the present application, in order to improve the parallel operation ability, multiple operation modules each including two multipliers and an adder can be configured in the high-radix Montgomery modular multiplication circuit. m groups of loop operation units are respectively configured to correspond to one operation module according to a preset number of groups. For example, one operation module including two multipliers and an adder is configured for every 4 groups of loop operation units to improve the parallel operation efficiency of multiple groups of loop operations in the loop control device.

[0038] As a preferred embodiment of the present application, as Figure 3 shown, the operation module further includes: a first register, a second register, and a third register; among them, the first register is arranged between the first multiplier and the adder, the first output terminal is connected to the input terminal of the first register, and the output terminal of the first register is connected to the seventh input terminal of the adder; the second register is arranged between the second multiplier and the adder, the second output terminal is connected to the input terminal of the second register, and the output terminal of the second register is connected to the eighth input terminal of the adder; the input terminal of the third register is connected to the third output terminal of the adder, and the output terminal of the third register serves as the output port of the operation module. In this embodiment, by arranging the first register between the first multiplier and the adder, the second register between the second multiplier and the adder, and the third register between the adder output and the loop operation device, the first multiplier / second multiplier and the adder are interrupted by the register, effectively improving the operation execution efficiency of the operation module.

[0039] As a preferred embodiment of the present application, as Figure 3 shown, the first register further includes a fourth output terminal, and the fourth output terminal is used as an output port of the operation module to output the data obtained by the multiplication operation of the first multiplier. Based on the operation steps in the pre-start loop operation of the high-base Montgomery modular multiplication algorithm that do not require the execution of addition operations, a fourth output terminal is provided on the first register, so that the fourth output terminal directly outputs the data obtained by the multiplication operation of the first multiplier outside the operation module, and is directly fed back to the loop operation device without passing through the adder, improving the feedback efficiency of the intermediate operation results and reducing the unnecessary occupation of operation resources.

[0040] As a preferred embodiment of the present application, as Figure 4 shown, each group of loop control units includes a loop controller, which is used to respond to the start control signal of the first-level loop operation of the core state machine module, read the data stored in the storage module, and perform the first-level loop operation by time-division multiplexing the first multiplier, the second multiplier, and the adder in the operation module based on the data stored in the storage module read, and is also used to control the start and end of each round of the second-level loop operation. Specifically, when the loop controller responds to the start control signal of the first-level loop operation of the core state machine module and controls this group of loop control units to execute a corresponding round of the first-level loop operation, the loop controller reads and calls the corresponding data from the storage module based on the execution progress of the first-level loop operation. For example, when performing the pre-start loop operation, the first set of results S(1) in the dual-port random static memory is read from the storage module, and the corresponding group of multipliers A(i) in the first single-port random static memory is read according to the execution round number i of the first-level loop operation, and the first set of multipliers B(1) in the second single-port random static memory is read. The way for the loop controller to control the start and end of each round of the second-level loop operation is based on the execution progress of the loop control unit performing the pre-start loop operation. When the execution of the pre-start loop operation is completed, the loop controller controls the start of the first round of the second-level loop operation. When the result of a round of the second-level loop operation is obtained, the loop controller controls the end of this round of the second-level loop operation and controls the start of the next round of the second-level loop operation. When the result of the k-th round of the second-level loop operation is obtained, the loop controller controls the end of the k-th round of the second-level loop operation. In the loop control unit provided in this embodiment, the loop controller is used to control the data transmission and reception between the loop control unit, the storage module, the core state machine, and the operation module.

[0041] As a preferred embodiment of the present application, as Figure 4As shown, each set of loop control units further includes: an intermediate data cache module, configured to cache the data read from the storage module and the intermediate operation data fed back by the operation module during the loop operation, and to be read and called by the loop controller; wherein, the intermediate data cache module includes: The first calculation result t register, configured to cache the first calculation result t fed back by the operation module; wherein, the first calculation result t refers to the first calculation result t obtained by the loop operation unit in the first calculation step of performing one round of pre-start loop operation.

[0042] The second calculation result u register, configured to cache the second calculation result u fed back by the operation module; wherein, the second calculation result u refers to the second calculation result u obtained by the loop operation unit in the second calculation step of performing one round of pre-start loop operation; The third calculation result d register, configured to cache the third calculation result d fed back by the operation module; wherein, the third calculation result d refers to the third calculation result d obtained by the loop operation unit in the third calculation step of performing one round of pre-start loop operation; the third calculation result d is used in the k rounds of secondary loop operations of the same round of primary loop operation.

[0043] The multiplier A register, configured to cache the multiplier A read by the loop control unit from the first single-port random static memory of the storage module. The multiplier B register, configured to cache the multiplier B read by the loop control unit from the second single-port random static memory of the storage module. The modulus N register, configured to cache the modulus N read by the loop control unit from the third single-port random static memory of the storage module. The intermediate result S register, configured to cache the result S read by the loop control unit from the dual-port random static memory of the storage module. In this embodiment, by setting the intermediate data cache module in each set of loop control units, on the one hand, it realizes the caching of the first calculation result t, the second calculation result u, and the third calculation result d fed back by the operation module, so as to facilitate the call in subsequent loop operation steps, and on the other hand, it realizes the call and caching of the multiplier A, multiplier B, modulus N, and result S in the storage module, enabling the data stored in the storage module to be read and called by multiple sets of parallel-running loop control units. Based on the caching of the data required during the loop operation by the intermediate data cache module, the parallel operation efficiency of the m sets of loop control units is improved.

[0044] Preferably, the first calculation result t register, the second calculation result u register, the third calculation result d register, the multiplier A register, the multiplier B register, the modulus N register, and the intermediate result S register included in the intermediate data cache module are all registers capable of caching k bit of data; when there is a new k bit of data to be read and written, the new data overwrites the old data for caching.

[0045] As a preferred embodiment of the present application, the high-base Montgomery modular multiplication circuit further includes: The first input terminal of the first multiplier is respectively connected to the multiplier A register in each group of loop control units to receive the multiplier A cached in the multiplier A register transmitted by the loop control unit; the first input terminal of the first multiplier is also respectively connected to the first calculation result t register in each group of loop control units to receive the first calculation result t cached in the first calculation result t register transmitted by the loop control unit; the second input terminal of the first multiplier is respectively connected to the multiplier B register in each group of loop control units to receive the multiplier B cached in the multiplier B register transmitted by the loop control unit.

[0046] The third input terminal of the second multiplier is respectively connected to the second calculation result u register in each group of loop control units to receive the second calculation result u cached in the second calculation result u register transmitted by the loop control unit; the fourth input terminal of the second multiplier is respectively connected to the modulus N register in each group of loop control units to receive the modulus N cached in the modulus N register transmitted by the loop control unit.

[0047] The fifth input terminal of the adder is respectively connected to the intermediate result S register in each group of loop control units to receive the result S cached in the intermediate result S register transmitted by the loop control unit; the sixth input terminal of the adder is respectively connected to the third calculation result d register in each group of loop control units to receive the third calculation result d cached in the third calculation result d register transmitted by the loop control unit. In this embodiment, by defining the registers in the loop control unit corresponding to the input terminals of the first multiplier, the second multiplier and the adder, the first multiplier is only used to perform the multiplication operation of the multiplier A and the multiplier B and the multiplication operation of the first calculation result t and the constant q, and the second multiplier is only used to perform the multiplication operation of the second calculation result u and the modulus N, so as to distinguish the multiplication operation steps of the first multiplier and the second multiplier in the loop operation, and avoid the situation that the execution of the first multiplier and the second multiplier is confused due to the parallel operation of multiple groups of loop control units.

[0048] As a preferred embodiment of the present application, a parallel operation method for high-base Montgomery modular multiplication is provided, which specifically includes: when the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the (r + 1)-th group of loop control units to start executing the next round of first-level loop operations when they are in the idle state; wherein, when r is equal to m, when the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the first group of loop control units to start executing the next round of first-level loop operations. Specifically, when a loop control unit starts to execute a round of first-level loop operations, it transitions from the idle state to the working state. Correspondingly, when a loop control unit ends a round of first-level loop operations, it transitions from the working state to the idle state. Since the operation steps for the third calculation result d in the pre-start loop operation in a round of loop operations require the first-round secondary loop result in the previous round of loop operations, and the secondary loop operations require the second-round to k-th round secondary loop results in the previous round of loop operations, therefore, in this embodiment, the completion of the first-round secondary loop operations in the first-level loop operation is used as the node to trigger the start of the next round of first-level loop operations, ensuring that when the next round of loop operations starts, at least the data required for its operation steps for the third calculation result d in the pre-start loop operation can be provided. While ensuring that a round of loop operations can be executed in a process, the parallel operation efficiency is improved.

[0049] As a preferred embodiment of the present application, the logic for the core state machine module to control the start of the loop control unit to execute the first-level loop operation is based on whether the first-round secondary loop operation in the first-level loop operation is completed; specifically, the core state machine module controls the first group of loop control units to start executing the first round of first-level loop operations. When the first group of loop control units completes the first-round secondary loop operation in the first round of first-level loop operations, the core state machine module controls the second group of loop control units to start executing the second round of first-level loop operations. When the second group of loop control units completes the first-round secondary loop operation in the second round of first-level loop operations, the core state machine module controls the third group of loop control units to start executing the second round of first-level loop operations, and so on. When the (m - 1)-th group of loop control units completes the first-round secondary loop operation in the (m - 1)-th round of first-level loop operations, the core state machine controls the m-th group of loop control units to start executing the m-th round of first-level loop operations.

[0050] As a preferred embodiment of the present application, a group of loop control units perform a round of first-level loop operations, specifically including: Step 1: Perform a pre-start loop operation to obtain the result of the pre-start loop operation; Step 2: Based on the result of the pre-start loop operation, perform k rounds of second-level loop operations. Wherein, the result of the pre-start loop operation refers to the obtained third calculation result d. In this embodiment, a round of first-level loop operations is divided into a pre-start loop operation and k rounds of second-level loop operations. The execution of the k rounds of second-level loop operations is based on the result of the pre-start loop operation, realizing that k rounds of second-level loop operations are nested in each round of first-level loop operations.

[0051] As a preferred embodiment of the present application, the method of performing the pre-start loop operation in Step 1 to obtain the result of the pre-start loop operation specifically includes: Step 11: The loop control unit reads the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, and the result S(1) stored in the dual-port random static memory and transmits them to the operation module to obtain the first calculation result t of the pre-start loop operation; Step 12: The loop control unit reads the constant q stored in the constant q register and combines it with the first calculation result t of the pre-start loop operation and transmits it to the operation module to obtain the second calculation result u of the pre-start loop operation; Step 13: The loop control unit reads the modulus N(1) stored in the third single-port random static memory, combines it with the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, the result S(1) stored in the dual-port random static memory, and the second calculation result u of the pre-start loop operation and transmits it to the operation module to obtain the third calculation result d of the pre-start loop operation.

[0052] Specifically, in the multiplier A(i), i refers to the number of rounds of the first-level loop operation currently executed by the loop control unit; the multiplier A(i) refers to the i-th group of multipliers A pre-stored in the first single-port random static memory; the multiplier B(1) refers to the first group of multipliers B pre-stored in the second single-port random static memory; the result S(1) refers to the result of the first round of the second-level loop operation in the previous round of loop operation. In particular, if the currently executed is the first round of loop operation, the result S(1) is the first group of initial results S pre-stored in the dual-port random static memory.

[0053] As a preferred embodiment of the present application, Step 11 specifically includes: The loop control unit initiates a read operation of multiplier A(i) to the first single-port random static memory of the storage module and a read operation of multiplier B(1) to the second single-port random static memory of the storage module based on the number of first-level loop rounds i it executes; wherein, the multiplier A(i) refers to the i-th group of multiplier A pre-stored in the first single-port random static memory; the multiplier B(1) refers to the first group of multiplier B pre-stored in the second single-port random static memory.

[0054] The loop control unit caches the read multiplier A(i) into the multiplier A register and the read multiplier B(1) into the multiplier B register. Meanwhile, the loop control unit initiates a read operation of result S(1) to the dual-port random static memory; wherein, the result S(1) refers to the result of the first-round secondary loop operation in the previous loop operation. In particular, if the currently executed is the first loop operation, the result S(1) is the first group of initial result S pre-stored in the dual-port random static memory.

[0055] The loop control unit caches the read result S(1) into the intermediate result S register. Meanwhile, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal and the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, enabling the first multiplier to perform the multiplication operation of multiplier A(i) and multiplier B(1). The first multiplier transmits the product of multiplier A(i) and multiplier B(1) to the adder through the seventh input terminal. The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, enabling the adder to execute the addition operation of the product of multiplier A(i) and multiplier B(1) and result S(1) and taking the addition operation result as the first calculation result t and transmitting it to the corresponding group of loop control units. The loop control unit caches the first calculation result t into the first calculation result t register; wherein, i is a positive integer less than or equal to k + 1. In this embodiment, by splitting the calculation step of the first calculation result of the pre-start loop in the high-radix Montgomery modular multiplication algorithm into two steps, the first step is to calculate the product of multiplier A(i) and multiplier B(1), and the second step is to calculate the sum of the product of multiplier A(i) and multiplier B(1) and result S(1), thereby achieving the acquisition of the first calculation result t.

[0056] As a preferred embodiment of the present application, step 12 specifically includes: The loop control unit transmits the first calculation result t cached in the first calculation result t register to the first multiplier through the first input terminal. The loop control unit reads the constant q from the constant q register and transmits it to the first multiplier through the second input terminal, so that the first multiplier performs the multiplication operation of the first calculation result t and the constant q. The first multiplier transmits the multiplication operation result of the first calculation result t and the constant q as the second calculation result u to the corresponding set of loop control units.

[0057] As a preferred embodiment of the present application, step 13 specifically includes: The loop control unit caches the second calculation result u into the second calculation result u register. At the same time, the loop control unit initiates a modulo N(1) read operation to the third single-port random static memory of the storage module; the loop control unit caches the read modulo N(1) into the modulo N register; wherein, the modulo N(1) refers to the first set of modulo N pre-stored in the third single-port random static memory.

[0058] The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs the multiplication operation of the multiplier A(i) and the multiplier B(1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal. The loop control unit transmits the modulo N(1) cached in the modulo N register to the second multiplier through the fourth input terminal, so that the second multiplier performs the multiplication operation of the second calculation result u and the modulo N(1).

[0059] The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal. The second multiplier transmits the product of the modulo N(1) and the second calculation result u to the adder through the eighth input terminal, so that the adder performs the addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(1), and the product of the modulo N(1) and the second calculation result u, and transmits the addition operation result as the third calculation result d to the corresponding set of loop control units; the loop control unit caches the third calculation result d in the third calculation result d register and obtains the pre-start loop operation result as the third calculation result d.

[0060] As a preferred embodiment of the present application, step 2 of performing k rounds of secondary loop operations based on the pre-start loop operation result specifically includes: The loop control unit initiates a read operation for the multiplier B(j + 1) to the second single-port random static memory of the storage module and a read operation for the modulus N(j + 1) to the third single-port random static memory of the storage module based on the number of rounds j of the secondary loop operation; where j is a positive integer less than or equal to k. The j + 1 is determined based on the number of rounds j of the secondary loop operation of the loop control unit. The multiplier B(j + 1) refers to the (j + 1)-th group of multipliers B pre-stored in the second single-port random static memory. The modulus N(j + 1) refers to the (j + 1)-th group of moduli N pre-stored in the third single-port random static memory.

[0061] The loop control unit caches the read multiplier B(j + 1) in the multiplier B register, caches the read modulus N(j + 1) in the modulus N register, and the loop control unit initiates a read operation for the result S(j + 1) to the dual-port random static memory; where the result S(j + 1) refers to the (j + 1)-th group of results S stored in the dual-port random static memory. Specifically, when the loop control unit executes the first round of the primary loop operation, the result S(j + 1) refers to the (j + 1)-th group of initial results S pre-stored in the dual-port random static memory. On the contrary, when the loop control unit does not execute the first round of the primary loop operation, the result S(j + 1) refers to the (j + 1)-th group of the k groups of results S stored in the dual-port random static memory based on the result of the secondary loop operation in the previous round of loop operation.

[0062] The loop control unit caches the read result S(j + 1) in the intermediate result S register. The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(j + 1) cached in the multiplier B register to the first multiplier through the second input terminal, so that the first multiplier performs the multiplication operation of the multiplier A(i) and the multiplier B(j + 1) to obtain the product of the multiplier A(i) and the multiplier B(j + 1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and transmits the modulus N(j + 1) cached in the modulus N register to the second multiplier through the fourth input terminal, so that the second multiplier performs the multiplication operation of the second calculation result u and the modulus N(j + 1) to obtain the product of the second calculation result u and the modulus N(j + 1).

[0063] The loop control unit transmits the result S(j + 1) cached in the intermediate result S register to the adder through the fifth input terminal, and transmits the third calculation result d cached in the third calculation result d register to the adder through the sixth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(j + 1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(j + 1) and the second calculation result u to the adder through the eighth input terminal, enabling the adder to perform the addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(j + 1), the product of the modulus N(j + 1) and the second calculation result u, and the third calculation result d of the pre-start loop operation, and caching the addition operation result as the result S(j) in the dual-port random static memory for reading in the next round of loop operation to complete one round of secondary loop operation.

[0064] Judge whether the number of execution rounds j of the secondary loop operation is equal to k. If so, end step 2. If not, increase the value of the number of execution rounds j of the secondary loop operation by 1 and enter the next round of secondary loop operation. In this embodiment, the loop controller controls the end of the secondary loop operation according to the number of execution rounds j of the secondary loop operation by judging whether the number of execution rounds j of the secondary loop operation is equal to k. When the number of execution rounds of the secondary loop operation does not reach k rounds, the loop control unit is controlled to keep executing the secondary loop operation.

[0065] As a preferred embodiment of the present application, the step 2 of performing k rounds of secondary loop operations based on the pre-start loop operation result further includes: when the number of execution rounds j of the secondary loop operation is equal to k, the loop control unit transmits the addition operation result to the result S register of the storage module as the result of the current round of primary loop operation for caching.

[0066] As a preferred embodiment of the present application, it is defined that there is an upper limit on the number of primary loop operations that each group of loop control units can complete; among them, the number of primary loop operations that the first group of loop control units can complete is equal to 1 + k / m rounds; the number of primary loop operations that the remaining groups of loop control units can complete is equal to k / m rounds; where k / m is a positive integer. For example: in some embodiments, the loop control module is configured to have 4 groups of loop control units, that is, m = 4, and the radix-32 Montgomery modular multiplication algorithm is implemented based on 4 groups of loop control units, that is, k = 32. In this loop control module, the first group of loop control units is used to implement 9 rounds of primary loop operations, and the second group to the fourth group of loop control units are used to implement 8 rounds of primary loop operations.

[0067] As a preferred embodiment of the present application, in order to ensure that the read, write, and call operations of the storage module and the arithmetic module by multiple groups of loop control units during parallel operation do not interfere with each other, the timing of the first multiplier, the second multiplier, and the adder in the arithmetic module multiplexed by m groups of loop control units in the loop control module is defined, the timing of reading the first single-port random static memory, the second single-port random static memory, the third single-port random static memory, the dual-port random static memory, and the constant q in the storage module is defined, and the timing of writing to the result S register in the storage module is defined. Thus, it is ensured that only one loop control module reads the first single-port random static memory / the second single-port random static memory / the third single-port random static memory in each clock cycle, and only one loop control module reads and writes the dual-port random static memory in each clock cycle. The specific definition logic can be but is not limited to: taking 4 time units as a group for the timing. In each group of timing, the first time unit allows reading the dual-port random static memory; the second time unit allows calling the multiplier; the third time unit allows calling the adder; the fourth time unit allows reading the first single-port random static memory, the second single-port random static memory, and the third single-port random static memory and writing to the dual-port random static memory. This embodiment controls the loop control unit to read, call, and transfer operations on each memory and register in a time-sharing manner, enabling the storage module to be read and written in a time-sharing manner and the arithmetic module to be multiplexed in a time-sharing manner, thereby improving the multiplexing efficiency of the storage module and the arithmetic module during the parallel operation of multiple groups of loop control units.

[0068] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. These programs can be stored in a computer-readable storage medium (such as various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs). When the program is executed, it performs the steps including the above method embodiments.

[0069] It should be noted that each of the aforementioned loop control modules, arithmetic modules, and storage modules can be but is not limited to a digital circuit module formed by a designer using the hardware description language Verilog HDL, or a digital circuit module formed by a designer through circuit drawing or compilation on software with circuit drawing or compilation functions. In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module.

[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A high-base Montgomery modular multiplication circuit, characterized in that, The high-radix Montgomery modular multiplication circuit includes: A loop control module, including m groups of loop control units, which are used to run loop operations in parallel to complete k + 1 rounds of loop operations, and are used to cache intermediate data during the loop operations, read the data required for the loop operations from the storage module, and transmit the corresponding data to the operation module according to the loop operation steps; A core state machine module, which is used to control the start and end of each round of first-level loop operations in the k + 1 rounds of loop operations run by the m groups of loop control units in the loop control module; An operation module, which is used to receive the data transmitted by the loop control module and perform multiplication and addition operations based on the received data; A storage module, which is used to store the data required for the loop operations for the loop control device to read and call; Among them, the k + 1 rounds of loop operations include a total of k + 1 rounds of first-level loop operations, and each round of first-level loop operations includes a pre-start loop operation and k rounds of second-level loop operations; the loop control device runs at most m rounds of loop operations in parallel; both k and m are positive integers, and k is a positive integer multiple of m.

2. The high-base Montgomery modular multiplication circuit according to claim 1, characterized in that, The storage module includes 3 single-port random static memories, 1 dual-port random static memory, and 2 registers; Among them, the 3 single-port random static memories include a first single-port random static memory for storing k + 1 groups of multiplicand A, a second single-port random static memory for storing k + 1 groups of multiplier B, and a third single-port random static memory for storing k + 1 groups of modulus N; The dual-port random static memory is used to store k + 1 groups of initial results S, and is also used to store the results of each round of second-level loop operations; the k + 1 groups of initial results S are used for the first round of loop operations; The 2 registers include a constant q register for storing 1 group of constant q and a result S register for storing the results of each round of first-level loop operations.

3. The high-base Montgomery modular multiplication circuit according to claim 1, characterized in that, The operation module includes a first multiplier, a second multiplier, and an adder; The first multiplier includes a first input terminal, a second input terminal, and a first output terminal; The second multiplier includes a third input terminal, a fourth input terminal, and a second output terminal; The adder includes a fifth input terminal, a sixth input terminal, a seventh input terminal, an eighth input terminal, and a third output terminal; Among them, the first input terminal, the second input terminal, the third input terminal, the fourth input terminal, the fifth input terminal, and the sixth input terminal are used as the input ports of the operation module to receive the data transmitted by the loop control device; The first output terminal of the first multiplier is connected to the seventh input terminal of the adder, and is used to transmit the data after the multiplication operation of the first multiplier to the adder for addition operation; The second output terminal of the second multiplier is connected to the eighth input terminal of the adder, and is used to transmit the data after the multiplication operation of the second multiplier to the adder for addition operation; The third output terminal of the adder is used as the output port of the operation module to output the data after the addition operation.

4. The high-base Montgomery modular multiplication circuit according to claim 3, characterized in that, The operation module further includes: a first register, a second register, and a third register; wherein, the first register is arranged between the first multiplier and the adder, the first output end is connected to the input end of the first register, and the output end of the first register is connected to the seventh input end of the adder; the second register is arranged between the second multiplier and the adder, the second output end is connected to the input end of the second register, and the output end of the second register is connected to the eighth input end of the adder; the input end of the third register is connected to the third output end of the adder, and the output end of the third register serves as the output port of the operation module.

5. The high-base Montgomery modular multiplication circuit according to claim 4, characterized in that, The first register further includes a fourth output end, and the fourth output end serves as the output port of the operation module for outputting the data obtained by the multiplication operation of the first multiplier.

6. The high-base Montgomery modular multiplication circuit according to claim 3, characterized in that, Each group of loop control units includes a loop controller, which is used to respond to the start control signal of the first-level loop operation of the core state machine module, read the data stored in the storage module, and perform the first-level loop operation by time-division multiplexing the first multiplier, the second multiplier, and the adder in the operation module based on the data read from the storage module, and is also used to control the start and end of each round of the second-level loop operation.

7. The high-base Montgomery modular multiplication circuit according to claim 6, characterized in that, Each group of loop control units further includes: an intermediate data cache module, which is used to cache the data read from the storage module and the intermediate operation data fed back by the operation module during the loop operation process, and is provided for the loop control unit to read and call; wherein, the intermediate data cache module includes: A first calculation result t register, which is used to cache the first calculation result t fed back by the operation module; A second calculation result u register, which is used to cache the second calculation result u fed back by the operation module; A third calculation result d register, which is used to cache the third calculation result d fed back by the operation module; A multiplier A register, which is used to cache the multiplier A read by the loop control module from the first single-port random static memory of the storage module; A multiplier B register, which is used to cache the multiplier B read by the loop control module from the second single-port random static memory of the storage module; A modulus N register, which is used to cache the modulus N read by the loop control module from the third single-port random static memory of the storage module; An intermediate result S register, which is used to cache the result S read by the loop control module from the dual-port random static memory of the storage module.

8. The high-base Montgomery modular multiplication circuit according to claim 7, characterized in that, The high-base Montgomery modular multiplication circuit further includes: The first input end of the first multiplier is respectively connected to the multiplier A register in each group of loop control units to receive the multiplier A cached in the multiplier A register transmitted by the loop control module; The first input end of the first multiplier is also respectively connected to the first calculation result t register in each group of loop control units to receive the first calculation result t cached in the first calculation result t register transmitted by the loop control module; The second input end of the first multiplier is respectively connected to the multiplier B register in each group of loop control units to receive the multiplier B cached in the multiplier B register transmitted by the loop control module; The third input end of the second multiplier is respectively connected to the second calculation result u register in each group of loop control units to receive the second calculation result u cached in the second calculation result u register transmitted by the loop control module; The fourth input terminal of the second multiplier is respectively connected to the modulo N register in each group of loop control units to receive the modulo N cached in the modulo N register transmitted by the loop control module; The fifth input terminal of the adder is respectively connected to the intermediate result S register in each group of loop control units to receive the result S cached in the intermediate result S register transmitted by the loop control module; The sixth input terminal of the adder is respectively connected to the third calculation result d register in each group of loop control units to receive the third calculation result d cached in the third calculation result d register transmitted by the loop control module.

9. A parallel operation method for a high-radix Montgomery modular multiplication circuit, characterized in that, The parallel operation method of the high-radix Montgomery modular multiplication circuit specifically includes: When the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the (r + 1)-th group of loop control units to start executing the next round of first-level loop operations when they are in the idle state; Among them, when r is equal to m, when the r-th group of loop control units completes the first round of secondary loop operations in the first-level loop operation, the core state machine controls the first group of loop control units to start executing the next round of first-level loop operations; when a loop control unit starts to execute a round of first-level loop operations, it changes from the idle state to the working state, and correspondingly, when a loop control unit ends a round of first-level loop operations, it changes from the working state to the idle state.

10. The parallel operation method for a high-radix Montgomery modular multiplication circuit according to claim 9, characterized in that, The execution of a round of first-level loop operations by a group of loop control units specifically includes: Step 1: Execute a pre-start loop operation to obtain the result of the pre-start loop operation; Step 2: Based on the result of the pre-start loop operation, execute k rounds of secondary loop operations to obtain the result of the first-level loop operation and end a round of first-level loop operations.

11. The parallel operation method for a high-radix Montgomery modular multiplication circuit according to claim 10, characterized in that, The method of executing the pre-start loop operation in Step 1 to obtain the result of the pre-start loop operation specifically includes: Step 11: The loop control unit reads the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, and the result S(1) stored in the dual-port random static memory and transmits them to the operation module to obtain the first calculation result t of the pre-start loop operation; Step 12: The loop control unit reads the constant q stored in the constant q register and combines the first calculation result t of the pre-start loop operation and transmits it to the operation module to obtain the second calculation result u of the pre-start loop operation; Step 13: The loop control unit reads the modulo N(1) stored in the third single-port random static memory, combines the multiplier A(i) stored in the first single-port random static memory, the multiplier B(1) stored in the second single-port random static memory, the result S(1) stored in the dual-port random static memory, and the second calculation result u of the pre-start loop operation and transmits them to the operation module to obtain the third calculation result d of the pre-start loop operation.

12. The parallel operation method for a high-radix Montgomery modular multiplication circuit according to claim 11, characterized in that, The specific content of Step 11 includes: The loop control unit initiates a read operation for the multiplier A(i) to the first single-port random static memory of the storage module and a read operation for the multiplier B(1) to the second single-port random static memory of the storage module based on the first-level loop round number it executes; The loop control unit caches the read multiplier A(i) into the multiplier A register and the read multiplier B(1) into the multiplier B register. Meanwhile, the loop control unit initiates a read operation for the result S(1) from the dual-port random static memory. The loop control unit caches the read result S(1) into the intermediate result S register. Meanwhile, the loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, enabling the first multiplier to perform the multiplication operation of multiplier A(i) and multiplier B(1). The first multiplier transmits the product of multiplier A(i) and multiplier B(1) to the adder. The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal, enabling the adder to perform the addition operation of the product of multiplier A(i) and multiplier B(1) and the result S(1), and taking the addition operation result as the first calculation result t and transmitting it to the corresponding group of loop control units. The loop control unit caches the first calculation result t for pre-starting the loop operation into the first calculation result t register. Here, i is a positive integer less than or equal to k + 1.

13. The parallel operation method for a high-radix Montgomery modular multiplication circuit according to claim 12, characterized in that, Step 12 specifically includes: The loop control unit transmits the first calculation result t cached in the first calculation result t register to the first multiplier through the first input terminal. The loop control unit reads the constant q from the constant q register and transmits it to the first multiplier through the second input terminal, enabling the first multiplier to perform the multiplication operation of the first calculation result t and the constant q. The first multiplier transmits the multiplication operation result of the first calculation result t and the constant q as the second calculation result u for pre-starting the loop operation to the corresponding group of loop control units.

14. The parallel operation method for a high-radix Montgomery modular multiplication circuit according to claim 13, characterized in that, Step 13 specifically includes: The loop control unit caches the second calculation result u into the second calculation result u register. Meanwhile, the loop control unit initiates a read operation for the modulus N(1) from the third single-port random static memory of the storage module. The loop control unit caches the read modulus N(1) into the modulus N register. The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(1) cached in the multiplier B register to the first multiplier through the second input terminal, enabling the first multiplier to perform the multiplication operation of multiplier A(i) and multiplier B(1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and transmits the modulus N(1) cached in the modulus N register to the second multiplier through the fourth input terminal, enabling the second multiplier to perform the multiplication operation of the second calculation result u and the modulus N(1). The loop control unit transmits the result S(1) cached in the intermediate result S register to the adder through the fifth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(1) and the second calculation result u to the adder through the eighth input terminal, enabling the adder to perform the addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(1), and the product of the modulus N(1) and the second calculation result u, and transmitting the addition operation result as the third calculation result d to the corresponding group of loop control units; The loop control unit caches the third calculation result d in the third calculation result d register and obtains the third calculation result d as the pre-start loop operation result.

15. The parallel operation method of the high-base Montgomery modular multiplication circuit according to claim 14, wherein, Performing k rounds of secondary loop operations based on the pre-start loop operation result described in step 2 specifically includes: The loop control unit initiates a read operation for the multiplier B(j + 1) to the second single-port random static memory of the storage module and a read operation for the modulus N(j + 1) to the third single-port random static memory of the storage module based on the number of rounds of the secondary loop operation; The loop control unit caches the read multiplier B(j + 1) in the multiplier B register, caches the read modulus N(j + 1) in the modulus N register, and the loop control unit initiates a read operation for the result S(j + 1) to the dual-port random static memory; The loop control unit caches the read result S(j + 1) in the intermediate result S register. The loop control unit transmits the multiplier A(i) cached in the multiplier A register to the first multiplier through the first input terminal, and transmits the multiplier B(j + 1) cached in the multiplier B register to the first multiplier through the second input terminal, enabling the first multiplier to perform the multiplication operation of the multiplier A(i) and the multiplier B(j + 1). The loop control unit transmits the second calculation result u cached in the second calculation result u register to the second multiplier through the third input terminal, and transmits the modulus N(j + 1) cached in the modulus N register to the second multiplier through the fourth input terminal, enabling the second multiplier to perform the multiplication operation of the second calculation result u and the modulus N(j + 1); The loop control unit transmits the result S(j + 1) cached in the intermediate result S register to the adder through the fifth input terminal, and transmits the third calculation result d cached in the third calculation result d register to the adder through the sixth input terminal. At the same time, the first multiplier transmits the product of the multiplier A(i) and the multiplier B(j + 1) to the adder through the seventh input terminal, and the second multiplier transmits the product of the modulus N(j + 1) and the second calculation result u to the adder through the eighth input terminal, enabling the adder to perform the addition operation of the result S(1), the product of the multiplier A(i) and the multiplier B(j + 1), the product of the modulus N(j + 1) and the second calculation result u, and the pre-start loop operation result, the third calculation result d, and caches the addition operation result as the result S(j) in the dual-port random static memory for reading in the next round of loop operation, completing one round of secondary loop operation; Determine whether the number of execution rounds of the secondary loop operation is equal to k. If so, end Step 2. If not, increment the value of the number of execution rounds of the secondary loop operation by 1 and enter the next round of the secondary loop operation; where j is a positive integer less than or equal to k.

16. The parallel operation method of the high-base Montgomery modular multiplication circuit according to claim 14, wherein, The step 2 of performing k rounds of secondary loop operations based on the pre-start loop operation result further includes: when the number of execution rounds of the secondary loop operation is equal to k, the loop control unit caches the addition operation result as the result of the current round of the primary loop operation in the result S register of the storage module.

Citation Information

Patent Citations

  • Optimized Montgomery modular multiplication method, optimized modular square method and optimized modular multiplication hardware

    CN103761068A

  • Montgomery modular-multiplication calculation method suitable for embedded system

    CN107665109A

  • System and method for quickly realizing Montgomery modular multiplication by using multiplier

    CN114840174A

  • Federal learning-based hardware acceleration data transmission method

    CN114880686A

  • Data processing method combining Karatsuba and Montgomery modular multiplication

    CN115344237A