Software-hardware coordinated segmented scanning Montgomery modular power calculation system and readable storage medium
By adopting a segmented scanning Montgomery modular power calculation system with soft and hard collaboration in the RSA algorithm, combining the 2k segmented scanning algorithm and the Montgomery algorithm, the problem of low modular power calculation in the existing technology is solved, and flexible and efficient modular power calculation is achieved.
Patent Information
- Application Number
- CN202111480141.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In the prior art, when performing modular exponentiation operations in RSA algorithms, software computing efficiency is low, while FPGA hardware acceleration is poor, making it difficult to take into account both.
A piecewise scanning Montgomery modular exponential computing system is proposed, combining 2k segmented scanning algorithm and Montgomery algorithm to perform data pre-calculation and result transmission through the ARM side, and the FPGA side completes calculation-intensive modular exponential computing.
It realizes flexible and efficient modular exponentiation operations of different modular N, making full use of the flexibility of the software and the acceleration effect of FPGA, and improving the efficiency of modular exponentiation calculation.
Smart Images

Figure CN114138235B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of encryption and decryption operations, and in particular to a software-hardware coordinated segmented scanning Montgomery modular power calculation system and a readable storage medium. Background Art
[0002] In recent years, with the development of computers and e-commerce, mankind has entered the information age. In today's information explosion, the demand for privacy security is increasing, and people have further demands for rapid encryption, secure transmission, and confidentiality preservation of information. The development of cryptography has a more urgent need in the field of electronic information. In cryptography, the most basic thing is that information encryption and decryption are divided into symmetric encryption and asymmetric encryption. The difference between the two is whether the same key is used for encryption and decryption.
[0003] The former uses the same secret key for encryption and decryption, which reduces the amount of calculation, increases the encryption speed, and improves the encryption efficiency. However, the management and distribution of the secret key are very difficult and not secure enough. The latter uses different secret keys for encryption and decryption, which is more secure. The public key is public, and the secret key is kept by oneself. The private key does not need to be given to others. However, encryption and decryption take a long time and are slow, so it is only suitable for encrypting a small amount of data.
[0004] The RSA encryption algorithm is an asymmetric encryption algorithm. RSA is widely used in public key encryption and electronic commerce. It is generally believed that the security of RSA comes from the decomposition of large numbers, and there is currently no suitable algorithm to process the decomposition of large numbers. The core operation of the RSA algorithm is the modular exponentiation operation in a finite field. The speed of software processing modular exponentiation operation has always been the biggest limitation of the RSA algorithm due to the huge amount of calculations. Therefore, hardware acceleration of modular exponentiation operation has broad application prospects.
[0005] In order to speed up the modular exponentiation process, many modular exponentiation optimization algorithms need to perform pre-calculations at the software level before modular exponentiation, and then pass the calculated values to the FPGA as input values. Therefore, pure software calculations and FPGA calculations cannot take both into account. Pure software calculations of large modular exponentiations are too inefficient, while pure FPGAs have poor flexibility and are subject to greater constraints on pre-calculation and result transmission. Summary of the invention
[0006] Purpose of the invention: To propose a segmented scanning Montgomery modular power calculation system with software and hardware collaboration. kThe segmented scanning algorithm and the Montgomery algorithm can independently complete the data pre-calculation of integer modular exponentiation operations and the transmission of the result data, transmit them to the specified FPGA address, and start the FPGA to perform specific modular exponentiation operations, making full use of the flexibility of the software and the acceleration effect of the FPGA, and flexibly and efficiently performing modular exponentiation operations with different moduli N, better meeting the needs of practical applications, thereby solving the above-mentioned problems existing in the prior art.
[0007] In the first aspect, a segmented scanning Montgomery modular power calculation system with software and hardware collaboration is proposed, which is specifically implemented by the following technical solutions:
[0008] The specific implementation of the present invention is based on SoC software and hardware collaborative implementation of 256-bit Montgomery modular exponentiation (M E modN), the ARM side is used to complete the overall task scheduling and generate 2 k The FPGA side completes the computationally intensive modular exponentiation operation to pre-calculate and correct the coefficients required by the binary segmented scanning algorithm.
[0009] ARM side generates 2 k The present invention adopts the 6-bit segmented scanning method, according to 2 k The specific process of the binary segment scanning algorithm requires (2 6 ) pre-calculated values and 1 correction factor.
[0010] In addition, a driver module needs to be developed on the ARM side to establish two-way communication with the data preprocessing module, drive the FPGA, and establish two-way data communication with the FPGA.
[0011] The FPGA end implements the modular inversion module, which uses the extended Euclidean algorithm to calculate the 64-bit modular inversion value and provides it to the Montgomery module as an additional preprocessing value. Find the modular inverse x of a with respect to m, that is, a*x≡1(mod m), and convert it into finding x and y so that ax+my=1 holds, where y is the additional value. The modular inversion uses a direct extended Euclidean algorithm and uses the properties of the Euclidean algorithm: ax+my=a1x1+m1y1=…=1, where a recursive operation is used to obtain the final result, and the recursive boundary is m. n = 0. The core recursive code is as follows:
[0012]
[0013] In the formula, mod represents the modulus operation, / represents integer division, (x n ,y n )(x n-1 y n-1 ) is the solution of two consecutive recursions, (an , m n )(a n-1 , m n-1 ) The source data is transformed between two adjacent recursive operations.
[0014] The data distribution module is a data exchange module between the ARM side and the FPGA side, and is used for ARM access data. Limited by the 32-bit operating system on the ARM side and the 128-bit Avalon interface, this module needs to process the bit width of the data. This module also completes the data readout function, reads out 256-bit result data from the SRAM, and splits the data.
[0015] The SRAM module stores source data, budget values, correction coefficients and result data and is designed to be 256 bits wide.
[0016] The Montgomery modular exponentiation module is used to implement the specific operation of 256-bit modular exponentiation. It mainly includes: 256-bit Montgomery modular multiplication (MontMult) module, state machine control module (fsm_ctrl), address generation module (addgent). The state machine control module completes the control flow of the entire modular exponentiation algorithm and calls the MontMult module in sequence;
[0017] The 256-bit Montgomery modular multiplication module is the core computing module, which is implemented based on the CIOS optimization algorithm. The original CIOS algorithm includes the following steps:
[0018] Step 1: Start the modular inversion module first and get N′0 represents the modular inverse value required for modular multiplication, which is used here to generate the correction coefficient. Represents the modular inverse operation of the lower 64 bits of N, R represents the modulus value of the 64-bit modular inverse, where R = 2 64 ;
[0019] Step 2: Divide the source data A, B, N into 4 segments according to 64 bits
[0020] Step 3: Traverse the four 64-bit data of source data A[63:0] and B, perform multiplication in sequence, and store them as t[0]~t[3] respectively.
[0021] Step 4: Under the finite field, obtain the correction coefficient m
[0022] m=t[0]×N′0
[0023] Where t[0] represents the lower 64 bits of the result of step 3, that is, the result data before correction.
[0024] Step 5: Multiply the correction coefficient by the four segments of N in turn and add the corresponding t[0] to t[3] to correct the result so that the corrected value can be divided by R, so that the shift operation can be used instead of the division operation.
[0025] Step 6: Determine whether the four segments of data of A have been traversed. If not, shift A right by 64 bits and return to step 3.
[0026] Step 7: Concatenate t[0] to t[3] from low to high, denoted as T. If T>N, return TN, otherwise return T.
[0027] According to the above algorithm flow, the original algorithm consists of two inner loops, which can be optimized in terms of pipeline. The segmented multiplication operation of step 3 and the result correction of step 5 can be pipelined. There is no need to wait for the loop of step 3 to end completely before starting the loop of step 5. The calculation of step 5 can be started in the next cycle after step 3 starts. Only a small number of cycles need to be added to complete the operation of the two inner loops.
[0028] Modular exponentiation is performed using 2 k The binary segmented scanning algorithm can be optimized to different degrees according to the different values of the segment number k, which is mainly reflected in the number of modular multiplication calls. Considering the actual execution efficiency and resource consumption, the present invention uses the 6-bit segmented scanning method, which can reduce the number of modular multiplications by 12.5% compared with the traditional modular exponentiation binary algorithm. The budget value and correction coefficient required by the 6-bit segmented algorithm are as follows:
[0029] where i∈{0,1,2,3…2 k -1}
[0030] Where n = 256;
[0031] In the formula, modN represents the modulo operation of N, M i Indicates the exponential operation on the source data, n represents the bit width of the modular multiplication, which is 256 here;
[0032] Montgomery Modulo 2 k The binary segmented scanning algorithm (k=6) specifically includes the following steps:
[0033] Step 1: Scan and segment the exponent E according to the rule of 6 bits per group. Here, it is 256 bits. Three zeros need to be added in the high position. Then, scan and segment it according to 6 bits per group to obtain 43 groups of 6-bit segments, recorded as R[0:42], and initialize r=0.
[0034] Step 2: If r=0, initialize S according to the value of R[r].
[0035] Step 3: If r! = 0, then S = MontMult(R[r], S)
[0036] Step 3: S = MontMult(S,S)
[0037] Step 4: Determine whether the process is completed 6 times. If not, return to step 4. If completed, go to step 5.
[0038] Step 5: Correct the result, S = MontMult(S, factor)
[0039] Step 6: Determine whether the scanning of 43 groups of 6-bit segments is completed. If the scanning is completed, output S; otherwise, r is incremented by 1 and return to step 3.
[0040] In a second aspect, a readable storage medium is proposed, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the computing system described in the first aspect is driven to execute predetermined computing instructions.
[0041] Beneficial effects:
[0042] The present invention realizes a segmented scanning Montgomery modular exponentiation algorithm based on software and hardware collaboration of SoC, can flexibly use ARM for pre-calculation, realize modular exponentiation operations of different moduli N, and perform modular exponentiation operations more efficiently.
[0043] The present invention adopts a modular design, and mainly realizes a modular inversion module and a Montgomery modular multiplication module, wherein the modular multiplication module adopts a pipelined CIOS algorithm to efficiently realize modular multiplication operations and improve efficiency.
[0044] The present invention adopts 2 k The binary segmented scanning algorithm is combined with the Montgomery algorithm to balance resource consumption and efficiency, effectively reducing the number of modular multiplication calls.
[0045] The present invention implements the modular inversion module of the extended Euclidean algorithm in hardware, making full use of the computing advantages of FPGA to improve computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is the overall architecture of the software and hardware implementation scheme of the present invention.
[0047] Figure 2 It is the Montgomery CIOS modular multiplication flow chart of the present invention.
[0048] Figure 3 It is a schematic diagram of the CIOS pipeline optimization of the present invention.
[0049] Figure 4 It is a schematic diagram of the optimized modular multiplication waveform designed by the present invention.
[0050] Figure 5 It is a 6-bit segmented scan modular exponentiation state machine of the present invention.
[0051] Figure 6 This is the overall process of the 6-bit segmented scanning modular exponentiation of the present invention. DETAILED DESCRIPTION
[0052] In the following description, a large number of specific details are provided to provide a more thorough understanding of the present invention. However, it is apparent to those skilled in the art that the present invention can be implemented without one or more of these details. In other examples, in order to avoid confusion with the present invention, some technical features well known in the art are not described.
[0053] The specific implementation of the present invention is based on SoC software and hardware collaborative implementation of 256-bit Montgomery modular exponentiation (M E modN), the ARM side is used to complete the overall task scheduling, and 2 k The pre-calculation and correction coefficients required by the binary segmented scanning algorithm are stored in the specified FPGA address, and the result data is moved out of the FPGA after the calculation is completed. The FPGA side completes the calculation-intensive tasks and accelerates the operation of modular exponentiation. The specific structure is shown in Figure 1 .
[0054] The most important module on the ARM side is the pre-calculation module, which is used to generate 2 k The present invention adopts the 6-bit segmented scanning method, according to 2 k The specific process of the binary segment scanning algorithm requires (2 6 ) pre-calculated values and 1 correction coefficient, and the pre-calculated value is only related to M and N, and the correction coefficient is related to the bit width and N. Therefore, in the case of the same modulus N, it can be fixed to a constant after one calculation, which also shows that the Montgomery algorithm is more suitable for the case of fixed modulus N. The specific required budget values and correction coefficients are listed in the algorithm introduction below.
[0055] In addition, a driver module needs to be developed on the ARM side to establish two-way communication with the data preprocessing module, drive the FPGA, and establish two-way data communication with the FPGA. The specific driver is mapped to the Avalon bus on the FPGA side, and specific data transmission is realized through the Avalon bus.
[0056] The FPGA end mainly implements the modular inverse module, which uses the extended Euclidean algorithm to calculate the 64-bit modular inverse value and provides it to the Montgomery module as an additional preprocessing value. Since the result of the modular inverse module is used as the input value of the modular exponentiation algorithm, the specific modular inverse x of a with respect to m, that is, a*x≡1(mod m), is converted into x and y to make ax+my=1, where y is the additional value. This module will not be called in large quantities, but only once at the beginning of the modular exponentiation algorithm, using the direct extended Euclidean algorithm and the properties of the Euclidean algorithm:
[0057] ax+my=a1x1+m1y1=…=1, where the recursive operation is used to get the final result, and the recursive boundary is m n = 0. The core recursive code is as follows:
[0058]
[0059] In the formula, mod represents the modulus operation, / represents integer division, (x n ,y n )(x n-1 y n-1 ) is the solution of two consecutive recursions, (a n ,m n )(a n-1 ,m n-1 ) The source data is transformed between two adjacent recursive operations.
[0060] The data distribution module, as the data handover module between the ARM side and the FPGA side, is used for the access of ARM. Limited by the fact that the ARM side is a 32-bit operating system and the Avalon interface is 128 bits, the actual interaction between ARM and Avalon is 32 bits, and the present invention is a 256-bit modular exponentiation operation, so two segments of data splicing are required in this module. First, the valid 32 bits must be extracted from the 128-bit Avalon data interface, divided into 4 beats, and then the valid data is spliced into 256 bits and transferred to SRAM. This module also completes the function of data reading, reads 256 bits of result data from SRAM, and splits the data.
[0061] The SRAM module stores source data, budget values, correction coefficients and result data and is designed to be 256 bits wide.
[0062] The Montgomery modular exponentiation module is used to implement the specific operation of 256-bit modular exponentiation, which mainly includes: 256-bit Montgomery modular multiplication (MontMult) module, state machine control module (fsm_ctrl), address generation module (addgent). The state machine control module completes the entire modular exponentiation algorithm process and calls the MontMult module in sequence; the address generation module is used to generate the address of the source data and the budget value, take out the corresponding number from the SRAM, and generate the address of the result data when the calculation is completed.
[0063] The 256-bit Montgomery modular multiplication module is the core computing module, which is implemented based on the CIOS optimization algorithm. The specific flowchart is shown in Figure 2 , the original CIOS algorithm includes the following steps:
[0064] Step 1: Start the modular inversion module first and get N′0 represents the modular inverse value required for modular multiplication, which is used here to generate the correction coefficient. Represents the modular inverse operation of the lower 64 bits of N, R represents the modulus value of the 64-bit modular inverse, where R = 2 64 ;
[0065] Step 2: Divide the source data A, B, N into 4 segments according to 64 bits
[0066] Step 3: Traverse the four 64-bit data of source data A[63:0] and B, perform multiplication in sequence, and store them as t[0]~t[3] respectively.
[0067] Step 4: Under the finite field, obtain the correction coefficient m
[0068] m=t[0]×N′0
[0069] Where t[0] represents the lower 64 bits of the result of step 3, that is, the result data before correction.
[0070] Step 5: Multiply the correction coefficient by the four segments of N in turn and add the corresponding t[0] to t[3] to correct the result so that the corrected value can be divided by R, so that the shift operation can be used instead of the division operation.
[0071] Step 6: Determine whether the four segments of data of A have been traversed. If not, shift A right by 64 bits and return to step 3.
[0072] Step 7: Concatenate t[0] to t[3] from low to high, denoted as T. If T>N, return TN, otherwise return T.
[0073] According to the above algorithm flow, the original algorithm consists of two inner loops, which can be optimized in terms of pipeline. The segmented multiplication operation in step 3 and the result correction in step 5 can be pipelined. There is no need to wait for the complete end of the loop of step 3 before starting the loop of step 5. Since it is segmented multiplication and correction, the correction of the result of this segment can be started directly after the first segment multiplication is completed in step 3. Only a small number of cycles need to be added to complete the operation of the two inner loops. See the schematic diagram. Figure 3 , the original two sequentially executed inner loops Inner Loop1 and Inner Loop2 are optimized into two inner loops of pipeline calculation. Theoretically, the two inner loops can be completed by adding one cycle.
[0074] According to the pipelined logic and flow chart, design the waveform, see Figure 4 As can be seen from the figure, the signals temp_cs and temp_after differ by only one beat, and the calculation is performed in a pipeline without waiting. Finally, the final modular multiplication result is output in the 23rd cycle.
[0075] Modular exponentiation is performed using 2 k The binary segmented scanning algorithm can be optimized to different degrees according to the different values of the segment number k, which is mainly reflected in the number of modular multiplication calls. The number of modular multiplications required for different segment numbers and the reduction ratio of modular multiplication compared with the traditional binary modular exponentiation algorithm are shown in the following table:
[0076] Table 1: Comparison of the number of modular multiplications required for different segment numbers and the reduction ratio of the number of modular multiplications compared with the traditional binary modular exponentiation algorithm
[0077] Number of segments / number of modular multiplications <![CDATA[2 k Binary segment scan (times)]]> Reduction ratio (%) Precomputed Values 4bit 378 1.56 17 5bit 357 7.03 33 6bit 336 12.50 65 7bit 324 15.62 129 8bit 320 16.67 257
[0078] Taking into account the actual execution efficiency and resource consumption, the present invention selects a 6-bit segmented scanning method.
[0079] Therefore, the budget value and correction factor required by the 6-bit segmentation algorithm are as follows:
[0080] R[i]=M i modN, where i∈{0, 1, 2, 3…2 k -1}
[0081] factor=2 65n modN, where n=256;
[0082] In the formula, modN represents the modulo operation of N, M i Indicates the exponential operation on the source data, n represents the bit width of the modular multiplication, which is 256 here;
[0083] Binary of i i in decimal R[i] 0000000 0 1 …… …… …… 0111111 63 <![CDATA[M 63 towards N]]> 1000000 64 <![CDATA[2 65×256 modeN]]>
[0084] Montgomery Modulo 2 k The specific implementation of the binary segment scanning algorithm (k = 6) adopts the state machine method, see Figure 5 The state machine has 6 states, namely (IDLE, FETCH, INITIAL, NEXT, LOOP, MODIFY, DONE), where FETCH is the data acquisition state; INITIAL is the initialization state, which is only executed once in the entire process; LOOP is the main loop body, and when k=6, the loop needs to be looped 6 times in each large loop; MODIFY is the state of result modification. See the detailed flowchart for details. Figure 6 , the specific steps are as follows:
[0085] Step 1: Scan and segment the exponent E according to the rule of 6 bits per group. Here, it is 256 bits. Three zeros need to be added in the high position. Then, scan and segment it according to 6 bits per group to obtain 43 groups of 6-bit segments, recorded as R[0:42], and initialize r=0.
[0086] Step 2: If r=0, initialize S according to the value of R[r].
[0087] Step 3: If r! = 0, then S = MontMult(R[r], S)
[0088] Step 3: S = MontMult(S,S)
[0089] Step 4: Determine whether the process is completed 6 times. If not, return to step 4. If completed, go to step 5.
[0090] Step 5: Correct the result, S = MontMult(S, factor)
[0091] Step 6: Determine whether the scanning of 43 groups of 6-bit segments is completed. If the scanning is completed, output S; otherwise, r is incremented by 1 and return to step 3.
[0092] So far, the modular exponentiation operation has been completed. The final S is the final modular exponentiation result. The design is simulated in real time to obtain the waveform of the completed modular exponentiation. Working is pulled high to calculate the modular inverse. The cycle of the modular inverse operation only occupies a small part of the entire modular exponentiation. rdy_all is pulled high to obtain the final result.
[0093] As mentioned above, although the present embodiment has been shown and described with reference to a specific preferred embodiment, it should not be interpreted as limiting the present embodiment itself. Various changes can be made to it in form and detail without departing from the spirit and scope of the present embodiment defined in the appended claims.
Claims
1. A segmented scanning Montgomery modular power calculation system with software and hardware collaboration, characterized in that: The system comprises: The ARM side is used to complete the overall task scheduling and pre-calculation; the ARM side includes: Pre-calculation module, used to generate 2 k The budget value required by the binary segmented scanning modular exponentiation algorithm and the correction factor factor; A driver module is used to establish two-way communication with the data preprocessing module, drive the FPGA end, and establish two-way data communication with the FPGA end; The FPGA end is electrically connected to the ARM end and is used to complete computationally intensive tasks and perform specific operations of modular exponentiation; the FPGA end includes: Montgomery modular exponentiation module, used to implement the specific operation of 256-bit modular exponentiation; The modular inverse module uses the Euclidean expansion algorithm to calculate the 64-bit modular inverse value and provides it to the Montgomery module as a preprocessing value; Data distribution module, as the data handover module between ARM and FPGA, is used for accessing data of ARM; An SRAM module is used for transmitting data on the ARM end and for fetching data from a Montgomery module and a modular inverse module; The result of the modular inverse module is used as the input value of the modular exponentiation algorithm to find the modular inverse x of a with respect to m, that is, a*x≡1(modm), which is converted into finding x and y so that ax+my=1 holds; the modular inverse adopts a direct extended Euclidean algorithm, using the properties of the Euclidean algorithm: ax+my=a1x1+m1y1=…=1, where a recursive operation is performed to obtain the final result; the core recursive code is as follows: In the formula, mod represents the modulus operation, / represents integer division, (x n ,y n )(x n-1 y n-1 ) is the solution of two consecutive recursions, (a n ,m n )(a n-1 ,m n-1 ) Two adjacent recursive source data transformations; The Montgomery modular exponentiation module comprises: 256-bit Montgomery modular multiplication module; The state machine control module sequentially calls the 256-bit Montgomery modular multiplication module to complete the control process of the entire modular exponentiation algorithm; An address generation module, used to generate addresses of source data and budget values, and fetch corresponding numbers from the SRAM module; The 256-bit Montgomery modular multiplication module is implemented based on the CIOS optimization algorithm, including the following steps: Step 1: Start the modular inversion module first and get N′0 represents the modular inverse value required for modular multiplication, which is used here to generate the correction coefficient. Represents the modular inverse operation of the lower 64 bits of N, R represents the modulus value of the 64-bit modular inverse, where R = 2 64 ; Step 2: Divide the source data A, B, N into 4 segments according to 64-bit units; Step 3, traverse the four 64-bit data of source data A[63:0] and B, perform multiplication in sequence, and store them as t[0]~[3] respectively; Step 4: Under the finite field, obtain the correction coefficient m: m=t[0]×N′0 Where t[0] represents the lower 64 bits of the result of step 3, that is, the result data before correction; Step 5: Multiply the correction coefficient by the four segments of N in turn and add the corresponding t[0] to [3] to correct the result so that the corrected value can be divided by R, so that the shift operation can be used instead of the division operation; Step 6: Determine whether the four segments of A have been traversed. If not, shift A right by 64 bits and return to step 3. Step 7: Concatenate t[0] to [3] from low to high, denoted as T. If T>N, return TN, otherwise return T.
2. The computing system according to claim 1, characterized in that The piecewise multiplication operation of step 3 and the result correction of step 5 are pipelined, and there is no need to wait for the loop of step 3 to be completely completed.
3. The computing system according to claim 1, characterized in that The ARM side performs pre-calculation and obtains 2 k The budget value required by the binary segment scanning algorithm, budget 2 6 +1 value and transmit it to SRAM through AvalonBus. The budget value includes the following two parts: M i modN, where i∈{0,1,2,3…2 k -1}; factor=2 65n modN, where n=256; In the formula, modN represents the modulo operation of N, M i Indicates the exponential operation on the source data, and n represents the bit width of the modular multiplication, which is 256 here.
4. The computing system according to claim 1, characterized in that: The Montgomery modular exponentiation module uses 2 k The implementation of binary segmented scanning algorithm includes the following steps: Step 1: Scan and segment the exponent E according to the rule of 6 bits per group. Here, it is 256 bits. Three zeros need to be added to the high position. Then, 43 groups of 6 bits are obtained according to the rule of 6 bits per group. They are recorded as R[0:42] and initialized to r=0. Step 2: If r=0, initialize S according to the value of R[r]; Step 3: If r!=0, then S=MontMult(R[r],S); Step 3, S = MontMult(S, S); Step 4: Determine whether the process is completed 6 times. If not, return to step 4. If completed, go to step 5. Step 5, correct the result, S = MontMult (S, factor); Step 6: Determine whether the scanning of 43 groups of 6-bit segments is completed. If the scanning is completed, output S; otherwise, r is incremented by 1 and return to step 3.
5. A readable storage medium, characterized in that: The readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the computing system according to any one of claims 1 to 4 is driven to execute predetermined computing instructions.
Citation Information
Patent Citations
A method and model for high-speed modular multiplication and modular exponentiation based on FPGA
CN109284085A
IC card decryption method based on improved Montgomery modular exponentiation circuit
CN112491543A