Device and method for processing modular multiplication
By dividing the multiplicand, multiplier and modulus into multiple digital elements, and calculating the updated carry result and sum result in parallel, a two-step method is adopted to solve the problem of low efficiency of modulus multiplication and achieve more efficient calculation.
Patent Information
- Application Number
- CN202111145915.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-07-30
- Filing Date
- 2021-09-28
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-09-28
AI Technical Summary
The implementation efficiency of modular multiplication is low, mainly because the calculation of the carry result and the sum result depends on the results of other characters, resulting in high computational complexity.
The multiplicand, multiplier and modulus are divided into multiple digital elements, and multiple processing components are used to parallelly calculate and update the carry result and the sum result. The remainder of the calculation result is calculated by simplifying the components, and a two-step method is used to solve the problem of iterative carry result and sum result.
Through parallel computing and a two-step approach, the computational efficiency of modular multiplication is improved and the computational complexity is reduced.
Smart Images

Figure CN114385112B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computing device, and more particularly to a modular operation device and a method for processing modular multiplication. Background Art
[0002] Modular multiplication with a large number of operands is widely used in public-key cryptography. For example, modular multiplication operations may include iteratively calculating a carry result and a sum result, and calculating a resulting remainder based on the sum result. However, the calculation of the carry result and sum result corresponding to a word depends on other carry results and other sum results corresponding to other words, resulting in inefficient implementation of modular multiplication. Therefore, efficient modular multiplication is an urgent problem to be solved. Summary of the Invention
[0003] Therefore, the present invention provides an apparatus and method for processing modular multiplication to solve the above-mentioned problem.
[0004] An embodiment of the present invention discloses a modular operation device for processing modular multiplication, comprising: a controller configured to divide a multiplicand into a plurality of multiplicand words, divide a multiplier into a plurality of multiplier words, and divide a modulus into a plurality of modulus words; a first plurality of processing elements coupled to the controller and configured to calculate a first plurality of updated carry results and a first plurality of updated sum results based on the plurality of multiplicand words, a multiplier word of the plurality of multiplier words, a first plurality of carry results, and a first plurality of sum results. result), wherein at least two processing components of the first plurality of processing components calculate at least two updated carry results of the first plurality of updated carry results in parallel based on the multiplier element and at least two multiplicand elements of the plurality of multiplicand elements, and the at least two processing components calculate at least two updated sum results of the first plurality of updated sum results in parallel based on the multiplier element and the at least two multiplicand elements; a second plurality of processing components, coupled to the controller, configured to calculate a second plurality of updated carry results and a second plurality of updated sum results based on the plurality of modulus elements, the first plurality of updated carry results and the first plurality of updated sum results; and a reduction element, coupled to the controller, configured to calculate a result remainder based on the second plurality of updated carry results and the second plurality of updated sum results.
[0005] The present invention also discloses a modular operation device for processing a modular multiplication, comprising: a controller configured to divide a multiplicand into a plurality of multiplicand blocks, a multiplier into a plurality of multiplier blocks, and a modulus into a plurality of modulus blocks; a processing unit coupled to the controller and configured to execute the following instructions: calculating a first plurality of summation results based on a first multiplicand block of the plurality of multiplicand blocks, a first multiplier block of the plurality of multiplier blocks, and a first modulus block of the plurality of modulus blocks; and calculating a second plurality of summation results and a plurality of delayed summation results based on a second multiplicand block of the plurality of multiplicand blocks, the first multiplier block, and a second modulus block of the plurality of modulus blocks. a first plurality of updated sum results according to the first plurality of sum results, the plurality of delayed sum results, the first multiplicand block, a second multiplier block of the plurality of multiplier blocks, and the first modulus block; and a second plurality of updated sum results and a plurality of updated delayed sum results according to the second plurality of sum results, the second multiplicand block, the second multiplier block, and the second modulus block; and a simplification component coupled to the controller and the processing component, configured to calculate a result remainder according to the first plurality of updated sum results, the second plurality of updated sum results, and the plurality of updated delayed sum results. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1 Schematic diagram of an analog-to-digital operation device according to an embodiment of the present invention.
[0007] Figure 2 Schematic diagram of the operation of a modular operation device according to an embodiment of the present invention.
[0008] Figure 3 A table illustrating parallel processing of processing components according to an embodiment of the present invention.
[0009] Figure 4 Schematic diagram of data flow for parallel processing of processing components according to an embodiment of the present invention.
[0010] Figure 5 Schematic diagram of processing components according to an embodiment of the present invention.
[0011] Figure 6 FIG. 4 is a schematic diagram of data dependency between a carry result and a sum result according to an embodiment of the present invention.
[0012] Figure 7This is a scheduling table for modular multiplication operations according to an embodiment of the present invention.
[0013] Figure 8 Flowchart of a process of embodiment 1 of the present invention.
[0014] Figure 9 Schematic diagram of the operation of an analog calculation device according to an embodiment of the present invention.
[0015] Figure 10 Schematic diagram of the operation of an analog calculation device according to an embodiment of the present invention.
[0016] Figure 11 This is a scheduling table for modular multiplication operations according to an embodiment of the present invention.
[0017] Figure 12 FIG. 4 is a schematic diagram of data dependency between a carry result and a sum result according to an embodiment of the present invention.
[0018] Figure 13 FIG. 4 is a schematic diagram of data dependency between a carry result and a sum result according to an embodiment of the present invention.
[0019] Figure 14 FIG. 4 is a schematic diagram of data dependency between a carry result and a sum result according to an embodiment of the present invention.
[0020] Figure 15 Flowchart of a process of embodiment 1 of the present invention.
[0021] Figure 16 Flowchart of a process of embodiment 1 of the present invention.
[0022] Figure 17 Flowchart of a process of embodiment 1 of the present invention.
[0023] Figure 18 A schematic diagram of the data flow of processing components according to an embodiment of the present invention.
[0024] The description of the accompanying drawings is as follows:
[0025] 10, 20, 90, 101: Modulus arithmetic device
[0026] 100: At least one processing circuit
[0027] 110, 220, 920, 1020: at least one storage device
[0028] 114: Program Code
[0029] 120: Communication interface device
[0030] 130: At least one cache memory
[0031] 140, 200, 900, 1000: Controller
[0032] C, Sc: carry result
[0033] S, Ss: total result
[0034] PE: Processing Component
[0035] 210, 910, 1010: Simplified components
[0036] A: Multiplicand
[0037] B: Multiplier
[0038] P: Modulus
[0039] AW: multiplied digit yuan
[0040] BW: Multiplier Yuan
[0041] PW: analog to digital element
[0042] Mc', Sc': Update carry result
[0043] Ms', Ss', S': Update the sum result
[0044] 30: Table
[0045] (S_s)~: shift sum result
[0046] q: additional quotient
[0047] (M_c)~': shift and update the carry result
[0048] 70, 110: Schedule
[0049] 80, 150, 160, 170: Process
[0050] 800-820, 1500-1524, 1600-1618, 1700-1716: Steps
[0051] 1040: Multiple cache memories
[0052] AB: Multiplicand block
[0053] BB: Multiplier Block
[0054] PB: Modular Block
[0055] L: delayed sum result
[0056] L': Update the delayed sum result
[0057] H: delayed carry result DETAILED DESCRIPTION
[0058] Figure 1FIG2 is a schematic diagram of a modular operation device 10 according to an embodiment of the present invention. The modular operation device 10 may include at least one processing circuit (e.g., unit or component) 100, at least one storage device 110, at least one communication interface device 120, at least one cache memory 130, and at least one controller 140. The at least one processing circuit 100 may be a (micro)processor, a multi-core processor, an application-specific integrated circuit (ASIC), or a central processing unit (CPU). The at least one storage device 110 may be any storage device for storing program code 114. The at least one processing circuit 100 may read and execute the program code 114 through the at least one storage device 110. The at least one storage device 110 may be any data storage device. For example, the at least one storage device 110 includes, but is not limited to, a Subscriber Identity Module (SIM), a Read-Only Memory (ROM), a flash memory, a Random-Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a Digital Versatile Disc Read-Only Memory (DVD-ROM), a Blu-ray Disc Read-Only Memory (BD-ROM), a magnetic tape, a hard disk, an optical data storage device, a non-volatile storage device, a non-transitory computer-readable medium (e.g., tangible media), and the like. The at least one communication interface device 120 may be a wireless transceiver that transmits and receives signals (e.g., data, messages, and / or packets) based on the processing results of the at least one processing circuit 100. The at least one cache memory 130 may be any type of cache memory (L1 / L2 / L3 / L4 / L5 / L#).The at least one cache memory 130 is read and executed by the at least one processing circuit 100 and may be directly connected to, adjacent to, or integrated as a part of the at least one processing circuit 100. The at least one controller 140 may control the components included in the analog-to-digital operation device 10.
[0059] According to a modular multiplication, the operation of performing the modular multiplication includes iteratively calculating a carry result and a sum result, and calculating a resulting remainder based on the sum result. For example, the carry result C corresponding to a less significant word (e.g., 32 bits or 64 bits) j-1 , Carry result C j And the sum result S corresponding to a character j , and the sum result S corresponding to a more significant word j+1 is calculated iteratively. However, the sum result S j is based on the carry result C j-1 In addition, the summation result S in the i-th iteration is j is based on the summation result S in the (i-1)th iteration j+1 The modular multiplication may be a Montgomery multiplication, but is not limited thereto.
[0060] Figure 2 Schematic diagram of the operation of the analog-to-digital operation device 20 according to an embodiment of the present invention. Figure 2 The analog-to-digital operation device 20 includes K processing elements PE0-PE K-1 , K processing components PE K ~PE 2K-1 , a controller 200, a reduction element 210 and at least one storage device (eg, main memory) 220, wherein K ≥ 2. The controller 200 is coupled to the processing elements PE0-PE K-1 , processing component PE K ~PE 2K-1 , a simplified component 210 and at least one storage device 220 .
[0061] Specifically, based on the Montgomery multiplication algorithm, the modular operation device 20 is used to calculate the modular multiplication (i.e., the remainder) of a multiplicand A and a multiplier B relative to a modulus P. The controller 200 divides the multiplicand A into (e+1) multiplicand words AW0-AW1. e , divide the multiplier B into (e+1) multiplier words BW0~BW e , and divide the modulus P into (e+1) modulus words PW0~PW e , where e≥1. According to (eg individual) multiplicand elements AW0-AW e 、(e+1) multiplier elements BW0~BW e A multiplication digit, (e+1) carry results Sc0~Sc e And (e+1) sum results Ss0~Ss e , processing components PE0~PE K-1 Calculate (e+1) updated carry results Mc0'~Mc e ' and (e+1) updated sum results Ms0' to Ms e '. According to (for example, individual) modular digital elements PW0~PW e 、(e+1) updated carry results Mc0'~Mc e ' and (e+1) updated sum results Ms0' to Ms e ', processing component PE K ~PE 2K-1 Calculate (e+1) update carry results Sc0'~Sc e ' and (e+1) updated sum results Ss0'~Ss e '. Then, according to the updated carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e ', the simplification component 210 calculates the result remainder S. That is, a two-step method is used to calculate the carry result and the sum result. Therefore, the problem of calculating the iterative carry result and the sum result can be solved.
[0062] In one embodiment, processing components PE0-PE K-1 and processing components PE K ~PE 2K-1 same.
[0063] In one embodiment, the multiplicand A is an n-bit integer. In one embodiment, the multiplicand elements AW0-AW e Each multiplicand element is a w-bit integer. In one embodiment, according to w, (e+1) multiplicand elements AW0~AW e The number of characters in . For example, in It should be noted that, according to the system requirements, the (e+1) multiplicand elements AW0~AW e 、(e+1) multiplier elements BW0~BW e and (e+1) modular digital elements PW0~PW e The number of characters and bit length can be changed, but are not limited to this.
[0064] In one embodiment, the multiplicand elements AW0-AW e , update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e ' have a one-to-one correspondence. In one embodiment, the modulus digital elements PW0 to PW e ', update the carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e ' has a one-to-one correspondence.
[0065] Figure 3 Table 30 is a table showing the parallel processing of processing components according to an embodiment of the present invention. Specifically, processing components PE0 to PE K-1 The number of components K is not greater than (e+1) multiplicand elements AW0~AW e That is, the number of processing elements PE0 to PE K-1 A processing component calculates and updates the carry results Mc0'~Mc e 'At least one updated carry result and updated sum result Ms0'~Ms e In one embodiment, according to the number of characters (e+1) and the number of components K, the updated carry results Mc0' to Mc e ' at least one updated carry result of a number f and updated sum results Ms0 '~Ms e '. For example, the processing element PE0 calculates the updated carry result Mc0', Mc K ',…,Mc f-K ',in In one embodiment, according to the number of components, corresponding to the processing components PE0 to PEK-1 The multiple component indices (element index) and the corresponding multiplicand elements AW0~AW e Multiple character indexes (word index), processing components PE0~PE K-1 Calculate and update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e In one embodiment, the processing element PE corresponding to the element index k k Calculate an updated carry result Mc corresponding to index j j 'And update the sum result Ms j ', where the component index k is equal to the remainder obtained by dividing the index j by the number of components, that is, k=j(mod K). In one embodiment, according to the multiplier element BW i and (e+1) multiplicand elements AW0~AW e At least two multiplicand elements, processing components PE0~PE K-1 At least two processing components of the device calculate and update the carry results Mc0' to Mc in parallel (eg, simultaneously). e 'At least two updated carry results, and according to the multiplier element BW i and the at least two multiplicand elements, processing components PE0 to PE K-1 At least two processing components of the parallel calculation update sum results Ms0'~Ms e In one embodiment, processing components PE0 to PE K-1 The number of at least two processing components, update the carry result Mc0'~Mc e 'The number of at least two updated carry results and the updated sum result Ms0'~Ms e 'The sum of at least two updated results is the same. It should be noted that the processing components PE0~PE K-1 The at least two processing components of the embodiment can perform at least one parallel calculation. For example, the processing components PE0-PE1 parallelly calculate and update the carry results Mc0'-Mc1', and parallelly calculate and update the carry result Mc f-K '~Mc f-K+1 '.
[0066] In one embodiment, the processing element PE K ~PE 2K-1 The number of components K is not greater than (e+1) modular digital elements PW0~PW e That is, the processing element PE K ~PE 2K-1 A processing component calculates and updates the carry result Sc0'~Sce 'At least one updated carry result and updated sum result Ss0'~Ss e In one embodiment, according to the number of characters (e+1) and the number of components K, the updated carry results Sc0' to Sc e ', a quantity f of at least one updated carry result and an updated sum result Ss0' to Ss e ' a number f of at least one sum result. For example, the processing element PE K Calculate and update the carry result Sc0', Sc K ',…,Sc f-K ',in In one embodiment, according to the number of components, corresponding to the processing components PE K ~PE 2K-1 Multiple component indicators and corresponding analog digital elements PW0~PW e Multiple character pointers, processing component PE K ~PE 2K-1 Calculate and update the carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e In one embodiment, the processing element PE corresponding to the element index k k Calculate an updated carry result Sc corresponding to index j j 'And update the sum result Ss j ', wherein the component index k is equal to the remainder obtained by dividing the index j by the number of components, that is, k=j(mod K). In one embodiment, according to the updated carry results Mc0'~Mc e 'At least two updated carry results and updated sum result Ms0'~Ms e 'The sum of at least two updates, processing component PE K ~PE 2K-1 At least two processing components of the processor calculate and update the carry results Sc0' to Sc in parallel (eg, simultaneously). e 'At least two updated carry results, and parallel calculation update sum results Ss0'~Ss e In one embodiment, the processing component PE K ~PE 2K-1 The number of at least two processing components, update the carry result Sc0'~Sc e 'The number of at least two updated carry results and the updated sum result Ss0'~Ss e 'The sum of at least two updated results is the same. Note that the processing component PE K ~PE 2K-1The at least two processing components of the processor can perform at least one parallel computation. K ~PE K+1 Parallel calculation updates the carry results Sc0'-Sc1', and parallel calculation updates the carry results Sc2'-Sc3'. It should be noted that, in order to simplify the embodiment, Figure 3 K processing components PE0~PE K-1 and K processing components PE K ~PE 2K-1 The assumption is the same, but not limited to this.
[0067] According to the above description, the data flow of the parallel processing of the processing components is as follows Figure 4 shown.
[0068] In one embodiment, according to (eg, individual) multiplier elements BW0-BW e 、Multiplied digital element AW0~AW e , carry result Sc0~Sc e And the sum of the results Ss0~Ss e (e+1) shifted sum result Processing components PE0 to PE K-1 Calculate and update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e In one embodiment, the summation results Ss0 to Ss e Divide by radix 2 w (i.e. a right shift of a character) (e.g. a base 2 w system), shift sum result For example, Note that the shift sum result The most significant shifted sum result In practice, the shift sum result is obtained according to at least one delay element, at least one flip-flop or at least one register.
[0069] In one embodiment, the update carry results Mc0'-Mc are calculated according to the following instructions. e Each updated carry result and updated sum result Ms0'~Ms e 'Each update sum result: a multiplicand element AW of multiple multiplicand elements j and the multiplier element BW iMultiply (multiply) to obtain a product (multiplication); carry the result Sc0 ~ Sc e The carry result Sc j and the shift sum result The result of the one-shift sum Add to the product to obtain a number; divide the number by a base 2 w To obtain a quotient and a remainder; determine the quotient as the updated carry result Mc0'~Mc e 'Update the carry result Mc j '; and determine the remainder to update the sum result Ms0'~Ms e 'Update the sum result Ms j '. Note that i is a multiplier word index. Therefore, the above description can be expressed as the following equation:
[0070] AW j ×BW i +Sc j +Ss j+1 =Mc j '2 w +Ms j '
[0071] (Formula 1)
[0072] In one embodiment, processing components PE0-PE K-1 It is also set according to the updated sum result Ms0'~Ms e 'A least significant result Ms0' and an inverse word, calculate an extra quotient q i The reverse character is the modulus element PW0~PW e The inverse of the least significant word PW0 divided by a base 2 w Therefore, the additional quotient q i The calculation can be presented as the following equation:
[0073] (Ms0′×(-PW0 -1 mod 2 w ))mod 2 w =q i (Formula 2)
[0074] In one embodiment, the additional quotient q i and multiplication element BW0~BW e A multiplication of digital element BW i There is a one-to-one correspondence between i , calculate the least significant result Ms0'. In one embodiment, at a corresponding multiplier element BW i In an iteration of (eg, during) the processing components PE0-PE K-1 Calculate the additional quotient q i In one embodiment, the carry results Sc0 to Sc e All carry results and sum results Ss0~Ss e All sum results are initialized to 0.
[0075] In one embodiment, according to the (eg individual) additional quotient q i , analog digital element PW0~PW e , update the carry result Mc0'~Mc e '(e+1) shifted updated carry result And update the sum result Ms0'~Ms e ', processing component PE K ~PE 2K-1 Calculate and update the carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e In one embodiment, the update carry results Mc0' to Mc e 'Multiply by base 2 w (i.e., a left shift of one character) to obtain the shift-update-carry result For example, Note that the shift update carries the result The least significant carry result In fact, the shift-update-carry result is obtained according to at least one delay element, at least one flip-flop or at least one register.
[0076] In one embodiment, the update carry results Sc0′ to Sc e Each updated carry result and updated sum result Ss0'~Ss e 'Each update sum result: Modulus digital element PW0~PW e The analog-to-digital element PW j With the additional quotient q i Multiply to obtain a product; update the carry result by shifting The carry result is updated by a shift And update the sum result Ms0'~Ms e The sum of the results of Ms j 'Add to the product to obtain a number; divide the number by base 2 w To obtain a quotient and a remainder; determine the quotient to update the carry result Sc0'~Sc e 'Update carry result Sc j '; and determine the remainder to update the sum result Ss0'~Ss e 'Update the sum result Ss j '. According to the updated sum result Ms0'~Ms e ' a least significant result and an inverse character, additional quotient q i is generated, and the inverse word is the modulus word PW0~PW e The reciprocal of the least significant character of divided by base 2 w Therefore, the above description can be expressed as the following equation:
[0077] PW j ×q i +Mc j-1 ′+Ms j ′=Sc j '2 w +Ss j ′ (Formula 3)
[0078] Figure 5 Schematic diagram of the processing components of the embodiment of the present invention. K-1 Each processing element contains a digital element AW to be multiplied j and the multiplier element BW i The multiplier (multiplier) for multiplying, and for carrying the result Sc j , shift sum result That is, the processing elements PE0 to PE K-1 Each processing element of the PE performs a multiply and accumulate (MAC) operation. K ~PE 2K-1 Each processing element contains a PW j and the additional quotient q i Multiplier for multiplication, and for updating the shift-carry result Total result Ms j 'And the adder for adding the products. That is, the processing element PEK ~PE 2K-1 Each processing element performs a multiplication-accumulation operation. Note that the additional quotient q in equation (2) i The calculation can be Figure 5 The multiplier in FIG. 1 may be executed, or may be executed by an additional multiplier (not shown in FIG. Figure 5 middle).
[0079] According to the above description, the following equation can be obtained:
[0080]
[0081]
[0082] In one embodiment, in the case corresponding to the multiplier elements BW0 to BW e The first character BW i In the i-th iteration, the processing component PE K ~PE 2K-1 Calculate the carry result Sc0~Sc e And the sum of the results Ss0~Ss e Then, in the corresponding multiplier elements BW0~BW e The second character (e.g. the next character) BW i+1 In the (i+1)th iteration, processing components PE0~PE K-1 Calculate and update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e ', and processing component PE K ~PE 2K-1 Calculate and update the carry result and update the sum result Ss0'~Ss e '.
[0083] According to the above description, the data dependency of the carry result and the sum result in the embodiment of the present invention is as follows: Figure 6 As shown. Figure 6 In the example, the operation executed by the processing element PE0 is represented by C task. K-1 The executed operation is represented by D task. K ~PE 2K-1 The executed operation is represented by H task. Note that, to simplify the embodiment, the number of multiplicand elements AW0-AW3, the number of multiplier elements BW0-BW3 and the number of processing elements (ie, 4) are assumed to be the same, but not limited thereto.
[0084] In one embodiment, in the processing element PEK ~PE 2K-1 In the corresponding multiplier element BW0~BW e The most significant word BW e In the last iteration, the updated carry results Sc0'~Sc e 'And update the sum result Ss0'~Ss e 'After that, the simplifying component 210 calculates the result remainder S. In one embodiment, according to the corresponding updated carry results Sc0'~Sc e ''s multiple weights and corresponding updated summation results Ss0'~Ss e ', the simplification component 210 calculates the result remainder S. For example, weight 2 jw Corresponding to the updated carry result Sc j 'The weight and corresponding update sum result Ss j+1 '. Therefore, the above description can be presented as the following equation:
[0085]
[0086] Figure 7 The scheduling table 70 of the modular multiplication operation according to the embodiment of the present invention is shown in FIG. e , carry result Sc0~Sc e And the sum of the results Ss0~Ss e Afterwards, processing components PE0~PE K-1 Calculate and update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e '. The analog digital elements PW0-PW in at least one storage device 220 are accessed. e , update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e 'After that, process the component PE K ~PE 2K-1 Calculate and update the carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e '. The updated carry results Sc0' to Sc0' in at least one storage device 220 are accessed. e 'And update the sum result Ss0'~Ss e 'After that, the simplifying component 210 calculates the result remainder S. In addition, the updated carry results Mc0' to Mc e 'And update the sum result Ms0'~Ms e'After that, process components PE0~PE K-1 The updated carry results Mc0′-Mc are stored (eg, written) in a first order of at least one storage device 220. e 'And update the sum result Ms0'~Ms e '. In the calculation of the update carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e 'After that, process the component PE K ~PE 2K-1 The updated carry results Sc0′ to Sc e 'And update the sum result Ss0'~Ss e '. Then, after calculating the result remainder S, the simplification component 210 stores the result remainder S in at least one storage device 220. It should be noted that, according to system requirements, Figure 7 The order of storing the iterative sum results and the iterative carry results in can be changed, but is not limited thereto.
[0087] The operation embodiment of the above-mentioned analog-to-digital operation device can be summarized as follows: Figure 8 The process 80 can be compiled into the program code 114. The process 80 includes the following steps:
[0088] Step 800: Start.
[0089] Step 802: The controller divides A into AW0~AW e , divide B into BW0~BW e , and divide P into PW0~PW e .
[0090] Step 804: The controller sets Sc0 to Sc e and Ss0~Ss e Initialized to 0.
[0091] Step 806: Based on the AW in the i-th outer iteration and the u-th inner iteration uK+j ×BW i +Sc uK+j +Ss uK+j+1 , processing components PE0~PE K-1 Each processing element PE j Calculate Mc uK+j and Ms. uK+j .
[0092] Step 808: Based on Ms0 in the i-th outer iteration and the 0-th inner iteration, processing component PE0 calculates qi .
[0093] Step 810: The controller determines whether f inner iterations are complete. If so, execute step 812; otherwise, execute step 806.
[0094] Step 812: Based on the PW in the i-th outer iteration and the v-th inner iteration vK+j ×q i +Mc vK+j-1 +Ms vK+j , each processing element PE j Calculate Sc vK+j and Ss vK+j .
[0095] Step 814: The controller determines whether the f inner iterations are complete. If so, execute step 816; otherwise, execute step 812.
[0096] Step 816: The controller determines whether the (e+1) outer iterations are complete. If so, execute step 818; otherwise, execute step 806.
[0097] Step 818: According to Sc0~Sc e and Ss0~Ss e , simplify the component calculation S.
[0098] Step 820: End.
[0099] The operational details and variations of process 80 can be found in the above description and will not be elaborated here.
[0100] Figure 9 FIG. 1 is a schematic diagram illustrating the operation of the analog-to-digital operation device 90 according to an embodiment of the present invention. Figure 9 In the embodiment, the analog-to-digital operation device includes two processing elements PE0-PE1, a controller 900, a simplification element 910 and at least one storage device 920. The controller 900 is coupled to the processing elements PE0-PE1, the simplification element 910 and the at least one storage device 920.
[0101] Specifically, based on the Montgomery multiplication algorithm, the modular operation device 90 is used to calculate the modular multiplication (i.e., remainder) of a multiplicand A and a multiplier B relative to a modulus P. The controller 900 divides the multiplicand A into (e+1) multiplicand elements AW0-AW1. e , divide the multiplier B into (e+1) multiplier elements BW0~BW e , and divide the modulus P into (e+1) modulus elements PW0~PW e , where e≥1. According to (for example, individual) (e+1) multiplier elements BW0~BW e A multiplication number element, the multiplied number element AW0~AWe 、(e+1) carry results Sc0~Sc e And (e+1) sum results Ss0~Ss e , processing component PE0 calculates the additional quotient q i 、(e+1) updated carry results Mc0'~Mc e ' and (e+1) updated sum results Ms0' to Ms e '. According to the (eg individual) additional quotient q i , analog digital element PW0~PW e , update the carry result Mc0'~Mc e 'And update the sum result Ms0'~Ms e ', the processing element PE1 is set to calculate (e+1) update carry results Sc0'~Sc e ' and (e+1) updated sum results Ss0'~Ss e '. Then, according to the updated carry result Sc0'~Sc e 'And update the sum result Ss0'~Ss e ', the simplification component 910 calculates the result remainder S. That is, a two-step method is used to calculate the carry result and the sum result. Therefore, the problem of calculating the iterative carry result and the sum result can be solved.
[0102] In one embodiment, the processing element PE0 and the processing element PE1 are the same processing element.
[0103] The operational details and variations of process 90 can be found in the above description and will not be elaborated here.
[0104] Figure 10 FIG. 1 is a schematic diagram illustrating the operation of the analog-to-digital operation device 101 according to an embodiment of the present invention. Figure 10 In the embodiment, the analog-to-digital operation device 101 includes a processing element PE, a controller 1000, a simplification element 1010, at least one storage device (e.g., main memory) 1020, and a plurality of cache memories 1040. The controller 1000 is coupled to the processing element PE, the simplification element 1010, the at least one storage device 1020, and the plurality of cache memories 1040.
[0105] Specifically, based on the Montgomery multiplication algorithm, the modular operation device 101 is used to calculate the modular multiplication (i.e., remainder) of a multiplicand A and a multiplier B relative to a modulus P. The controller 1000 divides the multiplicand A into K multiplicand blocks AB0-AB K-1 , divide the multiplier B into K multiplier blocks BB0~BBK-1 , and divide the modulus P into K modulus blocks PB0~PB K-1 , where K≥2. The processing element PE is configured to execute an instruction. The instruction includes the multiplicand blocks AB0 to AB K-1 The multiplicand block AB j , Modulus block PB0~PB K-1 Modulus block PB j and multiplier blocks BB0~BB K-1 Multiplier Block BB i , calculate f summation results S jf ~S (j+1)f-1 , where f≥2. The instruction includes the following steps: K-1 The multiplicand block AB j+1 , Modulus block PB0~PB K-1 Modulus block PB j+1 and multiplier block BB i , calculate f summation results S (j+1)f ~S (j+2)f-1 and f delayed sum results L jf ~L (j+1)f-1 The instruction includes the following steps: jf ~S (j+1)f-1 , delayed sum result L jf ~L (j+1)f-1 , multiplicand block AB j , modular block PB j and BB0~BB K-1 Multiplier Block BB i+1 , calculate the sum of f updates S jf '~S (j+1)f-1 '. The instruction contains the sum result S (j+1)f ~S (j+2)f-1 , multiplicand block AB j+1 , modular block PB j+1 and multiplier block BB i+1 , calculate the updated sum result S (j+1)f '~S (j+2)f-1 ' and updated delayed sum result (updated delayed sum result) jf '~L (j+1)f-1 '. Then, the simplified component 1010 is set to update the sum result S according to jf '~S (j+1)f-1 ', update the sum result S (j+1)f '~S (j+2)f-1 'And update the delayed sum result Ljf '~L (j+1)f-1 Calculate a result remainder S. That is, a post-processing method is used to calculate the carry result and the sum result. Therefore, the problem of calculating the iterative carry result and the sum result can be solved.
[0106] In one embodiment, according to the multiplicand block AB j+1 , modular block PB j+1 , Multiplier Block BB i+1 and f additional quotients q if ~q (i+1)f-1 , the processing component PE calculates the delay sum result L jf ~L (j+1)f-1 .
[0107] In one embodiment, the multiplicand A is an n-bit integer. In one embodiment, the multiplicand blocks AB0 to AB K-1 Each multiplicand block of is a b-bit integer and contains f multiplicand elements. In one embodiment, each multiplicand element is a w-bit integer, i.e. as well as It should be noted that, according to system requirements, the multiplicand blocks AB0 to AB K-1 Multiplier blocks BB0~BB K-1 And modular blocks PB0~PB K-1 The number of blocks and the bit length can be changed, but are not limited to this.
[0108] In one embodiment, the sum result S jf ~S (j+1)f-1 , update the sum result S jf '~S (j+1)f-1 ' and the multiplicand block AB j The f multiplicand elements AW jf ~AW (j+1)f-1 There is a one-to-one correspondence between (j+1)f ~S (j+2)f-1 , update the sum result S (j+1)f '~S (j+2)f-1 ' and the multiplicand block AB j+1 The f multiplicand elements AW (j+1)f ~AW (j+2)f-1 There is a one-to-one correspondence between them.
[0109] In one embodiment, the delayed summation result L jf ~L (j+1)f-1 And the sum result S jf ~S (j+1)f-1 In one embodiment, the delayed summation result Ljf ~L (j+1)f-1 The number and sum of the results S jf ~S (j+1)f-1 In one embodiment, the delayed summation result L jf ~L (j+1)f-1 The number of results is less than the sum of S jf ~S (j+1)f-1 In one embodiment, the delayed summation result L jf ~L (j+1)f-1 and multiplier block BB i f multiplier elements BW if ~BW (i+1)f-1 In one embodiment, the update delay summation result L jf '~L (j+1)f-1 ' and multiplier block BB i+1 f multiplier elements BW (i+1)f ~BW (i+2)f-1 There is a correspondence between them.
[0110] In one embodiment, the analog-to-digital operation device 101 further includes a loading and storing element coupled to the controller 1000, wherein the loading and storing element is configured to execute an instruction. The instruction includes calculating the sum result S in the processing element. jf ~S (j+1)f-1 Before, the multiplicand block AB is loaded (or copied) from at least one storage device 1020 j and modular block PB j to multiple cache memories 1040 (ie, multiplicand blocks AB j and modular block PB j The instruction includes calculating the sum result S in the processing element. (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 Before, the multiplicand block AB is loaded from at least one storage device 1020 j+1 and modular block PB j+1 to cache memory 1040. The instruction includes calculating the updated sum result S in the processing component. jf '~S (j+1)f-1 'Before, load the sum result S from at least one storage device 1020 jf ~S (j+1)f-1 , delayed sum result L jf ~L (j+1)f-1 , multiplicand block AB j and modular block PB jto multiple cache memories 1040. The instruction includes calculating the updated sum result S in the processing component. (j+1)f '~S (j+2)f-1 'And update the delayed sum result L jf '~L (j+1)f-1 'Before, load the sum result S from at least one storage device 1020 (j+1)f ~S (j+2)f-1 , multiplicand block AB j+1 and modular block PB j+1 In one embodiment, after accessing (eg, reading) the (eg, loaded) multiplicand block AB in the plurality of cache memories 1040, j Afterwards, the processing component calculates the sum result S jf ~S (j+1)f-1 . In accessing the multiplicand blocks AB in the plurality of cache memories 1040 j+1 Afterwards, the processing component calculates the sum result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 The sum result S in accessing multiple cache memories 1040 jf ~S (j+1)f-1 , delayed sum result L jf ~L (j+1)f-1 and multiplicand block AB j Afterwards, the processing component calculates the updated sum result S jf '~S (j+1)f-1 '. The sum result S in accessing multiple cache memories 1040 (j+1)f ~S (j+2)f-1 and multiplicand block AB j+1 Afterwards, the processing component calculates the updated sum result S (j+1)f '~S (j+2)f-1 That is, the data used for calculation in each block (e.g., the multiplicand elements of the multiplicand block) is only loaded into the cache memory once. In other words, the number of cache misses can be reduced.
[0111] In one embodiment, the instructions executed by the processing component include calculating the sum result S jf ~S (j+1)f-1 Then, the sum result S is stored (eg, written) in the plurality of cache memories 1040. jf ~S (j+1)f-1 The instruction includes calculating the sum result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1Then, the sum result S is stored in the plurality of cache memories 1040 (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 The instruction includes the following steps: jf '~S (j+1)f-1 'Afterwards, the updated sum result S is stored in multiple cache memories 1040 jf '~S (j+1)f-1 '. The instruction contains the following steps: (j+1)f '~S (j+2)f-1 'Afterwards, the updated sum result S is stored in multiple cache memories 1040 (j+1)f '~S (j+2)f-1 In one embodiment, the load and store components are configured to execute instructions. The instructions include calculating the sum result S in the processing component. jf ~S (j+1)f-1 After that, sum the result S jf ~S (j+1)f-1 The instruction includes calculating the sum result S in the processing component. (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 After that, sum the result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 The instruction includes calculating the updated sum result S in the processing component. jf '~S (j+1)f-1 'After that, the sum result S will be updated jf '~S (j+1)f-1 'Store in at least one storage device 1020. The instruction includes calculating the updated sum result S in the processing component (j+1)f '~S (j+2)f-1 'After that, the sum result S will be updated (j+1)f '~S (j+2)f-1 'Stored in at least one storage device 1020.
[0112] In one embodiment, the storage and loading component further executes instructions including loading the multiplicand block AB j and modular block PB j to a first cache memory of the plurality of cache memories 1040. The instruction includes when the processing component calculates the sum result S jf ~S (j+1)f-1 When the multiplicand block AB is loaded j+1 and modular block PBj+1 to a second cache memory of the plurality of cache memories 1040. The instruction includes when the processing component calculates the sum result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 When the sum result S is loaded jf ~S (j+1)f-1 , delayed sum result L jf ~L (j+1)f-1 , multiplicand block AB j and modular block PB j to the first cache memory. The instruction includes when the processing component calculates the updated sum result S jf '~S (j+1)f-1 ', load the sum result S (j+1)f ~S (j+2)f-1 , multiplicand block AB j+1 and modular block PB j+1 to the second cache memory.
[0113] Figure 11 The scheduling table 110 of the modular multiplication operation of the embodiment of the present invention is shown in FIG. jf ~S (j+1)f-1 Previously, the load and store component loaded the multiplicand block AB j and modular block PB j to the first cache memory. When the processing component calculates the sum result S jf ~S (j+1)f-1 When the load and store component loads the multiplicand block AB j+1 and modular block PB j+1 To the second cache memory. When the processing component calculates the sum result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 When the load and store component stores (eg stored) the sum result S jf ~S (j+1)f-1 To the first cache memory in at least one storage device 1020. The processing component calculates the updated sum result S jf '~S (j+1)f-1 'Before, the load and store components load the multiplicand block AB j and modular block PB j To the first cache memory. When the processing component calculates and updates the sum result S jf '~S (j+1)f-1 ', the loading and storage component loads the multiplicand block AB j+1 and modular block PBj+1 To the second cache memory. When the processing component calculates and updates the sum result S jf '~S (j+1)f-1 ', the load and store component stores the sum result S (j+1)f ~S (j+2)f-1 and the delayed sum result L jf ~L (j+1)f-1 to the second cache memory in at least one storage device 1020. When the processing component calculates and updates the sum result S (j+1)f '~S (j+2)f-1 ', the load and store component stores the updated sum result S jf '~S (j+1)f-1 ' to the first cache memory in at least one storage device 1020. The load and store component stores the updated sum result S (j+1)f '~S (j+2)f-1 to the second cache memory in at least one storage device 1020. That is, in the modular multiplication, a ping-pong cache memory (eg, a ping-pong buffer) is used to calculate the sum result.
[0114] In one embodiment, the processing component further executes an instruction including a block AB of the multiplicand j The most significant word AW (j+1)f-1 and multiplier block BB i , calculate f delayed carry results (delayed carry result) H if ~H (i+1)f-1 The instruction includes the most significant word AW (j+1)f-1 and multiplier block BB i+1 , calculate f delayed carry results H (i+1)f ~H (i+2)f-1 In one embodiment, the simplify component is further configured to execute the following instructions: update the sum result S jf '~S (j+1)f-1 ', update the sum result S (j+1)f '~S (j+2)f-1 ', update the delayed sum result L jf '~L (j+1)f-1 ', calculate the result remainder S. In one embodiment, according to the multiplicand block AB j+1 The least significant character AW (j+1)f and multiplier block BB i , the processing component calculates the sum of the delay results L jf ~L (j+1)f-1, and according to the least significant character AW (j+1)f and multiplier block BB i+1 , the processing component calculates the update delay sum result L jf '~L (j+1)f-1 In one embodiment, based on (eg, only based on) the least significant character AW (j+1)f , Multiplier Block BB i , modular block PB (j+1)f , f additional quotients q if ~q (i+1)f-1 and the first plurality of temporary carry results, the processing component calculates the delayed sum result L jf ~L (j+1)f-1 Based on (eg, only on) the least significant character AW (j+1)f , Multiplier Block BB i+1 , and the second plurality of temporary carry results, the processing component calculates and updates the delayed sum result L jf '~L (j+1)f-1 ', where according to the multiplicand block AB j+1 The least significant character AW (j+1)f+1 , calculate the first plurality of temporary carry results and the second plurality of temporary carry results. For example, the delayed carry result H if ~H (i+1)f-1 It is based on the multiplicand block AB j The most significant character AW (j+1)f-1 The calculation results show that the processing component does not follow the delayed carry result H if ~H (i+1)f-1 Calculate the delay sum result L jf ~L (j+1)f-1 . And, the delayed carry result H (i+1)f ~H (i+2)f-1 Based on the most significant character AW (j+1)f-1 The calculation results show that the processing component does not follow the delayed carry result H if ~H (i+1)f-1 Calculate the update delay sum L jf '~L (j+1)f-1 '.
[0115] In one embodiment, the processing component further executes an instruction including a block AB of the multiplicand j The most significant character AW (j+1)f-1 and multiplier block BB i The most significant character BW (i+1)f-1 , calculate the delayed carry result H (i+1)f-1 ; Determine the delayed carry result H (i+1)f-1 The sum result S jf ~S(j+1)f-1 The most significant result S (j+1)f-1 The instruction includes the following steps: j The most significant character AW (j+2)f-1 and multiplier block BB i The most significant character BW (i+1)f-1 , calculate the delayed carry result H (i+1)f-1 '; and determine the delayed carry result H (i+1)f-1 ' is the sum result S (j+1)f ~S (j+2)f-1 The most significant result S (j+2)f-1 .
[0116] In one embodiment, the updated sum result S is calculated jf '~S (j+1)f-1 'The instructions include: According to the sum result S jf ~S (j+1)f-1 , delayed sum result L jf ~L (j+1)f-1 , multiplicand block AB j and multiplier block BB i+1 , calculating a plurality of temporary sum results based on a first word of the first word; and calculating a plurality of temporary sum results based on the plurality of temporary sum results, the multiplicand block AB j and multiplier block BB i+1 The second character of the calculation updates the sum result S jf '~S (j+1)f-1 In one embodiment, the multiplier block BB i+1 The first word is the multiplier block BB i+1 The least significant character BW (i+1)f (i.e. iterating over the sum result i.e. iterating over the previous processing of the delayed sum result).
[0117] In one embodiment, the processing component further executes the following instructions: f-1 The least significant result S0, delayed summation results L0~L f-1 A least significant result L0 and corresponding to the multiplier block BB i The least significant word BW if+u An inverse word in an iteration of , calculate the additional quotient q if+u , wherein the inverse character is the modulus block PB0~PB K-1 The reciprocal of the least significant character PB0 divided by a base 2 w The remainder of .
[0118] In one embodiment, the additional quotient block qB iand multiplier blocks BB0~BB K-1 A multiplier block BB i There is a correspondence between the multiplier block BB i , calculate the minimum valid result S jf In one embodiment, in the block corresponding to the multiplier BB i In one iteration, the processing component calculates an additional quotient block qB i once.
[0119] Based on the above, an example of pseudo code is as follows:
[0120] Initialize S to 0.
[0121] Initialize carry out resultsL,L_n,Ca and Cb to 0.
[0122] for(i=0:k-1)begin
[0123] Perform a block_0process. / / Process a least significant block.
[0124] for(j=1:k-1)begin
[0125] Perform at least one block_1process. / / Process other blocks.
[0126] end
[0127] Determine L_nas L (k-1)f .
[0128] Determine Ca[k-1]as L_n.
[0129] Determine Cb[k-2:0]as Ca[k-2:0].
[0130] end
[0131] / / Post process of S+L.
[0132] Determine 0 as Cc and Ca[-1].
[0133] for(i=0:k-1)begin
[0134] Compute Cc and S j according to S j +L j +Ca[i-1]+Cc.
[0135] for(j=0:f-1)begin
[0136] Compute Cc and S j according to S j +L j +Cc.
[0137] end
[0138] end
[0139] Determine Cc+L_n as S fk .
[0140] Ca[k-1]is a(k-1)-th bit of Ca,and Cb[k-2:0]is a segment with a mostsignificant bit(k-2)and a least significant bit 0.
[0141] In addition,the block_0 process can be obtained as follows:
[0142] Initialize Cb[0]to 0.
[0143] for(v=0:f-1)begin
[0144] Compute Cb[0]and S v according to S v +L v +Cb[0].
[0145] end
[0146] for(u=0:f-1)begin
[0147] / / X task
[0148] Determine AW0×BW if+u +S0as T.
[0149] Determine T[w-1:0]as S0.
[0150] Determine T[2w-1:w]as Mc.
[0151] Determine S0×tas q jf+u .
[0152] Determine PW0×q if+u +S0as T′.
[0153] Determine T′[2w-1:w]as Sc.
[0154] for(v=1:f-1)begin / / Y task
[0155] Determine AW v ×BW if+u +S v +Mc as T.
[0156] Determine T[w-1:0]as S v .
[0157] Determine T[2w-1:w]as Mc.
[0158] Determine PW v ×q if+u +S v +S c as T′.
[0159] Determine T′[w-1:0]as S v-1 .
[0160] Determine T′[2w-1:w]as Sc.
[0161] end
[0162] Compute Cb[0]and S f-1 according to Cb[0]+Mc+Sc.
[0163] end
[0164] In addition,the block_1 process can be obtained as follows:
[0165] Initialize Cb[j]to Ca[j-1].
[0166] for(v=0:f-1)begin
[0167] Compute Cb[j]and S jf+v according to S jf+v +L jf+v +Cb[j].
[0168] end
[0169] for(u=0:f-1)begin
[0170] / / X task
[0171] Determine AW jf ×BW if+u +S jf as T.
[0172] Determine T[w-1:0]as S jf .
[0173] Determine T[2w-1:w]as Mc.
[0174] Determine PW jf ×q if+u +S jf as T′.
[0175] Determine T′[w-1:0]as L (j-1)f+u .
[0176] Determine T′[2w-1:w]as Sc.
[0177] for(v=1:f-1)begin / / Y task
[0178] Determine AW jf+v ×BW if+u +S jf+v +Mc as T.
[0179] Determine T[w-1:0]as S jf+v .
[0180] Determine T[2w-1:w]as Mc.
[0181] Determine PW jf+v ×q if+u +S jf+v+Sc as T′.Determine T′[w-1:0]as S jf+v-1 .
[0182] Determine T′[2w-1:w]as Sc.
[0183] end
[0184] Compute Cb[j]and S jf+f-1 according to Cb[j]+Mc+Sc.
[0185] end
[0186] According to the above description, the data dependency of the carry result and the sum result in the embodiment of the present invention is as follows: Figure 12 It should be noted that the number of multiplicand blocks AB0-AB3 and the number of multiplier blocks BB0-BB4 are not limited thereto.
[0187] Specifically, the data dependencies between the carry result and the sum result in a block_0 process are as follows: Figure 13 In addition, the data dependency between the carry result and the sum result in a block_1 process is as follows: Figure 14 shown.
[0188] The operation embodiment of the above-mentioned analog-to-digital operation device can be summarized as follows: Figure 15 The process 150 can be compiled into the program code 114. The process 1500 includes the following steps:
[0189] Step 1500: Start.
[0190] Step 1502: The controller divides A into AB0~AB K-1 , split B into BB0~BB K-1 , and divide P into PB0~PB K-1 .
[0191] Step 1504: In the i-th outer iteration, load and store components into BB i to cache memory.
[0192] Step 1506: In the i-th outer iteration and the j-th inner iteration, load and store components into L jf ~L (j+1)f-1 to cache memory.
[0193] Step 1508: In the i-th outer iteration and the j-th inner iteration, load and store components into AB j PB j and S jf ~S (j+1)f-1 .
[0194] Step 1510 : The controller determines whether j is 0. If so, execute step 1512 , otherwise execute step 1514 .
[0195] Step 1512: By executing the block_0 process, the processing component calculates S0~S f-1 , and in the i-th outer iteration and the j-th inner iteration, store S0~S f-1 . Execute step 1516.
[0196] Step 1514: By executing the block_1 process, the processing component calculates S jf ~S (j+1)f-1 and L (j-1)f ~L jf-1 , and in the i-th outer iteration and the j-th inner iteration, store S jf ~S (j+1)f-1 and L (j-1)f ~L jf-1 .
[0197] Step 1516: The controller determines whether the K inner iterations are complete. If so, the controller executes step 1520; otherwise, the controller executes step 1518.
[0198] Step 1518: The controller determines whether the (K-1) inner iterations are complete. If so, execute step 1508; otherwise, execute step 1506.
[0199] Step 1520: The controller determines whether the K outer iterations are complete. If so, execute step 1522; otherwise, execute step 1504.
[0200] Step 1522: According to S0~S Kf-1 and L0~L (K-1)f-1 , by performing post-processing, the simplified component calculation S.
[0201] Step 1524: End.
[0202] Figure 15 The block_0 process of step 1512 can be Figure 16 The process 160 is implemented as follows. The process 160 includes the following steps:
[0203] Step 1600: Start.
[0204] Step 1602: The controller initializes Mc to 0.
[0205] Step 1604: According to AW in the uth outer iteration and the vth inner iteration v ×BW if+u +S v +Mc, processing component calculates Sv and Mc.
[0206] Step 1606: Based on S0 in the uth outer iteration and the 0th inner iteration, the processing component calculates q if+u .
[0207] Step 1608: The controller determines whether the f inner iterations are complete. If so, execute step 1610; otherwise, execute step 1604.
[0208] Step 1610: The controller initializes Sc to 0.
[0209] Step 1612: Based on the PW in the uth outer iteration and the vth inner iteration v ×q if+u +S v +Sc, processing component calculates S v and Sc.
[0210] Step 1614: The controller determines whether the f inner iterations are complete. If so, execute step 1616; otherwise, execute step 1612.
[0211] Step 1616: The controller determines whether the f outer iterations are complete. If so, execute step 1618; otherwise, execute step 1602.
[0212] Step 1618: End.
[0213] Figure 15 The block_1 process of step 1514 in the Figure 17 The process 170 is implemented as follows. The process 170 includes the following steps:
[0214] Step 1700: Start.
[0215] Step 1702: The controller initializes Mc to 0.
[0216] Step 1704: According to AW in the uth outer iteration and the vth inner iteration jf+v ×BW if+u +S jf+v +Mc, processing component calculates S v and Mc.
[0217] Step 1706: The controller determines whether the f inner iterations are complete. If so, execute step 1708; otherwise, execute step 1704.
[0218] Step 1708: The controller initializes Sc to 0.
[0219] Step 1710: Based on the PW in the uth outer iteration and the vth inner iteration jf+v ×qif+u +S jf+v +Sc, processing component calculates S v and Sc.
[0220] Step 1712: The controller determines whether f inner iterations are complete. If so, execute step 1714; otherwise, execute step 1710.
[0221] Step 1714: The controller determines whether the f outer iterations are complete. If so, execute step 1716; otherwise, execute step 1702.
[0222] Step 1716: End.
[0223] The operational details and variations of processes 150 , 160 and 170 can be found in the above description and will not be elaborated here.
[0224] According to the above, the data flow of the processing component is as follows Figure 18 As shown. Figure 18 , t=-PW0 -1 mod2 w .
[0225] It should be noted that the modular multiplication provided in the present invention can be regarded as an improved and efficient Montgomery modular multiplication.
[0226] In the above operations, the word “determine” can be replaced by “compute / calculate,” “obtain,” “generate,” “output,” “use,” “choose / select,” or “decide.” The term “according to” can be replaced by “in response to.”
[0227] Those skilled in the art will be able to combine, modify, and / or vary the above-described embodiments in accordance with the spirit of the present invention. The aforementioned statements, steps, and / or processes (including suggested steps) may be implemented by a device, which may be hardware, software, firmware (a combination of a hardware device and computer instructions and data, where the computer instructions and data are read-only software on the hardware device), an electronic system, or a combination of the above. For example, the device may be analog-to-digital operation device 10.
[0228] The hardware may be analog circuits, digital circuits, and / or hybrid circuits. For example, the hardware may be an application-specific integrated circuit, a field programmable gate array (FPGA), a programmable logic device (PLD), coupled hardware components, or a combination of the foregoing. In other embodiments, the hardware may be a general-purpose processor, a microprocessor, a controller, a digital signal processor (DSP), or a combination of the foregoing.
[0229] Software may be a combination of program codes, instructions, and / or functions stored in a storage unit, such as a computer-readable medium. For example, the computer-readable medium may be a subscriber identity module (SIM), read-only memory (ROM), flash memory, random access memory (RAM), compact disc read-only memory (CD-ROM / DVD-ROM / BD-ROM), magnetic tape, hard disk, optical data storage device, non-volatile memory (NVSM), or a combination thereof. The computer-readable medium (e.g., storage unit) may be internally coupled to at least one processor (e.g., a processor integrated with the computer-readable medium) or externally coupled to at least one processor (e.g., a processor independent of the computer-readable medium). The at least one processor may include one or more modules for executing the software stored in the computer-readable medium. The combination of program codes, instructions, and / or functions may cause the at least one processor, one or more modules, hardware, and / or electronic system to perform the corresponding steps.
[0230] The electronic system may be a system on chip (SoC), a system in package (SiP), an embedded computer on module (CoM), a computer program product, a device, a mobile phone, a notebook computer, a tablet computer, an e-book, a portable computer system, and an analog-to-digital computing device 10 .
[0231] Based on the foregoing, the present invention provides an apparatus and method for processing modular multiplication. The operations performed by the modular arithmetic device are defined. A two-step method is employed to calculate the carry result and the sum result. Thus, the problem of calculating iterative carry results and sum results is solved.
[0232] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A modular operation device for processing a modular multiplication, comprising: a controller configured to divide a multiplicand into a plurality of multiplicand elements, divide a multiplier into a plurality of multiplier elements, and divide a modulus into a plurality of modulus elements; a first plurality of processing components coupled to the controller and configured to calculate a first plurality of updated carry results and a first plurality of updated sum results based on the plurality of multiplicand elements, a multiplier element of the plurality of multiplier elements, a first plurality of carry results, and a first plurality of sum results, wherein at least two processing components of the first plurality of processing components calculate at least two updated carry results of the first plurality of updated carry results in parallel based on the multiplier element and at least two multiplicand elements of the plurality of multiplicand elements, and the at least two processing components calculate at least two updated sum results of the first plurality of updated sum results in parallel based on the multiplier element and the at least two multiplicand elements; a second plurality of processing components coupled to the controller, configured to calculate a second plurality of updated carry results and a second plurality of updated sum results based on the plurality of modulus digital elements, the first plurality of updated carry results, and the first plurality of updated sum results; as well as a simplification component coupled to the controller and configured to calculate a result remainder according to the second plurality of updated carry results and the second plurality of updated sum results; The first plurality of processing components are used to execute equation AW j ×BW i +Sc j +Ss j+1 =Mc j '2 w +Ms j ', where j and j+1 are indices, i is a multiplier element index, AWj is one of the multiplicand elements, BW i is one of the multiple multiplier elements, Sc j is one of the first plurality of carry results, Ss j+1 is the shift sum result, Mc j ' is one of the first multiple update carry results, 2 w is the cardinality, 2 w W is the number of digits in the character, Ms j ' is one of the first plurality of updated sum results; The second plurality of processing components are configured to execute equation PW j ×q i +Mc j-1 '+Ms j '=Sc j '2 w +Ss j ', where j-1 is the index, PW j is one of the plurality of modular digital elements, q i is the additional quotient, Mc j-1 ' is the shift update carry result, Sc j ' is one of the second plurality of update carry results, Ss j ' is one of the second plurality of updated summation results.
2. The modular arithmetic device according to claim 1, wherein: At least two processing components of the second plurality of processing components calculate in parallel at least two updated carry results of the second plurality of updated carry results based on at least two updated carry results of the first plurality of updated carry results and at least two updated sum results of the first plurality of updated sum results, and calculate in parallel at least two updated sum results of the second plurality of updated sum results.
3. The modular arithmetic device according to claim 1, wherein: The first plurality of processing components calculates the first plurality of updated carry results and the first plurality of updated sum results based on the plurality of multiplicand elements, the plurality of multiplier elements, the first plurality of carry results, and a plurality of shifted sum results of the first plurality of sum results.
4. The modular operation device according to claim 3, wherein: Each updated carry result of the first plurality of updated carry results and each updated sum result of the first plurality of updated sum results are calculated according to the following instructions: multiplying a multiplicand element of the plurality of multiplicand elements by the multiplier element to obtain a product; adding a carry result of the first plurality of carry results and a shifted sum result of the plurality of shifted sum results to the product to obtain a number; Dividing the number by the base to obtain a quotient and a remainder; determining the quotient as an updated carry result of the first plurality of updated carry results; as well as The remainder is determined to be an updated sum result of the first plurality of updated sum results.
5. The analog-to-digital operation device according to claim 1, wherein: The first plurality of processing components are further configured to execute the following instructions: An additional quotient is calculated based on a least significant result of the first plurality of updated summation results and an inverse word, wherein the inverse word is a remainder of a reciprocal of a least significant word of the plurality of modulus words divided by the base.
6. The modular arithmetic device according to claim 1, wherein: The second processing components calculate the second updated carry results and the second updated sum results based on the plurality of modulus digital elements, the shifted updated carry results of the first updated carry results, and the first updated sum results.
7. The modular operation device according to claim 6, wherein: Each updated carry result of the second plurality of updated carry results and each updated sum result of the second plurality of updated sum results are calculated according to the following instructions: multiplying a modular digital element of the plurality of modular digital elements by the additional quotient to obtain a product; adding a shift-updated carry result of the plurality of shift-updated carry results and a sum result of the first plurality of updated sum results to the product to obtain a number; Dividing the number by the base to obtain a quotient and a remainder; determining the quotient as an updated carry result of the second plurality of updated carry results; as well as determining the remainder as an updated sum result of the second plurality of updated sum results; The additional quotient is generated according to a least significant result of the first plurality of updated summation results and an inverse word, and the inverse word is a remainder of a reciprocal of a least significant word of the plurality of modulus words divided by the base.
8. The modular arithmetic device according to claim 1, wherein: In a first iteration corresponding to a first word of the plurality of multiplier elements, the second plurality of processing components calculate the first plurality of carry results and the first plurality of sum results; In a second iteration corresponding to a second word of the plurality of multiplier elements, the first plurality of processing components calculate the first plurality of updated carry results and the first plurality of updated sum results; as well as In the second iteration, the second plurality of processing components calculate the second plurality of updated carry results and the second plurality of updated sum results.
9. A modular operation device for processing a modular multiplication, comprising: a controller configured to divide a multiplicand into a plurality of multiplicand blocks, divide a multiplier into a plurality of multiplier blocks, and divide a modulus into a plurality of modulus blocks; A processing component, coupled to the controller, configured to execute the following instructions: calculating a first plurality of summation results based on a first multiplicand block of the plurality of multiplicand blocks, a first multiplier block of the plurality of multiplier blocks, and a first modulus block of the plurality of modulus blocks; calculating a second plurality of summation results and a plurality of delayed summation results according to a second multiplicand block of the plurality of multiplicand blocks, the first multiplier block, and a second modulus block of the plurality of modulus blocks; calculating a first plurality of updated summation results based on the first plurality of summation results, the plurality of delayed summation results, the first multiplicand block, a second multiplier block of the plurality of multiplier blocks, and the first modulus block; as well as calculating a second plurality of updated sum results and a plurality of updated delayed sum results based on the second plurality of sum results, the second multiplicand block, the second multiplier block, and the second modulus block; as well as A simplification component is coupled to the controller and the processing component, and is configured to calculate a result remainder according to the first plurality of updated sum results, the second plurality of updated sum results, and the plurality of updated delayed sum results.
10. The modular operation device according to claim 9, wherein: A quantity of the plurality of delayed sum results is the same as a quantity of the first plurality of sum results.
11. The modular digital operation device according to claim 9, further comprising: at least one storage device; Multiple cache memories; and A loading and storing component is coupled to the controller and configured to execute the following instructions: Before the processing component calculates the first plurality of sum results, loading the first multiplicand block and the first modulus block from the at least one storage device into the plurality of cache memories; before the processing component calculates the second plurality of summation results and the plurality of delayed summation results, loading the second multiplicand block and the second modulus block from the at least one storage device into the plurality of cache memories; before the processing component calculates the first plurality of updated sum results, the plurality of delayed sum results, the first multiplicand block, and the first modulus block from the at least one storage device to the plurality of cache memories; as well as Before the processing component calculates the second plurality of updated sum results and the plurality of updated delayed sum results, the second plurality of sum results, the second multiplicand block, and the second modulus block are loaded from the at least one storage device into the plurality of cache memories.
12. The modular operation device according to claim 11, wherein: The loading and storing component is further configured to execute the following instructions: loading the first multiplicand block and the first modulus block into a first cache memory of the plurality of cache memories; loading the second multiplicand block and the second modulus block into a second cache memory of the plurality of cache memories when the processing component calculates the first plurality of sum results; loading the first plurality of summation results, the plurality of delayed summation results, the first multiplicand block, and the first modulus block into the first cache memory when the processing component calculates the second plurality of summation results and the plurality of delayed summation results; as well as When the processing component calculates the first plurality of updated sum results, the second plurality of sum results, the second multiplicand block, and the second modulus block are loaded into the second cache memory.
13. The modular operation device according to claim 9, wherein: The processing component is further configured to execute the following instructions: Calculating a first plurality of delayed-carry results according to a most significant word of the first multiplicand block and the first multiplier block; and A second plurality of delayed-carry results is calculated based on the most significant word and the second multiplier block.
14. The modular operation device according to claim 13, wherein: The simplified component is also configured to execute the following instructions: The result remainder is calculated according to the first plurality of updated sum results, the second plurality of updated sum results, the plurality of updated delayed sum results, the first plurality of delayed carry results, and the second plurality of delayed carry results.
15. The modular operation device according to claim 9, wherein: The processing component calculates the plurality of delayed sum results based on a least significant word of the second multiplicand block and the first multiplier block, and the processing component calculates the plurality of updated delayed sum results based on the least significant word and the first multiplier block.
16. The modular operation device according to claim 9, wherein: The processing component is further configured to execute the following instructions: calculating a first delayed-carry result according to a most significant word of the first multiplicand block and a most significant word of the first multiplier block; determining the first delayed-carry result as a most significant result of the first plurality of summation results; calculating a second delayed-carry result according to a most significant word of the second multiplicand block and the most significant word of the first multiplier block; as well as The second delayed-carry result is determined to be a most significant result of the second plurality of summation results.
17. The modular operation device according to claim 9, wherein: The instructions for calculating the first plurality of updated sum results include: calculating a plurality of temporary summation results according to the first plurality of summation results, the plurality of delayed summation results, the first multiplicand block, and a first word of the second multiplier block; as well as The first plurality of updated sum results are calculated according to the plurality of temporary sum results, the first multiplicand block, and a second word of the second multiplier block.
18. The modular operation device according to claim 17, wherein: The first word of the second multiplier block is a least significant word of the second multiplier block.
19. The modular operation device according to claim 9, wherein: The processing component is further configured to execute the following instructions: An additional quotient is calculated based on a least significant result of the first plurality of sum results, a least significant result of the plurality of delayed sum results, and an inverse word in an iteration corresponding to a least significant word of the second multiplier block, wherein the inverse word is a remainder of a reciprocal of a least significant word of the plurality of modulus blocks divided by a base.
Citation Information
Patent Citations
Apparatus and method for calculating an lnteger quotient
TW200400442A
Apparatus And Method For Calculating A Result Of A Modular Multiplication
TW200403584A
Device and method for performing multiple modulus conversion using inverse modulus multiplication
TW200416564A
Combined polynomial and natural multiplier architecture
TW200504583A