Low-delay two-dimensional differential addition chain reasoning method

By constructing a low-latency two-dimensional differential addition chain inference method with parallel windows and five flag signals, the problem of high latency and difficulty in balancing complexity in the traditional two-dimensional differential addition chain inference process is solved, and a low-latency and low-complexity hardware inference circuit design is realized.

CN120669954APending Publication Date: 2025-09-19BEIJING INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510774417.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The inference process of the traditional two-dimensional differential addition chain has the problem of high latency and complexity that is difficult to balance. Especially in hardware inference circuits, directly adopting a cross-clock domain strategy will increase the design complexity, while not adopting a cross-clock domain strategy will increase the calculation latency.

Method used

A low-latency two-dimensional differential addition chain inference method based on five flag signals is adopted. By constructing a parallel window to process data of multiple iterative rounds in parallel, and controlling data selection through the first to fourth flag signals and auxiliary flag signals, the number of clock cycles is reduced by combining parallel computing and interleaved bit technology.

Benefits of technology

It realizes low-latency and low-complexity two-dimensional differential addition chain reasoning, reduces the number of clock cycles in the reasoning process, and improves the efficiency of the hardware reasoning circuit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669954A_ABST
    Figure CN120669954A_ABST
Patent Text Reader

Abstract

The invention relates to a low-delay two-dimensional differential addition chain reasoning method, and relates to the technical field of digital integrated circuit design, and the method comprises the steps: constructing a parallel window to process data of a plurality of iteration rounds at the same time; and constructing and controlling data gating of each round of iteration process of the two-dimensional differential addition chain through a first mark signal, a second mark signal, a third mark signal, a fourth mark signal and an auxiliary mark signal so as to carry out multiple iterations on the two input array data and obtain a two-dimensional differential addition chain reasoning result. The invention provides a low-delay two-dimensional differential addition chain reasoning algorithm based on five mark signals, and solves the problem that a lot of clock cycles are overheads in the two-dimensional differential addition chain reasoning process by introducing a parallelization calculation reasoning window.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital integrated circuit design, and in particular to a low-latency two-dimensional differential addition chain reasoning method. Background Art

[0002] Traditionally, the inference process for missing elements in a two-dimensional differential addition chain is performed row by row. However, this strategy is not fully compatible with hardware inference circuits. The root cause of this problem is that the accumulation direction of the two-dimensional differential addition chain is completely opposite to the inference direction of the two-dimensional differential addition chain. Therefore, before starting the actual calculation process of the two-dimensional differential addition chain for regional client energy integrated demand-side management, the entire two-dimensional differential addition chain inference must be completed.

[0003] However, on the one hand, since the logic of each row of reasoning is relatively simple and far shorter than the critical path length of the modular multiplier in the two-dimensional differential addition chain accumulation process, in order to speed up the reasoning, the direct engineering strategy is to adopt a cross-clock domain approach to provide a dedicated high-speed clock for the reasoning circuit. However, this will inevitably increase the complexity of the design of the entire regional client energy integrated demand side management architecture; on the other hand, if the cross-clock domain strategy is not adopted, the clock cycle overhead of the entire reasoning process will be equivalent to the actual calculation process of the regional client energy integrated demand side management, which greatly increases the delay of the calculation process. Summary of the Invention

[0004] In view of this, the present invention aims to propose a low-latency two-dimensional differential addition chain reasoning method to solve the problem in the prior art that low latency and low complexity cannot be taken into account at the same time.

[0005] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0006] The present invention provides a low-latency two-dimensional differential addition chain reasoning method, comprising:

[0007] A parallel window is constructed to simultaneously process data from multiple iteration rounds; and data gating for each iteration process of a two-dimensional differential addition chain is constructed and controlled by a first flag signal, a second flag signal, a third flag signal, a fourth flag signal, and an auxiliary flag signal, so as to perform multiple iterations on the two input array data and obtain an inference result of the two-dimensional differential addition chain.

[0008] Furthermore, the first flag signal is used to indicate the data source of the doubling operation in the i-th iteration round;

[0009] The second flag signal is used to indicate the source of input data for the point addition operation associated with the second element in the addition chain element set of the i-th iteration round in the i-th iteration round;

[0010] The third flag signal is used to indicate that the calculation result is the difference between the horizontal coordinates of the two input values ​​of the point addition operation of the 0th element in the addition chain element set in the i-th iteration round;

[0011] The fourth flag signal is used to indicate that the calculation result is the difference between the horizontal coordinates of the two input values ​​of the point addition operation of the first element in the addition chain element set in the i-th iteration round;

[0012] The auxiliary flag signal is used to assist in the calculation of the second flag signal and the fourth flag signal.

[0013] Furthermore, the auxiliary flag signal is also used to indicate that the two initial values ​​of the second-bit element in the data set input in the 0th iteration round are 0, 1 or 1, 0.

[0014] Furthermore, after constructing the parallel window, the method further includes:

[0015] Obtain an interleaving bit sequence number, where the interleaving bit sequence number is configured as w-1, 2w-2, 3w-3...xw-x; where x is the total number of windows and w is the window width;

[0016] The interleaving bit is set at the first and last positions of each window, and the first interleaving bit of the current window is inherited from the last interleaving bit of the previous window.

[0017] Furthermore, after performing multiple iterations on the two input array data and obtaining a two-dimensional differential addition chain inference result, the method further includes:

[0018] The 0th element in the addition chain element set of the i-th iteration round is configured as the sum of the 0th element in the addition chain element set of the i-1th iteration round and the 1st element in the addition chain element set of the i-1th iteration round;

[0019] Selecting the addition chain element set of the i-1th iteration round based on the first flag signal to obtain the first element of the addition chain element set of the i-th iteration round;

[0020] Based on the second flag signal, the addition chain element set of the i-1th iteration round is selected to obtain the second element in the addition chain element set of the i-th iteration round and the second flag signal respectively;

[0021] The accumulated result of the two-dimensional difference addition chain is obtained.

[0022] Furthermore, before configuring the 0th element in the addition chain element set of the i-th iteration round as the sum of the 0th element in the addition chain element set of the i-1th iteration round and the 1st element in the addition chain element set of the i-1th iteration round, the method further includes:

[0023] Initialize the 0th element in the addition chain element set of the 0th iteration round to 1,1;

[0024] Initialize the first element of the addition chain element set in the 0th iteration round to 0,0;

[0025] Based on the auxiliary flag signal, the second element of the addition chain element of the 0th iteration round is obtained.

[0026] Furthermore, the original ITA addition chain is split into an I-type addition chain based on the ITA addition chain and an R-type addition chain based on modular square root operation.

[0027] According to the accumulated results of the two-dimensional differential addition chain, the addition chain is iterated in parallel through the I-type addition chain and the R-type addition chain to obtain the I-type addition chain iteration result and the R-type addition chain iteration result, and the modular inverse operation result is calculated.

[0028] Furthermore, the calculation to obtain the modular inverse operation result includes:

[0029] According to the formula

[0030]

[0031] Calculate the modular inverse operation result α -1 ,in, is the iterative result of type I addition chain, is the iterative result of the R-type addition chain, k is the preset scalar value; m is the length of the original ITA addition chain; α is the accumulated result of the two-dimensional difference addition chain.

[0032] Compared with the prior art, the present invention has the following advantages:

[0033] The present invention proposes a low-latency two-dimensional differential addition chain inference algorithm based on five flag signals, and solves the problem of a large number of overhead clock cycles in the two-dimensional differential addition chain inference process by introducing an inference window for parallel computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are provided to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:

[0035] Figure 1 This is a flow chart of the low-latency two-dimensional differential addition chain inference method of the present invention;

[0036] Figure 2 Schematic diagram of the low-latency two-dimensional differential addition chain reasoning process based on interleaved bits of the present invention;

[0037] Figure 3Schematic diagram of the structure of the low-latency multifunctional dual-point multiplication scheduling algorithm based on dual multipliers of the present invention;

[0038] Figure 4 This is a diagram showing the architecture of the low-latency multifunctional dual-point multiplication arithmetic logic unit of the present invention;

[0039] Figure 5 This is a diagram of the low-latency differential addition chain inference unit architecture of the present invention;

[0040] Figure 6 This is a diagram of the low-latency multi-function point multiplication architecture deployed on FPGA in the present invention. DETAILED DESCRIPTION

[0041] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0042] In the description of the present invention, it should be noted that the terms "upper," "lower," "inner," and "back" and other terms indicating orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0043] Furthermore, in the description of the present invention, unless otherwise expressly defined, the terms "mounted," "connected," "connect," and "connector" should be interpreted broadly. For example, these terms may refer to fixed, removable, or integral connections; mechanical or electrical connections; direct or indirect connections through an intermediary; or internal communication between two components. Those skilled in the art will appreciate the specific meanings of these terms in the present invention based on the specific circumstances.

[0044] The following will refer to the attached Figures 1 to 6 The present invention is described in detail with reference to the embodiments.

[0045] In general, the present invention provides a low-latency two-dimensional differential addition chain reasoning method, including:

[0046] Step S1: construct a parallel window to process data of multiple iteration rounds simultaneously;

[0047] Step S2: construct and control the data selection of each round of iteration process of the two-dimensional differential addition chain through the first flag signal, the second flag signal, the third flag signal, the fourth flag signal, and the auxiliary flag signal, so as to iterate the two input array data multiple times and obtain the two-dimensional differential addition chain inference result.

[0048] In one possible implementation, the first flag signal (PD i ), used to indicate the data source of the doubling operation in the i-th iteration round;

[0049] The second flag signal (PA i ), used to indicate the second element in the addition chain element set in the i-th iteration round The source of input data for the relevant point addition operation;

[0050] The third flag signal Used to indicate that the calculation result is the 0th element in the addition chain element set of the i-th iteration round The difference between the horizontal coordinates of the two input values ​​of the point addition operation;

[0051] Fourth flag signal Used to indicate that the calculation result is the first element in the addition chain element set of the i-th iteration round The difference between the horizontal coordinates of the two input values ​​of the point addition operation;

[0052] Auxiliary flag signal (F i ), used to assist the second flag signal (PA i ) and the fourth flag signal Calculation.

[0053] This paper proposes a low-latency two-dimensional differential addition chain inference algorithm based on five flag signals and introduces staggered bits to construct an internal parallelized inference window. This solves the problem of excessive clock cycle overhead in the two-dimensional differential addition chain inference process.

[0054] In a possible implementation, the auxiliary flag signal is further used to indicate that the two initial values ​​of the second-bit element in the data set input in the 0th iteration round are 0, 1 or 1, 0.

[0055] In one possible implementation, after building the parallel window, the method further includes:

[0056] Get the interleaving bit sequence number, which is configured as w-1, 2w-2, 3w-3...xw-x; where x is the total number of windows and w is the window width;

[0057] The interlacing bit is set at the first and last positions of each window and the first interlacing bit of the current window is inherited from the last interlacing bit of the previous window.

[0058] The present invention proposes a low-latency two-dimensional differential addition chain inference algorithm based on five flag signals, and solves the problem of a large number of overhead clock cycles in the two-dimensional differential addition chain inference process by introducing interleaved bits to construct an inference window for internal parallel calculation.

[0059] In one possible implementation, this embodiment further provides a low-latency multifunctional dot multiplication algorithm specifically as follows:

[0060] After performing multiple iterations on the two input array data and obtaining a two-dimensional differential addition chain inference result, the method further includes:

[0061] The 0th element in the addition chain element set of the i-th iteration round is configured as the 0th element in the addition chain element set of the i-1-th iteration round. and the first element in the addition chain element set of the i-1th iteration round of and;

[0062] The addition chain element set C of the i-1th iteration round based on the first flag signal i-1 Perform gating to obtain the first element in the addition chain element set of the i-th iteration round;

[0063] The addition chain element set C of the i-1th iteration round based on the second flag signal i-1 The second element in the addition chain element set of the i-th iteration round is selected and the second flag signal is

[0064] The accumulated result of the two-dimensional difference addition chain is obtained.

[0065] In one possible implementation, before configuring the 0th element in the addition chain element set of the i-th iteration round as the sum of the 0th element in the addition chain element set of the i-1th iteration round and the 1st element in the addition chain element set of the i-1th iteration round, the method further includes:

[0066] Initialize the 0th element in the addition chain element set of the 0th iteration round to 1,1;

[0067] Initialize the first element of the addition chain element set in the 0th iteration round to 0,0;

[0068] Based on the auxiliary flag signal, the second element of the addition chain element of the 0th iteration round is obtained.

[0069] In one possible implementation, this embodiment provides a low-latency fully parallel modular inversion algorithm based on dual multipliers, specifically:

[0070] The original Itoh-Tsujii algorithm (ITA) addition chain is split into an I-type addition chain based on the ITA addition chain and an R-type addition chain based on modular square root operation.

[0071] According to the accumulated results of the two-dimensional differential addition chain, the addition chain is iterated in parallel through the I-type addition chain and the R-type addition chain to obtain the I-type addition chain iteration result and the R-type addition chain iteration result, and the modular inverse operation result is calculated.

[0072] In one possible implementation, calculating the modular inverse operation result includes:

[0073] According to the formula

[0074]

[0075] Calculate the modular inverse operation result α -1 ,in, is the iterative result of type I addition chain, is the iterative result of the R-type addition chain, k is the preset scalar value; m is the length of the original ITA addition chain; α is the accumulated result of the two-dimensional difference addition chain.

[0076] This embodiment proposes a low-latency, fully parallel modular inversion algorithm based on dual multipliers. By fully utilizing two parallel, independent modular multipliers, the length of the addition chain during the modular inversion process is halved, significantly reducing the number of iterative modular squaring operations required in the final iteration. A low-latency dual-point multiplication algorithm compatible with both regional client energy integrated demand-side management operations and effective current source model operations is proposed. An operator scheduling scheme, arithmetic logic unit hardware architecture, and a supporting register-level scheduling scheme are also presented.

[0077] More specifically, this first embodiment provides a time-delayed two-dimensional differential addition chain inference algorithm:

[0078] The low-latency TDDAC inference algorithm proposed in this design does not use a cross-clock domain approach, but instead shares the same clock domain with the TDDAC accumulation process. By constructing parallel windows within the TDDAC inference process, this design resolves the mismatch between the critical path lengths of the TDDAC inference and TDDAC accumulation processes, reducing the clock cycle overhead of the TDDAC inference process and, consequently, the latency of the TDDAC inference process.

[0079] First, based on the original TDDAC research, in order to realize the pure hardware TDDAC accumulation logic, this design introduces five additional flag signals F i , PD i , P.A. i , and This design uses these five flag signals to control the data gating of each iteration during the TDDAC accumulation process. Therefore, in fact, the inference process of TDDAC is to calculate the five flag signals required for each iteration during the TDDAC accumulation process. The functions of these five flag signals are as follows:

[0080] PD i Used to indicate the data source for the doubling operation in the current round. Both scalars in are even, so The result of the doubling operation must always be stored, but at the same time there are three possibilities for the source of the doubling operation, namely and PD i The value range is 0, 1, 2.

[0081] PA i Used to indicate the current round Another source of input data for the related point addition operation. and The parity of is always (o,o) and (e,e), so in each round of iterative calculation All from the previous round and For this reason, the dot addition operation does not require an additional flag signal to indicate the source of the input data. The related point addition operation, due to The parity of may be (e,o) or (o,e), so the input of the other point addition operation may be (o,o) or (e,e). When the input of the other point addition operation is (e,e), The parity and Keep consistent; otherwise, its parity is Different.

[0082] Used to indicate that the calculation result is The difference x between the two input values ​​of the point addition operation diff0 .because The parity of is always (o,o) and is fixed as one input of this point addition operation, so when the other input is or When , according to the parity, the difference between the two point addition input values ​​exists in two cases: P+Q and PQ. Then the difference between the two is (P+Q), if Then the difference between the two is (PQ).

[0083] Used to indicate that the calculation result is The difference x between the two input values ​​of the point addition operation diff1 .because The parity of is one odd and one even, and the input of the other point addition operation is (o,o) or (e,e). According to the parity, the difference between the two point addition operation input values ​​exists in two cases: P and Q. Then the difference between the two is P, if The difference between the two is Q.

[0084] F i Used to assist the sign signal. Mainly used to assist PA i and The calculation of, for the first round of the accumulation process (first row), is used to indicate The two initial values ​​of are 0, 1 or 1, 0. At the same time, the difference required for the right point addition operation can also be determined, that is, value.

[0085] Based on the above analysis and the meaning of the five flag signals, the generation process of the five flag signals can be simplified to two loops that can be fully parallelized by the hardware architecture. The low-latency two-dimensional differential addition chain inference algorithm (LLW-TDDAC) proposed in this design is shown in Table 1.

[0086] As shown in the comments in Table 1, the reasoning process of the two-dimensional differential addition chain can be simplified into two large loops (lines 4 to 17 and 20 to 25 of Table 1). For the first loop, the two sets of data strobes (switch statements) contained in it have no data dependencies between each other and can be implemented as a purely parallel hardware computing architecture; for the second loop, only one set of data strobes is included, which can also be implemented as a purely parallel hardware computing architecture. However, due to the PA i and The calculation depends on F i , so there is a data dependency with the second set of data strobes in the first cycle. But considering k flag 、l flag The combinational logic path for the second set of data selections (rows 13 to 16 of Table 1) consists only of a bitwise XOR operation and a two-bit data comparison. Therefore, even if there are certain data dependencies, the combinational logic depth is still relatively low, and pure combinational logic can be used for parallel implementation.

[0087]

[0088] Note that in the algorithm in Table 1, only Depends on the two scalar values ​​​​of the bit k scanned in the current loop t(w-1)+i and k t(w-1)+i , and the other four flag signals all depend on the current bit and the previous bit of the two scalar values ​​at the same time. Therefore, this design introduces the concept of interleaved bits to adapt to the algorithm in Table 1, such as Figure 2The interleaved bits are numbered w-1, 2w-2, 3w-3, and so on, occupying the first and last bits of each window. The value of the first interleaved bit in the i-th window is inherited from the last interleaved bit in the (i-1)-th window. In the i-th window, this value is only used as an input value and is not modified by the inference algorithm of that window.

[0089] The second embodiment of the present invention provides a low-latency multifunctional dot multiplication algorithm:

[0090] Based on the low-latency two-dimensional differential addition chain inference algorithm proposed above, this paper further proposes a low-latency multifunctional point multiplication algorithm as shown in Table 2. This algorithm can be regarded as a TDDAC accumulation process based on the TDDAC inference results. The algorithm in Table 2 uses (1,1) pairs. Initialize with (0,0) Initialize. Then, the auxiliary flag signal F in the TDDAC inference algorithm is used. m Then it enters the iterative phase of the TDDAC accumulation process, where Fixed to and and Based on PD i and PA i It is important to note that the input data is selected in the LLW-TDDAC algorithm. and Separate instructions and The difference of the operands in the point addition operation is generated, but it is not reflected in the algorithm of Table 2. It is only used as a control signal in the data selection process of the hardware circuit.

[0091]

[0092] The following discusses the versatility of the algorithm in Table 2 and draws an analogy with the Montgomery ladder point multiplication algorithm. The Montgomery ladder point multiplication algorithm performs one point addition operation and one doubling operation per iteration. In comparison, the low-latency multifunctional point multiplication algorithm requires an additional point addition operation per iteration. At the same time, considering the data dependency analysis of the point addition operation and the doubling operation, the one point addition operation in each iteration of the Montgomery ladder point multiplication algorithm can be fixed as input and output, while only the doubling operation is retained with the flexibility of the input end. Therefore, in order to use the algorithm in Table 2 to be compatible with the Montgomery ladder point multiplication algorithm, step 7 of the algorithm in Table 2 can be retained, while steps 8 and 9 can be used to implement the data gating function of the doubling operation. In summary, the algorithm in Table 2, as an algorithm mainly for regional client energy comprehensive demand-side management (ECDSM) calculations, has the ability to be compatible with ECDSM calculations.

[0093] The third embodiment of this embodiment proposes a low-latency fully parallel modular inversion algorithm based on dual multipliers: based on the low-latency (LLW) two-dimensional differential addition chain (TDDAC) inference algorithm and the low-latency multifunctional ECDSM algorithm, the low-latency inference and accumulation of TDDAC can be realized. At the same time, the accumulation process of TDDAC is compatible with the calculation of ECSM. However, the accumulation process of TDDAC is still based on the projected coordinate system. Therefore, after the accumulation process is completed, a modular inversion operation is also required to complete the coordinate system conversion. Therefore, it is necessary to design a modular inversion algorithm compatible with the ECDSM algorithm. Aiming at the hardware architecture of dual multipliers, this design proposes a low-latency fully parallel modular inversion algorithm as shown in Algorithm 13. By introducing the modular square root operation to realize the fully parallel working mode of the dual multipliers, the traditional ITA addition chain with an original length of (m-1) is halved, so that the length of the new addition chain is only Significantly reduce the highest power in the iterative calculation process of modular inverse operations, thereby reducing the number of cycles required for iteration, and ultimately achieving the design goal of low latency.

[0094]

[0095] The necessary mathematical derivations of the algorithm in Table 3 are as follows. In order to realize the fully parallel modular inverse operation, the basic building blocks of the original ITA addition chain are retained. Based on this, a new building element designed using modular square root operation is introduced. in And α∈GF(2 m ). As in Chapter 3, There is a special mathematical relationship:

[0096]

[0097] analogy Similar mathematical relationships can be found.

[0098]

[0099] It is easy to find that there is an equivalence relationship between the two:

[0100] R k (α)=[I m-k (α)] -1

[0101] At the same time, the calculation process of traditional ITA makes full use of the relationship α -1 =[I m-1 (α)] 2 , considering R k (α) and I k (α) symmetric characteristics, it is also easy to find that the result of the modular inverse operation can be obtained only from R k (α) is expressed as:

[0102] α -1 =R m-1 (α)

[0103] Inspired by this, the result of the modular inverse operation can be represented by both, while taking into account R k (α) and I k (α) Symmetric property, the original addition chain is divided equally, and the final joint representation is:

[0104]

[0105] This means that the original ITA addition chain can be split and two new addition chains of type I and type R can be constructed, and the two new addition chains can be iteratively calculated at the same time. In the iterative calculation, since there is no data dependency between the two new addition chains of type I and type R, the two addition chains can be executed completely independently. At the same time, considering the symmetry of their calculation modes, the iterative processes of the two addition chains of type I and type R can be synchronized, thus forming a fully parallel modular inversion algorithm (the algorithm in Table 3). Using GF(2 163 ) as an example, the calculation steps of the low-latency fully parallel modular inversion algorithm based on dual multipliers are shown in Table 4.

[0106] Table 4

[0107]

[0108] Compared with the calculation method of the traditional ITA addition chain, the total number of modular multiplications of the low-latency fully parallel modular inversion algorithm based on dual multipliers proposed in this design has doubled. However, since the hardware architecture proposed in this design has a dual multiplier architecture, and no matter whether the multifunctional ECDSM architecture executes ECDSM or ECSM, only one modular inversion operation is required in each round of iterative calculation, the two multipliers can execute the low-latency fully parallel modular inversion algorithm in parallel and synchronously. Therefore, the number of clock cycles introduced by modular multiplication in the modular inversion part does not increase. At the same time, it is noted that in the traditional ITA addition chain, as the number of calculations increases, the number of modular exponentiation operations also rises rapidly. Similarly, in GF(2 163 ), the highest modular exponentiation operation required in the traditional ITA addition chain calculation process is In the proposed low-latency fully parallel modular inversion algorithm based on dual multipliers, this data is approximately halved to only In existing ECDSM / ECSM designs, the main hardware resource overhead and the primary optimization target are the iterative calculations of ECDSM / ECSM rather than the modular inverse operation. The cost of introducing a modular high-order power module specifically for modular inverse operation is too high, so modular high-order power operation can only be implemented through multiple iterations of modular square or modular fourth power modules. In this case, compared to This consumes a significant number of clock cycles to iterate and accumulate to the target value, increasing latency. However, the low-latency, fully parallel modular inversion algorithm proposed in this chapter significantly reduces the exponent of the modular exponentiation operation in the final step, significantly reducing the clock cycle overhead and ultimately achieving latency reduction.

[0109] The fourth embodiment of this embodiment proposes a hardware scheduling solution:

[0110] The original intention of this design is to develop a low-latency ECDSM architecture. Although the data dependencies of the algorithm in Table 2 can be easily satisfied based on a single multiplier, the clock cycle overhead of a single multiplier is high, which will cause the clock cycle overhead saved by the LLW-TDDAC inference algorithm to be consumed by the TDDAC accumulation process. Therefore, this chapter proposes a low-latency multifunctional ECDSM architecture based on dual multipliers. By properly scheduling the finite field operation unit, a low-latency TDDAC accumulation process can be achieved while satisfying the data dependencies. The scheduling scheme is as follows: Figure 3 As shown in Figure 3, the data path generated by this scheduling scheme is also compatible with the low-latency, fully parallel modular inversion algorithm proposed above (the algorithm in Table 3). In summary, this architecture achieves the design goals of low power consumption and multi-functionality at the three levels of TDDAC inference, TDDAC accumulation, and modular inversion.

[0111] like Figure 3As shown in the figure, the low-latency multifunctional point multiplication scheduling scheme uses gray and white background colors to divide different TDDAC accumulation rounds. Each round of iteration requires two point addition operations and one point doubling operation. Due to the introduction of DAC, the Y coordinate in the point addition and point doubling operations can be omitted, so the computational overhead of one point addition operation is 4 modular multiplication operations, while the computational overhead of one point doubling operation is only 2 modular multiplication operations. Based on two modular multipliers, a total of 10 modular multiplication operations in one iteration process are prioritized, and then other finite field operations are reasonably planned, which can be formed as follows Figure 3 The TDDAC accumulation process scheduling scheme with a single iteration cost of 5 clock cycles is shown. The dotted line in the figure indicates that the data needs to be cached in the register to accommodate the data dependency; the red symbol represents the calculation result generated by the current round (white background), which is directly passed to the next round (gray background). The figure introduces two subscript variables i and k, where i is controlled by PD i , which is used as the input value for the gate doubling operation. According to Algorithm 11, the possible values ​​of i are 1, 2, and 3. Similarly, k is controlled by PA i , used to select an input value of the right-side dot addition operation in the TDDAC accumulation process. The possible values ​​of i are 1 and 2. In the figure, the 4th and 5th clock cycles of the current iteration round need to indicate the affine coordinate value x corresponding to the difference between the two input data of the dot addition operation. diff0 and x diff1 , the values ​​are respectively determined by the flag signal and Control, when needed x diff0 and x diff1 The data preparation is completed in the first clock cycle.

[0112] A fifth embodiment of the present invention provides an arithmetic logic unit architecture:

[0113] Based on the low-latency multifunctional point multiplication scheduling scheme, this embodiment proposes a register-level hardware scheduling scheme as shown in Table 5, and the corresponding arithmetic logic unit architecture is as follows: Figure 4 As shown in the figure, this architecture comprises a modular multiplier-adder and a modular multiplier. The multiplication portion of each unit utilizes a hybrid KOM architecture, with a pipeline stage inserted within the multiplication unit to optimize timing. Furthermore, the architecture includes four modular squarer units, which are paired to form two modular exponentiation units. During ECDSM and ECSM calculations, the modular exponentiation units are used to quickly calculate modular fourth powers in point addition and doubling operations.

[0114] Table 5

[0115]

[0116] SQR_R: modular squarer result, QUA_R: modular quarticizer result, MUL_R: modular multiplier result, SUM_R: modular adder result, R: register storage value.

[0117] As for the modular inverse operation, since this design adopts a low-latency fully parallel modular inverse algorithm, a modular square root unit needs to be introduced at the same time. On finite fields of different sizes, the Hamming weight W(S N ) has a similar trend with N, that is, in the process of modular square accumulation (from left to right), W(S N ) rises at a slower rate, and the corresponding modular exponentiation circuit complexity increases less; but in the modular square accumulation process (from right to left), W(S N ) rises very fast, and the complexity of the modular circuit quickly increases to the highest level. N ) have an imbalanced characteristic on the left and right sides, whereas the accumulation progress of the I-type addition chain and the R-type addition chain in the low-latency fully parallel modular inversion algorithm is completely synchronized. The main clock cycle overhead of the low-latency multifunctional point multiplication architecture comes from the ECDSM / ECSM iteration process, and the main computational load of the iteration process comes from the modular multiplier. Therefore, the critical path must be designed in the modular multiplier rather than the modular squaring circuit. Guided by this idea, without introducing a critical path to the modular squaring unit, modular quartic circuits can be used on finite fields of various sizes. This design uses two modular quadratic units (SQRT) in a similar manner to the modular squarer unit (SQR). Both the modular squarer unit and the modular squarer unit can be cascaded to quickly calculate the high-order power cases in the algorithm shown in Table 3, further saving clock cycles in the modular inversion operation.

[0118] In addition, this architecture also includes an additional modular addition unit and six registers, five of which can control the data stored in them through the scheduling scheme based on Table 6.2 to store and cache intermediate variables in the ECDSM and ECSM calculation processes, ultimately satisfying the data dependencies in the scheduling scheme. In addition, another register prepares x in advance from four possible differential cases based on the inference results of TDDAC. diff0 and x diff1 to match the accumulation process of TDDAC.

[0119] The sixth embodiment of the present invention provides a low-latency differential addition chain inference architecture:

[0120] Based on the low-latency window two-dimensional differential addition chain inference algorithm proposed in the embodiment (the algorithm in Table 1), this paper proposes the LLW-TDDAC low-latency differential addition chain inference architecture as shown in the following example. Figure 5The architecture is based on a window divided by staggered bits, and simultaneously calculates the flag signal F required for all TDDAC accumulation processes within the window. i , PD i , P.A. i , and The logic for generating the auxiliary flag signal is located in Figure 5 Below, it generates F based on the input scalar sequence k and l i Directly used to calculate the PA in the current window i and The flag signal PD i and The generation logic is also based on the scalar sequences k and l. Since the low-latency differential addition chain inference architecture is a pure combinational logic circuit, the same considerations for the combinational logic depth of the modular squaring circuit apply to the design of the low-latency differential addition chain inference architecture. Furthermore, the low-latency differential addition chain inference architecture requires additional logic resources. As the window size increases, the combinational logic depth and resource usage of the inference architecture increase. Therefore, a comprehensive consideration of the ATP indicators of the overall architecture is required to ultimately determine the window size of the low-latency differential addition chain inference architecture.

[0121] The sixth embodiment of the present invention provides an overall FPGA architecture. The low-latency multi-function point multiplication architecture proposed in this embodiment is implemented in hardware on the FPGA. The overall architecture block diagram of this design is shown in FIG. Figure 6 As shown in the figure, the FPGA's overall architecture primarily consists of a phase-locked loop (PLL) module built into the FPGA, providing a global clock; a low-latency, multifunctional dot-multiplication arithmetic logic unit (ALU), proposed in this chapter, which uses dual-channel scalar inputs to perform computational tasks; and an instruction memory (IM) for storing the ALU's operating instructions. The ALU includes the LLW-TDDAC low-latency differential addition chain inference architecture proposed in this chapter, two modular multipliers, four modular squarers, two modular square roots, a register table, a series of multiplexers, and a small state machine for controlling the operating mode. The addresses of all multiplexers together constitute the ALU's microinstructions, which have a bit width of 29 bits. To simplify the ALU logic, only the small state machine for controlling the operating mode remains within the ALU. The ALU specifically performs the control required for ECDSM and ECSM inference, stored in the IIM as microinstructions. A large number of microinstructions can be retrieved one by one to control the data selection of the multiplexers, thereby controlling the ALU's operation. In addition, the LLW-TDDAC low-latency differential addition chain inference architecture is based on dual-path input scalars to infer the required flag signal F in the ECDSM process. i , PD i , P.A. i , and The flag signal is stored in the instruction memory as a microinstruction. When the arithmetic logic unit is calculating the ECSM in multi-function mode, the valid signal of one scalar is invalid. At this time, the LLW-TDDAC low-latency differential addition chain inference architecture is bypassed, and the only scalar is directly stored in the instruction memory as a microinstruction.

[0122] In order to solve the problem of a large number of overhead clock cycles in the TDDAC inference process, the present invention proposes a low-latency TDDAC inference algorithm based on five flag signals, and constructs an inference window for internal parallel calculation by introducing interleaving bits.

[0123] A low-latency fully parallel modular inversion algorithm based on dual multipliers is proposed. By fully utilizing two parallel independent modular multipliers, the length of the addition chain in the modular inversion process is halved, thereby significantly reducing the number of iterative modular squaring operations required in the last round of iterations.

[0124] A low-latency double-point multiplication algorithm that is compatible with both ECDSM and ECSM operations is proposed, and an operator scheduling scheme, arithmetic logic unit hardware architecture, and a corresponding register-level scheduling scheme are given.

[0125] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A low-latency two-dimensional differential addition chain inference method, characterized in that: include: Construct parallel windows to process data from multiple iterations simultaneously; Construct and control the data selection of each round of iteration process of the two-dimensional differential addition chain through the first flag signal, the second flag signal, the third flag signal, the fourth flag signal, and the auxiliary flag signal, so as to iterate the two input array data multiple times and obtain the two-dimensional differential addition chain inference result.

2. The low-latency two-dimensional differential addition chain inference method according to claim 1, characterized in that: The first flag signal is used to indicate the data source of the doubling operation in the i-th iteration round; The second flag signal is used to indicate the source of input data for the point addition operation associated with the second element in the addition chain element set of the i-th iteration round in the i-th iteration round; The third flag signal is used to indicate that the calculation result is the difference between the horizontal coordinates of the two input values ​​of the point addition operation of the 0th element in the addition chain element set in the i-th iteration round; The fourth flag signal is used to indicate that the calculation result is the difference between the horizontal coordinates of the two input values ​​of the point addition operation of the first element in the addition chain element set in the i-th iteration round; The auxiliary flag signal is used to assist in the calculation of the second flag signal and the fourth flag signal.

3. The low-latency two-dimensional differential addition chain inference method according to claim 1, characterized in that: The auxiliary flag signal is also used to indicate that the two initial values ​​of the second-bit element in the data set input in the 0th iteration round are 0, 1 or 1, 0.

4. The low-latency two-dimensional differential addition chain inference method according to claim 1, characterized in that: After constructing the parallel window, the method further includes: Obtain an interleaving bit sequence number, where the interleaving bit sequence number is configured as w-1, 2w-2, 3w-3...xw-x; where x is the total number of windows and w is the window width; The interleaving bit is set at the first and last positions of each window, and the first interleaving bit of the current window is inherited from the last interleaving bit of the previous window.

5. The low-latency two-dimensional differential addition chain inference method according to claim 1, characterized in that: After performing multiple iterations on the two input array data and obtaining a two-dimensional differential addition chain inference result, the method further includes: The 0th element in the addition chain element set of the i-th iteration round is configured as the sum of the 0th element in the addition chain element set of the i-1th iteration round and the 1st element in the addition chain element set of the i-1th iteration round; Selecting the addition chain element set of the i-1th iteration round based on the first flag signal to obtain the first element of the addition chain element set of the i-th iteration round; Based on the second flag signal, the addition chain element set of the i-1th iteration round is selected to obtain the second element in the addition chain element set of the i-th iteration round and the second flag signal respectively; The accumulated result of the two-dimensional difference addition chain is obtained.

6. The low-latency two-dimensional differential addition chain inference method according to claim 5, characterized in that: Before configuring the 0th element in the addition chain element set of the i-th iteration round as the sum of the 0th element in the addition chain element set of the i-1th iteration round and the 1st element in the addition chain element set of the i-1th iteration round, the method further includes: Initialize the 0th element in the addition chain element set of the 0th iteration round to 1,1; Initialize the first element of the addition chain element set in the 0th iteration round to 0,0; Based on the auxiliary flag signal, the second element of the addition chain element of the 0th iteration round is obtained.

7. The low-latency two-dimensional differential addition chain inference method according to claim 5, characterized in that: The original ITA addition chain is split into an I-type addition chain based on the ITA addition chain and an R-type addition chain based on modular square root operation; According to the accumulated results of the two-dimensional differential addition chain, the addition chain is iterated in parallel through the I-type addition chain and the R-type addition chain to obtain the I-type addition chain iteration result and the R-type addition chain iteration result, and the modular inverse operation result is calculated.

8. The low-latency two-dimensional differential addition chain inference method according to claim 7, characterized in that: The modular inverse operation result obtained by the calculation includes: According to the formula Calculate the modular inverse operation result α -1 ,in, is the iterative result of type I addition chain, is the iterative result of the R-type addition chain, k is the preset scalar value; m is the length of the original ITA addition chain; α is the accumulated result of the two-dimensional difference addition chain.