Lightweight elliptic curve cipher reconfigurable processor
By designing a lightweight elliptic curve cipher reconfigurable processor, the problems of complex hardware and insufficient attack resistance in the existing technology are solved, and the high security and low power consumption ECC algorithms are quickly switched, which is suitable for the fields of national cipher security, network security teaching and trade secrets.
Patent Information
- Application Number
- CN202510349696.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-04-28
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing reconfigurable elliptic curve cryptographic processors have problems such as complex hardware structure, large circuit scale and insufficient resistance to side channel attacks.
A lightweight elliptic curve password reconfigurable processor is designed, including a shared data area, a reconfigurable instruction memory, a master state machine logic module, a decoder, a parameter configuration manager, a computing array unit, a data channel router and a branch jump logic module. The shared data area is used to store initial data and global configuration parameters, and the decoder and a parameter configuration manager are driven by the master state machine logic module, to realize the flexible reconstruction of the computing unit array and the efficient computing result transmission of the data channel router.
It realizes high security, low power consumption and fast ECC algorithm switching, which is suitable for national password security, network security teaching and trading secrets, and has broad application prospects.
Smart Images

Figure CN120277027A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a lightweight elliptic curve cryptography reconfigurable processor. Background Art
[0002] Elliptic Curve Cryptography (ECC) is a public-key encryption technology based on the theory of elliptic curves, which can be widely applied to application scenarios such as public-key encryption, signature, and key exchange. Compared with traditional public-key cryptosystems (DSA, RSA, etc.), ECC is constructed based on the Elliptic Curve Discrete Logarithm Problem (ECDLP). Therefore, it can use shorter keys while providing the same level of security, making it have better performance, and thus gradually replacing traditional public-key cryptosystems.
[0003] For the hardware implementation of an elliptic curve cryptography processor, according to whether its result meets the definition of reconfigurability, it can be divided into processors implemented in a non-reconfigurable manner and those implemented in a reconfigurable manner. For the former, its representative method is Application Specific Integrated Circuit (ASIC). This method can design and manufacture integrated circuits according to the specific needs of users, so it has the advantages of small area, low power consumption, high reliability, and good performance. However, this method has the disadvantages of poor flexibility and single functionality. For the latter, its representative methods are Field Programmable Gate Array (FPGA) and a reconfigurable circuit for cryptographic algorithms implemented based on software-defined chip technology. Compared with FPGA, the reconfigurable circuit for cryptographic algorithms implemented based on software-defined chip technology can better balance flexibility, performance, and power consumption indicators. Therefore, this design method is gradually becoming an important development direction for future cryptographic chips.
[0004] For a reconfigurable elliptic curve cryptography processor, the patent application with the publication number CN 101826142 A discloses "a reconfigurable elliptic curve cryptography processor". This processor consists of a control unit, a data path unit, and an input / output unit, and can support the reconfigurable configuration of elliptic curve cryptography from the finite field layer to the curve layer to the group operation layer and then to the application layer. However, this reconfigurable processor has problems such as a complex hardware structure, a large circuit scale, and an unclear ability to resist side-channel attacks. Summary of the Invention
[0005] In view of the above problems, the present invention provides a lightweight elliptic curve cryptography reconfigurable processor.
[0006] To solve the above technical problems, the present invention provides a lightweight elliptic curve cryptography reconfigurable processor, which includes a shared data area, a reconfigurable instruction memory, a main control state machine logic module, a decoder, a parameter configuration manager, an arithmetic unit array, a data path router, and a branch jump logic module;
[0007] The shared data area is used for the upper-layer system to write initial data, global configuration parameters, two sets of moduli, two sets of pre-computed parameters Mu, and data to be operated. The initial data includes elliptic curve configuration parameters and a set of noise data. The global configuration parameters include reconfigurable data of the arithmetic unit array. The reconfigurable data of the arithmetic unit array includes base address sequences of point addition and double point functions, base addresses of scalar multiplication and coordinate conversion functions, modulus length, field element width, segmented values of modular inverse operands, and the starting address of the main program without initialization instructions;
[0008] The reconfigurable instruction memory is used for the upper-layer system to write reconfigurable instructions, and the reconfigurable instructions include initialization instructions, arithmetic instructions, and jump instructions;
[0009] The main control state machine logic module is used to receive the control word written by the upper-layer system and enable operation according to the control word, so as to send control signals to the decoder and the parameter configuration manager to make the decoder and the parameter configuration manager start working;
[0010] The decoder is used to read and decode instructions starting from the 0 address of the reconfigurable instruction memory under the drive of the main control state machine logic module, and synchronously feedback the decoded instruction type into the main control state machine logic module. If the reconfigurable instruction contains an initialization instruction, the decoder preferentially executes the initialization instruction. If it does not contain an initialization instruction or the initialization instruction has been executed, the remaining instructions are executed until the last instruction is executed; the decoder is also used to control the operation of the data path router to realize the gating of the data to be operated input and the operation result output;
[0011] The parameter configuration manager is used to extract the global configuration parameters and initial data of the shared data area under the drive of the main control state machine logic module to reconstruct the function of the modular arithmetic circuit of the arithmetic unit array, and send the noise data and the base address sequences of the point addition and double point functions to the main control state machine logic module. The main control state machine logic module outputs the scrambled function address to the decoder according to the noise data and the base address sequences of the point addition and double point functions;
[0012] The arithmetic unit array is used to output corresponding operation results and operation completion flags according to the data to be operated and operation instructions gated by the decoder;
[0013] The branch jump logic module is used to receive the jump instruction decoded and output by the decoder, execute the corresponding jump logic, and return the corresponding jump flag to the master state machine logic module and the decoder. The decoder is also used to determine the current instruction or whether the instruction sequence is completed and determine the fetch address of the next instruction according to the jump flag, the operation completion flag, and the out-of-order function address. The master state machine logic module exits the program according to the jump flag;
[0014] The data path router is used to send and store the operation result to the corresponding position in the shared data area according to the result output strobe result, and the shared data area feeds back the operation result to the upper-level system.
[0015] Further, the operation unit array includes a modular addition unit, a modular subtraction unit, four groups of modular multiplication units, a general adder, a general subtractor, and a modular inverse unit over a large prime field.
[0016] Further, the shared data area includes 16 vector registers, each vector register having a bit width of 600 bits, and the 16 vector registers support parallel 8-vector readout, 4-vector write, and data-level parallel operations.
[0017] Further, a serial-to-parallel / parallel-to-serial conversion logic module is connected to the front end of the shared data area. The serial-to-parallel / parallel-to-serial conversion logic module is used to implement the parallel-to-serial conversion of the current output vector register and the serial-to-parallel conversion of the current written data. The serial-to-parallel / parallel-to-serial conversion logic module automatically executes the corresponding conversion logic according to the bus address transmitted in order.
[0018] Further, the reconfigurable instruction further includes function call instructions, including top-level function call instructions and bottom-level function call instructions. The top-level function is a scalar multiplication function, which has a scalar encoding type and a scalar value address. When executing the scalar multiplication function call instruction, the next address of the instruction is written into the return address register 1, and the function parameter is loaded into the branch jump logic module. Then the decoder reads the address of the scalar multiplication function from the master state machine logic module as the fetch address of the next instruction, and then continues to execute the subsequent instructions. The bottom-level functions are dot addition, double dot, and coordinate transformation functions. When executing the bottom-level function call instruction, the next address of the instruction is written into the return address register 2, and then the decoder reads the out-of-order address of the dot addition and double dot functions or the address of the coordinate transformation function from the master state machine logic module as the fetch address of the next instruction, and then continues to execute the subsequent instructions.
[0019] Further, the branch jump logic module outputs jump flags according to four categories of situations. One is the index - 0 judgment logic for detecting whether the scalar multiplication loop is completed; the second is the unconditional jump logic, which is used in function return instructions or program return instructions; the third is the NAF - encoding dot - free addition judgment logic related to scalar encoding and the judgment logic for the current bit of the binary - encoded scalar being 0; the fourth is the constant comparison circuit.
[0020] Further, the upper - layer system polls whether the processor completion status is set in a polling manner or responds to the processor's interrupt. Once it detects that the processor has finished the operation, the upper - layer system directly reads the operation result from the specified address in the shared data area, and then clears the completion flag in the status register or clears the interrupt flag. When the main control state machine logic module detects the corresponding change, it enters the idle state and waits for the upper - layer system to enable it again.
[0021] Compared with the prior art, the beneficial effects of the present invention include: The present invention can cope with application scenarios such as high security, low power consumption, fast switching and reconstruction of ECC algorithms, and can thus be flexibly applied to national cryptography security, network security teaching, and commercial cryptography fields, with broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the principle block diagram of the lightweight elliptic curve cryptography reconfigurable processor according to the embodiment of the present invention;
[0023] Figure 2 is the working flowchart of the main control state machine logic module according to the embodiment of the present invention;
[0024] Figure 3 is the principle block diagram of the arithmetic array unit according to the embodiment of the present invention;
[0025] Figure 4 is the state machine schematic diagram of the decoder according to the embodiment of the present invention;
[0026] Figure 5 is the principle block diagram of the branch jump logic module according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following further clarifies the present invention with reference to the drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0028] Such as Figure 1As shown in the figure, an embodiment of the present invention provides a lightweight elliptic curve cryptography reconfigurable processor, which includes a shared data area 1, a reconfigurable instruction memory 2, a master control state machine logic module 3, a decoder 4, a parameter configuration manager 5, an arithmetic array unit 6, a data path router 7, and a branch jump logic module 8.
[0029] Among them, the shared data area 1 includes 16 vector registers, each vector register has a bit width of 600 bits, and the 16 vector registers support parallel 8-vector reading, 4-vector writing, and data-level parallel operations, which can significantly improve the ECC operation performance. The shared data area 1 is open to the upper-layer system for reading and writing operations, so as to allow the upper-layer system to write initial data, global configuration parameters, two sets of moduli, two sets of pre-computed parameters Mu, and data to be operated (operands). The above initial data includes elliptic curve configuration parameters and a set of noise data. The above global configuration parameters include the reconfigurable data of the arithmetic unit array, and the reconfigurable data of the arithmetic unit array includes the base address sequence of the point addition and double point functions, the base addresses of the scalar multiplication and coordinate conversion functions, the modulus length, the field element width, the modulus inverse operand segmentation value, and the starting address of the main program without initialization instructions.
[0030] The reconfigurable instruction memory 2 is only open to the upper-layer system for writing operations, so as to allow the upper-layer system to write reconfigurable instructions, and the above reconfigurable instructions include initialization instructions, arithmetic instructions, and jump instructions.
[0031] See Figure 2 , the master control state machine logic module 3 is used to receive the control word written by the upper-layer system, and enable work according to the above control word, so as to send control signals to the decoder 4 and the parameter configuration manager 5, so that the decoder 4 and the parameter configuration manager 5 start to work.
[0032] The decoder 4 is used to read and decode instructions starting from the 0 address of the reconfigurable instruction memory 2 under the drive of the master control state machine logic module 3, and synchronously feedback the decoded instruction type into the master control state machine logic module 3. If the reconfigurable instruction contains an initialization instruction, the decoder 4 preferentially executes the initialization instruction. If it does not contain an initialization instruction or the initialization instruction has been executed, the remaining instructions are executed until the last instruction is executed.
[0033] The parameter configuration manager 5 is used to extract the global configuration parameters and initial data of the shared data area 1 under the drive of the master control state machine logic module 3, so as to reconstruct the function of the modular arithmetic circuit of the arithmetic unit array, and send the noise data and the base address sequence of the point addition and double point functions to the master control state machine logic module 3. The master control state machine logic module 3 outputs the scrambled function address to the decoder 4 according to the noise data and the base address sequence of the point addition and double point functions. The decoder 4 is also used to control the operation of the data path router 7 to realize the gating of the data to be operated input and the operation result output.
[0034] The arithmetic unit array 6 is used to output corresponding arithmetic results and arithmetic completion flags according to the gated input data to be calculated and arithmetic instructions.
[0035] The branch jump logic module 8 is used to receive the jump instructions decoded and output by the decoder 4, execute corresponding jump logic, and return corresponding jump flags to the main control state machine logic module 3 and the decoder 4. The decoder 4 is also used to determine the current instruction or judge whether the instruction sequence is completed and determine the fetch address of the next instruction according to the jump flag, arithmetic completion flag and out-of-order function address. The main control state machine logic module 3 also completes program exit according to the jump flag.
[0036] The data path router 7 is used to send and store the arithmetic results to corresponding positions in the shared data area according to the gating results, and the shared data area 1 feeds back the arithmetic results to the upper-level system.
[0037] Since the common bit width of the on-chip bus of the SoC is 32 bits or 64 bits, and the bit width of the vector register is 600 bits, therefore, a serial-parallel / parallel-serial conversion logic module 9 is connected to the front end of the shared data area 1. The serial-parallel / parallel-serial conversion logic module 9 is used to implement the parallel-serial conversion of the current output vector register and the serial-parallel conversion of the currently written data. The serial-parallel / parallel-serial conversion logic module 9 automatically executes corresponding conversion logic according to the bus address transmitted in order. Since the bit width of the reconfigurable instruction of the present invention is 24 bits, it can be directly connected to the data interface of the on-chip bus.
[0038] The arithmetic unit array 6 of the embodiment of the present invention includes a modular addition unit, a modular subtraction unit, four groups of modular multiplications, a general adder, a general subtractor, and a modular inverse unit on a large prime number field. The block diagram of the arithmetic unit array 6 is as Figure 3 shown Figure 3The ai and bi signals are 600-bit operands for each path selected by the MUX circuit and fed into the arithmetic operation unit. The signals named in the en_xx mode are the module enable signals for each arithmetic operation unit. The signals named in the over_xx mode are the completion flag signals output after each arithmetic operation unit completes the corresponding operation, which are active high in a single cycle. Modules such as the modular multiplication unit, modular addition unit, modular subtraction unit, and modular inverse unit in the arithmetic operation unit array need to first load configuration parameters to reconstruct the required arithmetic functions. After reconstruction, under the control of the arithmetic operator, arithmetic operation units such as the modular multiplication unit, modular inverse unit, and ordinary adder are fed with operands through a 4-way MUX. However, only the enabled arithmetic operation units can perform operations. After the operation is completed, the operation result is sent out through the operation result MUX circuit controlled by the arithmetic operator, and a maximum of 4-way operation results can be output simultaneously. When multiple arithmetic operation units are enabled to work simultaneously, their operation completion signals may not be set to active high at the same time, and a decoder circuit is required for synchronization processing. Most of the arithmetic operation units in the arithmetic operation unit array 6 are the underlying logic units that constitute the elliptic curve cryptography algorithm. Therefore, these units largely determine the overall performance overhead and resource overhead of the reconfigurable system. Each arithmetic operation unit in this arithmetic operation unit array 6 will reconfigurably support finite field arithmetic with a large prime field bit width not exceeding 600 bits. First, for the modular multiplication operation in the large prime field, the present invention adopts a general Barrett reduction modular multiplication algorithm based on segmented scanning of operands to implement the large number modular multiplication operation with a prime field bit width below 600 bits. Different from the Montgomery modular multiplication algorithm, the Barrett modular multiplication algorithm does not require specific domain conversion and can efficiently process moduli with different bit widths. Therefore, the modular multiplication circuit unit designed based on the segmented scanning Barrett modular multiplication algorithm has the advantages of flexible call, small hardware scale, and the ability to preset the segmented parameters of operands (flexibly balancing the circuit scale and modular multiplication performance). Before using this modular multiplication unit, it is necessary to first perform pre-configuration of relevant parameters through the parameter configuration manager 5. Specifically, during the operation of the modular multiplication algorithm, a pre-computation parameter Mu is required to calculate the approximate quotient value in each round of Barrett reduction; a value n is used to represent the bit width of the finite field where the current curve is located; and a value p is used as the modulus for the modular multiplication operation. This modular multiplication algorithm performs modular multiplication calculations in a cyclic iterative manner. In each iteration, the algorithm first calculates the product of a and b i , and then uses the Barrett reduction method to perform modular reduction on the operation result, and finally feeds the result of the current round back to the next round. When the loop ends, the algorithm needs to determine whether the calculation result exceeds the modulus size and perform at most one correction operation on it.
[0039] Secondly, for the modular inverse operation in a large prime number field, the present invention adopts the extended Euclidean modular inverse algorithm. Since this extended Euclidean modular inverse algorithm does not rely on division, but only requires addition and subtraction operations, shift operations, and parity judgment operations, it calculates the inverse element by the method of successive division. Therefore, the hardware circuit scale is small and it has good operating efficiency. The operation unit array 6 will call the ordinary addition module and ordinary subtraction module implemented by the segmented scanning method of the operand, enabling the operation unit array 6 to efficiently process the modular inverse operation with any bit width below 600 bits in the prime number field.
[0040] Finally, the hardware designs of modular addition, modular subtraction, ordinary adder, and ordinary subtractor are relatively simple. The modular addition and modular subtraction algorithm modules will also call ordinary addition and subtraction units internally. Therefore, only the implementation circuit of ordinary addition and subtraction needs to be introduced in detail. The operands of addition and subtraction are segmented and executed for addition or subtraction operations according to a preset length, and then the carry or borrow is added to the previous segmented data. If there is still a carry or borrow generated, the addition operation with the previous segmented data is iteratively executed until no carry or borrow is generated, that is, the ordinary addition and subtraction operations are completed. In view of the fact that the addition and subtraction circuit adopts the processing method of segmented addition and subtraction of operands, the carry chain is interrupted, which is easy to extend to the addition and subtraction operations with ultra-large bit widths.
[0041] Once enabled, the decoder 4 will actively read the reconfigurable instructions in the reconfigurable instruction memory 2, and can read up to 4 instructions simultaneously at most, and start synchronous decoding for all the instructions read at one time. The decoder 4 will output corresponding decoding results for different instruction types, and these decoding results serve as the control signals for the branch jump logic module 8 and the operation unit array 6, the path selection signals for the data path router 7, and the instruction type data for the main control state machine logic module 3. Specifically, the decoder 4 determines the current instruction or judges whether the instruction sequence is completed and determines the fetch address of the next instruction according to the jump flag of the branch jump logic module 8, the completion flag of the operation unit array, and the out-of-order function address. The state machine of the instruction decoder 4 is as Figure 4 shown. Different from the traditional embedded processor decoding process, which is divided into five processes: fetch, decode, execute, memory access, and write-back, in order to simplify the decoding logic, this decoder 4 directly removes the memory access and write-back processes. The data loading of the operation unit array 6 and the writing of the operation result back to the shared data area 1 are responsible for the data path router 7. In order to adapt to the multi-cycle characteristics of the operation instructions and the characteristics that the operation cycles of different operation instructions may be different, the decoder 4 adds the detection of the operation flag status to determine whether the current instruction or multiple parallel instructions have all completed the operation. In order to support the jumps of the underlying functions of elliptic curve point addition, double point, and scalar multiplication functions, the decoder 4 adds the setting of the function call flag status to obtain the fetch address of the next instruction.
[0042] Regarding the reconfigurable instruction design, in order to simplify the instructions and reduce the instruction storage overhead, the VLIW type instructions commonly used in ECC reconfigurable accelerators are not adopted. However, to ensure the performance of the ECC reconfigurable processor and the flexibility of instruction encoding, the arithmetic instructions are designed to support up to 4-way parallel decoding and invocation. By analyzing the call frequencies and data dependencies of the point addition and double point function modular multiplication operations in the ECC algorithm, it is found that setting the maximum parallelism of the modular multiplication operation to 4 can ensure the fastest execution speed of the point addition and double point functions. At the same time, the parallelism of the supporting reconfigurable instructions should also be set to 4. After testing, 4-way parallel decoding and 4-way modular multiplication parallel invocation can significantly improve the arithmetic performance of the ECC reconfigurable processor. The decoder 4 supports up to 4 arithmetic instructions for parallel decoding and up to 4-way parallel invocation of the arithmetic operation unit, and naturally also supports the parallel decoding and invocation of 2 or 3 arithmetic instructions.
[0043] The present invention implements 4 types of reconfigurable instructions, namely initialization instructions, arithmetic instructions, jump instructions, and function call instructions. The function definitions of the 4 types of instructions are shown in Table 1. Among them, the arithmetic instructions support parallel processing of up to 4 instructions. The completion time of the parallel instructions is referenced by the slowest running instruction. The next instruction (or the next set of parallel instructions) can only be executed after the slowest instruction has finished running. The instruction encoding is up to 24 bits long and as short as 14 bits, which is much smaller than the length of VLIW type instructions, greatly saving the instruction storage overhead.
[0044] Table 1 Reconfigurable Instruction Definition
[0045]
[0046]
[0047] The block diagram of the branch jump logic module 8 is as Figure 5 shown. The decoded output of the jump instruction is used to control the operation of the branch jump logic module. Figure 5 In it, the index signal is the bitwise index value of the scalar K. The index value decreases from high to low and is set in combination with the field element width parameter and the scalar K during initialization. When the index drops to 0, it indicates that the ECC scalar multiplication is completed. As defined in Table 1 for the jump instruction, the scalar value encoding circuit is responsible for generating jump flags in both NAF encoding and binary encoding cases, facilitating the programming of the call to the point addition and double point functions when writing the scalar multiplication function. The constant comparison circuit has a simple structure and only judges whether the operation result is 0, 1, or negative and generates a jump flag. Finally, under the control of the decoded output result, the final jump flag is output for the decoding operation of the decoder module. Overall, the branch jump logic module has a simple structure, a small circuit scale, and is closely combined with the key arithmetic function of ECC - scalar multiplication operation, significantly improving the jump execution efficiency of the scalar multiplication instruction sequence.
[0048] The specific reconfigurable implementation and call process of the ECC algorithm are as follows:
[0049] a) In the initial stage of the first step, the upper-layer system writes initial data, global configuration parameters and other data to the shared data area 1 through the system bus. The upper-layer system continues to write the pre-compiled ECC reconfigurable instructions to the reconfigurable instruction memory 2. The starting part of the reconfigurable instruction is all initialization instructions, and the rest is a mixed instruction sequence of arithmetic instructions, function call instructions and jump instructions.
[0050] b) The upper-layer system writes a control word to the processor through the system bus to enable the main control state machine logic module 3 to work. The main control state machine logic module 3 drives the decoder 4. The decoder 4 starts reading instructions from the 0 address of the reconfigurable instruction memory 2 and decodes and executes them. The instruction type output by the decoding is synchronously fed back into the main control state machine logic module 3.
[0051] c) The decoder 4 first executes the initialization instruction sequence. Based on this, the main control state machine logic module 3 enables the parameter configuration manager 5. The parameter configuration manager 5 receives the routed global configuration data, two sets of modulus, two sets of pre-computed parameters Mu and a set of noise data. Except for the instruction address and noise data, which are fed back to the main control state machine logic module 3, other configuration items, modulus, Mu and other data are written into the arithmetic unit array 6. For example, for modular multiplication operations, when writing values such as modulus, modulus length, Mu, etc., modular multiplication operations with different operand lengths below 600 bits can be reconfigured and implemented.
[0052] d) Once the initialization instruction sequence is executed, the arithmetic unit array completes the arithmetic function reconfiguration. The decoder starts to execute the subsequent instructions, that is, a mixed instruction sequence composed of arithmetic instructions, function call instructions and jump instructions. The state machine executed by the decoder is as Figure 5 shown.
[0053] 1) When executing arithmetic instructions (supporting 4 instructions to be executed simultaneously), the operator sequence, operand address transformation sequence and operation enable are sent to the arithmetic unit array 6. The operand address transformation sequence is used to select a pair of operands (if it is a unary operation, select one operand) for the specified arithmetic unit from the maximum 4 groups of operands sent by the data path router 7. The arithmetic unit to be enabled will emit a high-level completion flag signal after completing the operation and output the corresponding operation result. The decoder 4 synchronizes the operation completion flag and accordingly drives the data path router 7 to write the operation result to the specified address unit of the shared data area 1.
[0054] 2) If a function call instruction is executed, the processing method varies according to the function type. The top-level function is a scalar multiplication function with parameters such as scalar encoding type and scalar value address. When a scalar multiplication function call instruction is executed, the address of the next instruction is written into the return address register 1, and function parameters such as scalar values are loaded into the branch jump logic module 8. Then, the decoder 4 reads the address of the scalar multiplication function from the main control state logic module 3 as the fetch address of the next instruction and continues to execute the subsequent instructions. The bottom-level functions are point addition, double point, and coordinate transformation functions. These functions have no parameters, and the required input data is provided by the address units at addresses 8 - 13 specified in the shared data area 1. When these function call instructions are executed, the address of the next instruction is written into the return address register 2. Then, the decoder 4 reads the out-of-order address of the point addition or double point function or the address of the coordinate transformation function from the main control state logic module 3 as the fetch address of the next instruction and continues to execute the subsequent instructions.
[0055] 3) If a jump instruction is executed, the decoded output of the jump instruction is sent to the branch jump logic module 8. For different jump modes, as Figure 5 shown, four types of jump flags are output. One is the index - 0 judgment logic for detecting whether the scalar multiplication loop is completed; the second is the unconditional jump logic, which is used in function return instructions or program return instructions; the third is the NAF - encoding - related no - point - addition judgment logic for scalar encoding and the judgment logic for the current bit of the binary - encoded scalar being 0; the fourth is the constant comparison circuit. The decoder 4 jumps to the address in the reconfigurable instruction memory 2 recorded in the return address register when the value of the jump flag is '1'. If the value of the jump flag is '0', the decoder 4 fetches the instruction from the next address of the current instruction address.
[0056] e) When the last instruction (which must be a program return instruction) is executed, the main control state machine logic module 3 is triggered. The main control state machine logic module 3 sets the status register and sets the completion flag and the interrupt flag. The upper - layer system can poll the status of whether the ECC reconfigurable processor completion status is set in a polling manner or respond to the interrupt of the ECC reconfigurable processor. Once it is detected that the ECC operation is completed, the upper - layer system can directly read the operation result from the specified address in the shared data area 1, and then clear the completion flag or the interrupt flag in the status register. When the main control state machine logic module 3 detects the corresponding change, it enters the idle state and waits for the upper - layer system to enable the ECC reconfigurable processor again.
[0057] The present invention has the advantages of small scale, low power consumption, and easy configuration and use. It can safely and efficiently support the hardware reconfiguration of three major types of elliptic curve cryptography algorithms over large prime number fields, and supports the dynamic configuration of elliptic curve cryptography algorithms with a bit width within 600 bits. Given that in the current ECC algorithms, the maximum field element width of the large prime number finite field recommended by NIST is 521 bits, therefore, the finite field operation of the ECC reconfigurable processor designed by the present invention supports operations on data with a maximum of 600 bits, which is sufficient to meet the current ECC cryptographic service requirements. In the design of reconfigurable instructions, short-encoded instructions are preferably used as much as possible, and in the design of the decoder 4, the branch jump logic module 8, and the main control state machine logic module 3, circuits with simple and efficient structures are preferably used, significantly reducing the circuit scale of the entire processor. The present invention can resist side-channel attacks against ECC algorithms. By introducing noise data and randomly calling the function sequences of ECC point addition and double point operations, the parallelism of the base field multiplier called by each function and the order of each arithmetic operation can be flexibly scrambled and programmed, that is, the reconfigurable instruction implementation methods of the point addition and double point function group sequences are different, which can make the energy trace envelope of ECC point addition and double point operations in a random state, significantly confusing the interface of the energy trace envelopes of the two operations and also confusing the energy traces of the two operations, resulting in the inability to distinguish the energy traces of point addition and double point operations, making the side-channel attack against the non-equilibrium state of the scalar value '0-1' in ECC multiple point operations twice as difficult or infeasible. Given that the combination of reconfigurable instructions and noise data can achieve the random call of point addition operations and double point operations, can confuse the energy trace envelopes of the two operations, and can resist side-channel attacks, it makes it possible to apply the NAF encoding of scalar values in the scalar multiplication algorithm. Otherwise, it is necessary to apply the Montgomery ladder modular exponentiation algorithm to implement the ECC scalar multiplication operation and perform '0-1' power consumption balancing processing on the scalar value. Usually, the scalar multiplication using NAF encoding has a performance improvement of about 1 / 3 compared to the ordinary binary scan algorithm. Therefore, the processor of the present invention is both safe and efficient.
[0058] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, the other parts not specifically described belong to the prior art or common general knowledge. Without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A lightweight elliptic curve cryptography reconfigurable processor, characterized in that, It includes a shared data area, a reconfigurable instruction memory, a master control state machine logic module, a decoder, a parameter configuration manager, an arithmetic array unit, a data path router, and a branch jump logic module; The shared data area is used for the upper-level system to write initial data, global configuration parameters, two sets of modulus, two sets of pre-computed parameters Mu, and data to be operated. The initial data includes elliptic curve configuration parameters and a set of noise data. The global configuration parameters include reconfigurable data of the arithmetic unit array. The reconfigurable data of the arithmetic unit array includes the base address sequences of point addition and double point functions, the base addresses of scalar multiplication and coordinate conversion functions, the modulus length, the field element width, the segmented value of the modular inverse operand, and the starting address of the main program without initialization instructions; The reconfigurable instruction memory is used for the upper-level system to write reconfigurable instructions. The reconfigurable instructions include initialization instructions, arithmetic instructions, and jump instructions; The master control state machine logic module is used to receive the control word written by the upper-level system and enable operation according to the control word, so as to send control signals to the decoder and the parameter configuration manager to make the decoder and the parameter configuration manager start working; The decoder is used to read and decode instructions starting from the 0 address of the reconfigurable instruction memory under the drive of the master control state machine logic module, and synchronously feedback the decoded instruction type into the master control state machine logic module. If the reconfigurable instruction contains an initialization instruction, the decoder preferentially executes the initialization instruction. If it does not contain an initialization instruction or the initialization instruction has been executed, the remaining instructions are executed until the last instruction is executed; The decoder is also used to control the operation of the data path router to realize the gating of the data to be operated input and the operation result output; The parameter configuration manager is used to extract the global configuration parameters and initial data of the shared data area under the drive of the master control state machine logic module to reconstruct the function of the modular arithmetic circuit of the arithmetic unit array, and send the noise data and the base address sequences of point addition and double point functions to the master control state machine logic module. The master control state machine logic module outputs the out-of-order function address to the decoder according to the noise data and the base address sequences of point addition and double point functions; The arithmetic unit array is used to output the corresponding operation result and operation completion flag according to the data to be operated and the operation instruction gated by the decoder; The branch jump logic module is used to receive the jump instruction decoded and output by the decoder to execute the corresponding jump logic, and return the corresponding jump flag to the master control state machine logic module and the decoder. The decoder is also used to determine the current instruction or judge whether the instruction sequence is completed and determine the fetch address of the next instruction according to the jump flag, the operation completion flag, and the out-of-order function address. The master control state machine logic module completes the program exit according to the jump flag; The data path router is used to send and store the operation result to the corresponding position of the shared data area according to the result output gating result, and the shared data area feeds the operation result back to the upper-level system.
2. The lightweight elliptic curve cryptography reconfigurable processor according to claim 1, characterized in that The arithmetic unit array includes modulo addition units, modulo subtraction units, four groups of modulo multiplication units, ordinary adders, ordinary subtractors, and modulo inverse units over a large prime field.
3. A lightweight elliptic curve cryptography reconfigurable processor according to claim 1, characterized in that, The shared data area includes 16 vector registers, each with a bit width of 600 bits, and the 16 vector registers support parallel 8-vector readout, 4-vector write-in, and data-level parallel operations.
4. A lightweight elliptic curve cryptography reconfigurable processor according to claim 3, characterized in that, A serial-parallel / parallel-serial conversion logic module is connected to the front end of the shared data area. The serial-parallel / parallel-serial conversion logic module is used to implement the parallel-serial conversion of the current output vector register and the serial-parallel conversion of the current write-in data. The serial-parallel / parallel-serial conversion logic module automatically executes the corresponding conversion logic according to the sequentially transmitted bus address.
5. A lightweight elliptic curve cryptography reconfigurable processor according to claim 1, characterized in that, The reconfigurable instructions further include function call instructions, including top-level function call instructions and bottom-level function call instructions. The top-level function is a scalar multiplication function, which has a scalar encoding type and a scalar value address. When the scalar multiplication function call instruction is executed, the address of the next instruction is written into the return address register 1, the function parameters are loaded into the branch jump logic module, and then the decoder reads the address of the scalar multiplication function from the main control state machine logic module as the fetch address of the next instruction, and then continues to execute the subsequent instructions. The bottom-level functions are dot addition, double dot, and coordinate transformation functions. When the bottom-level function call instruction is executed, the address of the next instruction is written into the return address register 2, and then the decoder reads the out-of-order address of the dot addition or double dot function or the address of the coordinate transformation function from the main control state machine logic module as the fetch address of the next instruction, and then continues to execute the subsequent instructions.
6. A lightweight elliptic curve cryptography reconfigurable processor according to claim 1, characterized in that, The branch jump logic module outputs jump flags according to four types of situations. One is the index-0 judgment logic for detecting whether the scalar multiplication loop is completed; the second is the unconditional jump logic, which is used in function return instructions or program return instructions; the third is the NAF encoding dot addition-free judgment logic related to scalar encoding and the binary encoding scalar current bit-0 judgment logic; The fourth is a constant comparison circuit.
7. A lightweight elliptic curve cryptography reconfigurable processor according to claim 1, characterized in that, The upper-level system polls in a polling manner whether the processor completion status is set, or responds to the interrupt of the processor. Once it detects that the processor has finished operating, the upper-level system directly reads the operation result from the specified address in the shared data area, and then clears the completion flag in the status register or clears the interrupt flag. When the main control state machine logic module detects the corresponding change, it enters the idle state and waits for the upper-level system to enable it again.
Citation Information
Patent Citations
Reconfigurable elliptic curve cipher processor
CN101826142A
Cited By
Programmable hardware architecture supporting elliptic curve point multiplication
CN122548805A