Apparatus and method for accelerating computation, and chip

By designing the computing acceleration device in the 64-bit RISC-V processor, using the bit width selection module and the corresponding processing module, the problem of low computing efficiency of large bit width zero knowledge is solved, and a significant acceleration effect is achieved.

WO2025118193A1PCT designated stage expired Publication Date: 2025-06-12SUNLUNE (SINGAPORE) PTE LTD

Patent Information

Application Number
PCT/CN2023/136848
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-06
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the 64-bit RISC-V architecture processor system, how to efficiently implement zero-knowledge proof operations with large bit width, especially large integer finite domain operations with 256-bit and 384-bit scales, is inefficient.

Method used

An operation acceleration device is designed, including an instruction decoding module, a bit width selection module, a small bit width processing module and a large bit width processing module. The bit width of the instruction is determined by instruction decoding, and the corresponding processing module is selected according to the bit width, and the large bit width register stack and arithmetic logic unit are used for processing.

Benefits of technology

The acceleration of large-bit wide computing is achieved in the small-bit wide processor architecture system, which significantly reduces the number of instructions and clock cycles, and improves the efficiency of zero-knowledge proof computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023136848_12062025_PF_FP_ABST
    Figure CN2023136848_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are an apparatus and method for accelerating computation, and a chip. The apparatus for accelerating computation comprises: an instruction decoding module, a bit width selection module, a small bit width processing module, and a large bit width processing module. The instruction decoding module is configured to decode an inputted instruction, determine the bit width of the instruction on the basis of a decoding result, generate a control signal comprising bit width information of the instruction, and send the control signal to the bit width selection module. The bit width selection module is configured to forward the control signal to the small bit width processing module when the bit width of the instruction is a small bit width, and forward the control signal to the large bit width processing module when the bit width of the instruction is a large bit width. The small bit width processing module is configured to process the instruction by using a register file having a small bit width and a corresponding arithmetic logic unit. The large bit width processing module is configured to process the instruction by using a register file having a large bit width and a corresponding arithmetic logic unit.
Need to check novelty before this filing date? Find Prior Art

Description

Computing acceleration device, method and chip Technical Field

[0001] This article relates to the field of cryptographic computing technology, and in particular to a computing acceleration device, method, and chip. Background Art

[0002] Zero-Knowledge Proof (ZKP) is a cryptographic algorithm that can prove the correctness of a proposition without revealing any other information, thereby addressing privacy and scalability issues in communication. The key theoretical foundation of ZKP is elliptic curves, whose basic operations are based on large integer arithmetic over finite fields. Furthermore, the number theoretic transformations that ZKP relies on also require extensive large integer arithmetic.

[0003] Although zero-knowledge proofs haven't yet reached a true industry standard, common zero-knowledge proof schemes like Groth16, PlonK, and STARK all make extensive use of large-integer finite field operations on 256-bit and 384-bit scales. Common 32-bit, 64-bit, and even 128-bit processors require a large number of small-bit-width instructions to implement these large-integer finite field operations, resulting in low efficiency. For example, the most commonly used 384-bit Montgomery modular multiplication requires only eight conventional instructions using 384-bit registers and arithmetic units. However, using a 64-bit RISC (Reduced Instruction Set Computer)-V architecture processor, the computation requires approximately 1,500 instructions.

[0004] Therefore, how to implement large-bit-width zero-knowledge proof operations in a 64-bit RISC-V architecture processor system is a problem that needs to be solved.

[0005] Summary of the Invention

[0006] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0007] The present disclosure provides an operation acceleration device, comprising: an instruction decoding module, a bit width selection module, a small bit width processing module and a large bit width processing module;

[0008] an instruction decoding module configured to decode an input instruction, determine the bit width of the instruction according to the decoding result, generate a control signal containing the bit width information of the instruction, and send the control signal to the bit width selection module;

[0009] a bit width selection module configured to obtain bit width information of the instruction from the control signal, forward the control signal to the small bit width processing module when the bit width of the instruction is small, and forward the control signal to the large bit width processing module when the bit width of the instruction is large; wherein the small bit width means that the data bit width is less than or equal to the first threshold, and the large bit width means that the data bit width is greater than the first threshold;

[0010] A small bit-width processing module configured to process instructions using a small bit-width register file and a corresponding arithmetic logic unit;

[0011] The large bit width processing module is configured to process instructions using a large bit width register file and a corresponding arithmetic logic unit.

[0012] The present disclosure provides a method for accelerating computing, comprising:

[0013] The instruction decoding module decodes the input instruction, determines the bit width of the instruction according to the decoding result, generates a control signal containing the bit width information of the instruction, and sends it to the bit width selection module;

[0014] The bit width selection module obtains the bit width information of the instruction from the control signal, and forwards the control signal to the small bit width processing module when the bit width of the instruction is small, and forwards the control signal to the large bit width processing module when the bit width of the instruction is large; wherein the small bit width means that the data bit width is less than or equal to the first threshold, and the large bit width means that the data bit width is greater than the first threshold;

[0015] The small bit-width processing module uses a small bit-width register file and a corresponding arithmetic logic unit to process instructions; the large bit-width processing module uses a large bit-width register file and a corresponding arithmetic logic unit to process instructions.

[0016] The present disclosure provides a chip including the above-mentioned computing acceleration device.

[0017] Compared with related technologies, the computing acceleration device, method and chip provided by the present disclosure have the following characteristics: the instruction decoding module decodes the input instruction, determines the bit width of the instruction based on the decoding result, generates a control signal containing the bit width information of the instruction and sends it to the bit width selection module; the bit width selection module obtains the bit width information of the instruction from the control signal, forwards the control signal to the small bit width processing module when the bit width of the instruction is small, and forwards the control signal to the large bit width processing module when the bit width of the instruction is large; the small bit width processing module uses the small bit width register stack and the corresponding arithmetic logic unit to process the instruction; the large bit width processing module uses the large bit width register stack and the corresponding arithmetic logic unit to process the instruction. The above computing acceleration device, method and chip can realize the acceleration of large bit width operations in a small bit width processor architecture system.

[0018] Other features and advantages of the present disclosure will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present disclosure. Other advantages of the present disclosure can be realized and obtained through the solutions described in the description and the drawings.

[0019] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description.

[0020] Summary of the Figures

[0021] The accompanying drawings are used to provide an understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation to the technical solution of the present disclosure.

[0022] FIG1 is a schematic structural diagram of a computing acceleration device provided by an embodiment of the present disclosure;

[0023] FIG2 is a schematic structural diagram of a large bit-width processing module provided by an embodiment of the present disclosure;

[0024] FIG3 is a schematic diagram of partitioning of a large-bit-width register file provided by an embodiment of the present disclosure;

[0025] FIG4 is a schematic diagram of a first preset bit width register group provided by an embodiment of the present disclosure;

[0026] FIG5 is a schematic diagram of a second preset bit width register group provided by an embodiment of the present disclosure;

[0027] FIG6 is a schematic structural diagram of a small bit-width processing module provided by an embodiment of the present disclosure;

[0028] FIG7 is a flow chart of a calculation acceleration method provided by an embodiment of the present disclosure;

[0029] FIG8 is a schematic structural diagram of a chip provided by an embodiment of the present disclosure;

[0030] FIG9 is a schematic structural diagram of a computing acceleration device based on a 64-bit RISC-V core provided by an embodiment of the present disclosure;

[0031] FIG10 is a schematic diagram of the distribution of a 128-bit basic register file provided by an embodiment of the present disclosure;

[0032] FIG11 is a schematic diagram of the structure of a 256-bit register file provided in an embodiment of the present disclosure;

[0033] FIG12 is a schematic diagram of the structure of a 384-bit register file provided in an embodiment of the present disclosure;

[0034] FIG13 is a schematic diagram of basic registers and ALU data flows from a register stack perspective provided by an embodiment of the present disclosure.

[0035] Details

[0036] The present disclosure describes a plurality of embodiments, but this description is exemplary rather than restrictive, and it will be apparent to those skilled in the art that there may be more embodiments and implementations within the scope of the embodiments described in the present disclosure. Although many possible feature combinations are shown in the drawings and discussed in the detailed description, many other combinations of the disclosed features are also possible. Unless specifically limited, any feature or element of any embodiment may be used in combination with any other feature or element in any other embodiment, or may replace any other feature or element in any other embodiment.

[0037] The present disclosure includes and contemplates combinations of features and elements known to those of ordinary skill in the art. The disclosed embodiments, features, and elements of the present disclosure may also be combined with any conventional features or elements to form a unique inventive solution defined by the claims. Any features or elements of any embodiment may also be combined with features or elements from other inventive solutions to form another unique inventive solution defined by the claims. Therefore, it should be understood that any feature shown and / or discussed in this disclosure may be implemented individually or in any appropriate combination. Therefore, the embodiments are not subject to other limitations except for the limitations set forth in the appended claims and their equivalents. In addition, various modifications and changes may be made within the scope of protection of the appended claims.

[0038] In addition, when describing representative embodiments, the specification may have presented the method and / or process as a specific sequence of steps. However, to the extent that the method or process does not rely on the specific order of the steps described herein, the method or process should not be limited to the steps in the specific order described. As will be understood by those skilled in the art, other orders of steps are also possible. Therefore, the specific order of the steps set forth in the specification should not be interpreted as limiting the claims. In addition, the claims to the method and / or process should not be limited to performing their steps in the order written, and those skilled in the art can readily understand that these orders can be changed without departing from the scope of this disclosure.

[0039] As shown in FIG1 , an embodiment of the present disclosure provides an operation acceleration device, comprising: an instruction decoding module 10 , a bit width selection module 20 , a small bit width processing module 30 , and a large bit width processing module 40 ;

[0040] an instruction decoding module configured to decode an input instruction, determine the bit width of the instruction according to the decoding result, generate a control signal containing the bit width information of the instruction, and send the control signal to the bit width selection module;

[0041] a bit width selection module configured to obtain bit width information of the instruction from the control signal, forward the control signal to the small bit width processing module when the bit width of the instruction is small, and forward the control signal to the large bit width processing module when the bit width of the instruction is large; wherein the small bit width means that the data bit width is less than or equal to the first threshold, and the large bit width means that the data bit width is greater than the first threshold;

[0042] A small bit-width processing module configured to process instructions using a small bit-width register file and a corresponding arithmetic and logic unit (ALU);

[0043] The large bit width processing module is configured to process instructions using a large bit width register file and a corresponding arithmetic logic unit.

[0044] The computing acceleration device provided by the embodiment of the present application comprises an instruction decoding module that decodes an input instruction, determines the bit width of the instruction based on the decoding result, generates a control signal containing the bit width information of the instruction and sends it to the bit width selection module; the bit width selection module obtains the bit width information of the instruction from the control signal, forwards the control signal to the small bit width processing module when the bit width of the instruction is small, and forwards the control signal to the large bit width processing module when the bit width of the instruction is large; the small bit width processing module processes the instruction using a small bit width register stack and a corresponding arithmetic logic unit; the large bit width processing module processes the instruction using a large bit width register stack and a corresponding arithmetic logic unit. The above computing acceleration device can realize accelerated computing of large bit width computing in a small bit width processor architecture system.

[0045] In an exemplary embodiment, the first threshold is 64 bits. In other embodiments, the first threshold may also be other values.

[0046] In an exemplary embodiment, as shown in FIG2 , the large bit-width processing module includes: a large bit-width register file 401 , a dual bit-width register file access unit 403 , a first large bit-width arithmetic logic unit 405 , and a second large bit-width arithmetic logic unit 407 ;

[0047] a dual-bit-width register file access unit configured to, when the bit width of the instruction is a first preset bit width, determine, based on register information included in the instruction, a position of a first target register segment that the instruction needs to access in the register file of the large bit width, and write or read data of the first preset bit width to the first target register segment; and, when the bit width of the instruction is a second preset bit width, determine, based on the register information included in the instruction, a position of a second target register segment that the instruction needs to access in the register file of the large bit width, and write or read data of the second preset bit width to the second target register segment;

[0048] a first large bit width arithmetic logic unit, configured to perform arithmetic logic operations on data of a first preset bit width;

[0049] The second large bit width arithmetic logic unit is configured to perform arithmetic logic operations on data of a second preset bit width.

[0050] In an exemplary embodiment, the first preset bit width is 256 bits, and the second preset bit width is 384 bits. In other embodiments, the first preset bit width and the second preset bit width may also be other values.

[0051] In an exemplary embodiment, the large bit-width register file includes N basic registers of a third preset bit width; the value of the third preset bit width is a common divisor of the value of the first preset bit width and the value of the second preset bit width; and N is a positive integer greater than 1.

[0052] In an exemplary embodiment, the first preset bit width is 256 bits, the second preset bit width is 384 bits, and the third preset bit width is 128 bits.

[0053] In an exemplary embodiment, as shown in FIG3 , the large-bit-width register file is divided into a first dedicated storage area, a common storage area, and a second dedicated storage area in ascending order of register addresses;

[0054] The first dedicated storage area is configured to store predetermined parameters of a first preset bit width; the second dedicated storage area is configured to store predetermined parameters of a second preset bit width; and the public storage area is configured to store data of the first preset bit width and data of the second preset bit width.

[0055] By dividing a large-bit-width register file into a dedicated first and second dedicated storage areas, commonly used calculation parameters can be pre-loaded into the corresponding fixed registers when the program starts. This ensures that parameters with the first preset bit width will not be corrupted by instructions with the second preset bit width, and vice versa. This improves calculation speed while ensuring reliability.

[0056] In an exemplary embodiment, the predetermined parameter of the first preset bit width and the predetermined parameter of the second preset bit width are constant parameters used in zero-knowledge proof related protocols, such as some constant parameters in elliptic curve operations.

[0057] Take the new zk-SNARK elliptic curve encryption algorithm BLS12-381 used in the field of zero-knowledge proof as an example: commonly used constant parameters include: the large prime number P corresponding to the basic field of the elliptic curve, the parameters R and INV in the Montgomery large number modular multiplication algorithm, and the constant R2 in the extended Euclidean algorithm.

[0058] P=0x1a0111ea397fe69a4b1ba7b6434bacd764774b84f38512bf6730d2a0f6b0f6241eabfffeb153ffffb9fefffffffffaaab;

[0059] R=2^256%P

[0060] =0x15f65ec3fa80e4935c071a97a256ec6d77ce5853705257455f48985753c758baebf4000bc40c0002760900000002fffd;

[0061] INV=P^-1%2^256

[0062] =0xceb06106feaafc9468b316fee268cf5819ecca0e8eb2db4c16ef2ef0c8e30b48286adb92d9d113e889f3ffffffffcfffd;

[0063] R2=R^2%P

[0064] =0x11988fe592cae3aa9a793e85b519952d67eb88a9939d83c08de5476c4c95b6d50a76e6a609d104f1f4df1f341c341746;

[0065] Among them, “^” is the exponentiation operator and “%” is the remainder operator.

[0066] In an exemplary embodiment, a dual-width register file access unit is configured to determine a position of a first target register segment to be accessed by the instruction in a large-width register file based on register information included in the instruction in the following manner: determining a register of a first preset bit width to be accessed based on a register identifier included in the instruction, determining a base register segment corresponding to the register of the first preset bit width to be accessed based on a composition rule of a register group of the first preset bit width, and using the base register segment as the first target register segment to be accessed by the instruction;

[0067] Among them, the first preset bit width register group is set in the first dedicated storage area and the public storage area in the large bit width register stack, including multiple registers of the first preset bit width, each register of the first preset bit width is composed of n1 basic registers of the third preset bit width; n1=W1 / W3; W1 is the value of the first preset bit width, and W3 is the value of the third preset bit width.

[0068] In an exemplary embodiment, a dual-width register file access unit is configured to determine a position of a second target register segment to be accessed by the instruction in a large-width register file based on register information included in the instruction in the following manner: determining a register of a second preset bit width to be accessed based on a register identifier included in the instruction, determining a base register segment corresponding to the register of the second preset bit width to be accessed based on a composition rule of a register group of the second preset bit width, and using the base register segment as a first target register segment to be accessed by the instruction;

[0069] Among them, the second preset bit width register group is set in the second dedicated storage area and the public storage area in the large bit width register stack, including multiple registers of the second preset bit width, each register of the second preset bit width is composed of n2 basic registers of the third preset bit width; n2=W2 / W3; W2 is the value of the second preset bit width, and W3 is the value of the third preset bit width.

[0070] In an exemplary embodiment, as shown in FIG4 , the first preset bit width register group includes a1 dedicated registers R1_i of the first preset bit width and a2 general registers R1_j of the first preset bit width, whose register addresses are arranged consecutively from small to large;

[0071] The dedicated register of the first preset bit width is set in a first dedicated storage area in a register file with a large bit width;

[0072] The general register of the first preset bit width is set in a common storage area of ​​a register file with a large bit width; a1=M1 / n1; a2=M2 / n1;

[0073] Wherein, M1 is the total number of basic registers included in the first dedicated storage area in the large bit-width register file, and M2 is the total number of basic registers included in the common storage area in the large bit-width register file; 0≤i≤a1-1; a1≤j≤a1+a2-1.

[0074] In an exemplary embodiment, as shown in FIG5 , the second preset bit width register group includes b1 second preset bit width dedicated registers R2_k and b2 second preset bit width general registers R2_t, whose register addresses are arranged consecutively from large to small;

[0075] The dedicated register of the second preset bit width is set in a second dedicated storage area in a register file with a large bit width;

[0076] The general register of the second preset bit width is set in the public storage area of ​​the register file with a large bit width; b1 = M3 / n2; b2 = M2 / n2;

[0077] Wherein, M3 is the total number of basic registers included in the second dedicated storage area in the large bit-width register file, M2 is the total number of basic registers included in the public storage area in the large bit-width register file; 0≤k≤b1-1; b1≤t≤b1+b2-1.

[0078] In an exemplary embodiment, as shown in FIG6 , the small bit-width processing module includes: a small bit-width register file 301 , a small bit-width register file access unit 303 , and a small bit-width arithmetic logic unit 305 ;

[0079] The small bit-width register file includes a plurality of registers of a fourth preset bit-width;

[0080] a small bit-width register file access unit, configured to determine, based on register information included in the instruction, a position of a target register to be accessed by the instruction in the small bit-width register file, and write or read data of a fourth preset bit-width to the target register;

[0081] a small bit-width arithmetic logic unit, configured to perform arithmetic logic operations on data of a fourth preset bit-width;

[0082] The fourth preset bit width is less than or equal to the first threshold.

[0083] In an exemplary embodiment, the fourth preset bit width is 64 bits. The RISC-V architecture processor can support 64 bits.

[0084] In an exemplary embodiment, the instruction decoding module is configured to determine the bit width of the instruction according to the decoding result in the following manner: determining the bit width of the instruction according to a value of a predetermined bit segment of the instruction opcode.

[0085] In an exemplary embodiment, the predetermined bit segment includes two bits.

[0086] In an exemplary embodiment, the bit width of the instruction includes: a first preset bit width, a second preset bit width and a fourth preset bit width; wherein the first preset bit width and the second preset bit width are large bit widths, and the fourth preset bit width is a small bit width.

[0087] For example, the instruction set can include instructions with three bit widths: 64-bit, 256-bit, and 384-bit. Taking RISC-V as an example, while general 64-bit instructions remain unchanged, new 256-bit and 384-bit dedicated acceleration instructions are added. For example, an opcode (operation code) of 0001011 represents a 64-bit addition instruction, an opcode of 0101011 represents a 256-bit addition instruction, and an opcode of 1001011 represents a 384-bit addition instruction. Therefore, the instruction bit width can be determined based on the special bit field (bits 5 and 6) of the opcode. An opcode with a special bit field of "00" represents a 64-bit instruction, a special bit field of "01" represents a 256-bit instruction, and a special bit field of "10" represents a 384-bit instruction.

[0088] As shown in FIG7 , an embodiment of the present disclosure provides a method for accelerating computing, including:

[0089] Step S10: The instruction decoding module decodes the input instruction, determines the bit width of the instruction according to the decoding result, generates a control signal containing the bit width information of the instruction, and sends the control signal to the bit width selection module;

[0090] Step S20: The bit width selection module obtains the bit width information of the instruction from the control signal, and forwards the control signal to the small bit width processing module when the bit width of the instruction is small, and forwards the control signal to the large bit width processing module when the bit width of the instruction is large; wherein the small bit width means that the data bit width is less than or equal to the first threshold, and the large bit width means that the data bit width is greater than the first threshold;

[0091] In step S30 , the small bit-width processing module processes instructions using a small bit-width register file and a corresponding arithmetic logic unit; the large bit-width processing module processes instructions using a large bit-width register file and a corresponding arithmetic logic unit.

[0092] The operation acceleration method provided by the embodiment of the present application is as follows: the instruction decoding module decodes the input instruction, determines the bit width of the instruction according to the decoding result, generates a control signal containing the bit width information of the instruction and sends it to the bit width selection module; the bit width selection module obtains the bit width information of the instruction from the control signal, forwards the control signal to the small bit width processing module when the bit width of the instruction is small, and forwards the control signal to the large bit width processing module when the bit width of the instruction is large; the small bit width processing module uses the small bit width register stack and the corresponding arithmetic logic unit to process the instruction; the large bit width processing module uses the large bit width register stack and the corresponding arithmetic logic unit to process the instruction. The above operation acceleration method can realize the acceleration of large bit width operations in a small bit width processor architecture system.

[0093] In an exemplary embodiment, the large bit-width processing module includes: a large bit-width register file, a dual bit-width register file access unit, a first large bit-width arithmetic logic unit, and a second large bit-width arithmetic logic unit;

[0094] The large bit width processing module utilizes a large bit width register file and a corresponding arithmetic logic unit to process instructions, including:

[0095] When the dual-width register file access unit determines that the bit width of the instruction is a first preset bit width, the dual-width register file access unit determines, based on the register information included in the instruction, a position of a first target register segment that the instruction needs to access in the register file with a large bit width, and writes or reads data of the first preset bit width to or from the first target register segment; when the dual-width register file access unit determines that the bit width of the instruction is a second preset bit width, the dual-width register file access unit determines, based on the register information included in the instruction, a position of a second target register segment that the instruction needs to access in the register file with a large bit width, and writes or reads data of the second preset bit width to or from the second target register segment;

[0096] The first large-bit-width arithmetic logic unit performs arithmetic logic operations on data of a first preset bit width; and the second large-bit-width arithmetic logic unit performs arithmetic logic operations on data of a second preset bit width.

[0097] In an exemplary embodiment, the large bit-width register file includes N basic registers of a third preset bit width; the value of the third preset bit width is a common divisor of the value of the first preset bit width and the value of the second preset bit width; and N is a positive integer greater than 1.

[0098] In an exemplary embodiment, the large-bit-width register file is divided into a first dedicated storage area, a public storage area, and a second dedicated storage area in ascending order of register addresses;

[0099] The first dedicated storage area is configured to store predetermined parameters of a first preset bit width; the second dedicated storage area is configured to store predetermined parameters of a second preset bit width; and the public storage area is configured to store data of the first preset bit width and data of the second preset bit width.

[0100] In an exemplary embodiment, the dual-width register file access unit determines, based on register information included in the instruction, a position of a first target register segment that the instruction needs to access in the large-width register file, including:

[0101] The dual-width register file access unit determines a register of the first preset bit width that needs to be accessed based on a register identifier included in the instruction, determines a base register segment corresponding to the register of the first preset bit width that needs to be accessed based on a composition rule of a register group of the first preset bit width, and uses the base register segment as a first target register segment that needs to be accessed by the instruction;

[0102] Among them, the first preset bit width register group is set in the first dedicated storage area and the public storage area in the large bit width register stack, including multiple registers of the first preset bit width, each register of the first preset bit width is composed of n1 basic registers of the third preset bit width; n1=W1 / W3; W1 is the value of the first preset bit width, and W3 is the value of the third preset bit width.

[0103] In an exemplary embodiment, the dual-width register file access unit determines, based on register information included in the instruction, a position of a second target register segment that the instruction needs to access in the large-width register file, including:

[0104] The dual-width register file access unit determines a register of the second preset bit width that needs to be accessed based on a register identifier included in the instruction, determines a base register segment corresponding to the register of the second preset bit width that needs to be accessed based on a composition rule of a register group of the second preset bit width, and uses the base register segment as a first target register segment that needs to be accessed by the instruction;

[0105] Among them, the second preset bit width register group is set in the second dedicated storage area and the public storage area in the large bit width register stack, including multiple registers of the second preset bit width, each register of the second preset bit width is composed of n2 basic registers of the third preset bit width; n2=W2 / W3; W2 is the value of the second preset bit width, and W3 is the value of the third preset bit width.

[0106] In an exemplary embodiment, the first preset bit width register group includes a1 dedicated registers R1_i of the first preset bit width and a2 general registers R1_j of the first preset bit width, whose register addresses are arranged consecutively from small to large;

[0107] The dedicated register of the first preset bit width is set in a first dedicated storage area in a register file with a large bit width;

[0108] The general register of the first preset bit width is set in a common storage area of ​​a register file with a large bit width; a1=M1 / n1; a2=M2 / n1;

[0109] Wherein, M1 is the total number of basic registers included in the first dedicated storage area in the large bit-width register file, and M2 is the total number of basic registers included in the common storage area in the large bit-width register file; 0≤i≤a1-1; a1≤j≤a1+a2-1.

[0110] In an exemplary embodiment, the second preset bit width register group includes b1 second preset bit width dedicated registers R2_k and b2 second preset bit width general registers R2_t, whose register addresses are arranged consecutively from large to small;

[0111] The dedicated register of the second preset bit width is set in a second dedicated storage area in a register file with a large bit width;

[0112] The general register of the second preset bit width is set in the public storage area of ​​the register file with a large bit width; b1 = M3 / n2; b2 = M2 / n2;

[0113] Wherein, M3 is the total number of basic registers included in the second dedicated storage area in the large bit-width register file, M2 is the total number of basic registers included in the public storage area in the large bit-width register file; 0≤k≤b1-1; b1≤t≤b1+b2-1.

[0114] In an exemplary embodiment, the small bit-width processing module includes: a small bit-width register file, a small bit-width register file access unit, and a small bit-width arithmetic logic unit;

[0115] The small bit-width register file includes a plurality of registers of a fourth preset bit-width;

[0116] The small bit-width processing module uses a small bit-width register file and the corresponding arithmetic logic unit to process instructions, including:

[0117] The small-bit-width register file access unit determines, based on register information included in the instruction, a position of a target register to be accessed by the instruction in the small-bit-width register file, and writes or reads data of a fourth preset bit width to the target register;

[0118] The small bit-width arithmetic logic unit performs arithmetic logic operations on data of a fourth preset bit-width;

[0119] The fourth preset bit width is less than or equal to the first threshold.

[0120] In an exemplary embodiment, the instruction decoding module determines the bit width of the instruction according to the decoding result, including:

[0121] The instruction decoding module determines the bit width of the instruction according to the value of the predetermined bit segment of the instruction operation code.

[0122] As shown in FIG8 , an embodiment of the present disclosure provides a chip including the computing acceleration device 100 described above.

[0123] Large-bitwidth operations use a unified register file to support both 256-bit and 384-bit bit widths. After the bit-width selection module determines the bit width corresponding to the instruction, it connects the corresponding register segment in the register file to the arithmetic logic unit (ALU) component of the corresponding bit width, and the result is updated via the data path to the corresponding register segment.

[0124] Figure 9 shows a schematic diagram of the structure of an operation acceleration device based on a 64-bit RISC-V core. The characteristics of the operation acceleration device are the large-bit-width accelerator part and the bit-width selection module behind the instruction decoding module. The large-bit-width accelerator part uses a unified large-bit-width register file (128 bits) to support two bit widths of 256 bits and 384 bits. After the bit-width selection module determines the bit width corresponding to the instruction, the data of the register segment in the large-bit-width register file needs to be connected to the ALU of the corresponding bit width through the first data path, and the result of the ALU operation is stored in the corresponding register segment via the first data path.

[0125] The large-bit-width register file is physically composed of 92 128-bit basic registers. Figure 10 shows a schematic diagram of the distribution of the 128-bit basic register file. Among them, Reg00 to Reg91 are 92 128-bit basic registers. In actual use, instructions of different bit widths correspond to physically different basic register segments. The large-bit-width accelerator is partially aimed at the calculation of zero-knowledge proof-related protocols in the field of cryptography, and mainly provides acceleration for the operation of elliptic curve points under the Galois field. For some constant parameters required for the calculation of zero-knowledge proof-related protocols, the consumption of memory access during the calculation process can be reduced by dividing dedicated registers in the large-bit-width register file to store constant parameters. As shown in Figure 10, in the large-bit-width register file, 8 128-bit registers (Reg00 to Reg07) are divided near the 256-bit ALU side to form 4 256-bit registers, which are used to store constant parameters used in the 256-bit Galois field calculation process. The side close to the 384-bit ALU is divided into 12 128-bit registers (Reg80 to Reg91), forming four 384-bit registers, which are used to store parameters commonly used in 384-bit Galois field calculations.

[0126] Figure 11 shows a schematic diagram of the structure of a 256-bit register file. The wide-bit-width register file is logically divided into 4+36 256-bit registers. Each 256-bit register is composed of two 128-bit basic registers. Among them, four 256-bit registers (R2_00 to R2_03) are special registers used to store constant parameters used in the calculation process of 256-bit instructions. Thirty-six 256-bit registers (R2_04 to R2_39) are general-purpose registers.

[0127] Taking the 256-bit addition instruction "ADD2R2_04,R2_05,R2_00" as an example, the execution process is as follows: the instruction decoder generates the opcode OPCODE = "0101011" based on "ADD2." The bit-width selector determines that this is a 256-bit instruction based on the 5th and 6th bits of the OPCODE being "01." Based on this, the registers involved in the instruction are mapped to the physical register segment (base register segment) of the large-bit-width register file. Destination register R2_04 corresponds to Reg08-Reg09, and source registers R2_05 and R2_00 correspond to Reg10-Reg11 and Reg00-Reg01, respectively. This generates the corresponding control signals, connecting the two double-128-bit data, Reg10-Reg11 and Reg00-Reg01, to the 256-bit adder (256-bit ALU) for calculation and storing the result in the physical register segment Reg08-Reg09.

[0128] Figure 12 shows a schematic diagram of the structure of a 384-bit register file. The wide-bit-width register file is logically divided into 4+24 384-bit registers. Each 384-bit register is composed of three 128-bit basic registers. Among them, four 384-bit registers (R3_00 to R3_03) are special-purpose registers used to store constant parameters used in the calculation process of 384-bit instructions. Twenty-four 384-bit registers (R3_04 to R3_27) are general-purpose registers.

[0129] Taking the 384-bit addition instruction "ADD3 R3_06, R3_02, R3_07" as an example, the execution process is as follows: the instruction decoder generates the opcode OPCODE = "1001011" based on "ADD3." The bit-width selector determines that this is a 384-bit instruction based on the fact that bits 5 and 6 of the OPCODE are "10." Based on this, the registers involved in the instruction are mapped to the physical register segment (base register segment) of the large-bit-width register file. Destination register R3_06 corresponds to Reg71-Reg73, and source registers R3_02 and R3_07 correspond to Reg83-Reg85 and Reg68-Reg70, respectively. This generates the corresponding control signals, connecting the triple-128-bit data, Reg83-Reg85 and Reg68-Reg70, to the 384-bit adder (384-bit ALU) for calculation, and storing the result in the physical register segment Reg71-Reg73.

[0130] Logically, 256-bit instructions acquire, calculate, and store corresponding input and output data based on the identifiers of 256-bit registers, while 384-bit instructions acquire, calculate, and store corresponding input and output data based on the identifiers of 384-bit registers. Logically, 256-bit and 384-bit registers can be mapped to physical 128-bit registers, and the 256-bit and 384-bit ALUs acquire and store data from the physical 128-bit base registers based on this mapping. Figure 13 shows a schematic diagram of the base register and ALU data flow from a register file perspective. For the 256-bit instruction shown in Figure 13, a logical representation could be "ADD2 R2_19, R2_20, R2_21." For the 384-bit instruction shown in Figure 13, a logical representation could be "SUB3 R3_14, R3_14, R3_15."

[0131] The aforementioned 256-bit and 384-bit acceleration instructions are widely used in the field of zero-knowledge proofs. In common zero-knowledge proof protocols, multi-scalar multiplication (MSM) and number theoretic transform (NTT) require 256-bit and 384-bit acceleration instructions.

[0132] Number Theoretic Transforms (NTT) are typically performed in 256-bit memory, so using 256-bit acceleration modules can achieve better performance and efficiency. However, using 384-bit acceleration modules to calculate the 256-bit addition, subtraction, and multiplication operations in NTT wastes physical registers and ALU area, and reduces performance and efficiency. Therefore, when dealing with a large number of NTT operations, using only 384-bit acceleration modules will be less effective.

[0133] While multiple scalar multiplication (MSM) typically uses 256-bit coefficients, it's primarily used to compute scalar multiplications of 384-bit elliptic curve points. This translates to 384-bit modular multiplications, additions, and subtractions. Specifically, a typical elliptic curve point addition typically consists of approximately 10 modular multiplications and 10 additions (this number varies slightly depending on the implementation). If a 384-bit acceleration module is used, each modular multiplication can be converted into 3 multiplications and 2 additions, resulting in a 384-bit elliptic curve point addition requiring approximately 30 384-bit multiplications and 30 384-bit additions. However, if a 256-bit acceleration module is used, each modular multiplication is converted into 10 multiplications and 12 additions, resulting in a 384-bit elliptic curve point addition requiring approximately 100 256-bit multiplications and 130 256-bit additions. Although 256-bit addition and multiplication require fewer clock cycles and smaller area and power consumption than 384-bit, the number of multiplication instructions becomes 3.3 times that of 384-bit, and the number of addition instructions is 4.3 times that of 384-bit. The sudden increase in the number of instructions shows that using a 256-bit acceleration module to support a large number of 384-bit operations has relatively poor performance and efficiency.

[0134] In the field of zero-knowledge proof, both 256-bit and 384-bit accelerated operations are required. Therefore, the solution of the embodiment of the present application reuses a large-bit-width register stack, that is, 256-bit and 384-bit use the same physical register stack, but logically they can be mapped to the underlying physical registers according to the 256-bit register identifier and the 384-bit register identifier respectively. Taking the most common Montgomery modular multiplication in the most complex 384-bit elliptic curve cryptography system as an example, a general 64-bit RISC-V system requires about 1500 instructions to complete, while the dual-bit-width accelerator of the embodiment of the present application only needs 3 multiplication instructions and 2 addition instructions to complete it, and the required clock cycle is also shortened from more than 1000 cycles of the general system to about 28 cycles. Taking the most commonly used modular multiplication among the most widely used 256-bit elliptic curve multi-scalar multiplication, number theory transformation, polynomial commitment, etc. as an example, the general 64-bit RISC-V system requires more than 600 instructions to complete, while the dual-bit width accelerator of the embodiment of the present application also only needs 3 multiplication instructions and 2 addition instructions to complete, and uses 256-bit registers and ALU components, without any waste in terms of efficiency.

[0135] The solution of the embodiment of the present application uses a dual-bit-width register architecture, a dual-bit-width ALU, and a multi-bit-width instruction set, so as to support various algorithms of the encryption system with the fastest speed and highest efficiency from the hardware perspective; dedicated registers for commonly used parameters are planned for 256-bit and 384-bit calculations respectively; for the 384-bit elliptic curve encryption algorithm, the calculation can be completed with the least number of instructions and the least number of cycles, thereby occupying the commanding heights of absolute performance; in a wider range of 256-bit encryption systems, it can have excellent performance and good versatility in the fields of elliptic curve multi-scalar multiplication, number theory transformation, polynomial commitment, etc.

[0136] It will be appreciated by those skilled in the art that the functional modules / units in the apparatus disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0137] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.

Claims

1. An operation acceleration device, comprising: an instruction decoding module, a bit width selection module, a small bit width processing module, and a large bit width processing module; The instruction decoding module is configured to decode the input instruction, determine the bit width of the instruction according to the decoding result, generate a control signal containing the bit width information of the instruction, and send it to the bit width selection module; The bit width selection module is configured to obtain the bit width information of the instruction from the control signal, forward the control signal to the small bit width processing module when the bit width of the instruction is a small bit width, and forward the control signal to the large bit width processing module when the bit width of the instruction is a large bit width; wherein, the small bit width means that the data bit width is less than or equal to the first threshold, and the large bit width means that the data bit width is greater than the first threshold; The small bit width processing module is configured to process the instruction by using a register bank with a small bit width and a corresponding arithmetic logic unit; The large bit width processing module is configured to process the instruction by using a register bank with a large bit width and a corresponding arithmetic logic unit.

2. The device according to claim 1, wherein: The large bit width processing module includes: a register bank with a large bit width, a double-bit width register bank access unit, a first large bit width arithmetic logic unit, and a second large bit width arithmetic logic unit; The double-bit width register bank access unit is configured to, when the bit width of the instruction is the first preset bit width, determine the position of the first target register segment to be accessed by the instruction in the register bank with a large bit width according to the register information included in the instruction, and write or read data with the first preset bit width to / from the first target register segment; when the bit width of the instruction is the second preset bit width, determine the position of the second target register segment to be accessed by the instruction in the register bank with a large bit width according to the register information included in the instruction, and write or read data with the second preset bit width to / from the second target register segment; The first large bit width arithmetic logic unit is configured to perform arithmetic logic operations on data with the first preset bit width; The second large bit width arithmetic logic unit is configured to perform arithmetic logic operations on data with the second preset bit width; The register bank with a large bit width includes N basic registers with a third preset bit width; the value of the third preset bit width is the greatest common divisor of the values of the first preset bit width and the second preset bit width; N is a positive integer greater than 1.

3. The device according to claim 2, wherein: The register bank with a large bit width is sequentially divided into a first dedicated storage area, a common storage area, and a second dedicated storage area in ascending order of register addresses; wherein, the first dedicated storage area is configured to store predetermined parameters with the first preset bit width; the second dedicated storage area is configured to store predetermined parameters with the second preset bit width; the common storage area is configured to store data with the first preset bit width and data with the second preset bit width.

4. The device according to claim 3, wherein: A double-width register file access unit is configured to determine the position of a first target register segment to be accessed by an instruction in a large-width register file according to the register information included in the instruction in the following manner: determining a register with a first preset width to be accessed according to the register identifier included in the instruction, determining a base register segment corresponding to the register with the first preset width according to the composition rule of the first preset-width register group, and using the base register segment as the first target register segment to be accessed by the instruction; Wherein, the first preset-width register group is arranged in a first dedicated storage area and a common storage area in the large-width register file, and includes a plurality of registers with the first preset width. Each register with the first preset width is composed of n1 base registers with a third preset width; n1 = W1 / W3; W1 is the value of the first preset width, and W3 is the value of the third preset width.

5. The apparatus according to claim 3, Wherein: A double-width register file access unit is configured to determine the position of a second target register segment to be accessed by an instruction in a large-width register file according to the register information included in the instruction in the following manner: determining a register with a second preset width to be accessed according to the register identifier included in the instruction, determining a base register segment corresponding to the register with the second preset width according to the composition rule of the second preset-width register group, and using the base register segment as the first target register segment to be accessed by the instruction; Wherein, the second preset-width register group is arranged in a second dedicated storage area and a common storage area in the large-width register file, and includes a plurality of registers with the second preset width. Each register with the second preset width is composed of n2 base registers with a third preset width; n2 = W2 / W3; W2 is the value of the second preset width, and W3 is the value of the third preset width.

6. The apparatus according to claim 4, Wherein: The first preset-width register group includes a1 first preset-width dedicated registers R1_i and a2 first preset-width general registers R1_j with consecutive register addresses in ascending order; The first preset-width dedicated registers are arranged in the first dedicated storage area in the large-width register file; The first preset-width general registers are arranged in the common storage area in the large-width register file; a1 = M1 / n1; a2 = M2 / n1; Wherein, M1 is the total number of base registers included in the first dedicated storage area in the large-width register file, M2 is the total number of base registers included in the common storage area in the large-width register file; 0 ≤ i ≤ a1 - 1; a1 ≤ j ≤ a1 + a2 - 1.

7. The apparatus according to claim 5, Wherein: The second preset-width register group includes b1 second preset-width dedicated registers R2_k and b2 second preset-width general registers R2_t with consecutive register addresses in descending order; The second preset-width dedicated registers are arranged in the second dedicated storage area in the large-width register file; The general register with the second preset bit width is set in the common storage area of the register bank with a large bit width; b1 = M3 / n2; b2 = M2 / n2; where M3 is the total number of base registers included in the second dedicated storage area in the register bank with a large bit width, M2 is the total number of base registers included in the common storage area in the register bank with a large bit width; 0 ≤ k ≤ b1 - 1; b1 ≤ t ≤ b1 + b2 - 1.

8. The apparatus according to claim 1, wherein: The small bit-width processing module includes: a register bank with a small bit width, a small bit-width register bank access unit, and a small bit-width arithmetic logic unit; The register bank with a small bit width includes multiple registers with a fourth preset bit width; The small bit-width register bank access unit is configured to determine the position of the target register to be accessed by the instruction in the register bank with a small bit width according to the register information included in the instruction, and write or read data with a fourth preset bit width to / from the target register; The small bit-width arithmetic logic unit is configured to perform arithmetic logic operations on the data with a fourth preset bit width; The instruction decoding module is configured to determine the bit width of the instruction according to the decoding result in the following manner: determine the bit width of the instruction according to the value of the preset bit segment of the instruction opcode; where the fourth preset bit width is less than or equal to the first threshold.

9. An operation acceleration method, including: The instruction decoding module decodes the input instruction, determines the bit width of the instruction according to the decoding result, generates a control signal including the bit width information of the instruction, and sends it to the bit width selection module; The bit width selection module obtains the bit width information of the instruction from the control signal, and forwards the control signal to the small bit-width processing module when the bit width of the instruction is a small bit width, and forwards the control signal to the large bit-width processing module when the bit width of the instruction is a large bit width; wherein, the small bit width means the data bit width is less than or equal to the first threshold, and the large bit width means the data bit width is greater than the first threshold; The small bit-width processing module processes the instruction by using the register bank with a small bit width and the corresponding arithmetic logic unit; the large bit-width processing module processes the instruction by using the register bank with a large bit width and the corresponding arithmetic logic unit.

10. A chip, including: The operation acceleration apparatus according to any one of claims 1 - 8.

Citation Information

Patent Citations

  • Arithmetic unit, method and device capable of supporting different bit width arithmetic data

    CN107688854A

  • RISC-V-based processor special for post-quantum cryptography algorithm

    CN116432765A

  • 8-bit RISC microcontroller with double arithmetic logic units

    CN1766834A

  • Asymmetric clustered processor architecture based on value content

    US20070294507A1

Cited By

  • Vector data processor, instruction processing method and system on chip

    CN121364892A

  • RISC-V processor-oriented finite field modular operation acceleration implementation method

    CN121966834A

  • Binary translation optimization method, binary translator and electronic equipment

    CN122018920A

  • Method for splicing and generating target SRAM (Static Random Access Memory), electronic equipment and medium

    CN122132328A

  • A method for splicing to generate target SRAM, an electronic device and a medium

    CN122132328B