Low-bit wide vector instruction extension method and device based on RISC-V

By introducing a low-bit-width vector instruction extension method in the RISC-V vector architecture, the problems of waste of computing resources and inefficient execution in the prior art are solved, and more efficient computing performance in DNN quantization inference tasks are achieved.

CN120029672APending Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510055622.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In deep neural network inference tasks, existing RISC-V vector extensions can only use high-bit width instructions to perform low-bit width operations, resulting in waste of computing resources and inefficient execution.

Method used

By introducing a low-bit width vector instruction extension method in the RISC-V vector architecture, it includes modifying the instructions to set the standard element width and vector register group multiplier, calculate the extended bit width of the low-bit width vector element, perform fixed-point number multiplication operation and intercept the accumulated result.

Benefits of technology

The RISC-V architecture is implemented to reduce the number of instructions, improve operational efficiency, and maintain low expansion overhead on workloads such as DNN quantization inference, and is compatible with the RISC-V vector architecture standard.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029672A_ABST
    Figure CN120029672A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer system structures, and discloses a low-bit wide vector instruction expansion method and device based on RISC-V. The method comprises the steps that before instruction calling, a modification instruction is used for conducting modification setting on the standard element width and a vector register block multiplier in an RISC-V vector architecture; calculating the bit width of the expanded low-bit-width vector element, and calculating the number of low-bit-width elements which can be contained in the standard element width; performing fixed-point number multiplication according to the standard element width to obtain an intermediate result; and intercepting and accumulating the intermediate result to obtain a calculation result of the corresponding instruction. According to the method, low bit width support is added for RISC-V vector extension, so that the number of instructions is reduced and the operation efficiency is improved on the workload of high parallelism and low data bit width, such as DNN quantitative reasoning, of an RISC-V framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer architecture technology, and in particular to a method and device for extending low-bit-width vector instructions based on RISC-V. Background Art

[0002] RISC-V is an open source instruction set architecture (ISA) that adopts the design concept of Reduced Instruction Set Computer (RISC). It gives priority to the implementation of frequently used simple instructions and avoids the use of complex instructions, thereby increasing the execution speed of instructions and making hardware implementation and optimization easier. The RISC-V instruction set architecture adopts a modular design, allowing users to customize and expand according to specific application requirements, so that it can flexibly respond to different application scenarios.

[0003] In the reasoning scenario of Deep Neural Network (DNN), there is no high requirement for the data bit width of weights and activations, so it is often combined with quantization technology to use low-bit-width weights or activations to reduce storage space and increase reasoning speed. Existing quantized models use 4-bit or even 2-bit low-bit-width fixed-point numbers.

[0004] The existing RISC-V vector extension supports a minimum vector operation bit width of 8 bits. When faced with deep neural network inference tasks, only high-bitwidth instructions can be used to perform low-bitwidth operations, resulting in a waste of computing resources and increasing the number of instructions required to complete the task, which seriously affects execution efficiency.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0006] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0007] The low-bit-width vector instruction extension method based on RISC-V provided in the embodiment of the present disclosure includes:

[0008] Use modification instructions to modify the standard element width and vector register group multiplier in the RISC-V vector architecture before the instruction call;

[0009] Calculate the extended bit width cw of the low bit width vector element;

[0010] Based on the extended bit width cw of the low-bit width vector elements, calculate the number of low-bit width elements that can be contained in the standard element width, and reverse the order of the low-bit width elements from operand 1 in one standard element width;

[0011] Expand and store the original low-bit width operand;

[0012] Perform fixed-point multiplication according to the standard element width to obtain intermediate results;

[0013] The intermediate results are intercepted and accumulated to obtain the calculation results of the corresponding instructions.

[0014] In some embodiments, the bit width cw of the low bit width vector element after expansion satisfies:

[0015]

[0016] Among them, wd2 and wd1 are the actual bit widths of the two input vector elements in the lbvmacc instruction, sew is the standard element width, and Indicates rounding up and rounding down respectively.

[0017] In some embodiments, the original low-bit width operand is expanded and stored, including:

[0018] For operand 2, it is extended to cw bits in sequence and stored in vector register vs2; for operand 1, it is extended and flipped and stored in vector register vs1.

[0019] In some embodiments, intercepting the intermediate result includes:

[0020] Take the cwth bit to 2cw-1th bit of the intermediate result, and calculate the multiplication and accumulation result of the corresponding low-bit width element group.

[0021] In some embodiments, the method comprises:

[0022] Based on the RISC-V vector instruction standard, it is extended. According to the RISC-V vector architecture standard, general vector registers are used to store operands and results, and vtype vector registers store control information.

[0023] In the RISC-V vector architecture standard, the number of bits of a vector register, VLEN, indicates the number of bits of 0 and 1 data contained in a vector register; the standard element width, sew, indicates the bit width of a vector element to be operated on, and this value is configured and modified through instructions in the RISC-V vector architecture; the vector register group multiplier, LMUL, enables one or more vector registers to form a vector register group, and the vector register group is used as a single operand of a vector instruction and is set through instructions; the number of vector instruction elements, VLMAX, indicates the maximum number of elements that can be operated using a single vector instruction given the current sew and LMUL settings, VLMAX = LMUL*VLEN / sew.

[0024] In some embodiments, the method further comprises defining an lbvmacc extension instruction based on the RISC-V vector architecture standard;

[0025] In the lbvmacc extended instruction, 6 consecutive bits are used to distinguish the lbvmacc instruction from other instructions;

[0026] Use 1 bit as the mask control bit; use 5 consecutive bits to define the vector register where the lbvmacc instruction operand 2 is located;

[0027] We also use 5 consecutive bits to define the vector register where operand 1 of the lbvmacc instruction is located;

[0028] Use 3 consecutive bits to define the data bit width actually operated in operand 2 of the lbvmacc instruction;

[0029] We also use 5 consecutive bits to define the vector register to which the result of the lbvmacc instruction is written back;

[0030] It also uses three consecutive bits to define the data bit width actually operated on in operand 1 of the lbvmacc instruction;

[0031] Use 4 consecutive bits to distinguish the lbvmacc instruction from other instructions.

[0032] In some embodiments, if the mask control bit is 1, it means that the calculation result of the instruction is controlled by the value of the mask register; if the mask control bit is 0, it means that the calculation result of the instruction is not controlled by the value of the mask register;

[0033] When using 5 consecutive bits to define the vector register where operand 2 of the lbvmacc instruction is located, when LMUL is greater than 1, the vector register sequence required for a single operand is doubled, and the low-width operands are filled first and then arranged in intervals in the vector register group according to specific rules;

[0034] When using three consecutive bits to define the data width actually performed in operand 1 of the lbvmacc instruction, a three-bit encoding is used to represent different data widths from 1 to 8. This value determines the size of the padding for operand 1 and the order of arrangement in the register.

[0035] The low-bit-width vector instruction expansion device based on RISC-V provided in the embodiment of the present disclosure includes:

[0036] A configuration module for modifying the standard element width and vector register group multiplier in the RISC-V vector architecture using a modification instruction before the instruction is called;

[0037] A processing module, used for calculating the bit width cw of the low bit width vector element after expansion;

[0038] The processing module is further used to calculate the number of low-bit width elements that can be included in the standard element width based on the bit width cw after the low-bit width vector element is expanded, and reverse the order of the low-bit width elements from operand 1 in a standard element width;

[0039] A storage module, used for expanding and storing the original low-bit width operands;

[0040] The processing module is further used to perform fixed-point multiplication operations according to the standard element width to obtain an intermediate result;

[0041] The processing module is further used to intercept and accumulate the intermediate results to obtain the calculation results of the corresponding instructions.

[0042] An embodiment of the present disclosure provides an electronic device, the device comprising:

[0043] at least one processor;

[0044] and a memory communicatively coupled to the at least one processor;

[0045] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute the low-bit-width vector instruction extension method based on RISC-V in the above embodiment.

[0046] An embodiment of the present disclosure provides a storage medium storing program instructions, which, when running, execute the low-bit-width vector instruction extension method based on RISC-V in the above embodiment.

[0047] The RISC-V-based low-bit-width vector instruction extension method, apparatus, device, and storage medium provided in the embodiments of the present disclosure can achieve the following technical effects:

[0048] The low-bitwidth support is added to the RISC-V vector extension, which reduces the number of instructions and improves the operating efficiency of the RISC-V architecture on high-parallelism and low-bitwidth workloads such as DNN quantized reasoning. The additional overhead brought by the extension is low, and the high-bitwidth fixed-point multiplication unit in the existing RISC-V vector extension is fully utilized to perform low-bitwidth multiplication and accumulation operations. Less additional control logic is required, and the original frequency, power consumption, and average instruction cycle are effectively maintained. This low-bitwidth extension is compatible with the existing RISC-V vector architecture, complies with the relevant provisions of the RISC-V vector architecture standard on standard element width and vector register group multiplier, and does not conflict with existing vector extension instructions.

[0049] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0051] Figure 1 It is a flowchart of a low-bit-width vector instruction expansion method based on RISC-V provided by an embodiment of the present disclosure;

[0052] Figure 2 is a schematic diagram of an lbvmacc instruction encoding format provided by an embodiment of the present disclosure;

[0053] Figure 3 is a schematic diagram of an lbvmacc instruction execution provided by an embodiment of the present disclosure;

[0054] Figure 4 It is a structural schematic diagram of a low-bit-width vector instruction extension device based on RISC-V provided in an embodiment of the present disclosure;

[0055] Figure 5 It is a block diagram of an exemplary electronic device provided by an embodiment of the present disclosure and capable of implementing an embodiment of the present disclosure. DETAILED DESCRIPTION

[0056] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0057] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0058] Unless otherwise stated, the term "plurality" means two or more.

[0059] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0060] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0061] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0062] The following describes the RISC-V-based low-bit-width vector instruction extension method, device, and storage medium provided by the embodiments of the present disclosure in conjunction with the accompanying drawings.

[0063] Figure 1 is a flow chart of a low-bit-width vector instruction expansion method based on RISC-V provided by an embodiment of the present disclosure, such as Figure 1 As shown, the low-bit-width vector instruction extension method based on RISC-V may include:

[0064] S01, using a modification instruction to modify and set the standard element width and vector register group multiplier in the RISC-V vector architecture before the instruction call;

[0065] S02, calculating the bit width cw of the low bit width vector element after expansion;

[0066] S03, based on the bit width cw after the low-bit-width vector element is expanded, calculating the number of low-bit-width elements that can be included in the standard element width, and reversing the order of the low-bit-width elements from operand 1 in one standard element width;

[0067] S04, expanding and storing the original low-bit width operand;

[0068] S05, performing fixed-point multiplication operation according to the standard element width to obtain an intermediate result;

[0069] S06, intercepting and accumulating the intermediate results to obtain the calculation results of the corresponding instructions.

[0070] In this disclosure, low bit width support is added to RISC-V vector extensions, which reduces the number of instructions and improves the operating efficiency of the RISC-V architecture on high parallelism and low data bit width workloads such as DNN quantization reasoning. It is possible to perform more low bit width vector data multiplication and accumulation operations with a single instruction at the cost of low extension overhead.

[0071] In some embodiments, the bit width cw of the low bit width vector element after expansion satisfies:

[0072]

[0073] Among them, wd2 and wd1 are the actual bit widths of the two input vector elements in the lbvmacc instruction, sew is the standard element width, and Indicates rounding up and rounding down respectively.

[0074] In some embodiments, the above-mentioned expansion and storage of the original low-bit width operand may include:

[0075] For operand 2, it is extended to cw bits in sequence and stored in vector register vs2; for operand 1, it is extended and flipped and stored in vector register vs1.

[0076] In some embodiments, intercepting the intermediate result may include:

[0077] Take the cwth bit to 2cw-1th bit of the intermediate result, and calculate the multiplication and accumulation result of the corresponding low-bit width element group.

[0078] In some embodiments, Figure 1 The method in can be extended based on the vector instruction standard of RISC-V. According to the RISC-V vector architecture standard, the general vector registers are used to store operands and results, and the vtype vector registers store control information;

[0079] In the RISC-V vector architecture standard, the number of bits of a vector register, VLEN, indicates the number of bits of 0 and 1 data contained in a vector register; the standard element width, sew, indicates the bit width of a vector element to be operated on, and this value is configured and modified through instructions in the RISC-V vector architecture; the vector register group multiplier, LMUL, enables one or more vector registers to form a vector register group, and the vector register group is used as a single operand of a vector instruction and is set through instructions; the number of vector instruction elements, VLMAX, indicates the maximum number of elements that can be operated using a single vector instruction given the current sew and LMUL settings, VLMAX = LMUL*VLEN / sew.

[0080] In some embodiments, Figure 1 The method in can define lbvmacc extension instructions based on the RISC-V vector architecture standard;

[0081] In the lbvmacc extended instruction, 6 consecutive bits are used to distinguish the lbvmacc instruction from other instructions;

[0082] Use 1 bit as the mask control bit; use 5 consecutive bits to define the vector register where the lbvmacc instruction operand 2 is located;

[0083] We also use 5 consecutive bits to define the vector register where operand 1 of the lbvmacc instruction is located;

[0084] Use 3 consecutive bits to define the data bit width actually operated in operand 2 of the lbvmacc instruction;

[0085] We also use 5 consecutive bits to define the vector register to which the result of the lbvmacc instruction is written back;

[0086] It also uses three consecutive bits to define the data bit width actually operated on in operand 1 of the lbvmacc instruction;

[0087] Use 4 consecutive bits to distinguish the lbvmacc instruction from other instructions.

[0088] In some embodiments, Figure 1 In the method, if the mask control bit is 1, it means that the calculation result of the instruction is controlled by the value of the mask register; if the mask control bit is 0, it means that the calculation result of the instruction is not controlled by the value of the mask register;

[0089] When using 5 consecutive bits to define the vector register where operand 2 of the lbvmacc instruction is located, when LMUL is greater than 1, the vector register sequence required for a single operand is doubled, and the low-width operands are filled first and then arranged in intervals in the vector register group according to specific rules;

[0090] When using three consecutive bits to define the data width actually performed in operand 1 of the lbvmacc instruction, a three-bit encoding is used to represent different data widths from 1 to 8. This value determines the size of the padding for operand 1 and the order of arrangement in the register.

[0091] Figure 2 is a schematic diagram of an lbvmacc instruction encoding format provided by an embodiment of the present disclosure, Figure 3 is a schematic diagram of lbvmacc instruction execution provided by an embodiment of the present disclosure, combined with Figure 2 and Figure 3 ,right Figure 1 The low-bit-width vector instruction extension method based on RISC-V is further described in.

[0092] In a specific example, based on the RISC V vector instruction standard, according to the RISC-V vector architecture standard, 32 general vector registers v0-v31 are used to store operands and results, and the vtype vector register stores control information. At the same time, there are the following two configurable concepts in the RISC-V vector architecture standard.

[0093] The number of bits of a vector register, VLEN, indicates the number of bits of 0 and 1 data contained in a vector register.

[0094] Standard element width (sew) represents the bit width of a vector element being operated on. In the RISC-V vector architecture, this value must be a power of 2 and greater than or equal to 8, and can be configured and modified through specific instructions.

[0095] The vector register group multiplier (LMUL) allows one or more vector registers to form a vector register group, which can be used as a single operand of a vector instruction and can be set through specific instructions.

[0096] The number of vector instruction elements (VLMAX) indicates the maximum number of elements that can be operated using a single vector instruction given the current sew and LMUL settings, VLMAX = LMUL*VLEN / sew.

[0097] Based on it, the lbvmacc instruction extension is defined. The encoding format of the extended instruction is as follows Figure 2 shown.

[0098] like Figure 2 As shown, (1) instruction can use 6 consecutive bits to distinguish lbvmacc instruction from other instructions, such as Figure 2 The func6 field at bits 26-31.

[0099] (2) Instructions can use 1 bit as a mask control bit. According to the RISC-V vector extension standard, if this bit is 1, it means that the calculation result of the instruction is controlled by the value of the mask register. If this bit is 0, it means that the calculation result of the instruction is not controlled by the value of the mask register. Figure 2 The 25-bit vm field.

[0100] (3) The instruction can use 5 consecutive bits to define the vector register where the lbvmacc instruction operand 2 is located. When LMUL is greater than 1, the vector register sequence required for a single operand is doubled. For example, when LMUL is 2 and the vector register encoded in the vs2 field is v3, it means that vector register v3 and vector register v4 together store operand 2 of the lbvmacc instruction. Low-width operands need to be padded first and then arranged in intervals within the vector register group according to specific rules. For example Figure 2 The vs2 field of bits 20-24.

[0101] (4) The instruction can use 5 consecutive bits to define the vector register where the operand 1 of the lbvmacc instruction is located. It is affected by LMUL as in (3). Figure 2 The vs1 field of bits 15-19.

[0102] (5) The instruction can use three consecutive bits to define the data width of the actual operation in operand 2 of the lbvmacc instruction. The three-bit encoding represents different data widths from 1 to 8. This value determines the size of the padding for operand 2 and the order of arrangement in the register. Figure 2 The wd2 field of bits 12-14.

[0103] (6) The instruction can use 5 consecutive bits to define the vector register to which the result of the lbvmacc instruction is written back, such as Figure 2 The vd field of bits 7 to 11.

[0104] (7) Instructions can use three consecutive bits to define the actual data width of operand 1 of the lbvmacc instruction. Using three bits to encode different data widths from 1 to 8 determines the size of the padding for operand 1 and the order of arrangement in the register. Figure 2 The wd1 field of bits 4-6.

[0105] (8) Instructions can use 4 consecutive bits to distinguish lbvmacc instructions from other instructions. Figure 2 Bits 0-3 of the func3 field.

[0106] Combination Figure 3 The example introduces the execution process of the instruction, where array A is a vector containing 4 elements, each vector element has a bit width of 3 bits, and is used as operand 2 in the instruction. Array B is a vector containing 4 elements, each vector has a bit width of 2 bits, and is used as operand 1 in the instruction. The result of the instruction operation is to calculate the product of the corresponding element positions of the two arrays and accumulate them, 3*1+6*0+5*3+7*2=32. The execution steps of the instruction are as follows:

[0107] S11, before the instruction is called, you can choose to use related instructions to modify and set the standard element width and vector register group multiplier in the RISC-V vector architecture. Figure 2 The standard element width is 16 bits, the register bank multiplier is 1, and the number of vector registers is 32 bits.

[0108] S12, calculate the bit width after the low-bit width vector element is expanded. Assume that the expanded bit width cw, to ensure that the result does not overflow after multiplication and accumulation of all low-bit width vector elements, cw should meet the following conditions:

[0109]

[0110] Among them, wd2 and wd1 are the actual bit widths of the two input vector elements in the lbvmacc instruction, sew is the standard element width, and Indicates rounding up and rounding down respectively.

[0111] While satisfying the above constraints, the value of cw should be as small as possible to avoid wasting computing resources.

[0112] Substitute specific values, such as Figure 2 , we can get a reasonable value of cw as 8.

[0113] S13, calculate the number of low-width elements that can be included in the standard element width And reverse the order of the low-width elements from operand 1 in a standard element width. Figure 2 , we can calculate that the number of low-width elements is 2, and reverse the order of the elements in array B to get B' and B".

[0114] S14, expand and store the original low-bit width operand. For operand 2, expand it to cw bits in sequence and store it in vector register vs2. For operand 1, expand it and flip it and store it in vector register vs1.

[0115] like Figure 2 , as shown, 3 and 6 of operand 2 are stored in bits 16 to 31 of vector register vs2, and the extended value is 774, 5 and 7 are stored in bits 0 to 15 of vector register vs1, and the extended value is 1287; 1 and 0 of operand 1 are stored in bits 16 to 31 of vector register vs1, and the extended value is 1, 3 and 2 are stored in bits 0 to 15 of vector register vs1, and the extended value is 515.

[0116] S15, performing fixed-point multiplication operation according to the standard element width to obtain an intermediate result.

[0117] like Figure 2 , the standard element width is 16, so the 16 to 31 bits of the vs2 register are multiplied by the 16 to 31 bits of the vs1 register, and the result is 774; the 0 to 15 bits of the vs2 register are multiplied by the 0 to 15 bits of the vs1 register, and the result is 662805;

[0118] S16, result truncation, for the result calculated in S15 according to the standard element width, take the cwth to 2cw-1th bits, and obtain the multiplication and accumulation result of the corresponding low-bit width element group.

[0119] like Figure 2 , take the 8th to 15th bits of the first intermediate result 774, and get the result of the multiplication and accumulation of the first two bits of array A and the first two bits of array B, which is 3; take the 8th to 15th bits of the second intermediate result 662805, and get the result of the multiplication and accumulation of the last two bits of array A and the last two bits of array B, which is 29.

[0120] S17, result accumulation, the result obtained in the previous step is accumulated again to obtain the final calculation result of the instruction, and is written back to the result register vd specified in the instruction.

[0121] like Figure 2 , the multiplication and accumulation results obtained in the previous step are 3 and 29, and the calculation result 32 of this instruction is accumulated and written back.

[0122] The instruction extension implementation device needs to include a memory, a processor, etc. For the processor, based on the processor based on the RISC-V vector extension, it additionally includes components for executing the value acquisition and result accumulation functions in S16 and S17.

[0123] In the present disclosure, low-bit-width support is added to the RISC-V vector extension, so that the RISC-V architecture reduces the number of instructions and improves the operating efficiency on high-parallelism and low-data-bit-width workloads such as DNN quantized reasoning. The additional overhead brought by the extension is low, and the high-bit-width fixed-point multiplication operation unit in the existing RISC-V vector extension is fully utilized to perform low-bit-width multiplication and accumulation operations, requiring less additional control logic to be added, and effectively maintaining the original frequency, power consumption, and average instruction cycle. The low-bit-width extension is compatible with the existing RISC-V vector architecture, complies with the relevant provisions of the RISC-V vector architecture standard on the standard element width and vector register group multiplier, and has no conflict with the existing vector extension instructions.

[0124] Figure 4 is a schematic diagram of the structure of a low-bit-width vector instruction expansion device based on RISC-V provided by an embodiment of the present disclosure, such as Figure 4 As shown, a low-bit-width vector instruction extension device based on RISC-V may include:

[0125] Configuration module 401, used for modifying and setting the standard element width and vector register group multiplier in the RISC-V vector architecture using modification instructions before instruction call;

[0126] A processing module 402 is used to calculate the bit width cw of the low bit width vector element after expansion;

[0127] The processing module 402 is further configured to calculate the number of low-bit-width elements that can be included in the standard element width based on the bit width cw after the low-bit-width vector element is expanded, and to reverse the order of the low-bit-width elements from operand 1 in a standard element width;

[0128] The storage module 403 is used to expand and store the original low-bit width operand;

[0129] The processing module 402 is further used to perform fixed-point multiplication operations according to the standard element width to obtain an intermediate result;

[0130] The processing module 402 is further used to intercept and accumulate the intermediate results to obtain the calculation results of the corresponding instructions.

[0131] In some embodiments, the bit width cw of the low bit width vector element after expansion satisfies:

[0132]

[0133] Among them, wd2 and wd1 are the actual bit widths of the two input vector elements in the lbvmacc instruction, sew is the standard element width, and Indicates rounding up and rounding down respectively.

[0134] In some embodiments, the original low-bit width operand is expanded and stored, including:

[0135] For operand 2, it is extended to cw bits in sequence and stored in vector register vs2; for operand 1, it is extended and flipped and stored in vector register vs1.

[0136] In some embodiments, intercepting the intermediate result includes:

[0137] Take the cwth bit to 2cw-1th bit of the intermediate result, and calculate the multiplication and accumulation result of the corresponding low-bit width element group.

[0138] In some embodiments, in a RISC-V based low bit width vector instruction extension device, the vector instruction standard based on RISC-V is extended, and according to the RISC-V vector architecture standard, the general vector registers are used to store operands and results, and the vtype vector registers store control information;

[0139] In the RISC-V vector architecture standard, the number of bits of a vector register, VLEN, indicates the number of bits of 0 and 1 data contained in a vector register; the standard element width, sew, indicates the bit width of a vector element to be operated on, and this value is configured and modified through instructions in the RISC-V vector architecture; the vector register group multiplier, LMUL, enables one or more vector registers to form a vector register group, and the vector register group is used as a single operand of a vector instruction and is set through instructions; the number of vector instruction elements, VLMAX, indicates the maximum number of elements that can be operated using a single vector instruction given the current sew and LMUL settings, VLMAX = LMUL*VLEN / sew.

[0140] In some embodiments, in a RISC-V based low bit width vector instruction extension device, a lbvmacc extension instruction is defined based on the RISC-V vector architecture standard;

[0141] In the lbvmacc extended instruction, 6 consecutive bits are used to distinguish the lbvmacc instruction from other instructions;

[0142] Use 1 bit as the mask control bit; use 5 consecutive bits to define the vector register where the lbvmacc instruction operand 2 is located;

[0143] We also use 5 consecutive bits to define the vector register where operand 1 of the lbvmacc instruction is located;

[0144] Use 3 consecutive bits to define the data bit width actually operated in operand 2 of the lbvmacc instruction;

[0145] We also use 5 consecutive bits to define the vector register to which the result of the lbvmacc instruction is written back;

[0146] It also uses three consecutive bits to define the data bit width actually operated on in operand 1 of the lbvmacc instruction;

[0147] Use 4 consecutive bits to distinguish the lbvmacc instruction from other instructions.

[0148] In some embodiments, if the mask control bit is 1, it means that the calculation result of the instruction is controlled by the value of the mask register; if the mask control bit is 0, it means that the calculation result of the instruction is not controlled by the value of the mask register;

[0149] When using 5 consecutive bits to define the vector register where operand 2 of the lbvmacc instruction is located, when LMUL is greater than 1, the vector register sequence required for a single operand is doubled, and the low-width operands are filled first and then arranged in intervals in the vector register group according to specific rules;

[0150] When using three consecutive bits to define the data width actually performed in operand 1 of the lbvmacc instruction, a three-bit encoding is used to represent different data widths from 1 to 8. This value determines the size of the padding for operand 1 and the order of arrangement in the register.

[0151] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0152] Figure 5 A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown.

[0153] Combination Figure 5As shown, an embodiment of the present disclosure provides a low-bit-width vector instruction extension device 500 based on RISC-V, including a processor 504 and a memory 501. Optionally, the device may also include a communication interface 502 and a bus 503. Among them, the processor 504, the communication interface 502, and the memory 501 can communicate with each other through the bus 503. The communication interface 502 can be used for information transmission. The processor 504 can call the logic instructions in the memory 501 to execute the low-bit-width vector instruction extension method based on RISC-V in the above embodiment.

[0154] In addition, the logic instructions in the memory 501 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0155] The memory 501, as a computer-readable storage medium, can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 504 executes the functional application and data processing by running the program instructions / modules stored in the memory 501, that is, implements the low-bit-width vector instruction extension method based on RISC-V in the above embodiment.

[0156] The memory 501 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 501 may include a high-speed random access memory and may also include a non-volatile memory.

[0157] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute a RISC-V-based low-bit-width vector instruction extension method.

[0158] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0159] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, and other media that can store program codes, or a transient storage medium.

[0160] The above description and the accompanying drawings sufficiently illustrate the embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Portions and features of some embodiments may be included in or replace portions and features of other embodiments. As used in the description of the embodiments, the singular forms "a", "an" and "one" shall be used interchangeably unless the context clearly indicates otherwise.

[0161] (an) and "(the) are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the presence of other identical elements in the process, method or device including the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0162] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and simplicity of description, the specific working process of the systems, devices and units described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0163] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0164] The flowchart and block diagram in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A low-bit-width vector instruction extension method based on RISC-V, characterized in that: The method comprises: Use modification instructions to modify the standard element width and vector register group multiplier in the RISC-V vector architecture before the instruction call; Calculate the extended bit width cw of the low bit width vector element; Based on the extended bit width cw of the low-bit width vector elements, calculate the number of low-bit width elements that can be contained in the standard element width, and reverse the order of the low-bit width elements from operand 1 in one standard element width; Expand and store the original low-bit width operand; Perform fixed-point multiplication according to the standard element width to obtain intermediate results; The intermediate results are intercepted and accumulated to obtain the calculation results of the corresponding instructions.

2. The method according to claim 1, characterized in that The extended bit width cw of the low bit width vector element satisfies: Among them, wd2 and wd1 are the actual bit widths of the two input vector elements in the lbvmacc instruction, sew is the standard element width, and Indicates rounding up and rounding down respectively.

3. The method according to claim 1, characterized in that The method of expanding and storing the original low-bit width operand includes: For operand 2, it is extended to cw bits in sequence and stored in vector register vs2; for operand 1, it is extended and flipped and stored in vector register vs1.

4. The method according to claim 1, characterized in that: The intercepting of the intermediate result comprises: Take the cwth bit to 2cw-1th bit of the intermediate result, and calculate the multiplication and accumulation result of the corresponding low-bit width element group.

5. The method according to claim 1, characterized in that: The method comprises: Based on the RISC-V vector instruction standard, it is extended. According to the RISC-V vector architecture standard, general vector registers are used to store operands and results, and vtype vector registers store control information. In the RISC-V vector architecture standard, the number of bits of a vector register VLEN indicates the number of bits of 0 and 1 data contained in a vector register; the standard element width sew indicates the bit width of a vector element to be operated on, and in the RISC-V vector architecture, the value of sew is configured and modified through instructions; the vector register group multiplier LMUL makes one or more vector registers form a vector register group, and the vector register group is used as a single operand of a vector instruction and set through instructions; the vector instruction element number VLMAX indicates the maximum number of elements that can be operated using a single vector instruction given the current sew and LMUL settings, VLMAX = LMUL*VLEN / sew.

6. The method according to claim 5, characterized in that The method further includes defining an lbvmacc extension instruction based on the RISC-V vector architecture standard; In the lbvmacc extended instruction, 6 consecutive bits are used to distinguish the lbvmacc instruction from other instructions; Use 1 bit as the mask control bit; use 5 consecutive bits to define the vector register where the lbvmacc instruction operand 2 is located; We also use 5 consecutive bits to define the vector register where operand 1 of the lbvmacc instruction is located; Use 3 consecutive bits to define the data bit width actually operated in operand 2 of the lbvmacc instruction; We also use 5 consecutive bits to define the vector register to which the result of the lbvmacc instruction is written back; It also uses three consecutive bits to define the data bit width actually operated on in operand 1 of the lbvmacc instruction; Use 4 consecutive bits to distinguish the lbvmacc instruction from other instructions.

7. The method according to claim 5, characterized in that If the mask control bit is 1, it means that the calculation result of the corresponding instruction is controlled by the value of the mask register; if the mask control bit is 0, it means that the calculation result of the corresponding instruction is not controlled by the value of the mask register; When using 5 consecutive bits to define the vector register where operand 2 of the lbvmacc instruction is located, when LMUL is greater than 1, the vector register sequence required for a single operand is doubled, and the low-width operands are filled first and then arranged in intervals in the vector register group according to specific rules; When using three consecutive bits to define the data width actually performed in operand 1 of the lbvmacc instruction, a three-bit encoding is used to represent different data widths from 1 to 8, which determines the size of the padding for operand 1 and the order of arrangement in the register.

8. A low-bit width vector instruction extension device based on RISC-V, characterized in that: The device comprises: A configuration module for modifying the standard element width and vector register group multiplier in the RISC-V vector architecture using a modification instruction before the instruction is called; A processing module, used for calculating the bit width cw of the low bit width vector element after expansion; The processing module is further used to calculate the number of low-bit width elements that can be included in the standard element width based on the bit width cw after the low-bit width vector element is expanded, and reverse the order of the low-bit width elements from operand 1 in a standard element width; A storage module, used for expanding and storing the original low-bit width operands; The processing module is further used to perform fixed-point multiplication operations according to the standard element width to obtain an intermediate result; The processing module is further used to intercept and accumulate the intermediate results to obtain the calculation results of the corresponding instructions.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; It is characterized in that the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the RISC-V based low-bit-width vector instruction extension method described in any one of claims 1 to 7.

10. A storage medium storing program instructions, characterized in that: When the program instructions are running, they execute the RISC-V based low-bit-width vector instruction extension method as described in any one of claims 1 to 7.