Instruction execution method and device, electronic equipment and readable storage medium
By using the first and second data paths in the operation array to perform data operations and updating the number of operations in the local control unit, the problems of low execution efficiency and high power consumption of multiple microinstructions are solved, and more efficient instruction execution and lower power consumption are achieved.
Patent Information
- Application Number
- CN202511178336.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
In the prior art, multiple microinstructions generated after instruction decoding have data dependencies when executed in a computing array, resulting in low execution efficiency and high power consumption.
In the operation array, the first data of the target microinstruction is obtained through the first data path, and a data operation operation is performed with it based on the second data path. The number of operations completed is updated using a local control unit. When the number of operations completed meets a preset condition, the execution result of the microinstruction is determined, thereby reducing the repeated transmission of data between operation units.
The performance of instruction execution is improved and power consumption is reduced. By reducing the number of times data is transferred between computing units, processing efficiency is improved and overall power consumption is reduced.
Smart Images

Figure CN120670030A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to an instruction execution method, device, electronic device, and readable storage medium. Background Art
[0002] In the instruction pipeline, instructions are decoded into microinstructions, and a single instruction may generate multiple microinstructions. For microinstructions with data dependencies, data interaction occurs during their execution within the computational array. For example, a single instruction may be split into multiple microinstructions, each of which defines a portion of its operands. The execution result of a single microinstruction can only be determined after the operands defined in that microinstruction are combined with the operands defined in each of the multiple microinstructions to perform data operations.
[0003] Therefore, how to achieve the execution of these multiple microinstructions with higher performance and lower power consumption has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] Embodiments of the present invention provide an instruction execution method, apparatus, electronic device, and readable storage medium, which can solve the problem of how to execute microinstructions with higher performance and lower power consumption.
[0005] In order to solve the above problems, an embodiment of the present invention discloses an instruction execution method applied to a computing array, the method comprising: For any operation unit in the operation array, when first data of a target microinstruction is obtained from a first data path, a data operation operation is performed on the first data based on second data input along a second data path each time, and the second data is sent to a next operation unit through the second data path; After performing the data operation once, the local control unit of the operation unit updates the number of times the first data has been operated once; In a case where the number of operations performed satisfies a preset number condition, the current result of the data operation is determined as the execution result of the target microinstruction.
[0006] On the other hand, an embodiment of the present invention discloses an instruction execution device, applied to a computing array, comprising: a first processing module configured to, for any operation unit in the operation array, perform a data operation on the first data of the target microinstruction obtained from the first data path based on second data input along the second data path each time, and send the second data to the next operation unit through the second data path; an updating module, configured to update, by the local control unit of the computing unit, the number of times the first data has been computed after performing the data computing operation once; The first determining module is configured to determine the current result of the data operation as the execution result of the target microinstruction when the number of operations performed satisfies a preset number condition.
[0007] On the other hand, an embodiment of the present invention discloses an electronic device, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the aforementioned method.
[0008] An embodiment of the present invention further discloses a readable storage medium having executable instructions stored thereon. When executed by one or more processors, the processors are enabled to execute the method described above.
[0009] The embodiment of the present invention includes the following advantages: for any operation unit in the operation array, when the first data of the target microinstruction is obtained from the first data path, a data operation operation is performed on the first data based on the second data input along the second data path each time, and the second data is sent to the next operation unit through the second data path. After performing a data operation operation, the local control unit of the operation unit updates the number of times the first data has been operated. When the number of operations meets the preset number condition, the current result of the data operation operation is determined as the execution result of the target microinstruction. In this way, the operation units can share the data required for processing the microinstruction by sequentially passing the second data downward based on the second data path. In this way, the operation of a GATHER instruction can be completed by obtaining the required data from the physical register stack only once. Therefore, the overall performance of the solution is better and the power consumption is lower. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0011] Figure 1 It is a data rearrangement schematic diagram in the prior art; Figure 2 This is a flowchart of a method for executing an instruction provided by an embodiment of the present invention; Figure 3 This is a schematic diagram of data transmission provided by an embodiment of the present invention; Figure 4 This is another data transmission schematic diagram provided by an embodiment of the present invention; Figure 5 is a schematic diagram of a computing array provided by an embodiment of the present invention; Figure 6 This is a schematic diagram of a processing process provided by an embodiment of the present invention; Figure 7 is another processing process schematic diagram provided by an embodiment of the present invention; Figure 8 is another processing process schematic diagram provided by an embodiment of the present invention; Figure 9 is another processing process schematic diagram provided by an embodiment of the present invention; Figure 10 is another processing process schematic diagram provided by an embodiment of the present invention; Figure 11 This is a processing diagram of a computing unit provided by an embodiment of the present invention; Figure 12 This is a schematic diagram of generating a selection signal provided by an embodiment of the present invention; Figure 13 This is a selection diagram provided by an embodiment of the present invention; Figure 14 is a block diagram of an instruction execution device provided by an embodiment of the present invention; Figure 15 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0012] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0013] First, the application scenarios involved in the embodiments of the present invention are described. Taking the permutation instruction as an example, when the on-chip network processes the instruction, decoding may generate multiple microinstructions (UOPs). Among them, the permutation instruction can also be called a data query instruction. The permutation instruction is an instruction used to rearrange vector data according to a rule. For example, vector instructions such as the GATHER instruction in the Vector Instruction Extension (RVV) instruction set are used to perform permutation operations on data in vector registers. The GATHER instruction can include the vrgather.vv instruction and the vrgatherei16.vv instruction, and their instruction formats are expressed as vrgather.vvvd, vs2, vs1 and vrgatherei16.vv vd, vs2, vs1 respectively.
[0014] Instructions in RVV use independent vector configuration instructions to set the vector configuration register (VTYPE CSR). The vector configuration register is a register defined in RVV that controls vector configuration. The VTYPE register includes the vector length multiplier (LMUL) field and the selected element width (SEW) field. LMUL and SEW are both 3 bits. SEW is used to specify the bit width of the operand of a vector instruction, and LMUL is used to specify the number of vector registers for the operand of a vector instruction. Vector registers are user-programmable registers defined in RVV, and RVV defines 32 vector registers for user use.
[0015] In the RVV instruction set, vector register numbers in a vector instruction use LMUL consecutively numbered vector registers. For example, in the instruction "vrgather.vv v24,v8,v16" with LMUL=8, the vs1 register group contains eight vector registers (v16 to v23), the vs2 register group contains eight vector registers (v8 to v15), and the vd register group contains eight vector registers (v24 to v31). For the GATHER instruction, the data in the vs1 register group can be considered an index vector, the data in the vs2 register group can be considered a data table vector, and the data in the vd register group can be considered a destination vector. The maximum number of elements represented by a single vector register number that a vector instruction can operate on is expressed as VLMAX. VLMAX = VLEN × LMUL / SEW, where VLEN represents the vector register length (in bits) and DLEN represents the datapath length (in bits) in a specific implementation. DLEN is the maximum bit width of data that the ALU can access from a datapath in a specific implementation, measured in bits. DLEN can be the same as or different from VLEN, but DLEN can always be multiplied or divided by VLEN. If VLEN / DLEN = n, then data from a vector register can be transferred through the datapath n times before reaching the ALU. DLEN typically determines the number of microinstructions that are generated by splitting a vector instruction. In most implementations, DLEN and VLEN are equal.
[0016] Specifically, the vrgather.vv instruction is a data reordering instruction in the vector extension for batch querying the values of index vectors in the data table vector. The bit width of a single index element in the vs1 register group of the vrgather.vv instruction, the bit width of a single data element provided by the vs2 register group, and the bit width of a single destination element provided by the vd register group are specified by SEW in VTYPE. For example, when SEW is 32, the bit widths of the index element, data element, and vd register group are all 32 bits, a vector register in the vs1 register group includes VLEN / SEW index elements, the number of index elements included in the vs1 register group is VLEN×LMUL / SEW, and all the index elements included in the vs1 register group constitute an index vector. A vector register in the vs2 register group includes VLEN / SEW data elements, the number of data elements included in the vs2 register group is VLEN×LMUL / SEW, and all the data elements included in the vs2 register group constitute a data table vector. A vector register in the vd register group includes VLEN / SEW data elements. The number of destination elements included in the vd register group is VLEN×LMUL / SEW. All destination elements included in the vd register group constitute a destination vector.
[0017] The vrgather.vv instruction is used to query the data elements in vs2 using the index elements in vs1, and store the query results in the corresponding position in vd and vs1. The elements in vd, vs2, and vs1 are split into several elements according to SEW. vd[i], vs2[i], and vs1[i] represent the i-th element in the vd register group, the i-th element in the vs2 register group, and the i-th element in the vs1 register group, respectively. If vs1[i]>=VLMAX, the value of vd[i] is set to 0. Otherwise, the value of vd[i] is set to vs2[vs1[i]].
[0018] Take vrgather.vv v6,v2,v4 with VLEN=128, SEW=32, LMUL=2 as an example. Figure 1 FIG. 1 is a schematic diagram of data rearrangement according to an embodiment of the present invention. Figure 1 As shown in the figure, the vs1 register group (v4 and v5) contains 8 index elements: vs1[0] to vs1[7], and the specific values of vs1[0] to vs1[7] are 1, 8, 7, 19, 2, 3, 0, and 5 respectively. The vs2 register group (v2 and v3) contains 8 data elements: vs2[0] to vs2[7], and the specific values of vs2[0] to vs2[7] are 10, 32, 54, 76, 98, BA, DC, and FE respectively. When this instruction is executed, each index element in v4 and v5 is used to find the element at the corresponding position in the v2 and v3 registers. In this example, since vs1[1] and vs1[3] meet the condition of >= VLMAX, the corresponding elements vd[1] and vd[3] in the destination register group are both set to 0. The positions of vd[0], vd[2], vd[4], vd[5], vd[6], and vd[7] are filled with the specific values of vs2[1], vs2[7], vs2[2], vs2[3], vs2[0], and vs2[5] respectively: 32, FE, 54, 76, 10, and BA.
[0019] The vrgatherei16.vv instruction is a special version of the vrgather.vv instruction. It fixes the bit width of the index element in the vs1 register group to 16 bits, and SEW is only used to specify the data bit width in the vs2 and vd register groups. For the vrgatherei16.vvv6,v2,v4 instruction with VLEN=128, SEW=32, and LMUL=2, according to the RISC-V specification, the number of registers operated by a register number must meet the principle of equal number of elements. The continuous vector registers represented by register number v6 contain a total of 8 elements, and the continuous vector registers represented by register number v4 contain 8 elements. In this example, v4 contains 128 / 16=8 elements. Therefore, register v4 includes vs1[0] to vs1[7], that is, the register group represented by register number v4 only includes register v4. The vrgatherei16.vv instruction is executed in the same way as the vrgather.vv instruction and will not be repeated here.
[0020] In combination with the above example, it can be seen that each index vector needs to be searched for LMUL indexed vs2 registers. Therefore, the larger the LMUL, the larger the query amount. In an embodiment of the present invention, for a GATHER instruction, multiple UOPs can be generated, and the number of index elements, data elements and destination elements defined in a UOP is the same. For example, the number of index elements, data elements and destination elements defined in a UOP are all the number of data elements included in a single vs2 register. For example, for the vrgather.vv v6,v2,v4 instruction in the above example, two UOPs0 and UOP1 can be generated. The operands defined in UOP0 are v4, v2 and v6, and the operands defined in UOP1 are v5, v3 and v7, that is, the index elements defined in UOP0 are the 4 elements included in v4, and the data elements are the 4 elements included in v2. After querying the corresponding elements from the vs2 register group using the values of the 4 index elements in v4 as indexes, the values of the 4 destination elements in the v6 register are obtained, that is, the execution result is obtained. Accordingly, the execution result can be written back to complete the execution of the UOP. That is, in an embodiment of the present invention, the UOP can be split according to the granularity of DLEN. In this way, the number of UOPs obtained by instruction decoding can be reduced, thereby simplifying the UOP counting circuit and improving the execution efficiency. Among them, splitting the UOP is the process of decoding the instruction to obtain the UOP. It should be noted that any UOP split out of the instruction only contains one data table vector, one index vector and one result vector. However, each index element needs to query LMUL indexed vs2 registers, or each data element needs to be queried by LMUL vs1 registers as indexes. The number of UOPs split out by this splitting method is small, but data transfer is required between UOPs to complete the operation.
[0021] In the existing implementation method, only one element is queried per cycle. When SEW is small, for example, SEW=8, in the above example, VLEN / SEW=128 / 8=32 rounds of operations are required to obtain all the required elements from the memory storing vs2, which results in low processing efficiency. How to conveniently and efficiently implement the execution of multiple UOPs has become the key to affecting the execution efficiency of the commonly used instruction GATHER. For example, in a related technology, several operation units that do not transfer data to each other execute the UOPs split out of the GATHER instruction. In this way, it is necessary to split many UOPs in the instruction decoding stage to realize that any combination of index in the vs1 register and any data in the vs2 register enters the operation unit for operation at the same time, that is, a UOP is required between each group of vs1 registers and vs2 registers. Therefore, a GATHER instruction needs to be split into as many as (VLEN / DLEN×LMUL) 2UOPs are needed to complete the operation (LMUL when DLEN=VLEN 2 UOPs). Assuming a specific implementation has two available arithmetic pipelines, with DLEN and VLEN equal, then when LMUL is the maximum value of 8, each pipeline must execute up to 32 UOPs to complete the execution of all 64 UOPs, and thus complete the execution of a GATHER instruction. Therefore, this solution has low execution efficiency. Furthermore, in this approach, each register in the vs1 or vs2 register bank must be read by up to eight UOPs. Repeatedly reading the same data also leads to high power consumption.
[0022] To this end, an embodiment of the present invention provides an instruction execution method.
[0023] Reference Figure 2 , shows a flowchart of the steps of an instruction execution method provided by an embodiment of the present invention, which can be applied to a computing array, such as Figure 2 As shown, the method may specifically include the following steps: Step 101: For any operation unit in the operation array, when first data of a target microinstruction is obtained from a first data path, a data operation operation is performed on the first data based on second data input along a second data path each time, and the second data is sent to the next operation unit through the second data path.
[0024] Step 102: After performing the data calculation operation once, the local control unit of the calculation unit updates the number of times the first data has been calculated.
[0025] Step 103: When the number of operations performed satisfies a preset number condition, the current result of the data operation is determined as the execution result of the target microinstruction.
[0026] Among them, the method can be specifically applied to an electronic device including the operation array. After the UOP in the instruction pipeline passes through the renaming unit, the dispatch unit, the emission unit, and the physical register stack, it enters the operation array, and the operation unit in the operation array executes the UOP. In an embodiment of the present invention, the operation array includes multiple operation units, and the number of operation units included in the operation array can be set as needed, and the embodiment of the present invention does not limit this. Furthermore, these multiple operation units adopt a preset connection method. For example, these multiple operation units can be connected in series, or these multiple operation units can also be interconnected. For example, one operation unit is connected to other operation units. The embodiment of the present invention does not limit this.
[0027] The target microinstruction is a microinstruction obtained by decoding the target instruction, and the target instruction is the above-mentioned data query instruction. When the operation unit is used to execute the microinstruction, a source operand defined in the microinstruction, that is, the above-mentioned first data, will be input to the operation unit. Specifically, the first data can be actively obtained by the operation unit, or it can be issued by the processing unit responsible for providing operands in the electronic device. For example, it can be issued by a control unit located outside the operation unit. Assuming that a GATHER instruction generates 4 UOPs, each of the 4 operation units that process these 4 UOPs will obtain the first data defined in one of the UOPs.
[0028] The first data path is a data transmission path connecting the physical register file and each arithmetic unit in the arithmetic array. This data transmission path is used to obtain the first data defined in each microinstruction obtained by decoding the target instruction from the physical register file and pass it only to the arithmetic units related to the order of microinstructions that enter the arithmetic array after being split from the same instruction. The second data path is a data transmission path connecting the physical register file and each arithmetic unit in the arithmetic array. This data transmission path is used to obtain the second data defined in each microinstruction obtained by decoding the target instruction from the physical register file and pass it through the arithmetic units in the arithmetic array, sharing it in sequence within the arithmetic array. The microinstruction processed by the next arithmetic unit is decoded from the same instruction as the arithmetic unit. Specifically, for each arithmetic unit processing a target microinstruction split from the target instruction, the second data from the second data path is passed to the next arithmetic unit, allowing each arithmetic unit to obtain the second data defined in other UOPs belonging to the same MOP as the microinstruction it is processing. In this way, each second data required for the operation only needs to be obtained once from the physical register file to complete the operation of an instruction, eliminating the need for the data path to feed the second data defined in other UOPs multiple times to each arithmetic unit. Therefore, the overall power consumption of the solution is lower.
[0029] In the method of the present invention, it is only necessary to split a GATHER instruction into VLEN / DLEN×LMUL UOPs, rather than splitting it into (VLEN / DLEN×LMUL) 2 This reduces the number of microinstruction splits. When DLEN = VLEN and LMUL is 2, 4, or 8, this reduces the number of uops by 2, 12, and 56, respectively, compared to the related art approach. The reduction in uops is significant when LMUL is large, resulting in a more efficient overall solution.
[0030] In an embodiment of the present invention, each arithmetic unit is locally provided with a control unit (hereinafter referred to as the local control unit). The local control unit is responsible for recording the number of times the first data has queried the second data, i.e., the aforementioned number of operations. The local control unit may use an iteration counter to record the number of operations. The initial value of the number of operations is 0. Each update of the number of operations may be performed by adding 1 to the current value of the number of operations, thereby achieving statistics on the second data that has been queried. Referring to the aforementioned example, it can be seen that for the microinstruction generated by the GATHER instruction, the index element of the microinstruction must be used to query the data elements defined in the microinstruction as well as the data elements defined in other microinstructions. In other words, multiple queries are required to obtain the query results corresponding to the index elements of the microinstruction. Compared to a method in which the data path repeatedly feeds the arithmetic unit with the second data defined in other UOPs, and the central control unit controls the arithmetic unit to write back the current result as the execution result without transferring data between UOPs, in an embodiment of the present invention, the local control unit maintains the number of operations. This allows for convenient determination of whether the query is currently complete based on the number of operations. The execution result is directly obtained and written back when the preset number of times is met, further improving overall execution efficiency to a certain extent.
[0031] Specifically, if the number of operations completed satisfies a preset number condition, the query is determined to be currently completed, and accordingly, the current result of the data operation can be determined as the execution result of the target microinstruction. Conversely, if the number of operations completed does not meet the preset number condition, the query is determined to be currently incomplete, and accordingly, the data operation can be continued by waiting for the second data subsequently transmitted. The data operation can be a query operation.
[0032] In summary, in the instruction execution method provided by the embodiment of the present invention, for any operation unit in the operation array, when the first data of the target microinstruction is obtained from the first data path, a data operation operation is performed with the first data based on the second data input along the second data path each time, and the second data is sent to the next operation unit through the second data path. After performing a data operation operation, the local control unit of the operation unit updates the number of times the first data has been operated. When the number of operations meets the preset number condition, the current result of the data operation operation is determined as the execution result of the target microinstruction. In this way, the operation units can share the data required for processing the microinstruction by sequentially transferring the second data downward based on the second data path. In this way, the operation of a GATHER instruction can be completed by obtaining the required data from the physical register stack only once. Therefore, the overall performance of the solution is better and the power consumption is lower.
[0033] In an embodiment of the present invention, the local control unit maintains the number of operations completed, and can directly end the data operation and obtain the execution result when the number of operations completed meets the preset number conditions, so that the result can be written back in time, which can further improve the overall execution efficiency to a certain extent, and thus improve the overall processing performance.
[0034] Optionally, the embodiment of the present invention further includes: Step S21, the entry operation unit receives the second data sent by the operand providing unit; wherein, the operand providing unit sends the second data defined in the microinstruction to the entry operation unit each time it receives a microinstruction of a data query instruction, and sends the first data defined in the microinstruction to the operation units in the operation array based on the first data path in the order of unit numbers.
[0035] The entry arithmetic unit may be a predefined arithmetic unit. The operand providing unit is directly connected to each arithmetic unit through the first data path. When the first data is input in the order of unit numbers from small to large, the entry arithmetic unit is the arithmetic unit with the smallest unit number. When the first data is input in the order of unit numbers from large to small, the entry arithmetic unit is the arithmetic unit with the largest unit number. That is, in the embodiment of the present invention, the entry arithmetic unit is the arithmetic unit that starts processing the microinstructions obtained by decoding the target instruction the earliest, and the microinstruction obtained by decoding the target instruction that is transmitted the earliest enters the arithmetic unit for processing first.
[0036] The operand providing unit includes a physical register file. For the GATHER instruction processing scenario, the operand providing unit specifically includes a vector physical register file, and the operation array is a GATHER array. This ensures that the operand providing unit is capable of providing first and second data. In an embodiment of the present invention, the target microinstruction is the microinstruction obtained by decoding a data query instruction. Accordingly, upon receiving each microinstruction for a data query instruction, the operand providing unit sequentially assigns microinstructions to the operation units in the operation array in ascending order of unit number, and inputs the first data of the assigned microinstruction to each operation unit. This microinstruction is the target microinstruction for that operation unit. For the i-th UOP obtained by decoding the target instruction, the first data defined in the i-th UOP is read from the physical register file and sent to the j-th operation unit. Where j represents the order in which the i-th UOP was received. In the case of in-order issuance, j=i. In the case of out-of-order issuance, j is not necessarily equal to i. At the same time, the second data of all microinstructions obtained by decoding the data query instruction are uniformly transmitted starting from the entry operation unit. That is, the operand providing unit first inputs the second data into the entry operation unit, and then transmits the second data to the next operation unit in the operation array starting from the entry operation unit. Figure 3This is a data transmission diagram provided by an embodiment of the present invention. Figure 3 As shown, the first data of D1 bit defined in the i-th UOP of the target instruction can be read from the vector physical register file first, and then transferred to the operation unit j. Figure 4 This is another data transmission diagram provided by an embodiment of the present invention. Figure 4 As shown, the second data of D2bit defined in the i-th UOP of the target instruction can be read from the vector physical register file first, and then transferred to the entry operation unit.
[0037] Where D1 represents the sum of the lengths of all index elements included in the first data, and D2 represents the sum of the lengths of all data elements included in the second data, that is, the data bit width used by each operation unit for each round of operation. For the vrgather.vv instruction, D1 = D2 = DLEN. For the vrgatherei16.vv instruction, since the element bit width of the index vector of the vrgatherei16.vv instruction may be inconsistent with the element bit width of the data table vector, and the RISC-V specification constrains the data table vector and index vector of the vrgatherei16.vv instruction to each contain at most 8 vector registers, the data transfer process can be split according to the longer vector of the data table vector and index vector, that is, D1 and D2 may not be equal. Specifically, the division principle of the total size of the index elements and the total size of the data elements carried in a UOP obtained by splitting the GATHER instruction is: the maximum bit width of the total size of the index elements carried by each UOP or the total size of the data elements reaches DLEN, and the number of index elements carried is equal to the number of data elements.
[0038] For the vrgatherei16.vv instruction, assuming VLEN = DLEN = 128, SEW = 64, LMUL = 8, the length of the data table vector of the vrgatherei16.vv instruction is VLEN × LMUL = 1024 bit, and the length of the index vector is VLEM × LMUL / 64 × 16 = 256 bit. Each UOP split out carries two data elements (a total of 128 bit) and two index elements (a total of 32 bit). For the GATHER instruction, the bit width of the data element is SEW, and EEW represents the bit width of the index element. In the vrgatherei16.vv instruction, EEW = 16, and in the vrgather.vv instruction, EEW = SEW. Then D1 = EEW < SEW? DLEN / (SEW / EEW) : DLEN, that is, if EEW is less than SEW, D1 is determined to be DLEN / (SEW / EEW). Otherwise, if EEW is not less than SEW, D1 is determined to be DLEN. D2 = EEW > SEW? DLEN / (EEW / SEW) : DLEN, that is, if EEW is greater than SEW, D2 is determined to be DLEN / (EEW / SEW). Otherwise, if EEW is not greater than SEW, D2 is determined to be DLEN. Assuming SEW = 64 and EEW = 16, D1 = DLEN / (64 / 16) = DLEN / 4, D2 = DLEN.
[0039] For different values of SEW and EEW, the sizes of D1 and D2 can be as shown in Table 1 below:
[0040] Table 1 Assume that the arithmetic array includes 8 cascaded arithmetic units: arithmetic unit 0 to arithmetic unit 7. Figure 5 It is a schematic diagram of an arithmetic array provided by an embodiment of the present invention. As Figure 5 shown, the arithmetic array is connected to the vector physical register bank through three groups of independent data paths. Among them, the first data path and the second data path are respectively used to obtain data from the vector physical register bank, and the third data path is used to write back data to the vector physical register bank. The first data can be stored in the first buffer, the second data is stored in the second buffer, and the current result of the data operation is stored in the third buffer. Only the second buffer between the arithmetic units can transfer data to the second buffer of the next arithmetic unit, and there is no data path between other buffers.
[0041] Assume that after decoding a GATHER instruction, UOP0-1 is obtained, and UOP0 and UOP1 are issued in sequence. Then, UOP0 and UOP1 are assigned to operation unit 0 and operation unit 1 for processing, respectively. That is, in the embodiment of the present invention, for the microinstructions obtained by decoding the target instruction, they are assigned to the operation units in the operation array in the order of issuance of the microinstructions, starting from the operation unit with the smallest unit number. Accordingly, the operand providing unit inputs the first data defined in UOP0 to operation unit 0 through the first data path, and simultaneously inputs the second data defined in UOP0 to operation unit 0 through the second data path. One cycle or several cycles later (depending on the operand readiness), the operand providing unit inputs the first data defined in UOP1 to operation unit 1 through the first data path, and simultaneously inputs the second data defined in UOP1 to operation unit 0 (entry operation unit) through the second data path. The second data defined by UOP0 is then input to operation unit 1 from the second buffer of operation unit 0.
[0042] The data path widths of the first data path, the second data path, and the third data path are all DLEN. MAX Indicates the maximum value of LMUL supported by a specific implementation. The maximum number of UOPs obtained by decoding a GATHER instruction can be expressed as VLEN / DLEN×LMUL MAX , so VLEN / DLEN×LMUL needs to be set in the operation array MAX arithmetic units to ensure that the arithmetic array can meet the processing requirements of any GATHER instruction. MAX It can be 8. Accordingly, when DLEN=VLEN, the number of operation units can be 8.
[0043] By setting two data paths for input and one data path for output, the first data and the second data can arrive in the same clock cycle, and the operation result can be written back using the third data path in the same clock cycle, so the operation efficiency is higher. Of course, it is also possible to set only one data path for reading and writing to transmit data, and use one data path for reading and writing to complete the reception of the first data, the reception of the second data, and the writing back of the third data in multiple cycles. The embodiment of the present invention does not limit this. And by setting DLEN=VLEN, the number of operation units in the operation array can be made to reach the minimum number of units LMUL MAX , which can make the operation array reach the maximum throughput, and calculate 1 / LMUL per cycle MAX LMUL=LMUL MAX The GATHER instruction has higher computing performance. MAX= 8, DLEN is half of VLEN, requiring 16 arithmetic units. The throughput of the arithmetic array then becomes 1 / 16 of the GATHER instructions per cycle for LMUL = 8. This approach doesn't reduce the hardware circuit size, but it does halve the computational efficiency and result in lower performance.
[0044] In an embodiment of the present invention, for all microinstructions that a data query instruction is split into, the second data that needs to be shared is uniformly sent from the entry operation unit to the operation array, so that the second data of each microinstruction starts from the entry operation unit and is passed down in sequence. The second data defined by each microinstruction is always passed to the entry operation unit, and then passed to the next operation unit, ensuring that each operation unit in the operation array can obtain the second data defined in all microinstructions that a data query instruction is split into, thereby ensuring that the microinstructions are executed correctly. Among them, the second data input along the second data path includes the second data defined in all UOPs obtained by decoding the target instruction.
[0045] Optionally, a unit status bit is provided in the arithmetic unit in the embodiment of the present invention, and the embodiment of the present invention further includes: Step S31: The control unit sets the unit status bit to a first value when receiving the first data, and restores the unit status bit to a second value after detecting that the number of operations meets the preset number condition; the first value indicates that the operation unit is in an activated state, and the second value indicates that the operation unit is in an inactivated state.
[0046] Correspondingly, the sending of the second data to the next operation unit through the second data path includes: sending the second data to the next operation unit through the second data path when the unit status bit is the first value.
[0047] Wherein, each operation unit maintains a 1-bit unit status bit. The first value can be 1, and the second value can be 0. Accordingly, for any operation unit, if the operation unit obtains the first data, that is, the UOP enters the operation unit and starts execution, then the local control unit of the operation unit sets the unit status bit to 1 to indicate that it has entered the activation state. Wherein, the default value of the unit status bit is 0, indicating that no microinstructions are executed in the operation unit. Furthermore, after each update of the number of times the first data has been operated, if the number of times the number of times the first data has been operated meets the preset number condition, the unit status bit is restored to 0 to indicate that the activation state has been exited.
[0048] If the unit status bit is the second value, it means that the second data defined in all UOPs into which the target instruction is split has passed through this operation unit, that is, it has been passed to the next operation unit by this operation unit. Therefore, each operation unit can only send the second data to the next operation unit when the unit status bit is the first value, that is, in the active state. In this way, unnecessary sending operations can be avoided.
[0049] Optionally, if the microinstruction carries the same instruction sequence number (i.e., the sequence number uniquely indicating the MOP corresponding to the microinstruction) as the previously sent microinstruction, i.e., both carry the instruction sequence number of the target instruction (indicating that the microinstruction and the previously sent microinstruction were decoded from the same instruction), the operand providing unit directly sends the second data defined in the microinstruction to the entry arithmetic unit. If the microinstruction carries a different instruction sequence number than the previously sent microinstruction (indicating that the microinstruction and the previously sent microinstruction were decoded from different instructions), and if the unit status bit of the entry arithmetic unit is the second value, the second data defined in the microinstruction is sent to the entry arithmetic unit. In this case, if the unit status bit of the entry arithmetic unit is the first value, it indicates that the entry arithmetic unit has not yet completed processing the UOP in the target instruction. Therefore, the operation can continue to wait until the unit status bit of the entry arithmetic unit reaches the second value, i.e., the entry arithmetic unit is idle, and then continue to send the second data defined in the UOP obtained by decoding the new instruction, i.e., begin executing the new instruction. This avoids execution conflicts. That is, when the instruction sequence number carried by the microinstruction is different from that carried by the microinstruction sent last time, the microinstruction is received starting from the entry operation unit.
[0050] In a first implementation, the first data is an index element defined by the first source operand of the target microinstruction, and the second data is a data element defined by the second source operand of the target microinstruction. Accordingly, in this method, sending the second data to the next operation unit through the second data path specifically includes: only sending the second data to the next operation unit through the second data path. In this implementation, the second buffer stores the data elements (referred to as data element groups) defined by the second source operand of a microinstruction each time, that is, the data element groups are shared along the second data path each time, so that each operation unit that processes the decoded target instruction obtains the groups defined in all microinstructions decoded by the target instruction. In this method, the amount of data required to be transferred is smaller.
[0051] In a second implementation, the first data is a data element defined by a second source operand of the target microinstruction, and the second data is an index element defined by the first source operand of the target microinstruction. Accordingly, in this implementation, sending the second data to the next arithmetic unit via the second data path specifically includes sending the second data, a current result of the data operation, and the number of operations performed to the next arithmetic unit via the second data path.
[0052] In this implementation, the second buffer stores the index element (referred to as the index element group) defined by the first source operand of each microinstruction. Each time the second data is shared along the second data path, the current result of the data operation and the number of operations maintained by the control unit are first obtained from the third buffer (equivalent to the control signal maintained by the local operation unit). The index element group, current result, and number of operations are then sent to the next operation unit. This ensures that each operation unit can hold the data required for the data query in each clock cycle. The next operation unit can be an operation unit with a number one greater than the current operation unit.
[0053] In actual applications, the index element group, the current operation result, and the number of operations completed may contain more than DLEN × 2 bits of data, while the data vector group only contains DLEN bits of data. Therefore, the first implementation requires less data to be transferred, which can reduce hardware circuitry and power consumption.
[0054] Optionally, the step of performing a data operation on the first data based on the second data input along the second data path each time specifically includes: Step 1011: For any received second data, based on each index element in the first data, query operations are performed on multiple data elements in the second data to obtain query results corresponding to each index element.
[0055] Step 1012: Write the query results corresponding to the index elements into the destination elements corresponding to the index elements in the destination buffer.
[0056] Specifically, an index element in the first data can be used to query multiple data elements in the second data to obtain the query result corresponding to the index element. Let vs1[i] represent the index element. If vs1[i]>=VLMAX, then the query result corresponding to the index element is 0, and the result of this query operation is the final query result corresponding to the index element. If vs1[i] is less than VLMAX, and the second data received this time includes vs2[vs1[i]], then the query result of this time is the final query result corresponding to the index element. Otherwise, the query result of this time is the temporary query result corresponding to the index element. Accordingly, the query result corresponding to vs1[i] can be written into vd[i] in the destination buffer. The destination buffer is the above-mentioned third buffer.
[0057] In the embodiment of the present invention, the second data received at any one time is queried based on the index element. Since the second data defined in each microinstruction obtained by decoding the target instruction will pass through each arithmetic unit that processes the microinstruction obtained by decoding the target instruction, performing the query each time ensures that the second data defined in each microinstruction obtained by decoding the target instruction can be queried, thereby ensuring the accuracy of the query.
[0058] It should be noted that, in an embodiment of the present invention, after the final query result corresponding to the index element is obtained, a queried flag will be set for the index element. Exemplarily, this control sets a queried status bit for each index element, and after the final query result corresponding to the index element is obtained, the queried status bit of the index element is set to 1 to achieve the setting of a queried flag for the index element. The queried status bit defaults to 0. Accordingly, in the case where the queried flag is not set for the index element, an operation of performing a query operation on multiple data elements in the second data based on the index element is performed. In this way, the correct query result is avoided from being overwritten, resulting in inaccurate instruction execution results.
[0059] Furthermore, when adopting the aforementioned second implementation method, when the index element group is transmitted, the queried status bit of each index element is synchronously transmitted to the next operation unit. The operation unit writes the received current operation result and the number of operations completed into the third buffer and the local control unit, respectively, and writes the queried status bit of each index element into the local control unit. Then, based on each index element in the second data whose queried status bit is 0, a query operation is performed on multiple data elements in the first data to obtain the query results corresponding to each index element; the query results corresponding to each index element are written to the destination element corresponding to each index element in the destination buffer (i.e., the third buffer). Thereafter, the number of operations completed and the queried status bit of each index element are updated. Then, the second data, the destination element (i.e., the current operation result), the number of operations completed, and the queried status bit of each index element in the third buffer are continuously sent to the next operation unit.
[0060] Optionally, when the number of operations performed satisfies a preset number condition, determining the current result of the data operation as the execution result of the target microinstruction specifically includes: Step 1031 : When the number of operations performed satisfies the preset number condition, read each destination element of the destination buffer as the execution result of the target microinstruction.
[0061] Accordingly, the embodiment of the present invention further includes: step S41, writing the execution result back to the physical register file through the third data path.
[0062] In the embodiment of the present invention, if the number of operations completed satisfies the preset number condition, it indicates that the data elements defined in all microinstructions obtained by decoding the target instruction have been queried. At this time, the value of each destination element in the destination buffer is the final query result. Therefore, each destination element in the destination buffer can be used as the execution result of the target microinstruction.
[0063] Furthermore, a third data path connects each operation unit with the vector physical register file. Writing back can be performed through the third data path, that is, writing the execution result back to the above-mentioned vector physical register file to complete the execution of the target microinstruction.
[0064] In this embodiment of the present invention, each destination element in the destination buffer is read as the execution result of the target microinstruction only when the number of operations completed meets a preset number condition. The execution result is then written back to the physical register file via a third data path. This ensures that the final query result can be written back to the physical register file, ensuring the accuracy of the execution result.
[0065] Optionally, the embodiment of the present invention further includes: Step S51: Calculate the sum of the number of operations and the data length of the second data to obtain the queried length.
[0066] Step S52: When the searched length reaches the total length of the second source operand of the data arrangement instruction, determine that the number of operations completed satisfies the preset number condition.
[0067] The data length of the second data may be the sum of the lengths of the multiple data elements included in the second data. For example, assuming that one second data element includes four data elements, each of which is 32 bits, the sum of the data lengths of the second data is 128 bits. It should be noted that, when adopting the second implementation method described above and transferring the index element group, the product of the number of operations performed and the sum of the data lengths of the first data (i.e., the sum of the lengths of the multiple data elements included in the first data) is used as the queried length.
[0068] The total length of the second source operand of the data permutation instruction can be expressed as LMUL × VLEN. With n representing the number of operations completed and D2 representing the sum of the data lengths of the second data, D2 × n represents the queried length. Accordingly, if D2 × n = LMUL × VLEN, it is determined that the number of operations completed satisfies the preset number condition, i.e., the preset number condition can be expressed as n = LMUL × VLEN / D2. Furthermore, if D2 × n = LMUL × VLEN, it indicates that the data elements defined in all microinstructions obtained by decoding the target instruction have been queried using the index element, and therefore, it can be determined that the preset number condition has been met.
[0069] In the embodiment of the present invention, the queried length is calculated based on the number of operations, and whether the preset number condition is currently met can be determined based on the queried length and the total length of the second source operand. In this way, the determination efficiency is high.
[0070] In this embodiment of the present invention, the index element group in the first buffer of each operation unit only queries the data element group currently stored in the second buffer of that operation unit, and stores the current query result (i.e., the current result of the data operation) in the third buffer. In this embodiment of the present invention, the query process in each operation unit can be completed within a single clock cycle. The number of operations recorded by the local control unit of each operation unit can indicate how many times the index element group in the first buffer of that operation unit has queried the data element group in the second buffer.
[0071] Take the vrgather.vv v24,v8,v16 instruction with DLEN=VLEN and LMUL=8 as an example. In this vrgather.vv instruction, v24-v31 are the vd register group, v8-v15 are the vs2 register group, and v16-v23 are the vs1 register group. This instruction is decoded to produce UOP0-UOP7. The first data defined in UOP0 is the index element represented by the data in the v16 register, the second data is the data element represented by the data in the v8 register, and the destination element is the element in the v24 register. The first data defined in UOP1 is the index element represented by the data in the v17 register, the second data is the data element represented by the data in the v9 register, and the destination element is the element in the v25 register. The first data defined in UOP2 is the index element represented by the data in the v18 register, the second data is the data element represented by the data in the v10 register, and the destination element is the element in the v26 register. The first data defined in UOP3 is the index element represented by the data in the v19 register, the second data is the data element represented by the data in the v11 register, and the destination element is the element in the v27 register. The first data defined in UOP4 is the index element represented by the data in the v20 register, the second data is the data element represented by the data in the v12 register, and the destination element is the element in the v28 register, ..., the first data defined in UOP7 is the index element represented by the data in the v23 register, the second data is the data element represented by the data in the v15 register, and the destination element is the element in the v31 register. That is, the vs1 register group and vs2 register group of this instruction are divided into 8 rounds and passed to the operation array in sequence. And because DLEN=VLEN, each round just passes the elements of one vector register each of the vs1 and vs register groups into the operation array.
[0072] Figure 6 This is a schematic diagram of a processing process provided by an embodiment of the present invention. Figure 6 As shown, in the first round, the data in register v8 of the vs2 register bank is used as the second data of UOP0 and is stored in the second buffer of operation unit 0. The data in register v16 of the vs1 register bank is used as the first data of UOP0 and is stored in the first buffer of operation unit 0. Operation unit 0 enters the active state. Simultaneously, a query operation is performed in operation unit 0 (i.e., the index element in register v16 is used to query the data element in register v8). The result of this operation (represented by v24(v8)) is stored in the third buffer. The content stored in the third buffer can also be called a result vector. The control unit in operation unit 0 updates the number of operations completed = 1 (not shown in the figure), indicating that the index element of operation unit 0 has completed a query of the data vector. In this example, the result of the destination register v24 has completed the query of the data element in v8.
[0073] Figure 7 This is another processing diagram provided by an embodiment of the present invention. Figure 7 As shown, at the beginning of the second round (i.e., at the next clock cycle), only ALU 0 is active, and the data in its second buffer is sent to ALU 1 for storage in its second buffer. The data in register v9 is stored in ALU 0's second buffer as the second data for UOP 1, and the data in register v17 is stored in ALU 1's first buffer as the first data for UOP 1. ALU 1 also becomes active. In the second round, the two active ALUs complete their respective lookups (using the index element in register v16 to query the data element in register v9, and using the index element in register v17 to query the data element in register v8), respectively, and store the results in their respective third buffers. Simultaneously, the control units in ALUs 0 and 1 update their counts of operations to 2 and 1, respectively (not shown). The updated counts indicate that ALUs 0 and 1 have completed the lookups for the data elements in registers v8-9 and v8, respectively.
[0074] Figure 8 This is another processing diagram provided by an embodiment of the present invention. Figure 8 As shown, at the beginning of the third round (i.e., at the next clock cycle), operations units 0 and 1 are active, and the data in their second buffers is stored in the second buffers of operations units 1 and 2, respectively. The data in register v10 is stored in the second buffer of operations unit 0 as the second data of UOP 2, and the data in register v18 is stored in the first buffer of operations unit 2 as the first data of UOP 2. Operation unit 2 also becomes active. In the third round, the three active operations units complete their respective query operations (using the index element in register v16 to query the data element in register v10, using the index element in register v17 to query the data element in register v9, and using the index element in register v18 to query the data element in register v8), and store the results in their respective third buffers. Simultaneously, the control units in operations units 0, 1, and 2 update their counts of operations to 3, 2, and 1, respectively (not shown). The updated counts indicate that operations units 0, 1, and 2 have completed queries on the data elements in registers v8-10, v8-9, and v8, respectively.
[0075] Figure 9 This is another processing diagram provided by an embodiment of the present invention. Figure 9As shown, when the eighth round is reached, the data in register v15 is stored in the second buffer of unit 0 as the second data of UOP7, and the data in register v23 is stored in the first buffer of unit 7 as the first data of UOP7. All eight units in the operation array are activated. In the eighth round, unit 0 completes its query operation (using the index element in register v16 to query the data element in register v15) and stores the result in the third buffer of unit 0. At this point, the control unit of unit 0 records the number of operations completed as 8 (not shown in the figure), indicating that the query of the data elements in registers v8-15 has been completed, that is, the number of operations completed meets the preset number condition. Therefore, the control unit of unit 0 resets the unit status bit to 0 and transfers each destination element in the third buffer of unit 0 to the vector physical register file for storage. If no new microinstructions enter the operation array in the next round, unit 0 will deactivate.
[0076] Figure 10 This is another processing diagram provided by an embodiment of the present invention. Figure 10 As shown, in the ninth round, since no new microinstructions enter the operation array, no new data in the data register enters operation unit 0. However, the operation unit that was active in the previous round will still pass the data in the second buffer to the next operation unit. It should be noted that in this example, operation unit 7 does not have a next operation unit, so the data in the second buffer within this unit is directly overwritten by the data passed from operation unit 6. When operation unit 0, which has completed the operation in the eighth round, exits the active state, the data in the three buffers becomes invalid, and the data in the first buffer area, the second buffer area, and the third buffer area in operation unit 0 can be cleared.
[0077] In the ninth round, operation unit 1 completes its query operation (using the index element in register v17 to query the data element in register v15) and stores the result in its third buffer. At this point, the control unit of operation unit 1 records a count of 8 operations (not shown), indicating that the query of the data elements in registers v8-15 has been completed, and the number of operations meets the preset count requirement. Therefore, the control unit of operation unit 1 resets the unit status bit to 0 and transfers each destination element in the third buffer of operation unit 1 to the vector physical register file for storage. If no new microinstructions enter the operation array in the next round, operation unit 1 will deactivate. This continues in this manner. After six more rounds, or a total of 15 rounds of operations, the operation of a vrgather.vv instruction is completed.
[0078] Figure 11This is a processing diagram of a computing unit provided by an embodiment of the present invention, such as Figure 11 As shown, the operation unit can first initialize the number of operations completed to 0. If no first data enters the operation unit, it continues to wait for one clock cycle before continuing to make a judgment. If first data enters the operation unit, it receives second data transmitted on the second data path. The second data may come from the vector physical register file or from the previous operation unit. A query operation is performed based on the first and second data. Next, the number of operations completed is incremented by 1. Then, a determination is made as to whether a preset number condition is met. If so, the destination element in the third buffer is written back to the vector physical register file, and the judgment is continued after waiting for one clock cycle. If not, the second data in the second buffer is sent to the next operation unit, and the newly entered second data is received. In this embodiment of the present invention, the operation array only needs to read the first data and the second data once in groups to obtain all operands required by the UOP. After the UOP operation is completed, the result only needs to be stored once in the vector physical register file. In this way, the number of times the read and write ports of the vector physical register file are occupied is only proportional to the LMUL, thereby minimizing the impact on the execution of other instructions. The second data defined in the microinstruction obtained by decoding the target instruction passes through each operation unit in sequence. It should be noted that for an operation unit, performing the query operation, sending the second data to the next operation unit, and receiving the second data sent by the previous operation unit, these three operations can be completed within one clock cycle or within multiple clock cycles, and the embodiment of the present invention does not limit this.
[0079] A vrgatherei16.vv instruction requires LMUL / SEW×EEW index vector registers and LMUL data table vector registers to participate in the operation, producing LMUL vector register results. In this embodiment of the present invention, a D1-bit index element and a D2-bit data element are transferred to the operation array during each clock cycle. After LMUL cycles (where the number of cycles equals the number of uops resulting from instruction decomposition), the D2-bit target element is output to the vector physical register file during each clock cycle. This continuous transfer of LMUL cycles completes the operation of the instruction.
[0080] In this embodiment of the present invention, the second data of the first UOP sent to the computation array by different instructions all enter and participate in the computation from the entry computation unit. In this embodiment of the present invention, the computation array can complete the computation of the GATHER instruction within a linear time complexity proportional to the LMUL, resulting in greater efficiency. Specifically, a round represents one clock cycle. Computational unit 0 receives data and begins computation in round 1, completes computation in round LMUL, and can then receive new microinstructions for computation in round LMUL+1. Similarly, computational unit i begins computation in round i+1, completes computation in round LMUL+i, and can then receive new microinstructions for computation in round LMUL+i+1. This means that the computation array does not need to wait for all microinstructions of the current instruction to complete computation. While computational unit 0 is in an inactive state, it can begin receiving decoded microinstructions for the next instruction from computational unit 0. Consequently, processing efficiency is improved. Furthermore, the computation array can initiate computation upon receiving only the partial index vector and partial data table vector required by a computation unit, without having to wait for the entire index vector and data table vector to enter the computation unit, or without having to wait for the entire data table vector to enter the computation unit before performing computation. Therefore, the processing efficiency can be further improved.
[0081] For example, assuming that the 8 UOPs decoded from instruction 1 are currently being processed, after operation unit 0 exits the active state, a UOP decoded from instruction 2 is sent to operation unit 0 to start processing. In subsequent cycles, other operation units will complete the execution of the UOP instructions decoded from instruction 1 in turn, and exit the active state in turn. Accordingly, in subsequent cycles, the second data defined in the UOP decoded from instruction 2 can be directly fed into operation unit 0. That is, in subsequent cycles, the UOP decoded from instruction 2 can be directly sent to each operation unit in the order of unit numbers for execution by the operation unit. Ensure that instruction 2 can be executed normally.
[0082] It should be noted that, when the GATHER instruction is decoded to obtain 4 UOPs, 4 consecutive units can be used. That is, in the embodiment of the present invention, according to LMUL MAX Setting the number of operation units included in the operation array can ensure that the operation array can adapt to processing any GATHER instruction. In an embodiment of the present invention, multiple operation arrays can also be set for the several UOP numbers that will be generated by the decoding of the GATHER instruction. Assuming that the decoding of the GATHER instruction may generate three numbers of 2, 4, and 8, then an operation array including these three numbers of 2, 4, and 8 can be set. For any GATHER instruction, a corresponding number of operation arrays is used, and execution begins from the operation unit with the smallest unit number as the entry. The embodiment of the present invention does not impose any restrictions on this.
[0083] Furthermore, in an embodiment of the present invention, for ease of presentation, three buffers are used to store index vectors, data table vectors, and result vectors, respectively. In practical applications, one buffer can also be used to simultaneously satisfy the requirements for storing index vectors and operation results, that is, the third buffer and the first buffer are the same buffer. For an index element in the buffer, after obtaining the final query result corresponding to the index element, the final query result of the index element (that is, the target element that was queried) is used to overwrite the index element, so that the index for which the final query result has been obtained is no longer valid for subsequent operations. In this way, by reusing the buffer for storing index elements to store operation results, the area of the storage unit in the operation unit can be reduced. For example, the size of each buffer is DLEN. When three buffers are set, the storage space requirement of each operation unit is DLEN×3. By multiplexing, when only two buffers are set, the storage space requirement of each operation unit is reduced to DLEN×2.
[0084] In an embodiment of the present invention, multiple groups of arithmetic units are connected to form an arithmetic array to accelerate the calculation of the GATHER instruction. First data (D1 bits) and second data (D2 bits) are sent sequentially and in pairs to the arithmetic array via a data path. The paired first data (D1 bits) and second data (D2 bits) can be sent to the arithmetic units simultaneously, or in different clock cycles, which is not limited in this embodiment of the present invention. Neither D1 nor D2 is greater than DLEN. The first data (D1 bits) defined in different UOPs are sent sequentially to the memories of the corresponding arithmetic units in the order in which they were received, marking the arithmetic units as activated. The second data (D2 bits) are stored in the memory of the first arithmetic unit. Whenever new paired data enters the arithmetic array, the active arithmetic unit passes the second data in its second buffer to the next arithmetic unit. Simultaneously, all active arithmetic units use their internal index elements to perform query operations on the data elements entering the arithmetic unit, and the operation results are stored as temporary results in the third buffer within the arithmetic unit. At the same time, each operation unit will record and update the number of operations n. When the preset number condition D2×n=VLEN×LMUL is met, it is determined that all query operations have been completed in the operation unit, and the results in the third buffer are written back to the vector physical register stack.
[0085] Optionally, the query operation is performed on multiple data elements in the second data based on each index element in the first data to obtain query results corresponding to each index element, specifically including: Step 1011a: for any index element in the first data, divide the index element into a first index component, a second index component, and a third index component in ascending order based on the number of the data elements; Step 1011b: When the third index component is 0 and the second index component is equal to the microinstruction sequence number of the target microinstruction, generate a selection signal based on the first index component; select a target data element from the data elements of the second data based on the selection signal as the query result corresponding to the index element.
[0086] Step 1011c: when the third index component is not 0, and / or the second index component is not equal to the microinstruction sequence number, set the query result corresponding to the index element to 0.
[0087] The microinstruction sequence number is the UOP number of the target microinstruction, indicating the microinstruction's order among the multiple microinstructions obtained by decoding the target instruction. The first, second, and third index components can be referred to as the low-order index, the middle index, and the high-order index, respectively. That is, an index element in the first buffer includes three parts: the high-order index, the middle index, and the low-order index.
[0088] Specifically, when the queried status bit of the index element is not 1, the execution starts from the step of dividing the index element into a first index component, a second index component and a third index component in order from low to high based on the number of data elements. The bit width of the low-order index is determined by the number of data elements (that is, the number of data elements included in the second data (second buffer)). Assuming that the total size of the data elements defined in a UOP is DLEN, the number of data elements is DLEN / SEW. Accordingly, the bit width Width of the low-order index is low =log2(DLEN / SEW). Assuming DLEN = 128 and SEW = 32, the bit width of the low-order index is log2(128 / 32) = 2. Assuming DLEN = 128 and SEW = 16, the bit width of the low-order index is log2(128 / 16) = 3.
[0089] Accordingly, the lower Width of the index element can be low As the first index component, the 0th to Width low bit as the first index component. low +1~Width low +1+r bits as the second index component. Among them, r represents the width of the middle index bit Width middle=log2(VLEN / DLEN×LMUL MAX ). The part of the index element other than the first index component and the second index component is used as the third index component. That is, the high index width Width high =max(EEW-Width low -Width middle ,0). The high-order index width may be 0. In the case where the high-order index width is 0, the third index component is considered to be always equal to 0. For example, when VLEN=DLEN=256, SEW=8, EEW=8, Width low =log2(DLEN / SEW)=4,Width middle =log2(VLEN / DLEN×8)=4, at this time Width high =0.
[0090] For the i-th UOP split from the target instruction, the first data defined in the i-th UOP is the index element included in the i-th group index vector, and the second data defined in the i-th UOP is the data element included in the i-th group data table vector. If VLEN = DLEN, the i-th UOP will obtain the i-th group of VLEN bits in the data table vector and the index vector. Accordingly, when the i-th UOP enters the operation array, the i-th group index vector and the i-th group data table vector are sent to the operation array.
[0091] Assume that the middle index and low index are INDEX MID , INDEX LOW , according to the index width rule, INDEX LOW The value can be 0~2^WIDTH LOW -1, corresponds to the number of all elements in a data table vector (i.e., a second data). Among them, 2^WIDTH LOW Indicates a width of 2 LOW Power, 2^WIDTH LOW -1 means WIDTH of 2 LOW The power minus 1. Therefore, the low-order index can be converted into a one-hot code as a selection signal. When the index element is less than VLMAX, the data element corresponding to the index element is in the INDEX MID Therefore, the index element is less than VLMAX and the UOP number and INDEX MID If they are equal, it means that the middle index component represents the INDEX MID The data table vectors are stored in the second buffer. Conversely, when the UOP number and INDEX MID If they are not equal, it means that the middle index component represents the INDEXMID The data table vector does not exist in the second buffer. In this case, the query result corresponding to the index element can be directly set to 0.
[0092] When UOPs are issued sequentially, the number of queries equals the UOP number of the UOP. When UOPs are issued out of order, the number of queries equals the UOP number of the UOP. Accordingly, the UOP number of the UOP corresponding to the second data (i.e., the UOP defined by the second data) can be sent to the next arithmetic unit along with the second data. In this embodiment of the present invention, the UOP number in the arithmetic unit (i.e., the above-mentioned i) always corresponds to the number of the data table vector in the second buffer (i.e., both are i).
[0093] Furthermore, corresponding to this partitioning method, the high-order index is used to determine whether the index element is less than VLMAX, that is, whether the index element is within [0, VLMAX). If the third index component is not 0, it means that the index element is not less than VLMAX. Therefore, according to the RVV instruction set specification, the query result corresponding to the index element can be set to 0. The selection signal can be an s-bit target one-hot encoding, where s is the number of data elements included in the second data. The initial one-hot encoding is all 0, and the t-th bit in the initial one-hot encoding represents the t-th data element included in the second data. Accordingly, assuming that the low-order index is equal to u, then the u-th bit in the initial one-hot encoding is set to 1 to obtain the target one-hot encoding. When using the selection signal for selection, the v-th bit of each data element in the second buffer is combined into a bit sequence, and the selection signal is used to perform an AND operation with each bit sequence to obtain the v-th bit in the query result corresponding to the index element. Assuming that a data element consists of 32 bits, then the 32-bit sequence of each data element is ANDed to obtain a 32-bit query result. Of course, when the third index component is not 0 and / or the second index component is not equal to the microinstruction sequence number, the selection signal is all 0 and the selected data is also 0. In the embodiment of the present invention, when the third index component is 0 and the second index component is equal to the microinstruction sequence number of the target microinstruction, the query result corresponding to the index element is determined as the final query result of the index element.
[0094] Assuming VLEN=128, DLEN=128, and SEW=32, the second buffer contains four 32-bit data. Figure 12 FIG. 1 is a schematic diagram of generating a selection signal provided by an embodiment of the present invention, such as Figure 12As shown, index[31:5], index[4:2], and index[1:0] represent the high-order index, intermediate index, and low-order index, respectively, and uopIdx[2:0] represents the UOP number. When the electronic device uses sequential issuance, uopIdx[2:0] represents the number of operations performed, iterIdx. This number also factors into data selection. If the second index component equals the microinstruction sequence number of the target microinstruction, a 1 is output to indicate that the index query is complete. Accordingly, the query result corresponding to the index element is determined as the final query result. The conversion circuit is used to convert index[1:0] into the target one-hot encoding. The intermediate index and the UOP number are compared. Let w represent the number of data elements included in a second data element. The data elements included in the iterIdx-th second data element are: iterIdx×w data elements to iterIdx×w+w-1 data elements. Correspondingly, if the two are equal, it means that the range of the element index in the second buffer is within the interval [iterIdx×w,iterIdx×w+w-1], and the AND gate outputs 1. Otherwise, it outputs 0. If the third index component is 0, the AND gate outputs 1. Otherwise, it outputs 0. Finally, the output of the AND gate is the selection signal. For example, assuming index = 0x0000_0019 and iterIdx = 0x6, we can get index[31:5] = 0, index[4:2] = 0x6, and index[1:0] = 0x1. Since index[4:2] == iterIdx, the AND gate outputs 1, and the conversion circuit converts index[1:0] = 0x1 into the one-hot code b0010.
[0095] Figure 13 This is a selection diagram provided by an embodiment of the present invention, such as Figure 13 As shown, taking the selection signal 0100 as an example, the selection signal and the values of each element in the bit sequence serve as inputs to multiple AND gates. The final output of an OR gate connected to these multiple AND gates is the selection result from the bit sequence. In a specific implementation, a corresponding number of selection circuits can be provided according to the number of bits in a data element.
[0096] In an embodiment of the present invention, by dividing the index elements into the first index component, the second index component and the third index component, the query can be conveniently implemented based on the division of the index elements into the first index component, the second index component and the third index component, so the processing efficiency is high.
[0097] Reference Figure 14 , shows a block diagram of an instruction execution device provided by an embodiment of the present invention, such as Figure 14 As shown, the device may specifically include: a first processing module 201 configured to, for any operation unit in the operation array, perform a data operation on the first data of the target microinstruction obtained through the first data path based on second data input along the second data path each time, and send the second data to the next operation unit through the second data path; An updating module 202 is configured to update the number of operations performed on the first data by a local control unit of the computing unit after performing the data computing operation once; The first determining module 203 is configured to determine the current result of the data operation as the execution result of the target microinstruction if the number of operations performed satisfies a preset number condition.
[0098] Optionally, the device further comprises: a receiving module, configured to receive, by the entry operation unit, the second data sent by the operand providing unit; In which, each time the operand providing unit receives a microinstruction of a data query instruction, it sends the second data defined in the microinstruction to the entry operation unit, and sends the first data defined in the microinstruction to the operation units in the operation array based on the first data path in the order of unit numbers.
[0099] Optionally, a unit status bit is provided in the arithmetic unit; and the device further comprises: a second processing module in the control unit, configured to set the unit status bit to a first value upon receiving the first data, and restore the unit status bit to a second value after detecting that the number of operations completed satisfies the preset number condition; the first value indicating that the operation unit is in an activated state, and the second value indicating that the operation unit is in an inactivated state; The first processing module 201 is specifically configured to: when the unit status bit is the first value, send the second data to the next operation unit through the second data path.
[0100] Optionally, the target microinstruction is a microinstruction obtained by decoding a data query instruction; the first data is an index element defined by a first source operand of the target microinstruction, and the second data is a data element defined by a second source operand of the target microinstruction.
[0101] Optionally, the first processing module 201 is specifically configured to: For any received second data, based on each index element in the first data, query operations are performed on multiple data elements in the second data to obtain query results corresponding to each index element; The query results corresponding to the index elements are written into the destination elements corresponding to the index elements in the destination buffer.
[0102] Optionally, the first determining module 203 is specifically configured to: When the number of operations completed satisfies the preset number condition, reading each destination element of the destination buffer as the execution result of the target microinstruction; The apparatus further includes a write-back module configured to write the execution result back to the physical register file via a third data path.
[0103] Optionally, the device further comprises: a calculation module, configured to calculate the sum of the number of operations performed and the data length of the second data to obtain a queried length; The second determining module is configured to determine, when the queried length reaches the total length of the second source operand of the data query instruction, that the number of operations performed satisfies the preset number condition.
[0104] Optionally, the first processing module 201 is specifically configured to: For any index element in the first data, dividing the index element into a first index component, a second index component, and a third index component in ascending order based on the number of the data elements; If the third index component is 0 and the second index component is equal to the microinstruction sequence number of the target microinstruction, generating a selection signal based on the first index component; selecting a target data element from the data elements of the second data based on the selection signal as a query result corresponding to the index element; When the third index component is not 0, and / or the second index component is not equal to the microinstruction sequence number, the query result corresponding to the index element is set to 0.
[0105] Optionally, the target microinstruction is a microinstruction obtained by decoding a data query instruction; the first data is a data element defined by a second source operand of the target microinstruction, and the second data is an index element defined by a first source operand of the target microinstruction; The first processing module 201 is specifically configured to send the second data, the current operation result of the data operation, and the number of operations to the next operation unit through the second data path.
[0106] In summary, in the instruction execution device provided by the embodiment of the present invention, for any operation unit in the operation array, when the first data of the target microinstruction is obtained, a data operation operation is performed with the first data based on the second data input along the second data path each time, and the second data is sent to the next operation unit through the second data path. After performing a data operation operation, the local control unit of the operation unit updates the number of times the first data has been operated. When the number of operations meets the preset number condition, the current result of the data operation operation is determined as the execution result of the target microinstruction. In this way, the operation units can share the data required for processing the microinstruction by sequentially transferring the second data downward based on the second data path. In this way, the operation of a GATHER instruction can be completed by obtaining the required data from the physical register stack only once. Therefore, the overall performance of the solution is better and the power consumption is lower.
[0107] Reference Figure 15 , is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 15 As shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other via the communication bus. The memory is used to store executable instructions, which enable the processor to execute the instruction execution method of the aforementioned embodiment. The executable instructions can constitute a program.
[0108] An embodiment of the present invention provides a readable storage medium having executable instructions stored thereon. When executed by one or more processors, the processors are enabled to execute the instruction execution method of the aforementioned embodiment.
[0109] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0110] Those skilled in the art will appreciate that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. It should be noted that all actions of acquiring signals, information, or data in this application are performed in compliance with the relevant data protection laws and policies of the country of residence and with authorization from the owner of the corresponding device. The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0111] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing terminal device to operate in a predictable manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0112] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0113] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0114] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0115] The above is a detailed introduction to an instruction execution method, an instruction execution device, an electronic device and a readable storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for executing an instruction, characterized in that: Applied to a computing array, the method comprises: For any operation unit in the operation array, when first data of a target microinstruction is obtained from a first data path, a data operation operation is performed on the first data based on second data input along a second data path each time, and the second data is sent to a next operation unit through the second data path; After performing the data operation once, the local control unit of the operation unit updates the number of times the first data has been operated once; In a case where the number of operations performed satisfies a preset number condition, the current result of the data operation is determined as the execution result of the target microinstruction.
2. The method according to claim 1, characterized in that The method further includes: receiving, by the entry arithmetic unit, the second data sent by the operand providing unit; Among them, each time the operand providing unit receives a microinstruction of a data query instruction, it sends the second data defined in the microinstruction to the entry operation unit, and sends the first data defined in the microinstruction to the operation units in the operation array based on the first data path in the order of unit numbers.
3. The method according to claim 1, characterized in that The arithmetic unit is provided with a unit status bit; the method further comprises: The control unit sets the unit status bit to a first value upon receiving the first data, and restores the unit status bit to a second value after detecting that the number of operations completed satisfies the preset number condition; the first value indicates that the operation unit is in an activated state, and the second value indicates that the operation unit is in an inactivated state; The sending of the second data to the next operation unit through the second data path includes: sending the second data to the next operation unit through the second data path when the unit status bit is the first value.
4. The method according to any one of claims 1 to 3, characterized in that: The target microinstruction is a microinstruction obtained by decoding the data query instruction; The first data is an index element defined by a first source operand of the target microinstruction, and the second data is a data element defined by a second source operand of the target microinstruction.
5. The method according to claim 4, characterized in that The performing a data operation on the first data based on the second data input along the second data path each time includes: For any received second data, based on each index element in the first data, query operations are performed on multiple data elements in the second data to obtain query results corresponding to each index element; The query results corresponding to the index elements are written into the destination elements corresponding to the index elements in the destination buffer.
6. The method according to claim 5, characterized in that The step of determining the current result of the data operation as the execution result of the target microinstruction when the number of operations performed satisfies a preset number condition includes: When the number of operations completed satisfies the preset number condition, reading each destination element of the destination buffer as the execution result of the target microinstruction; The method further includes: writing the execution result back to a physical register file through a third data path.
7. The method according to claim 1, characterized in that The method further comprises: Calculating the sum of the number of operations and the data length of the second data to obtain the queried length; In a case where the searched length reaches the total length of the second source operand of the data search instruction, it is determined that the number of operations completed satisfies the preset number condition.
8. The method according to claim 5, characterized in that The querying operation is performed on the plurality of data elements in the second data based on the index elements in the first data to obtain query results corresponding to the index elements, including: For any index element in the first data, dividing the index element into a first index component, a second index component, and a third index component in ascending order based on the number of the data elements; If the third index component is 0 and the second index component is equal to the microinstruction sequence number of the target microinstruction, generating a selection signal based on the first index component; selecting a target data element from the data elements of the second data based on the selection signal as a query result corresponding to the index element; When the third index component is not 0, and / or the second index component is not equal to the microinstruction sequence number, the query result corresponding to the index element is set to 0.
9. The method according to any one of claims 1 to 3, characterized in that: The target microinstruction is a microinstruction obtained by decoding the data query instruction; the first data is a data element defined by the second source operand of the target microinstruction, and the second data is an index element defined by the first source operand of the target microinstruction; The sending of the second data to the next operation unit through the second data path includes: sending the second data, the current operation result of the data operation, and the number of operations to the next operation unit through the second data path.
10. An instruction execution device, characterized in that: Applied to a computing array, the device comprises: a first processing module configured to, for any operation unit in the operation array, perform a data operation on the first data of the target microinstruction obtained from the first data path based on second data input along the second data path each time, and send the second data to the next operation unit through the second data path; an updating module, configured to update, by the local control unit of the computing unit, the number of times the first data has been computed after performing the data computing operation once; The first determining module is configured to determine the current result of the data operation as the execution result of the target microinstruction when the number of operations performed satisfies a preset number condition.
11. An electronic device, characterized in that: include: A processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus; The memory is used to store executable instructions, and the executable instructions enable the processor to execute the method according to any one of claims 1 to 9.
12. A readable storage medium, characterized in that: Executable instructions are stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Calculation method and related products
CN107992329A
An instruction scheduling method and processor including instruction scheduling unit
CN112379928A
Instruction processing device and method, processor, electronic equipment and storage medium
CN118733115A
Instruction transmitting unit, instruction execution unit, and related apparatus and method
US20220147351A1