RISC-V vector sliding method and execution unit
By using the RISC-V vector sliding method, a cyclic shifter is used to slide the concatenated data, which solves the problem of implementing element sliding in the vector register, and realizes parallel data processing and improves system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENGDU QUNXIN MICROELECTRONICS TECHNOLOGY CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-12
AI Technical Summary
There is no existing technology that provides a specific implementation scheme for sliding elements in a vector register based on vector sliding instructions.
A RISC-V vector sliding method is provided. By receiving sliding configuration information and micro-operation requests sent by the instruction scheduling module, the method uses a cyclic shifter to perform sliding processing on the concatenated data to realize the sliding of vector register elements. Parallel processing of data is achieved through the concatenation module and the cyclic shifter.
It reduces circuit area, improves system performance, reduces data loading/storage times, reduces memory bandwidth pressure, and enables data sliding for different micro-operation requests to be completed within one execution cycle.
Smart Images

Figure CN122018987A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of RISC-V architecture and provides a RISC-V vector sliding method and execution unit. Background Technology
[0002] The vector sliding instruction in the RISC-V architecture indicates that elements are slid up or down in the vector register group, but no specific implementation scheme for sliding elements in the vector register based on the vector sliding instruction has been found in the existing technology. Summary of the Invention
[0003] This application provides a RISC-V vector sliding method and execution unit to address the lack of specific implementation schemes in the prior art for sliding elements in a vector register based on vector sliding instructions.
[0004] The first aspect of this application provides a RISC-V vector sliding method applied to an execution unit, comprising: The system receives sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; the micro-operation requests are obtained by splitting the vector sliding instruction by the instruction decoding module and are used to instruct the first vector register and the second vector register. The concatenated data is obtained by concatenating the data in the first vector register of all micro-operation request indications. The spliced data is processed by a cyclic shifter according to the sliding configuration information to obtain a cyclic shift result. The cyclic shift result is then split and stored in the second vector register indicated by the micro-operation request.
[0005] A second aspect of this application provides an execution unit, comprising: The receiving module is used to receive sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; the micro-operation requests are obtained by splitting the vector sliding instructions by the instruction decoding module and are used to instruct the first vector register and the second vector register. The splicing module is used to splice the data in the first vector register of all micro-operation request indications to obtain spliced data; A cyclic shifter is used to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result, and then split the cyclic shift result and store it in the second vector register indicated by the micro-operation request.
[0006] This application achieves RISC-V vector sliding by receiving sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; concatenating the data in the first vector register indicated by the micro-operation request to obtain concatenated data; using a cyclic shifter to perform sliding processing on the concatenated data according to the sliding configuration information to obtain a cyclic shift result; splitting the cyclic shift result and storing it in the second vector register indicated by the micro-operation request. Furthermore, using a cyclic shifter to achieve RISC-V vector sliding reduces circuit area, enables parallel data processing, and improves system performance. Regardless of the number of micro-operation requests, only one execution cycle is required.
[0007] To make the above and other objects, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A first flowchart of the RISC-V vector sliding method according to an embodiment of this application is shown; Figure 2 A flowchart of the first modification process according to an embodiment of this application is shown; Figure 3 A flowchart illustrating the first mask calculation process according to an embodiment of this application is shown; Figure 4 A flowchart illustrating the third modification process of an embodiment of this application is shown; Figure 5 A flowchart illustrating the calculation process of the second and third masks according to an embodiment of this application is shown; Figure 6 This illustration shows a schematic diagram of two upward offsets according to an embodiment of this application; Figure 7 This diagram illustrates two downward offsets according to an embodiment of this application. Figure 8 This illustration shows a schematic diagram of sliding up by one offset when the initial sliding index is 0, according to an embodiment of this application. Figure 9 This illustration shows a schematic diagram of sliding up by an offset when the initial sliding index is not 0, according to an embodiment of this application. Figure 10This illustration shows a schematic diagram of sliding down by one offset when the initial sliding index is 0, according to an embodiment of this application. Figure 11 This illustration shows a schematic diagram of sliding down by an offset when the starting sliding index is not 0, according to an embodiment of this application. Figure 12 A structural diagram of an execution unit according to an embodiment of this application is shown; Figure 13 A structural diagram of the first correction module according to an embodiment of this application is shown; Figure 14 A structural diagram of the second modification module according to an embodiment of this application is shown; Figure 15 A structural diagram of the third modification module according to an embodiment of this application is shown; Figure 16 Another structural diagram of the execution unit of this application embodiment is shown; Figure 17 This paper illustrates a schematic diagram of the conversion of the upward and downward offsets according to an embodiment of this application. Figure 18 A schematic diagram of a cyclic shifter according to an embodiment of this application is shown; Figure 19 A schematic diagram of the input module according to an embodiment of this application is shown.
[0010] Explanation of symbols in the attached drawings: 121. Receiving module; 122. Splicing module; 123. Cyclic shifter; 124. First Correction Module; 125. Second Correction Module; 126. Third Correction Module; 131. First mask generation logic; 132. AND logic; 1311. The First Calculator; 1312. Left shifter; 1313. Inverter; 141. First selector; 142. Splicer; 151. Second mask generation logic; 152. First mask processing logic; 153. Second mask processing logic; 154, OR logic; 1511. The Second Calculator; 1512, First right shift logic; 1513, Second Right Shift Logic; 1514. XOR logic; 161. Second selector; 162. Output Selector; 181. Controller; 182. Input module; 183. Third selector; 184. Shifter; 1901, First conversion circuit; 1902, Fourth Selector; 1903, Second conversion circuit. Detailed Implementation
[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0012] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0013] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.
[0014] The abbreviations involved in this application are defined as follows: bit: a unit of information in binary. slide: to slide; vlen: The bit width of the vector register, which is 128 bits by default in this application; sew: selected element width (element width), with values such as 8, 16, 32, and 64; lmul: vector length multiplier, with values of 1 / 8, 1 / 4, 1 / 2, 1, 2, 4, 8, etc. vl: The actual number of valid elements in the data to be shifted, where the data to be shifted is the concatenated data in the following text; vlmax: The maximum number of effective elements in the cyclic shifter; tail: Elements that are outside the valid element range; offset: The offset is calculated by multiplying the offset by the element width sew. src: source, indicating the source vector register (used for sliding), i.e., the first vector register in the following embodiments; dst: destination, which refers to the target vector register, i.e. the second vector register in subsequent embodiments.
[0015] The various logics described in this application are specifically hardware circuits.
[0016] In the RISC-V architecture, vector sliding instructions are used to instruct elements to move up and down in a vector register group, but no specific implementation scheme for sliding elements in a vector register based on vector sliding instructions has been found in the existing technology.
[0017] To address the aforementioned technical problems, some embodiments of this application provide a RISC-V vector sliding method applied to an execution unit, such as... Figure 1 As shown, the RISC-V vector sliding method includes: Step 101: Receive sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module.
[0018] The micro-operation request is obtained by splitting the vector sliding instruction by the instruction decoding module, and is used to instruct the first vector register and the second vector register. The micro-operation instruction obtained by the instruction decoding module is sent to the instruction scheduling module; the instruction scheduling module then sends the micro-operation instruction to the execution unit.
[0019] The vector sliding instructions included in RISC-V are shown in Table 1.
[0020] Table 1
[0021] In RISC-V, vector sliding instructions include two types: one type slides an element up or down from the vector register by a certain offset; the other type slides all elements up or down by one address bit, and the freed-up bit is placed into the value of the integer / floating-point register.
[0022] The sliding configuration information includes: element bit width, offset, and sliding type. Element bit width includes values such as 8-bit, 16-bit, 32-bit, and 64-bit, used to indicate the bit width of elements in the vector register. The offset indicates the number of elements to slide; multiplying the offset by the element bit width determines the shift amount, i.e., the number of bits to slide. The sliding type indicates the sliding direction, including upward and downward sliding. An upward sliding direction is to the left (moving upwards). A downward sliding direction is to the right (moving downwards).
[0023] The vector register has a bit width of 128 bits and a length multiplier of 8, meaning that when there are 8 micro-operation requests, it can simultaneously slide 1024 bits of data.
[0024] In practice, the number of micro-operation requests can be determined based on the number of sliding bits supported by the cyclic shifter.
[0025] Step 102: Concatenate the data in the first vector register of all micro-operation request indications to obtain concatenated data.
[0026] This step allows us to obtain the full concatenated data, thereby enabling the synchronous sliding of data in each first vector register.
[0027] Step 103: Use a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain the cyclic shift result.
[0028] The specific implementation process of this step can be found in subsequent embodiments. Through the use of a cyclic shifter, the splicing data sliding processing can be achieved.
[0029] Step 104: Split the cyclic shift result and store it in the second vector register of the micro-operation request indication.
[0030] In this step, the cyclic shift result is split according to the order of the micro-operation instructions, and the splitting order is the same as the splicing order in step 102.
[0031] This embodiment receives sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; it concatenates the data in the first vector register indicated by the micro-operation requests to obtain concatenated data; it uses a cyclic shifter to perform sliding processing on the concatenated data according to the sliding configuration information to obtain a cyclic shift result; and it splits the cyclic shift result and stores it in the second vector register indicated by the micro-operation requests. This enables RISC-V vector sliding, and using a cyclic shifter to implement RISC-V vector sliding reduces circuit area, enables parallel data processing, and improves system performance. It requires only one execution cycle regardless of the number of micro-operation requests.
[0032] Furthermore, this application reduces the number of data loading / storage operations and lowers memory bandwidth pressure by sliding data within the vector register. For example, in signal processing, directly sliding time-domain / frequency-domain data within the vector register can reduce the frequency of memory read / write operations.
[0033] In some embodiments of this application, step 103 above utilizes a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result, including: Based on the sliding configuration information, determine the element width, offset, and sliding type; Input the element width, offset, sliding type, and splicing data into the cyclic shifter so that the cyclic shifter can perform sliding processing on the splicing data according to the element width, offset, and sliding type to obtain the cyclic shift result.
[0034] In detail, the cyclic shifter determines the effective offset based on the input offset and the sliding type; it calculates the shift amount based on the effective offset and the element bit width; and it slides the spliced data according to the shift amount to obtain the cyclic shift result.
[0035] Specifically, determining the effective offset based on the input offset and the sliding type includes: when the sliding type is an upward slide, the input offset is determined to be the effective offset; when the sliding type is a downward slide, the input offset is inverted and incremented by 1 to obtain the effective offset.
[0036] In some embodiments of this application, for the upward slide, the cyclic shifter places the overflowing data on the left side onto the right side in sequence. Since the low-order data of the upward slide instruction itself needs to remain unchanged, no additional processing is required.
[0037] For the downward shift, the cyclic shifter places the overflowing data on the right side onto the left side in sequence. At this time, there will be excess data on the left side. Since the sum of the index value and the offset is greater than or equal to vlmax, the default output should be 0. Therefore, a mask value needs to be constructed to mask the excess data on the left side.
[0038] Specifically, step 103 above uses a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain the cyclic shift result, and also includes: when the sliding type is downward sliding, correcting the cyclic shift result.
[0039] like Figure 2 As shown, the process of correcting the cyclic shift result includes: Step 201: Calculate the first mask based on the maximum number of valid elements, the offset, and the amount to be shifted.
[0040] The first mask, denoted as slide_down_mask, is used to zero out numbers whose sum of index value and offset is greater than the maximum number of valid elements.
[0041] The value to be shifted is an all-1 value, and the amount of data to be shifted is equal to the maximum number of valid elements. For example, when the maximum number of valid elements is 128 bits, the value to be shifted is a 128-bit all-1 value.
[0042] Step 202: Perform a masking operation on the cyclic shift result using the first mask to obtain the first corrected shift result.
[0043] In some embodiments of this application, such as Figure 3 As shown, step 201 above calculates the first mask based on the maximum number of valid elements, the offset, and the amount to be shifted, including: Step 301: Subtract the offset from the maximum number of valid elements to obtain the left shift amount, that is, calculate the left shift amount using vlmax-offset.
[0044] Step 302: Move the quantity to be shifted to the left to obtain the first data.
[0045] Step 303: Invert the first data to obtain the first mask.
[0046] In some embodiments of this application, step 103 above uses a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result, and further includes: when the sliding type is upward sliding, the offset is 1, and the starting sliding index is 0, the cyclic shift result is corrected.
[0047] The process of correcting the cyclic shift result includes: The smallest index element in the cyclic shift result is replaced with an integer or floating-point value to obtain the second corrected shift result. The smallest index element refers to the element at index 0.
[0048] In some embodiments of this application, step 103 above uses a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result, and further includes: when the sliding type is downward and the offset is 1, the cyclic shift result is corrected.
[0049] like Figure 4 As shown, the process of correcting the cyclic shift result includes: Step 401: Calculate the second mask and the third mask based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted.
[0050] The second mask, denoted as mask_real, indicates the index position after removing the largest index among the actual valid elements. The third mask, denoted as mask_bit, indicates the largest index position among the actual valid elements.
[0051] Step 402: Perform a masking operation on the second mask and the cyclic shift result to obtain the second data.
[0052] Step 403: Perform a masking operation on the third mask and the integer or floating-point value to obtain the third data.
[0053] Step 404: Perform an OR operation on the second data and the third data to obtain the third corrected shift result.
[0054] In some embodiments of this application, such as Figure 5 As shown, step 401 above calculates the second mask and the third mask based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted, including: Step 501: Subtract the actual number of effective elements from the maximum number of effective elements to obtain the right shift amount, that is, calculate the right shift amount using vlmax-vl.
[0055] Step 502: Move the quantity to be shifted to the right by the right shift amount to obtain the fourth data.
[0056] Step 503: Shift the fourth data one bit to the right to obtain the second mask.
[0057] Step 504: Perform an XOR operation on the fourth data and the second mask to obtain the third mask.
[0058] To more clearly illustrate the RISC-V vector sliding process, the sliding process is explained below using six vector sliding instructions, where offset is the offset, vstart is the starting sliding index, vl is the actual number of valid elements, vlmax is the maximum number of valid elements, src is the source vector register, and dst is the destination vector register.
[0059] (1) An instruction to slide up by a certain offset, wherein the offset is determined by an immediate number or an integer.
[0060] like Figure 6 As shown, Figure 6 This diagram illustrates two upward offsets in an embodiment of this application. vlmax=10, offset=2, vstart=0, vl=6. Figure 6The topmost element is the index. When the index is less than max(vstart, offset), dst remains unchanged. When the index is greater than or equal to max(vstart, offset) and less than vl, dst is the value of src after sliding up by offset. When the index is greater than or equal to vl, dst remains unchanged.
[0061] (2) The instruction to slide down by a certain offset.
[0062] like Figure 7 As shown, Figure 7 This diagram illustrates the downward movement of two offsets according to an embodiment of this application. vlmax=10, offset=2, vstart=2, vl=6. Figure 7 The topmost element is the index. When the index is less than vstart, dst remains unchanged. When the index is greater than or equal to vstart and less than vl, dst is the value of src after sliding down by offset. The value after sliding needs to be adjusted; that is, if the sum of the index and offset is greater than or equal to vlmax, it is considered that src has exceeded the range and there is no valid element in dst, so it will default to outputting 0. When the index is greater than or equal to vl, dst remains unchanged.
[0063] (3) The instruction to slide up by one offset includes two cases: one is vstart=0, such as Figure 8 As shown; another is vstart≠0, as... Figure 9 As shown.
[0064] like Figure 8 As shown, Figure 8 This diagram illustrates an embodiment of the present application where the initial sliding index is 0 and the slider slides up by one offset. Figure 8 The topmost element is the index value of the element, vstart=0. At this time, the first element of dst will be stored as either the integer x[rs1] or the floating-point value f[rs1]. When the index value is greater than or equal to 1 and less than vl, dst is the value of src after sliding up 1; when the index value is greater than or equal to vl, dst remains unchanged.
[0065] like Figure 9 As shown, Figure 9 This diagram illustrates a scenario where the initial sliding index is non-zero, and the slider moves up by an offset. Figure 9 The topmost element is the index, where vstart ≠ 0. When the index is less than vstart, dst remains unchanged. When the index is greater than or equal to vstart and less than vl, dst is the value of src after sliding up 1. When the index is greater than or equal to vl, dst remains unchanged.
[0066] (4) The instruction to slide down by one offset includes two cases: one is vstart=0, such as Figure 10 As shown; another is vstart≠0, as... Figure 11 As shown.
[0067] like Figure 10 As shown, Figure 10 This diagram illustrates a scenario where the sliding index of an embodiment of this application starts at 0 and slides down by one offset. When the index value is greater than or equal to vstart and less than vl-1, dst is the value of src after sliding down by 1. When the index value is equal to vl-1, the vl-th element of dst is stored as either an integer x[rs1] or a floating-point value f[rs1]. When the index value is greater than or equal to vl, dst remains unchanged.
[0068] like Figure 11 As shown, Figure 11 This diagram illustrates a scenario where the sliding index of an embodiment of this application slides down by one offset when it is not zero. When the index value is less than vstart, dst remains unchanged. When the index value is greater than or equal to vstart and less than vl-1, dst is the value of src after sliding down by 1. When the index value is equal to vl-1, the vl-th element of dst is stored as either an integer x[rs1] or a floating-point value f[rs1]. When the index value is greater than or equal to vl, dst remains unchanged.
[0069] In some embodiments of this application, an execution unit is also provided, such as... Figure 12 As shown, it includes: The receiving module 121 is used to receive sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; the micro-operation requests are obtained by splitting the vector sliding instruction by the instruction decoding module and are used to instruct the first vector register and the second vector register.
[0070] The splicing module 122 is used to splice the data in the first vector register of all micro-operation request indications to obtain spliced data.
[0071] The cyclic shifter 123 is used to perform sliding processing on the spliced data according to the sliding configuration information to obtain the cyclic shift result, and then split the cyclic shift result and store it into the second vector register indicated by the micro-operation request.
[0072] This embodiment can implement RISC-V vector sliding, and uses a cyclic shifter to implement RISC-V vector sliding, which can reduce circuit area, achieve parallel data processing, and improve system performance. It only requires one execution cycle for different numbers of micro-operation requests.
[0073] In some embodiments of this application, for the upward slide, the cyclic shifter places the overflowing data on the left side onto the right side in sequence. Since the low-order data of the upward slide instruction itself needs to remain unchanged, no additional processing is required.
[0074] For the downward shift, the cyclic shifter places the overflowing data on the right side onto the left side in sequence. At this time, there will be excess data on the left side. Since the sum of the index value and the offset is greater than or equal to vlmax, the default output should be 0. Therefore, a mask value needs to be constructed to mask the excess data on the left side.
[0075] To construct mask values, such as Figure 13 As shown, the execution unit further includes a first correction module 124. The first correction module 124 includes a first mask generation logic 131 and an AND logic 132.
[0076] The first mask generation logic 131 is used to calculate the first mask based on the maximum number of valid elements, the offset, and the amount to be shifted when the sliding type is downsliding. The amount to be shifted is an all-1 value, and the amount of data to be shifted is equal to the maximum number of valid elements. For example, when the maximum number of valid elements is 128, the amount to be shifted is {128{1'b1}}.
[0077] The AND logic 132 is used to perform an AND operation on the first mask and the cyclic shift result to obtain the first corrected shift result.
[0078] In some implementations, such as Figure 13 As shown, the first mask generation logic 131 includes: The first calculator 1311 is used to obtain the left shift by subtracting the offset from the maximum number of valid elements.
[0079] The left shifter 1312 is used to shift the amount to be shifted to the left by the left shift amount to obtain the first data.
[0080] Inverter 1313 is used to invert the first data to obtain the first mask. The first mask is denoted as slide_down_mask.
[0081] For example, with an upward sliding type, a maximum number of valid elements (vlmax) of 10, an offset of 2, and a selected element width (sew) of 8, the amount to be shifted for 1024 bits of data is 10'b11_1111_1111. The left shift is vlmax - offset = 8 bits. Shifting 10'b11_1111_111 left by 8 bits yields 10'b11_0000_0000, which is then inverted to obtain the first mask 10'b00_1111_1111.
[0082] like Figure 13 As shown, Figure 13 The shift amount is a 128-bit all-1 value {128{1'b1}}. Correspondingly, the first mask is also 128 bits. At this point, the maximum number of valid elements in the corresponding circular shift register is (sew=8, lmul=8). Therefore, the first mask is fully valid when sew=8. When sew=16, the first mask [63:0] (slide-down mask) is valid. When sew=32, the first mask [31:0] is valid. When sew=64, the first mask [15:0] is valid.
[0083] For the downward sliding case, the first mask is used to mask the cyclic shift result (i.e., bits with a value of 1 in the first mask will keep the corresponding value of the cyclic shift result unchanged, while bits with a value of 0 in the first mask will set the corresponding value of the cyclic shift result to zero). Thus, based on the cyclic shifter and the first mask constructed during right sliding, the initial upward or downward sliding result can be obtained.
[0084] In some embodiments of this application, such as Figure 14 As shown, the execution unit also includes a second correction module 125.
[0085] The second correction module 125 is used to replace the smallest index element in the cyclic shift result with an integer or floating-point value when the sliding type is upward, the offset is 1, and the starting sliding index is 0, so as to obtain the second correction shift result.
[0086] In some implementations, such as Figure 14 As shown, the second correction module 125 includes a first selector 141 and a splicer 142.
[0087] One input to the first selector 141 is the smallest index element in the cyclic shift result, and the other input is an integer or floating-point value. The smallest index element is the first index element.
[0088] The inputs to splicer 142 are the element data after removing the smallest index element from the cyclic shift result and the output of the first selector 141. Splicer 142 splices the two input data to obtain the second corrected shift result.
[0089] In some embodiments of this application, such as Figure 15 As shown, the execution unit also includes: a third correction module 126.
[0090] The third correction module 126 includes: a second mask generation logic 151, a first mask processing logic 152, a second mask processing logic 153, and an OR logic 154.
[0091] The second mask generation logic 151 is used to calculate the second mask (abbreviated as mask_real) and the third mask (abbreviated as mask_bit) based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted when the sliding type is down and the offset is 1. The second mask is used to indicate the index position after removing the maximum index among the actual valid elements. The third mask is used to indicate the maximum index position of the actual valid elements.
[0092] The first masking logic 152 is used to perform a masking operation on the second mask and the cyclic shift result to obtain the second data. This masking operation is used to ensure that the non-tail portion after removing the largest index of the actual valid elements retains its original calculated value.
[0093] The second mask processing logic 153 is used to perform a masking operation on the third mask and integer or floating-point values to obtain the third data. This masking operation can modify the data at the third mask position to a positive number or floating-point value.
[0094] OR logic 154 is used to perform an OR operation on the second and third data to obtain the third corrected shift result.
[0095] In some embodiments of this application, such as Figure 15 As shown, the second mask generation logic 151 includes: The second calculator 1511 is used to obtain the right shift by subtracting the actual number of effective elements from the maximum number of effective elements.
[0096] The first right shift logic 1512 is used to shift the value to the right by the shift amount to obtain the fourth data.
[0097] The second right shift logic 1513 is used to shift the fourth data one bit to the right to obtain the second mask.
[0098] The XOR logic 1514 is used to perform an XOR operation on the fourth data and the second mask to obtain the third mask.
[0099] For example, if vlmax = 10 and vl = 6, then the value to be shifted is 128 bits of all 1s, i.e., 10'b11_1111_1111. Based on the above calculations of the second and third masks, the value to be shifted, 10'b11_1111_1111, needs to be right-shifted by 4 bits (calculated from vlmax - vl) to obtain the fourth data, 10'b00_0011_1111. Then, the fourth data is right-shifted by 1 bit to obtain the second mask, 10'b00_0001_1111. Finally, the fourth data and the second mask are XORed to obtain the third mask, 00_0010_0000.
[0100] Since vl=6 and it is a down-1 instruction, the second mask 10'b00_0001_1111 is used as the mask value for the cyclic shift result. That is, the lower 5 elements will use the cyclic shift result; while the last valid element, that is, the 6th element from low to high, needs to be selected through the third mask 00_0010_0000.
[0101] In some embodiments of this application, an execution unit is also provided, such as... Figure 16 As shown, all vector sliding instructions can be implemented through this execution unit. Furthermore, this application employs highly parallel processing, enabling instructions to be completed within a single execution cycle. Specifically, the execution unit includes: The receiving module 121 is used to receive sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; the micro-operation requests are obtained by splitting the vector sliding instructions by the instruction decoding module and are used to instruct the first vector register and the second vector register.
[0102] The splicing module 122 is used to splice the data in the first vector register of all micro-operation request indications to obtain spliced data src.
[0103] The left-shifting cyclic shifter 123 is used to perform sliding processing on the spliced data src according to the shift amount to obtain the cyclic shift result rotate_result. In this embodiment, the cyclic shifter 123 is a left-rotating cyclic shifter. In specific implementation, the cyclic shifter 123 determines the shift amount according to the sliding configuration information.
[0104] The first correction module 124 includes: a first calculator 1311, a left shifter 1312, an inverter 1313, and an AND logic 132. The first calculator 1311 is used to obtain the left shift amount by subtracting the offset from the maximum number of valid elements. The left shifter 1312 is used to shift the amount to be shifted {128{1'b1}} to the left by the left shift amount to obtain the first data. The inverter 1313 is used to invert the first data to obtain the first mask slide_down_mask. The AND logic 132 is used to perform an AND operation on the first mask slide_down_mask and the cyclic shift result rotate_result to obtain the first corrected shift result.
[0105] The second selector 161 is connected to the cyclic shifter 123 and the AND logic 132, and is used to select one from the cyclic shift result rotate_result output by the cyclic shifter 123 and the first corrected shift result output by the AND logic 132 to obtain the cyclic shift result rotate_result_real.
[0106] The second correction module 125 includes a first selector 141 and a concatenator 142. One input to the first selector 141 is the smallest index element in the cyclic shift result, and the other input is an integer or floating-point value. The concatenator 142 takes the elements remaining after removing the smallest index element from the cyclic shift result and the output of the first selector 141 as inputs, and concatenates the input data to obtain the second corrected shift result.
[0107] The third correction module 126 includes: a second calculator 1511, a first right shift logic 1512, a second right shift logic 1513, an XOR logic 1514, a first mask processing logic 152, a second mask processing logic 153, and an OR logic 154. The second calculator 1511 is used to obtain the right shift amount by subtracting the actual number of valid elements from the maximum number of valid elements. The first right shift logic 1512 is used to shift the amount to be shifted to the right by the shift amount to obtain the fourth data, where the amount to be shifted is {128{1'b1}}. The second right shift logic 1513 is used to shift the fourth data one bit to the right to obtain the second mask mask_real. The XOR logic 1514 is used to perform an XOR operation on the fourth data and the second mask mask_real to obtain the third mask mask_bit. The first mask processing logic 152 is used to perform a mask operation on the second mask mask_real and the cyclic shift result to obtain the second data. The second mask processing logic 153 is used to perform a masking operation on the third mask_bit and an integer or floating-point value to obtain the third data. The OR logic 154 is used to perform an OR operation on the second data and the third data to obtain the third corrected shift result.
[0108] The second correction module 125 and the third correction module 126 obtain the final cyclic shift result slide_final_result through the output selector 162.
[0109] In some embodiments of this application, the cyclic shifts of the upward sliding type (i.e., left rotation) and the downward sliding type (i.e., right rotation) are interchangeable. The following example illustrates this using a left rotation cycle. Figure 17 As shown. From Figure 17 As can be seen, taking sew=8 as an example, if offset=2, the corresponding shift amount is 2x8=16 bits. At this time, the left rotation amount is represented as 6'b010000, and the corresponding right rotation amount is {(~3'b010+1), 3'b000}=6'b110000.
[0110] Based on this, a cyclic shifter compatible with both upward and downward sliding can be designed. For example... Figure 18As shown, the cyclic shifter includes: a controller 181, multiple input modules 182, a third selector 183, and a shifter 184. The element bit widths of each input module 182 are different; for example, the element bit widths of the multiple input modules 182 from left to right are 8 bits, 16 bits, 32 bits, and 64 bits, respectively.
[0111] The controller 181 is used to generate a first signal and send it to the input module 182 according to the element bit width and sliding type in the sliding configuration information; generate a second signal and send it to the third selector 183 according to the element bit width; and send an offset to the input module 182. The first signal is used to indicate the valid input module 182 and the sliding type; the second signal is used to indicate the valid input module 182.
[0112] Input module 182 is used to perform steering processing on the offset; when the first signal indicates that the input module is valid and the sliding type is upward, the offset is converted into a shift amount according to the element bit width of the input module; when the first signal indicates that the input module is valid and the sliding type is downward, the steering offset is converted into a shift amount according to the element bit width of the input module; and the shift amount is output.
[0113] The third selector 183 connects to multiple input modules 182 and is used to determine and output the shift amount of the valid input modules based on the second signal.
[0114] The shifter 184 is connected to the third selector 183 and is used to perform cyclic sliding processing on the spliced data according to the shift amount to obtain the cyclic shift result.
[0115] In some embodiments of this application, such as Figure 19 As shown, Figure 19 The diagram shows an input module 182 with an element bit width of 8 bits. The input module 182 includes a first conversion circuit 1901, a fourth selector 1902, and a second conversion circuit 1903; each second selector has a different element bit width.
[0116] The first conversion circuit 1901 includes a first branch and a second branch. The input terminal of the first branch is connected to the input terminal of the second branch and is used to input the offset shift_num. The first branch is used to perform a redirection process on the offset; the second branch is used to directly output the offset.
[0117] In some embodiments of this application, such as Figure 19 As shown, the first branch includes an inverter and an adder. The inverter is used to reverse the shift amount. The adder is connected to the inverter and is used to add 1 to the reversed data to obtain the downward offset rshift_num. In some embodiments of this application, the first branch is a signal line.
[0118] The two input terminals of the fourth selector 1902 are connected to the output terminals of the first branch and the second branch, respectively; the output terminal of the fourth selector 1902 is connected to the second conversion circuit 1903. The control terminal of the fourth selector 1902 is connected to the controller, and is used to output the offset of the second branch, i.e., the offset lshift_num, when the first signal indicates that this input module is valid and the sliding type is upward; and to output the offset of the first branch, i.e., the offset rshift_num, when the first signal indicates that this input module is valid and the sliding type is downward.
[0119] The second conversion circuit 1903 is connected to the fourth selector 1902 and is used to convert the offset output by the fourth selector 1902 into a shift value according to the element bit width of this input module. Taking sew8 as an example, the shift value output by the second conversion circuit 1903 is represented as {shift_num_sew8, 3'h0}. For sew16, the shift value output by the second conversion circuit 1903 is represented as {shift_num_sew16, 4'h0}. For sew32, the shift value output by the second conversion circuit 1903 is represented as {shift_num_sew32, 5'h0}. For sew64, the shift value output by the second conversion circuit 1903 is represented as {shift_num_sew64, 6'h0}.
[0120] In some embodiments of this application, a specific RISC-V vector sliding method is also provided, including: (1) The instruction decoding module receives the vector sliding instruction, splits the vector sliding instruction into multiple micro-operation requests, and sends the micro-operation requests to the instruction scheduling module.
[0121] During this step, out-of-order sending of micro-operation requests to the instruction scheduling module is supported. Each micro-operation request contains a number, which the instruction scheduling module uses to determine the order of micro-operations.
[0122] For example, when luml=8, the number of micro-operations (UOPs) that need to be split is the largest. Taking vslideup.vx as an example, the data width of a single micro-operation is 128 bits. With luml=8, the vslideup_vx instruction needs to process 1024 bits of data. The corresponding micro-operation UOP splitting method is as follows: Uop0: dst0 = vslideup_vx(src0, rs1, dst0); Uop1: dst1 = vslideup_vx(src1, rs1, dst1); Uop2: dst2 = vslideup_vx(src2, rs1, dst2); Uop3: dst3 = vslideup_vx(src3, rs1 dst3); Uop4: dst4 = vslideup_vx(src4, rs1, dst4); Uop5: dst5 = vslideup_vx(src5, rs1, dst5); Uop6: dst6 = vslideup_vx(src6, rs1, dst6); Uop7: dst7 = vslideup_vx(src7, rs1, dst7).
[0123] (2) The instruction scheduling module receives multiple micro-operation requests sent by the instruction decoding module and schedules the multiple micro-operation requests to the execution unit.
[0124] (3) The execution unit receives multiple micro-operation requests and stores them in the cache. After all micro-operations are collected, the order of the micro-operation requests is determined, and the data in the first vector register indicated by the micro-operation requests is concatenated in order to obtain concatenated data; configuration information is sent to the circular shifter, wherein the configuration information includes: sew=8, offset=2 and the sliding type is upward; the circular shifter is used to perform sliding processing on the concatenated data according to the configuration information to obtain the circular shift result.
[0125] By using the above shifting, 1024 bits of data can be slid to obtain the sliding result within one execution cycle.
[0126] (4) The execution unit splits the 1024-bit sliding result into 128-bit segments and writes the split 128-bit results into the relevant second vector register. This process requires 8 writes.
[0127] The RISC-V vector sliding method and execution unit provided in this application utilize a cyclic shifter to implement the sliding of vector elements, which solves the problem of the lack of specific implementation schemes for RISC-V vector sliding instructions in the prior art. Furthermore, the cyclic shifter reduces circuit area, enables parallel data processing, and improves system performance; under different numbers of micro-operation requests, the sliding can be completed in only one execution cycle.
[0128] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0129] It should also be understood that, in the embodiments of this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this application generally indicates that the preceding and following related objects have an "or" relationship.
[0130] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0131] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0132] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, apparatuses, or units, or they may be electrical, mechanical, or other forms of connection.
[0133] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0134] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0135] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0136] This application uses specific embodiments to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A RISC-V vector sliding method, characterized in that, Applied to the execution unit, including: The system receives sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; the micro-operation requests are obtained by splitting the vector sliding instruction by the instruction decoding module and are used to instruct the first vector register and the second vector register. The concatenated data is obtained by concatenating the data in the first vector register of all micro-operation request indications. The spliced data is processed by a cyclic shifter according to the sliding configuration information to obtain a cyclic shift result. The cyclic shift result is then split and stored in the second vector register of the micro-operation request indication.
2. The method according to claim 1, characterized in that, The cyclic shifter is used to perform sliding processing on the spliced data according to the sliding configuration information to obtain the cyclic shift result, including: Based on the sliding configuration information, determine the element width, offset, and sliding type; The element width, offset, sliding type, and splicing data are input into the cyclic shifter, so that the cyclic shifter performs sliding processing on the splicing data according to the element width, offset, and sliding type to obtain the cyclic shift result.
3. The method according to claim 2, characterized in that, The method of using a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result also includes: When the sliding type is downward, the first mask is calculated based on the maximum number of valid elements, the offset, and the amount to be shifted; wherein, the amount to be shifted is an all-1 value, and the data amount of the amount to be shifted is equal to the maximum number of valid elements; The first mask is used to perform a masking operation on the cyclic shift result to obtain the first corrected shift result.
4. The method according to claim 3, characterized in that, Calculate the first mask based on the maximum number of valid elements, the offset, and the amount to be shifted, including: The left shift amount is obtained by subtracting the offset from the maximum number of effective elements. The first data is obtained by shifting the quantity to be shifted to the left by the left shift amount; The first mask is obtained by inverting the first data.
5. The method according to claim 3, characterized in that, The method of using a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result also includes: When the sliding type is upward, the offset is 1, and the starting sliding index is 0, the smallest index element in the cyclic shift result is replaced with an integer or floating-point value to obtain the second corrected shift result.
6. The method according to claim 3, characterized in that, The method of using a cyclic shifter to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result also includes: When the sliding type is down and the offset is 1, the second mask and the third mask are calculated based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted; wherein, the second mask is used to indicate the index position of the maximum index among the actual valid elements; the third mask is used to indicate the maximum index position of the actual valid elements. The second data is obtained by performing a masking operation on the second mask and the cyclic shift result; The third data is obtained by performing a masking operation on the third mask and the integer or floating-point value; A third corrected shift result is obtained by performing an OR operation on the second data and the third data.
7. The method according to claim 6, characterized in that, Calculate the second and third masks based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted, including: The right shift amount is obtained by subtracting the actual effective element amount from the maximum effective element amount. The fourth data is obtained by shifting the amount to the right by the rightward shift amount; Shifting the fourth data one bit to the right yields the second mask; The third mask is obtained by performing an XOR operation on the fourth data and the second mask.
8. An execution unit, characterized in that, include: The receiving module is used to receive sliding configuration information and multiple micro-operation requests sent by the instruction scheduling module; The micro-operation request is obtained by splitting the vector sliding instruction by the instruction decoding module, and is used to instruct the first vector register and the second vector register; The splicing module is used to splice the data in the first vector register of all micro-operation request indications to obtain spliced data; A cyclic shifter is used to perform sliding processing on the spliced data according to the sliding configuration information to obtain a cyclic shift result, and then split the cyclic shift result and store it in the second vector register of the micro-operation request indication.
9. The execution unit according to claim 8, characterized in that, The execution unit further includes: a first correction module; The first correction module includes: a first mask generation logic and an AND logic; The first mask generation logic is used to calculate the first mask based on the maximum number of valid elements, the offset, and the amount to be shifted when the sliding type is downward; wherein, the amount to be shifted is an all-1 value, and the data amount of the amount to be shifted is equal to the maximum number of valid elements; The AND logic is used to perform an AND operation on the first mask and the cyclic shift result to obtain the first corrected shift result.
10. The execution unit according to claim 9, characterized in that, The first mask generation logic includes: A first calculator is used to obtain the left shift by subtracting the offset from the maximum number of effective elements; A left shifter is used to shift the amount to be shifted to the left by the left shift amount to obtain the first data; An inverter is used to invert the first data to obtain a first mask.
11. The execution unit according to claim 9, characterized in that, The execution unit further includes: a second correction module; The second correction module is used to replace the smallest index element in the cyclic shift result with an integer or floating-point value when the sliding type is an upward slide, the offset is 1, and the starting sliding index is 0, so as to obtain the second corrected shift result.
12. The execution unit according to claim 11, characterized in that, The second correction module includes: a first selector and a splicer; One input to the first selector is the smallest index element in the cyclic shift result, and the other input is an integer or floating-point value; The input to the splicer is the element after removing the smallest index element from the cyclic shift result and the output of the first selector. The input data is spliced to obtain the second corrected shift result.
13. The execution unit according to claim 9, characterized in that, The execution unit further includes: a third correction module; The third correction module includes: a second mask generation logic, a first mask processing logic, a second mask processing logic, and / or logic; The second mask generation logic is used to calculate the second mask and the third mask based on the maximum number of valid elements, the actual number of valid elements, and the amount to be shifted when the sliding type is down and the offset is 1; wherein, the second mask is used to indicate the index position of the maximum index among the actual valid elements; the third mask is used to indicate the maximum index position of the actual valid elements. The first mask processing logic is used to perform a masking operation on the second mask and the cyclic shift result to obtain the second data; The second mask processing logic is used to perform masking operations on the third mask and integer or floating-point values to obtain the third data; The OR logic is used to perform an OR operation on the second data and the third data to obtain a third corrected shift result.
14. The execution unit according to claim 13, characterized in that, The second mask generation logic includes: The second calculator is used to obtain the right shift by subtracting the actual number of effective elements from the maximum number of effective elements. The first right shift logic is used to shift the quantity to be shifted to the right by the right shift amount to obtain the fourth data; The second right shift logic is used to shift the fourth data one bit to the right to obtain the second mask; The XOR logic is used to perform an XOR operation on the fourth data and the second mask to obtain the third mask.
15. The execution unit according to claim 9, characterized in that, The execution unit further includes: a second selector; The second selector is connected to the cyclic shifter and the first mask generation logic, and is used to select a cyclic shift result from the output of the cyclic shifter and the output of the first mask generation logic.
16. The execution unit according to claim 8, characterized in that, The cyclic shifter includes: a controller, multiple input modules, a third selector, and a shifter; each input module has a different element bit width; The controller is configured to generate a first signal and send it to the input module based on the element bit width and sliding type in the sliding configuration information; generate a second signal and send it to the third selector based on the element bit width; and send an offset to the input module; the first signal is used to indicate a valid input module and sliding type; the second signal is used to indicate a valid input module; The input module is used to perform a steering process on the offset; when the first signal indicates that the input module is valid and the sliding type is upward, the offset is converted into a shift amount according to the element width of the input module; when the first signal indicates that the input module is valid and the sliding type is downward, the steering offset is converted into a shift amount according to the element width of the input module; and the shift amount is output. The third selector is connected to the plurality of input modules and is used to determine and output the shift amount of the valid input module according to the second signal; The shifter is connected to the third selector and is used to perform a cyclic sliding process on the spliced data according to the shift amount to obtain a cyclic shift result.
17. The execution unit according to claim 16, characterized in that, The input module includes: a first conversion circuit, a fourth selector, and a second conversion circuit; each fourth selector has a different element bit width. The first conversion circuit includes a first branch and a second branch. The input terminal of the first branch is connected to the input terminal of the second branch and is used to input an offset. The first branch is used to perform a reversal process on the offset. The second branch is used to directly output the offset. The two input terminals of the fourth selector are respectively connected to the output terminals of the first branch and the second branch; the output terminal of the fourth selector is connected to the second conversion circuit; the control terminal of the fourth selector is connected to the controller, and is used to output the offset of the second branch when the first signal indicates that the input module is valid and the sliding type is upward; and to output the offset of the first branch when the first signal indicates that the input module is valid and the sliding type is downward. The second conversion circuit is connected to the fourth selector and is used to convert the offset output by the fourth selector into a shift amount according to the element bit width of this input module.
18. The execution unit according to claim 17, characterized in that, The first branch includes: an inverter and an adder; The inverter is used to reverse the shift amount; The adder is connected to the inverter and is used to add 1 to the inverted data.