Vector shifting method, processor and computer equipment
The vector shift instructions are received and spliced through the vector execution module, sharing the same shift logic circuit, solving the problem that vector shift instructions require multiple shift logic circuits in the prior art, and realizing area saving and layout simplification.
Patent Information
- Application Number
- CN202311800890.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
In the prior art, vector shift instructions need to design different shift logic circuits separately, resulting in a large number of shift logic circuits in the execution module, expanding the area, and making it difficult to layout and wiring.
The vector shift instruction is received through the vector execution module, and the vector indicating the shift request is identified by data, and a splicing operation is performed for the one-vector and two-vector shift instructions to share the same shift logic circuit.
Reduces the number of shift logic circuits required to be deployed in the execution module, saves the area of the execution module, and simplifies the layout and routing process.
Smart Images

Figure CN120216023A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computer technologies, and in particular, to a vector shift method, a processor, and a computer device. Background Art
[0002] Vector shift instructions are widely used in devices with vector processing acceleration engines, such as AI (Artificial Intelligence) processors. The shift logic circuit of the vector shift instruction is integrated in the execution module of the processor. Usually, vector shift instructions include single-vector right shift instructions, single-vector left shift instructions, and double-vector shift instructions. Therefore, different shift logic circuits need to be designed separately for different functional vector shift instructions, which leads to a relatively large number of shift logic circuits to be integrated in the execution module, the area of the execution module expands, and it is difficult to perform layout and wiring in the execution module. Summary of the Invention
[0003] The embodiments of the present application provide a vector shift method, a processor, and a computer device, which can enable different types of shift instructions to share the same shift logic circuit, reduce the number of shift logic circuits required to be deployed in the execution module, and save the area of the execution module. The technical solutions are as follows:
[0004] On the one hand, a vector shift method is provided, which is executed by a processor. The processor includes a vector execution module. The method includes:
[0005] The vector execution module receives a vector shift instruction. The vector shift instruction carries a data identifier, and the data identifier indicates the vector for which a shift is requested. The vector shift instruction is a single-vector shift instruction or a double-vector shift instruction;
[0006] When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains a first vector indicated by the data identifier, obtains a preset vector, and splices the first vector and the preset vector to obtain a first spliced vector. Or when the vector shift instruction belongs to a double-vector shift instruction, the vector execution module obtains a first vector and a second vector indicated by the data identifier, and splices the first vector and the second vector to obtain the first spliced vector. The value in the preset vector is 0;
[0007] The vector execution module shifts the first spliced vector to obtain a shifted vector, and determines the execution result of the vector shift instruction based on the shifted vector.
[0008] On the other hand, a processor is provided. The processor includes a vector execution module;
[0009] The vector execution module is configured to receive a vector shift instruction, where the vector shift instruction carries a data identifier indicating the vector for which a shift is requested, and the vector shift instruction is a single-vector shift instruction or a two-vector shift instruction;
[0010] The vector execution module is configured to, when the vector shift instruction is a single-vector shift instruction, obtain a first vector indicated by the data identifier, obtain a preset vector, and splice the first vector and the preset vector to obtain a first spliced vector; or, when the vector shift instruction is a two-vector shift instruction, obtain a first vector and a second vector indicated by the data identifier, and splice the first vector and the second vector to obtain the first spliced vector, where the values in the preset vector are 0;
[0011] The vector execution module is configured to shift the first spliced vector to obtain a shifted vector, and determine an execution result of the vector shift instruction based on the shifted vector.
[0012] Optionally, the number of vector execution modules is multiple, and each vector execution module includes a shift unit and a register;
[0013] For any shift unit, when the vector shift instruction is a single-vector shift instruction, the shift unit is configured to respectively read first sub-vectors indicated by the data identifier from registers in each vector execution module, splice the read multiple first sub-vectors to obtain the first vector, and splice the first vector and the preset vector to obtain the first spliced vector; or
[0014] For any shift unit, when the vector shift instruction is a two-vector shift instruction, the shift unit is configured to respectively read first sub-vectors and second sub-vectors indicated by the data identifier from registers in each vector execution module, splice the read multiple first sub-vectors to obtain the first vector, splice the read multiple second sub-vectors to obtain the second vector, and splice the first vector and the second vector to obtain the first spliced vector.
[0015] Optionally, each vector execution module corresponds to a different position identifier indicating a position in the vector; the shift unit is configured to:
[0016] Determine the position identifier of each first sub-vector, where the position identifier of the first sub-vector is the position identifier of the vector execution module to which the register where the first sub-vector is located belongs;
[0017] Concatenate the multiple first sub-vectors according to the position identifiers of the multiple first sub-vectors to obtain the first vector, so that the position of each first sub-vector in the first vector is the position indicated by the position identifier of the first sub-vector.
[0018] Optionally, for any vector execution module, the shift unit in the vector execution module is used to determine the sub-vector at the position indicated by the position identifier in the vector execution module in the execution result, and write the determined sub-vector into the register of the vector execution module.
[0019] Optionally, the number of vector execution modules is multiple, and each vector execution module corresponds to a different position identifier, and the position identifier indicates the position in the vector; the vector shift instruction further includes shift parameters, and the shift parameters include sub-parameters for n-level shift operations, and the sub-parameters are used to indicate whether to perform a shift operation, n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted;
[0020] The vector execution module is used to, when the sub-parameter of the first-level shift operation indicates to perform a shift operation, move the first concatenated vector by the number of bits required for the first-level shift operation in the moving direction indicated by the vector shift instruction to obtain the first initial shifted vector; when the sub-parameter of the first-level shift operation indicates not to perform a shift operation, use the first concatenated vector as the first initial shifted vector;
[0021] The vector execution module is used to filter the first initial shifted vector based on the position identifier of the vector execution module to obtain the first target shifted vector, and the first target shifted vector includes the value at the position indicated by the position identifier in the first initial shifted vector;
[0022] The vector execution module is used to perform the second-level shift operation to the n-level shift operation on the first target shifted vector to obtain the nth target shifted vector, and determine the nth target shifted vector as the execution result of the vector shift instruction.
[0023] Optionally, the vector execution module is used to:
[0024] Determine a second quantity, where the second quantity is equal to the sum of the number of bits to be shifted for the second-level shift operation to the n-level shift operation;
[0025] Determine the target values at the multiple positions indicated by the position identifier in the first initial shifted vector;
[0026] Construct the first target shift vector from the second quantity of pre-order values, multiple target values, and the second quantity of post-order values in the first initial shift vector, and pad with zeros when the quantity of the pre-order values or the post-order values is less than the second quantity; wherein, the pre-order values refer to the values in the first initial shift vector before the multiple target values, and the post-order values refer to the values in the first initial shift vector after the multiple target values.
[0027] Optionally, the vector execution module is configured to:
[0028] When the sub-parameter of the k-th level shift operation indicates to perform a shift operation, shift the (k - 1)-th target shift vector by the number of bits required for the k-th level shift operation in the moving direction indicated by the vector shift instruction to obtain the k-th initial shift vector, and filter the k-th initial shift vector to obtain the k-th target shift vector; when the sub-parameter of the k-th level shift operation indicates not to perform a shift operation, filter the (k - 1)-th target shift vector to obtain the k-th target shift vector; where k is a positive integer greater than 1 and not greater than n.
[0029] Optionally, the vector execution module is configured to:
[0030] Determine a third quantity, where the third quantity is equal to the number of bits required for the k-th level shift operation;
[0031] Remove the first third quantity of values and the last third quantity of values in the k-th initial shift vector to obtain the k-th target shift vector.
[0032] On the other hand, a computer device is provided, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the vector shift method as described in the above aspect.
[0033] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the vector shift method as described in the above aspect.
[0034] On the other hand, a computer program product is provided, including a computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the vector shift method as described in the above aspect.
[0035] For the solution provided in the embodiments of the present application, if the vector shift instruction is a single-vector shift instruction, a preset vector is concatenated with a vector to be shifted, the concatenated vector is shifted, and the execution result is determined based on the shifted vector. If the vector shift instruction is a two-vector shift instruction, two vectors to be shifted are concatenated, the concatenated vector is shifted, and the execution result is determined based on the shifted vector. Therefore, whether it is a single-vector shift instruction or a two-vector shift instruction, the logic of concatenation, shifting, and determining the execution result needs to be executed. That is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module. Description of the Drawings
[0036] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0037] Figure 1 is a schematic structural diagram of a processor provided by an embodiment of the present application;
[0038] Figure 2 is a flowchart of a vector shift method provided by an embodiment of the present application;
[0039] Figure 3 is a flowchart of another vector shift method provided by an embodiment of the present application;
[0040] Figure 4 is a schematic diagram of a multi-level shift operation provided by an embodiment of the present application;
[0041] Figure 5 is a flowchart of another vector shift method provided by an embodiment of the present application;
[0042] Figure 6 is a schematic diagram of another multi-level shift operation provided by an embodiment of the present application;
[0043] Figure 7 is a flowchart of another vector shift method provided by an embodiment of the present application;
[0044] Figure 8 is a schematic diagram of the input and output signals of a shift unit provided by an embodiment of the present application;
[0045] Figure 9It is a schematic diagram of a shift unit provided by an embodiment of the present application;
[0046] Figure 10 It is a flowchart of another vector shift method provided by an embodiment of the present application;
[0047] Figure 11 It is a schematic diagram of another multi - stage shift operation provided by an embodiment of the present application;
[0048] Figure 12 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present application;
[0049] Figure 13 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Detailed implementation manners
[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0051] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first vector may be called the second vector, and similarly, the second vector may be called the first vector.
[0052] Among them, at least one means one or more than one. For example, at least one vector execution module may be one vector execution module, two vector execution modules, three vector execution modules, etc., any integer greater than or equal to one. A plurality means two or more than two. For example, a plurality of vector execution modules may be two vector execution modules, three vector execution modules, etc., any integer greater than or equal to two. Each means each of at least one. For example, each vector execution module refers to each vector execution module among a plurality of vector execution modules. If there are 3 vector execution modules in a plurality of vector execution modules, then each vector execution module refers to each of the 3 vector execution modules.
[0053] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in the present application are all fully authorized by users or relevant parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0054] The vector shift method provided by the embodiments of the present application is executed by a processor, which includes a vector execution module for executing vector shift instructions. In some embodiments, the processor is disposed in a computer device. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle terminal, etc., but is not limited thereto.
[0055] Figure 1 is a schematic structural diagram of a processor provided by the embodiments of the present application. As Figure 1 shown, the processor includes a plurality of vector execution modules 101 (two are taken as an example in the figure), and each vector execution module 101 includes a shift unit 1011 and a register 1012. Among them, there is a communication connection between the respective vector execution modules 101.
[0056] Taking a single vector shift instruction as an example, for the first vector to be shifted, the memory divides the first vector into a plurality of first sub-vectors, and loads the plurality of first sub-vectors into the registers 1012 of the plurality of vector execution modules 101 respectively. The shift unit 1011 in each vector execution module 101 is used to read the first sub-vectors in the respective registers 1012, splice the read plurality of first sub-vectors to obtain the first vector, then splice the first vector with a preset vector to obtain a first spliced vector, and shift the first spliced vector based on the single vector shift instruction. Subsequently, each shift unit 1011 only needs to write the vector at its respective responsible position in the execution result of the single vector shift instruction into the register 1012 of its own vector execution module 101.
[0057] Taking the two - vector shift instruction as an example, for the first vector and the second vector to be shifted, the memory divides the first vector into multiple first sub - vectors, loads the multiple first sub - vectors into the registers 1012 of multiple vector execution modules 101 respectively, divides the second vector into multiple second sub - vectors, and loads the multiple second sub - vectors into the registers 1012 of multiple vector execution modules 101 respectively. The shift unit 1011 in each vector execution module 101 is used to read the first sub - vector and the second sub - vector in each register 1012, splice the read multiple first sub - vectors to obtain the first vector, splice the read multiple second sub - vectors to obtain the second vector, then splice the first vector and the second vector to obtain the first spliced vector, and perform a shift on the first spliced vector based on this two - vector shift instruction. Subsequently, each shift unit 1011 only needs to write the vector at its respective responsible position in the execution result of this two - vector shift instruction into the register 1012 of its own vector execution module 101.
[0058] Figure 2 is a flowchart of a vector shift method provided by an embodiment of the present application. The embodiment of the present application is executed by a processor, and the processor includes a vector execution module. Refer to Figure 2 This method includes:
[0059] 201. The vector execution module receives a vector shift instruction. The vector shift instruction carries a data identifier, and the data identifier indicates the vector requested to be shifted. The vector shift instruction is a single - vector shift instruction or a two - vector shift instruction.
[0060] In the embodiment of the present application, the processor includes a vector execution module, and the vector execution module is used to process vectors, such as vector operations, vector shifts, and vector splicing. Among them, the processor can be a processor with a vector calculation engine (that is, a vector execution module), such as an AI processor or a CPU (Central Processing Unit) that supports the RISC - V Vector extension.
[0061] When the vector execution module receives a vector shift instruction, it can perform a vector shift according to the indication of the vector shift instruction. The vector shift instruction carries a data identifier, and the data identifier is used to indicate the vector requested to be shifted. For example, the data identifier is the storage address of the vector. Among them, the vector shift instruction includes multiple types. The vector shift instruction can be a single - vector shift instruction or a two - vector shift instruction. The single - vector shift instruction can also be a single - vector left - shift instruction or a single - vector right - shift instruction, etc. The two - vector shift instruction can be a two - vector left - shift instruction, etc. Among them, the single - vector shift instruction indicates a shift on a single vector, and the two - vector shift instruction indicates a shift after splicing two vectors.
[0062] When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains the first vector indicated by the data identifier, obtains a preset vector, and splices the first vector and the preset vector to obtain a first spliced vector. Or when the vector shift instruction belongs to a two-vector shift instruction, the vector execution module obtains the first vector and the second vector indicated by the data identifier, and splices the first vector and the second vector to obtain a first spliced vector. The value in the preset vector is 0.
[0063] If the vector shift instruction received by the vector execution module belongs to a single-vector shift instruction, the data identifier in the single-vector shift instruction only indicates one first vector. In addition to obtaining the first vector, the vector execution module also obtains a preset vector, and the values in the preset vector are all 0. The vector execution module splices the first vector and the preset vector to obtain a first spliced vector.
[0064] If the vector shift instruction received by the vector execution module belongs to a two-vector shift instruction, the data identifier in the two-vector shift instruction indicates two vectors, that is, the first vector and the second vector. Then the vector execution module obtains the first vector and the second vector, and splices the first vector and the second vector to obtain a first spliced vector.
[0065] 203. The vector execution module shifts the first spliced vector to obtain a shifted vector, and determines the execution result of the vector shift instruction based on the shifted vector.
[0066] After obtaining the first spliced vector, the vector execution module shifts the first spliced vector according to the vector shift instruction to obtain a shifted vector. For example, if the shift direction indicated by the vector shift instruction is to the right, the first spliced vector is shifted to the right; if the shift direction indicated by the vector shift instruction is to the left, the first spliced vector is shifted to the left.
[0067] Since when the vector shift instruction belongs to a single-vector shift instruction, the first vector to be shifted is spliced with the preset vector, the shifted vector is not necessarily the final execution result of the vector shift instruction. It is also necessary to screen the values to be retained in the shifted vector. Therefore, after obtaining the shifted vector, it is necessary to determine the execution result of the vector shift instruction based on the shifted vector.
[0068] Since the single-vector shift instruction indicates processing of one vector and the double-vector shift instruction indicates processing of two vectors, considering that the processing logics of the single-vector shift instruction and the double-vector shift instruction are different and different shift logic circuits need to be deployed respectively to implement them, resulting in resource waste. Therefore, in the embodiments of the present application, for the single-vector shift instruction, an additional preset vector is spliced, which is equivalent to padding zeros to the first vector to be shifted, so that both the single-vector shift instruction and the double-vector shift instruction need to perform the splicing process, unifying the processing logics of the two vector shift instructions, enabling the single-vector shift instruction and the double-vector shift instruction to share the same shift logic circuit, achieving resource sharing and being beneficial to saving resources.
[0069] In the method provided by the embodiments of the present application, if the vector shift instruction is a single-vector shift instruction, the preset vector is spliced with one vector to be shifted, the vector obtained after splicing is shifted, and the execution result is determined based on the shifted vector. If the vector shift instruction is a double-vector shift instruction, the two vectors to be shifted are spliced, the vector obtained after splicing is shifted, and the execution result is determined based on the shifted vector. Therefore, whether it is a single-vector shift instruction or a double-vector shift instruction, the logics of splicing, shifting, and determining the execution result need to be performed, that is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module.
[0070] The above Figure 2 embodiment is only a brief description of the vector shift method. In some embodiments, the process of shifting a vector includes multi-level shift operations. For the detailed process of shifting a vector, reference can be made to the following Figure 3 embodiment. Figure 3 is a flowchart of another vector shift method provided by the embodiments of the present application. The embodiments of the present application are executed by a processor, and the processor includes a vector execution module. Refer to Figure 3 and the method includes:
[0071] 301. The vector execution module receives a vector shift instruction, and the vector shift instruction carries a data identifier, and the data identifier indicates the vector requested to be shifted. The vector shift instruction is a single-vector shift instruction or a double-vector shift instruction.
[0072] In the embodiments of the present application, the vector shift instruction further includes shift parameters, and the shift parameters include sub-parameters of n-level shift operations. The sub-parameters are used to indicate whether to perform a shift operation. n is a positive integer, and the number of bits to be shifted also corresponds to each level of shift operation.
[0073] In a possible implementation, the shift parameter is a vector parameter with n bits, and the i-th value in the vector parameter is the sub-parameter for the i-th level of shift operation. Here, i is a positive integer not greater than n. Optionally, when the sub-parameter is equal to 1, it indicates that the shift operation is to be performed, and when the sub-parameter is equal to 0, it indicates that the shift operation is not to be performed. Optionally, the number of bits to be shifted for the i-th level of shift operation is M * 2 n-i . Here, M is a preset value, and when shifting is performed at the byte granularity, M is equal to 8.
[0074] For example, n is equal to 5, M is equal to 8, and the shift parameter is rs1
[10111] . The sub-parameter for the first level of shift operation is 1, indicating that the first level of shift operation needs to be performed, and the number of bits to be shifted is 8 * 2 5-1 = 128; the sub-parameter for the second level of shift operation is 0, indicating that the second level of shift operation does not need to be performed; the sub-parameter for the third level of shift operation is 1, indicating that the third level of shift operation needs to be performed, and the number of bits to be shifted is 8 * 2 5-3 = 32; the sub-parameter for the fourth level of shift operation is 1, indicating that the fourth level of shift operation needs to be performed, and the number of bits to be shifted is 8 * 2 5-4 = 16; the sub-parameter for the fifth level of shift operation is 1, indicating that the fifth level of shift operation needs to be performed, and the number of bits to be shifted is 8 * 2 5-5 = 8.
[0075] In a possible implementation, the single-vector shift instruction includes a single-vector left shift instruction and a single-vector right shift instruction, and the two-vector shift instruction includes a two-vector left shift instruction. Optionally, the single-vector left shift instruction indicates to left-shift a single vector, the single-vector right shift instruction indicates to right-shift a single vector, and the two-vector left shift instruction indicates to splice two vectors and then left-shift, outputting the upper half of the data after left-shifting.
[0076] Optionally, taking the vector bit width of 256 as an example, vs1 and vs2 represent the vectors for which shifting is requested (i.e., the source operands), vd represents the execution result of the vector shift instruction (i.e., the destination operand), and rs1[4:0] represents the shift parameter with a bit width of 5. Then, the formats of the single-vector left shift instruction, the single-vector right shift instruction, and the two-vector left shift instruction are as follows.
[0077] (1) Single-vector left shift instruction: vshfl vd,vs1,rs1.
[0078] The logical expression is: vd[255:0] = vs1[255:0] << {rs1[4:0],3’b0}.
[0079] Among them, this logical expression represents shifting the vector vs1[255:0] to the left by {rs1[4:0], 3’b0} to obtain the vector vd[255:0]. {rs1[4:0], 3’b0} means concatenating three 0s on the right side (i.e., the lower bits) of rs1[4:0].
[0080] (2) Single vector right shift instruction: vshfr vd, vs1, rs1.
[0081] The logical expression is: vd[255:0] = vs1[255:0] >> {rs1[4:0], 3’b0}.
[0082] Among them, this logical expression represents shifting the vector vs1[255:0] to the right by {rs1[4:0], 3’b0} to obtain the vector vd[255:0]. {rs1[4:0], 3’b0} means concatenating three 0s on the right side (i.e., the lower bits) of rs1[4:0].
[0083] (3) Double vector left shift instruction: vshflc vd, vs1, vs2, rs1.
[0084] The logical expression is: tmp[511:0] = {vs1[255:0], vs2[255:0]} << {rs1[4:0], 3’b0}, vd[255:0] = tmp[511:256].
[0085] Among them, this logical expression represents concatenating the vectors vs1[255:0] and vs2[255:0], shifting the concatenated vector to the left by {rs1[4:0], 3’b0} to obtain the vector tmp[511:0], and retaining the upper half of the data tmp[511:256] in the vector tmp[511:0] to obtain the vector vd[255:0]. {rs1[4:0], 3’b0} means concatenating three 0s on the right side (i.e., the lower bits) of rs1[4:0].
[0086] In the embodiments of the present application, the vector shift instruction received by the vector execution module can be any one of the above single vector left shift instruction, single vector right shift instruction, and double vector left shift instruction.
[0087] 302. When the vector shift instruction belongs to a single vector shift instruction, the vector execution module obtains the first vector indicated by the data identifier, obtains the preset vector, concatenates the first vector and the preset vector to obtain the first concatenated vector, or when the vector shift instruction belongs to a double vector shift instruction, the vector execution module obtains the first vector and the second vector indicated by the data identifier, concatenates the first vector and the second vector to obtain the first concatenated vector, and the value in the preset vector is 0.
[0088] That is, regardless of whether the vector shift instruction is a single-vector shift instruction or a two-vector shift instruction, the vector execution module needs to perform the process of vector splicing to obtain the first spliced vector. The difference lies in that if the vector shift instruction is a single-vector shift instruction, the first vector to be shifted is spliced with a preset vector, and if the vector shift instruction is a two-vector shift instruction, the first vector to be shifted is spliced with the second vector. It can be understood that for a single-vector shift instruction, an additional preset vector is provided to participate in the splicing process.
[0089] 303. When the sub-parameter of the i-th level shift operation in the vector execution module indicates that a shift operation is to be performed, according to the shift direction indicated by the vector shift instruction, the (i - 1)-th initial shifted vector is shifted by the number of bits required for the i-th level shift operation to obtain the i-th initial shifted vector.
[0090] Where i is a positive integer not greater than n. That is, the i-th level shift operation mentioned in steps 303 and 304 can be any level of shift operation in the n-level shift operation. Optionally, the 0-th initial shifted vector is the first spliced vector.
[0091] For the i-th level shift operation, the vector execution module obtains the value parameter of the i-th level shift operation in the shift parameter. If the sub-parameter of the i-th level shift operation indicates that a shift operation is to be performed, then according to the shift direction indicated by the vector shift instruction, the (i - 1)-th initial shifted vector is shifted by the number of bits required for the i-th level shift operation to obtain the i-th initial shifted vector. Among them, the (i - 1)-th initial shifted vector is the shift result of the (i - 1)-th level shift operation, that is, the shift result of the previous level shift operation of the i-th level shift operation.
[0092] For example, n is equal to 5, M is equal to 8, the shift parameter is rs1
[10111] , and the number of bits required for the i-th level shift operation is 8 * 2 5-i , and the shift direction indicated by the vector shift instruction is to the right. Taking i = 2 as an example, the sub-parameter of the first level shift operation is equal to 1, indicating that a shift operation needs to be performed. Then the first spliced vector is shifted 128 bits to the right to obtain the 1-st initial shifted vector.
[0093] 304. When the sub-parameter of the i-th level shift operation in the vector execution module indicates that no shift operation is to be performed, the (i - 1)-th initial shifted vector is used as the i-th initial shifted vector.
[0094] For the i-th level shift operation, the vector execution module obtains the value parameter of the i-th level shift operation from the shift parameters. If the sub-parameter of the i-th level shift operation indicates that no shift operation is to be performed, there is no need to perform the i-th level shift operation, and the (i - 1)-th initial shift vector can be directly used as the i-th initial shift vector. Among them, the (i - 1)-th initial shift vector is the shift result of the (i - 1)-th level shift operation, that is, the shift result of the previous level shift operation of the i-th level shift operation.
[0095] For example, n is equal to 5, M is equal to 8, the shift parameter is rs1
[10111] , and the number of bits to be shifted in the i-th level shift operation is 8 * 2 5-i , and the shift direction indicated by the vector shift instruction is to the right. Taking i = 2 as an example, the sub-parameter of the second level shift operation is equal to 0, indicating that no shift operation needs to be performed. Then, there is no need to perform the second level shift operation on the first initial shift vector, and the first initial shift vector can be directly determined as the second initial shift vector.
[0096] It should be noted that during the shifting process in steps 303 - 304 above, the vector is not filtered. Therefore, the bit width of each obtained initial shift vector is equal to the bit width of the first spliced vector, that is, the bit width of the processed vector is always the initial bit width.
[0097] Figure 4 is a schematic diagram of a multi-level shift operation provided by an embodiment of the present application. As Figure 4 shown, n is equal to 5, M is equal to 8, the bit width of the first spliced vector is 512, denoted as the first spliced vector[511:0]. The number of bits to be shifted in the first, second, third, fourth, and fifth level shift operations are 128, 64, 32, 16, and 8 respectively. The shift sub-parameter is used to indicate whether to perform a shift operation, and the instruction type identifier is used to indicate the type of the vector shift instruction. The shift direction of the vector shift instruction can be determined according to the instruction type identifier.
[0098] As Figure 4As shown, first, a first-level shift operation is performed on the first splicing vector [511:0] to obtain the first initial shift vector [511:0]. A second-level shift operation is performed on the first initial shift vector [511:0] to obtain the second initial shift vector [511:0]. A third-level shift operation is performed on the second initial shift vector [511:0] to obtain the third initial shift vector [511:0]. A fourth-level shift operation is performed on the third initial shift vector [511:0] to obtain the fourth initial shift vector [511:0]. A fifth-level shift operation is performed on the fourth initial shift vector [511:0] to obtain the fifth initial shift vector [511:0]. It should be noted that when the shift sub-parameter is equal to 1, it means that a shift operation needs to be performed, and when the shift sub-parameter is equal to 0, it means that no shift operation needs to be performed.
[0099] 305. The vector execution module determines the execution result of the vector shift instruction based on the nth initial shift vector.
[0100] The nth initial shift vector is also the shift result corresponding to the last-level shift operation.
[0101] In a possible implementation manner, during the shifting process in the above steps 303 - 304, the vector is not filtered. Therefore, the bit width of the nth initial shift vector is equal to the bit width of the first splicing vector. The vector execution module determines the execution result of the vector shift instruction in the following manner.
[0102] When the vector shift instruction belongs to a single-vector shift instruction, the vector formed by the values in the first position of the nth initial shift vector is determined as the execution result of the single-vector shift instruction. The first position is the position of the first vector in the first splicing vector. If the vector shift instruction belongs to a single-vector shift instruction, it means that a preset vector is additionally spliced during vector shifting. Then, after obtaining the nth initial shift vector, the additional spliced part needs to be filtered, and the part in the first position of the nth initial shift vector is retained to obtain the execution result of the single-vector shift instruction. For example, the bit width of the first vector is 256, the bit width of the preset vector is 256, the preset vector is spliced behind the first vector to obtain the first splicing vector, and the bit width of the first splicing vector is 512. Then, the position of the first vector in the first splicing vector is [511:256], and the position of the preset vector in the first splicing vector is [255:0]. The bit width of the nth initial shift vector is 512, and the part of the nth initial shift vector located at [511:256] is determined as the execution result of the single-vector shift instruction.
[0103] When the vector shift instruction belongs to a two-vector shift instruction, the vector execution module determines the vector formed by the values at the second position in the nth initial shift vector as the execution result of the two-vector shift instruction, where the second position is the position indicated by the two-vector shift instruction to be retained. Among them, the two-vector shift instruction indicates the position to be retained. For example, if the bit width of the nth initial shift vector is 512, the position to be retained by the two-vector left shift instruction introduced in step 301 above is [511:256], that is, half of the data at the high position needs to be retained. Therefore, after obtaining the nth initial shift vector, only the part to be retained is used as the execution result of the two-vector shift instruction. Optionally, the two-vector shift instruction carries a target position identifier, which indicates the position to be retained in the vector.
[0104] In the method provided by the embodiments of the present application, if the vector shift instruction is a single-vector shift instruction, a preset vector is concatenated with the vector to be shifted, the concatenated vector is shifted, and the execution result is determined based on the shifted vector. If the vector shift instruction is a two-vector shift instruction, the two vectors to be shifted are concatenated, the concatenated vector is shifted, and the execution result is determined based on the shifted vector. Therefore, whether it is a single-vector shift instruction or a two-vector shift instruction, the logic of concatenation, shifting, and determining the execution result needs to be executed. That is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module.
[0105] Moreover, for the three types of vector shift instructions, namely the single-vector left shift instruction, the single-vector right shift instruction, and the two-vector left shift instruction, by unifying the processing logics of these three vector shift instructions, three different shift logic circuits can be merged into a set of shared shift logic circuits, realizing an optimization method of sharing circuit resources, which is beneficial to reducing resource occupation and further reducing the area of the processor.
[0106] In the above Figure 3 embodiment, after the multi-level shift operation, the last obtained initial shift vector needs to be filtered. It can be seen that there is redundant data processing during the multi-level shift operation, that is, there is still room for optimization during the multi-level shift operation. Based on this, the embodiments of the present application also propose a scheme for filtering the vector to be shifted in advance during the multi-level shift operation. For the detailed process, please refer to the following Figure 5 embodiment. Figure 5 is a flowchart of another vector shift method provided by the embodiments of the present application. The embodiments of the present application are executed by a processor, and the processor includes a vector execution module. Refer to Figure 5, the method includes:
[0107] 501. The vector execution module receives a vector shift instruction. The vector shift instruction carries a data identifier, and the data identifier indicates the vector for which the shift is requested. The vector shift instruction is a single-vector shift instruction or a two-vector shift instruction.
[0108] In an embodiment of the present application, the vector shift instruction further includes a shift parameter. The shift parameter includes sub-parameters for an n-level shift operation. The sub-parameters are used to indicate whether to perform the shift operation. n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted.
[0109] Among them, this step 501 is the same as the above step 301, and will not be elaborated here one by one.
[0110] 502. When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains the first vector indicated by the data identifier, obtains a preset vector, and splices the first vector and the preset vector to obtain a first spliced vector. Or when the vector shift instruction belongs to a two-vector shift instruction, the vector execution module obtains the first vector and the second vector indicated by the data identifier, and splices the first vector and the second vector to obtain a first spliced vector. The value in the preset vector is 0.
[0111] This step 502 is the same as the above step 302, and will not be elaborated here one by one.
[0112] 503. When the sub-parameter of the i-th level of shift operation indicates to perform the shift operation, the vector execution module shifts the (i - 1)-th target shift vector by the number of bits required for the i-th level of shift operation in the moving direction indicated by the vector shift instruction to obtain the i-th initial shift vector, and filters the i-th initial shift vector to obtain the i-th target shift vector.
[0113] Among them, i is a positive integer not greater than n. That is to say, the i-th level of shift operation mentioned in this step 503 and step 504 can be any level of the n-level shift operation. Optionally, the 0-th target shift vector is the first spliced vector.
[0114] For the i-th level shift operation, the vector execution module obtains the value parameter of the i-th level shift operation from the shift parameters. If the sub-parameters of the i-th level shift operation indicate that a shift operation is to be performed, then, in accordance with the moving direction indicated by the vector shift instruction, the (i - 1)-th target shift vector needs to be shifted by the number of bits required for the i-th level shift operation to obtain the i-th initial shift vector. This process is the same as that of step 303 described above. The difference is that the (i - 1)-th target shift vector is the result of filtering the shift vector obtained from the (i - 1)-th level shift operation. And in this step 503, after obtaining the i-th initial shift vector, the i-th initial shift vector also needs to be filtered to obtain the i-th target shift vector. It can be understood that after filtering the parts of the i-th initial shift vector that are not required for subsequent shift operations, the remaining part is used as the i-th target shift vector. The i-th target shift vector is a part of the i-th initial shift vector, that is, the bit width of the i-th target shift vector is smaller than the bit width of the i-th initial shift vector.
[0115] In the embodiments of the present application, during the multi-level shift operation, after each level of shift operation is completed, the currently obtained shift vector needs to be filtered, that is, the parts that are not required for subsequent shift operations are filtered out, and only the useful parts are retained, thereby reducing the bit width of the vector to be processed, reducing the amount of data to be processed, saving processing resources, and facilitating the improvement of processing efficiency.
[0116] In a possible implementation manner, the vector execution module filters the i-th initial shift vector to obtain the i-th target shift vector, including: determining a target number, where the target number is the sum of a first number and a preset number. The first number is equal to the number of bits required for the shift operations from the (i + 1)-th level shift operation to the n-th level shift operation, and the preset number is equal to the number of bits to be retained for the execution result of the vector movement instruction. The vector formed by the first target number of values in the i-th initial shift vector is determined as the i-th target shift vector.
[0117] Optionally, the bit widths of the first vector, the preset vector, and the second vector are all 256, and the bit width of the first concatenated vector is 512. n is equal to 5, the shift parameter is rs1
[10111] , and the number of bits required for the i-th level shift operation is 8 * 2 5-iThen, the number of bits to be shifted for the first, second, third, fourth, and fifth level shift operations are 128, 64, 32, 16, and 8 respectively. Taking i = 1 as an example, after the first level shift operation, the first initial shift vector is obtained. The number of bits to be shifted for the remaining four level shift operations is 8 + 16 + 32 + 64 = 120. The number of bits to be retained for the execution result of the vector shift instruction is 256, that is, the first quantity is equal to 120, the preset quantity is equal to 256, so the target quantity is equal to 376. Then, after the first level shift operation, the number of bits to be retained is 376. Therefore, the vector formed by the first 376 values in the first initial shift vector is determined as the first target shift vector. Among them, the first 376 values refer to the first 376 values with higher bits in the first initial shift vector.
[0118] For example, taking the first level shift operation as an example, the above first concatenated vector is denoted as shf_din[511:0]. Then, the first target shift vector is as follows. When the sub-parameter of the first level shift operation indicates to perform a shift operation and the shift direction of the vector shift instruction is to the right, the first target shift vector can be denoted as {128’b0, shf_din[511:264]}, where 128’b0 represents 128 0s filled on the left after shifting 128 bits to the right, and shf_din[511:264] represents the first 248 values with higher bits in the first concatenated vector shf_din[511:0]. When the sub-parameter of the first level shift operation indicates to perform a shift operation and the shift direction of the vector shift instruction is to the left, the first target shift vector can be denoted as shf_din[383:8], where shf_din[383:8] represents the values from the 8th bit to the 383rd bit in the first concatenated vector shf_din[511:0]. Since the shift direction of the vector shift instruction is to the left, the first 128 values with higher bits in the first concatenated vector shf_din[511:0] have been shifted out.
[0119] Optionally, the vector execution module determines the vector formed by the first target number of values in the i-th initial shift vector as the i-th target shift vector, including: when the shift direction indicated by the vector shift instruction is to the right, obtaining the first target number of values in the i-th initial shift vector; replacing the last first quantity of the obtained values with 0, and determining the vector formed by the current target number of values as the i-th target shift vector.
[0120] In the embodiments of the present application, if the shift direction indicated by the vector shift instruction is to the right, then the number of bits that still need to be shifted to the right in the remaining multi-stage shift operations is the first quantity. That is, among the first target quantity of values, the last first quantity of values will ultimately be shifted out. That is, the last first quantity of values are the values to be used, but they need to participate in the shift operation. Therefore, the last first quantity of values can be uniformly replaced with 0, thereby simplifying the processing process and facilitating the improvement of processing efficiency.
[0121] For example, the bit widths of the first vector, the preset vector, and the second vector are all 256, the bit width of the first concatenated vector is 512, and the number of bits to be shifted in the first, second, third, fourth, and fifth stage shift operations are 128, 64, 32, 16, and 8 respectively, and the first quantity is equal to 120. Taking the first stage shift operation as an example, the first concatenated vector is denoted as shf_din[511:0]. When the sub-parameter of the first stage shift operation indicates to perform the shift operation and the shift direction of the vector shift instruction is to the right, the first target shift vector can be denoted as {128’b0, shf_din[511:384], 120’b0}, where 128’b0 represents 128 0s supplemented on the left after shifting 128 bits to the right, and shf_din[511:384] represents the values of the first 128 high bits in the first concatenated vector shf_din[511:0]. 120’b0 represents 120 0s replaced by the last 120 values in shf_din[511:264].
[0122] 504. When the sub-parameter of the vector execution module in the i-th stage shift operation indicates not to perform the shift operation, the (i - 1)-th target shift vector is filtered to obtain the i-th target shift vector.
[0123] For the i-th stage shift operation, the vector execution module obtains the value parameter of the i-th stage shift operation in the shift parameter. If the sub-parameter of the i-th stage shift operation indicates not to perform the shift operation, there is no need to perform the i-th stage shift operation. Just filter the (i - 1)-th target shift vector directly to obtain the i-th target shift vector. It can be understood that after filtering the part of the (i - 1)-th target shift vector that is not required for subsequent shift operations, the remaining part is used as the i-th target shift vector. The i-th target shift vector is a part of the (i - 1)-th target shift vector, that is, the bit width of the i-th target shift vector is smaller than the bit width of the (i - 1)-th target shift vector.
[0124] In a possible implementation, the vector execution module filters the (i - 1)-th target shift vector to obtain the i-th target shift vector, including: determining a target quantity, where the target quantity is the sum of a first quantity and a preset quantity, the first quantity is equal to the number of bits to be shifted for the (i + 1)-th to n-th shift operations, and the preset quantity is equal to the number of bits to be retained for the execution result of the vector shift instruction. The vector formed by the first target quantity of values in the (i - 1)-th target shift vector is determined as the i-th target shift vector. Among them, this process is the same as the process of filtering the i-th initial shift vector to obtain the i-th target shift vector in step 503 above, and will not be elaborated here one by one.
[0125] For example, the bit widths of the first vector, the preset vector, and the second vector are all 256, the bit width of the first concatenated vector is 512, and the number of bits to be shifted for the first, second, third, fourth, and fifth shift operations are 128, 64, 32, 16, and 8 respectively, and the target quantity is equal to 376. Taking the first-level shift operation as an example, the first concatenated vector is denoted as shf_din[511:0]. When the sub-parameter of the first-level shift operation indicates that no shift operation is to be performed, the 1st target shift vector can be denoted as shf_din[511:136], and shf_din[511:136] represents the values of the first 376 high bits in the first concatenated vector shf_din[511:0].
[0126] Figure 6 is a schematic diagram of another multi-level shift operation provided by an embodiment of the present application. As Figure 6 shown, n is equal to 5, the bit widths of the first vector, the preset vector, and the second vector are all 256, the bit width of the first concatenated vector is 512, denoted as the first concatenated vector [511:0]. The number of bits to be shifted for the first, second, third, fourth, and fifth shift operations are 128, 64, 32, 16, and 8 respectively. The shift sub-parameter is used to indicate whether to perform a shift operation, and the instruction type identifier is used to indicate the type of the vector shift instruction. The shift direction of the vector shift instruction can be determined according to the instruction type identifier.
[0127] As Figure 6As shown, first, a first-level shift operation is performed on the first splicing vector [511:0] to obtain the first initial shift vector [511:0]. The first initial shift vector [511:0] is filtered to obtain the first target shift vector [375:0] with 376 bits. A second-level shift operation is performed on the first target shift vector [375:0] to obtain the second initial shift vector [375:0]. The second initial shift vector [375:0] is filtered to obtain the second target shift vector [311:0] with 312 bits. A third-level shift operation is performed on the second target shift vector [311:0] to obtain the third initial shift vector [311:0]. The third initial shift vector [311:0] is filtered to obtain the third target shift vector [279:0] with 280 bits. A fourth-level shift operation is performed on the third target shift vector [279:0] to obtain the fourth initial shift vector [279:0]. The fourth initial shift vector [279:0] is filtered to obtain the fourth target shift vector [263:0] with 264 bits. A fifth-level shift operation is performed on the fourth target shift vector [263:0] to obtain the fifth initial shift vector [263:0]. The fifth initial shift vector [263:0] is filtered to obtain the fifth target shift vector [255:0] with 256 bits.
[0128] It should be noted that when the shift sub-parameter is equal to 1, it means that a shift operation needs to be performed. When the shift sub-parameter is equal to 0, it means that no shift operation needs to be performed.
[0129] 505. The vector execution module determines the nth target shift vector as the execution result of the vector shift instruction.
[0130] The nth target shift vector is also the result of filtering the shift vector obtained from the last-level shift operation.
[0131] In the embodiment of the present application, when the vector shift instruction belongs to a single-vector shift instruction, since a preset vector is additionally spliced behind the first vector, only the upper half of the data needs to be retained subsequently. When the vector shift instruction belongs to a two-vector shift instruction, the two-vector shift instruction instructs to retain the upper half of the data. Therefore, regardless of the type of the vector shift instruction, only the upper half of the data needs to be retained. And since in the above steps 503 and 504, the vector has been filtered during the multi-level shift operation, that is, the useless lower half of the data has been filtered, the bit width of the nth target shift vector is equal to the bit width of the first vector. Therefore, there is no need to filter the nth target shift vector again, and the nth target shift vector is the final execution result of the vector shift instruction.
[0132] For the method provided by the embodiments of the present application, whether it is a single vector shift instruction or a double vector shift instruction, it is necessary to execute the logic of splicing, shifting, and determining the execution result. That is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module.
[0133] Moreover, during the multi-level shift operation, after each level of shift operation is completed, it is necessary to filter the currently obtained shift vector, that is, filter out the parts that are not required for subsequent shift operations, and only retain the useful parts, thereby reducing the bit width of the vector to be processed, reducing the amount of data to be processed, saving processing resources, and facilitating the improvement of processing efficiency.
[0134] In some embodiments, the number of vector execution modules in the processor is multiple, and each vector execution module has a shift unit and a register. Each vector execution module is only responsible for processing the data at some positions in the vector. In this case, the detailed process of the vector shift method can be referred to the following Figure 7 embodiments. Figure 7 FIG. Figure 7 is a flowchart of another vector shift method provided by the embodiments of the present application. The embodiments of the present application are executed by a processor, and the processor includes multiple vector execution modules. Refer to
[0135] 701. For any one of the multiple vector execution modules, the vector execution module receives a vector shift instruction, and the vector shift instruction carries a data identifier, where the data identifier indicates the vector for which the shift is requested, and the vector shift instruction is a single vector shift instruction or a double vector shift instruction.
[0136] In the embodiments of the present application, multiple vector execution modules cooperate to process a vector shift instruction. Therefore, each vector execution module will receive the vector shift instruction. The vector shift instruction is the same as the vector shift instruction in step 301 above, and will not be elaborated here one by one.
[0137] In a possible implementation manner, each vector execution module corresponds to a different position identifier, and the position identifier indicates a position in the vector. Each vector execution module is responsible for processing the data at the position indicated by the corresponding position identifier.
[0138] In the case where the vector shift instruction is a single-vector shift instruction, the data identifier in the single-vector shift instruction indicates a first vector, where the first vector is stored in a memory. When the memory receives the single-vector shift instruction, it divides the first vector into multiple first sub-vectors according to the position identifiers of multiple vector execution modules. The multiple first sub-vectors are respectively located at the positions indicated by the position identifiers of the multiple vector execution modules in the first vector, and the first sub-vector at the position indicated by any position identifier is passed into the register of the vector execution module corresponding to the position identifier. That is to say, the memory will pass each part of the first vector into the registers of the vector execution modules responsible for each part respectively.
[0139] In the case where the vector shift instruction is a two-vector shift instruction, the data identifier in the two-vector shift instruction indicates a first vector and a second vector. The manner in which the second vector is passed into each register is the same as that of the first vector passed into each register as described above, and will not be elaborated here.
[0140] 702. When the vector shift instruction in the vector execution module is a single-vector shift instruction, the shift unit in each vector execution module reads the first sub-vector indicated by the data identifier from the register, splices the multiple read first sub-vectors to obtain a first vector, and splices the first vector and a preset vector to obtain a first spliced vector.
[0141] For any vector execution module, when the vector shift instruction is a single-vector shift instruction, the shift unit in the vector execution module reads the first sub-vector indicated by the data identifier from the register in each vector execution module, so as to splice the multiple read first sub-vectors into a first vector. Optionally, the shift unit reads the first sub-vector from the register in the same vector execution module, sends the read first sub-vector to other shift units, and receives the first sub-vectors sent by other shift units, and splices the obtained multiple first sub-vectors to obtain the first vector.
[0142] In a possible implementation manner, each vector execution module corresponds to a different position identifier, and the position identifier indicates the position in the vector. Then, the shift unit splicing the multiple read first sub-vectors to obtain a first vector includes: determining the position identifier of each first sub-vector, where the position identifier of the first sub-vector is the position identifier of the vector execution module to which the register where the first sub-vector is located belongs, and splicing the multiple first sub-vectors according to the position identifiers of the multiple first sub-vectors to obtain a first vector, so that the position of each first sub-vector in the first vector is the position indicated by the position identifier of the first sub-vector.
[0143] Taking the example that the multiple vector execution modules include vector execution module 1 and vector execution module 2, vector execution module 1 includes a shift unit 11 and a register 12, and vector execution module 2 includes a shift unit 21 and a register 22. The bit width of the first vector is 256, denoted as vs1[255:0]. The first vector can be divided into a first sub-vector a and a first sub-vector b. The first sub-vector a is vs1[255:128], and the first sub-vector b is vs1[127:0]. That is, the first sub-vector a is the upper half of the data in the high bits of the first vector, and the first sub-vector b is the lower half of the data in the low bits of the first vector. Vector execution module 1 is responsible for processing the upper half of the data in the vector. Therefore, the register 12 stores the first sub-vector a. Vector execution module 2 is responsible for processing the lower half of the data in the vector. Therefore, the register 22 stores the first sub-vector b. After the shift unit 11 and the shift unit 21 obtain the first sub-vector a and the first sub-vector b, they both need to splice the first sub-vector b behind the first sub-vector a to obtain the first vector.
[0144] 703. When the vector shift instruction in the vector execution module belongs to a two-vector shift instruction, the shift unit in each vector execution module reads the first sub-vector and the second sub-vector indicated by the data identifier in the register respectively, splices the read multiple first sub-vectors to obtain the first vector, splices the read multiple second sub-vectors to obtain the second vector, and splices the first vector and the second vector to obtain the first spliced vector.
[0145] For any vector execution module, when the vector shift instruction belongs to a two-vector shift instruction, the shift unit in this vector execution module needs to read the first sub-vector and the second sub-vector respectively, and splice the multiple first sub-vectors into the first vector and the second sub-vectors into the second vector.
[0146] In a possible implementation manner, each vector execution module corresponds to a different position identifier, and the position identifier indicates a position in the vector. The shift unit splices the read multiple first sub-vectors to obtain a first vector, and splices the read multiple second sub-vectors to obtain a second vector, including: determining the position identifier of each first sub-vector, where the position identifier of the first sub-vector is the position identifier of the vector execution module to which the register where the first sub-vector is located belongs; splicing the multiple first sub-vectors according to the position identifiers of the multiple first sub-vectors to obtain a first vector, so that the position of each first sub-vector in the first vector is the position indicated by the position identifier of the first sub-vector. Determining the position identifier of each second sub-vector, where the position identifier of the second sub-vector is the position identifier of the vector execution module to which the register where the second sub-vector is located belongs; splicing the multiple second sub-vectors according to the position identifiers of the multiple second sub-vectors to obtain a second vector, so that the position of each second sub-vector in the second vector is the position indicated by the position identifier of the second sub-vector.
[0147] Wherein, the processing manner of the first sub-vector in step 703 is the same as the processing manner of the first sub-vector in step 702 above, and the processing manner of the second sub-vector in step 703 is the same as the processing manner of the first sub-vector in step 702 above by the same token, and will not be elaborated here.
[0148] 704. The shift unit in the vector execution module shifts the first spliced vector to obtain a shifted vector, and determines the execution result of the vector shift instruction based on the shifted vector.
[0149] This step 704 can adopt the shift manner of steps 303 - 305 in the above Figure 3 embodiment, or adopt the shift manner of steps 503 - 505 in the above Figure 5 embodiment, and will not be elaborated here.
[0150] 705. The shift unit in the vector execution module determines the sub-vector at the position indicated by the position identifier in the vector execution module in the execution result, and writes the determined sub-vector into the register of the vector execution module.
[0151] In the embodiment of the present application, since each vector execution module is only responsible for part of the data in the vector, that is, only responsible for the data at the position indicated by the corresponding position identifier, therefore, after the shift unit obtains the entire execution result of the vector shift instruction, it only needs to write the sub-vector at the position indicated by the position identifier in the execution result into the corresponding register. Furthermore, different parts of the execution result are respectively stored in the registers of each vector execution module.
[0152] For example, multiple vector execution modules include vector execution module 1 and vector execution module 2. The bit width of the execution result is 256. Vector execution module 1 is responsible for the data in [255:128] of the 256 bits, that is, it is responsible for the upper half of the data. Vector execution module 2 is responsible for the data in [127:0] of the 256 bits, that is, it is responsible for the lower half of the data. Then, the shift unit 11 in vector execution module 1 needs to write the sub-vector in [255:128] of the execution result into the register 12 in vector execution module 1, and the shift unit 21 in vector execution module 2 needs to write the sub-vector in [127:0] of the execution result into the register 22 in vector execution module 2.
[0153] For the method provided by the embodiments of the present application, whether it is a single-vector shift instruction or a two-vector shift instruction, it is necessary to execute the logic of splicing, shifting, and determining the execution result. That is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module.
[0154] Moreover, the processor includes multiple vector execution modules. By increasing the number of vector execution modules, the number of components required to be laid out in a single vector execution module can be reduced, which is beneficial to reducing the area of a single vector execution module and reducing the difficulty of layout and wiring for a single vector execution module.
[0155] Taking the example that the processor includes two vector execution modules, the vector execution module includes a shift unit and a register. Figure 8 It is a schematic diagram of the input signal and output signal of a shift unit provided by the embodiments of the present application. As Figure 8 shown, the input signals of the shift unit include vs1_0[127:0], vs1_1[127:0], vs2_0[127:0], vs2_1[127:0], rs1[4:0], shfr_en, shfl_en, shflc_en, and valu_idx, and the output signal of the shift unit includes vd_dout[127:0]. Among them, the meanings of each signal are as follows.
[0156] vs1_0[127:0] represents the first sub-vector read by the shift unit from the register of the current vector execution module. vs1_1[127:0] represents the first sub-vector read by the shift unit from the register of another vector execution module. vs2_0[127:0] represents the second sub-vector read by the shift unit from the register of the current vector execution module. vs2_1[127:0] represents the second sub-vector read by the shift unit from the register of another vector execution module. rs1[4:0] represents the shift parameter. shfr_en represents the enable signal of the single-vector right shift instruction. shfl_en represents the enable signal of the single-vector left shift instruction. shflc_en represents the enable signal of the double-vector left shift instruction. valu_idx represents the number of the vector execution module where the shift unit is located, and can also be understood as the position identifier corresponding to the vector execution module. Based on valu_idx, the position responsible by the vector execution module can be determined. vd_dout[127:0] represents the data written by the shift unit into the register of the current vector execution module.
[0157] Figure 9 is a schematic diagram of a shift unit provided by an embodiment of the present application, as Figure 9 shown. The shift unit includes a splicing component 1, a splicing component 0, a discrimination component 0, a multi-stage shift component, and a discrimination component 1.
[0158] The splicing component 1 is used to splice vs1_0[127:0] and vs1_1[127:0] according to valu_idx to obtain vs1[255:0], and vs1[255:0] represents the first vector. The splicing component 0 is used to splice vs2_0[127:0] and vs2_1[127:0] according to valu_idx to obtain vs2[255:0], and vs2[255:0] represents the second vector.
[0159] The discrimination component 0 is used to determine whether to set vs2_mux to 0 through the shflc_en signal. When shflc_en is equal to 1, it represents a double-vector shift instruction, and there is no need to set vs2_mux to 0. When shflc_en is equal to 0, it represents a single-vector shift instruction, and it is necessary to set vs2_mux to 0.
[0160] The multi-stage shift component is used to perform step-by-step shifting according to shfr_en. Each stage of the shift operation is determined whether to execute according to the corresponding sub-parameter in rs1, and the shift direction is determined according to shfr_en. When shfr_en is equal to 1, it represents a right shift. When shfr_en is equal to 0, it represents a left shift. The final execution result of the multi-stage shift component is shf_res[255:0].
[0161] The discrimination component 1 is used to determine which data in shf_res[255:0] to retain according to valu_idx, and write the retained vd_out[127:0] into the register. When valu_idx is equal to 1, vd_out[127:0]=shf_res[255:128]; when valu_idx is equal to 0, vd_out[127:0]=shf_res[127:0].
[0162] In some embodiments, as described in the above Figure 7 embodiment, when the processor includes multiple vector execution modules, since each vector execution module is only responsible for the data at some positions in the vector, only the data at some positions is retained after obtaining the final execution result. Therefore, there is redundant data processing during the multi-level shift operation, that is, there is still room for optimization during the multi-level shift operation. Based on this, on the basis of the above Figure 7 embodiment, the embodiment of the present application also proposes a solution to retain only the data at the positions responsible by the current vector execution module during the multi-level shift operation. For the detailed process, see the following Figure 10 embodiment.
[0163] Figure 10 FIG. is a flowchart of another vector shift method provided by the embodiment of the present application. The embodiment of the present application is executed by a processor, and the processor includes multiple vector execution modules. Each vector execution module includes a shift unit and a register. See Figure 10 This method includes:
[0164] 1001. For any one of the multiple vector execution modules, the vector execution module receives a vector shift instruction. The vector shift instruction carries a data identifier, and the data identifier indicates the vector for which the shift is requested. The vector shift instruction is a single-vector shift instruction or a two-vector shift instruction.
[0165] In the embodiment of the present application, each vector execution module corresponds to a different position identifier, and the position identifier indicates the position in the vector. Among them, the vector shift instruction further includes shift parameters, and the shift parameters include sub-parameters for n-level shift operations. The sub-parameters are used to indicate whether to perform a shift operation. n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted.
[0166] 1002. When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains the first vector indicated by the data identifier, obtains a preset vector, and splices the first vector and the preset vector to obtain a first spliced vector. Or when the vector shift instruction belongs to a two-vector shift instruction, the vector execution module obtains the first vector and the second vector indicated by the data identifier, and splices the first vector and the second vector to obtain a first spliced vector. The value in the preset vector is 0.
[0167] In a possible implementation, the vector execution module includes a shift unit and a register. This step 1002 is the same as steps 702 and 703 above, and will not be elaborated here.
[0168] 1003. When the sub-parameter of the first-level shift operation in the vector execution module indicates to perform a shift operation, the first spliced vector is shifted by the number of bits required for the first-level shift operation in the moving direction indicated by the vector shift instruction to obtain the first initial shifted vector; when the sub-parameter of the first-level shift operation indicates not to perform a shift operation, the first spliced vector is used as the first initial shifted vector.
[0169] This step 1003 is the same as the processes of steps 303 and 304 above, and will not be elaborated here.
[0170] 1004. The vector execution module filters the first initial shifted vector based on the position identifier of the vector execution module to obtain the first target shifted vector. The first target shifted vector includes the values at the positions indicated by the position identifier in the first initial shifted vector.
[0171] After obtaining the first initial shifted vector, filter the first initial shifted vector according to the positions indicated by the position identifier of the vector execution module to obtain the first target shifted vector, so that the first target shifted vector includes the values at the positions indicated by the position identifier in the first initial shifted vector.
[0172] In the embodiments of the present application, after completing the first-level shift operation, data at the positions required to be processed by the current vector execution module is selected in advance in the first initial shifted vector, thereby reducing the bit width of the vector to be shifted, simplifying the subsequent shift process, and being beneficial to improving the shift efficiency.
[0173] In a possible implementation, the vector execution module filters the first initial shifted vector based on the position identifier of the vector execution module to obtain the first target shifted vector, including: determining a second quantity, where the second quantity is equal to the sum of the number of bits required for the second-level shift operation to the nth-level shift operation, and determining the target values at multiple positions indicated by the position identifier in the first initial shifted vector. The first target shifted vector is formed by the second quantity of pre-order values, multiple target values, and the second quantity of post-order values in the first initial shifted vector, and padding with zeros when the number of pre-order values or post-order values is less than the second quantity. Here, the pre-order values refer to the values before the multiple target values in the first initial shifted vector, and the post-order values refer to the values after the multiple target values in the first initial shifted vector.
[0174] Optionally, the bit widths of the first vector, the preset vector, and the second vector are all 256, and the bit width of the first concatenated vector is 512. n is equal to 5, and the number of bits to be shifted for the i-th level of shift operation is 8 * 2 5-i . Then the number of bits to be shifted for the first, second, third, fourth, and fifth levels of shift operations are 128, 64, 32, 16, and 8 respectively. After the first level of shift operation, the second quantity is 8 + 16 + 32 + 64 = 120. The multiple vector execution modules include vector execution module 1 and vector execution module 2. Vector execution module 1 is responsible for the data on [255:128] in the 256 bits, that is, it is responsible for the upper half of the data, and vector execution module 2 is responsible for the data on [127:0] in the 256 bits, that is, it is responsible for the lower half of the data. Denote the first initial shifted vector as lv1_res[511:0].
[0175] For vector execution module 1, the target values at multiple positions it is responsible for are lv1_res[511:384]. Since there are no previous values for lv1_res[511:384], 120 zeros need to be supplemented in front of lv1_res[511:384]. The 120 subsequent values of lv1_res[511:384] are lv1_res[383:264]. Therefore, the first target shifted vector with 368 bits that vector execution module 1 needs to retain can be denoted as {120’b0, lv1_res[511:264]}.
[0176] For vector execution module 2, the target values at multiple positions it is responsible for are lv1_res[383:256]. The 120 previous values of lv1_res[383:256] are lv1_res[503:384], and the 120 subsequent values of lv1_res[383:256] are lv1_res[255:136]. Therefore, the first target shifted vector with 368 bits that vector execution module 2 needs to retain can be denoted as lv1_res[503:136].
[0177] 1005. The vector execution module performs the second-level to the n-th level of shift operations on the first target shifted vector to obtain the n-th target shifted vector.
[0178] After obtaining the first target shifted vector, the vector execution module can perform the second-level to the n-th level of shift operations on the first target shifted vector.
[0179] In a possible implementation, when the sub-parameter of the k-th level shift operation in the vector execution module indicates to perform a shift operation, the (k - 1)-th target shifted vector is shifted by the number of bits required for the k-th level shift operation in the moving direction indicated by the vector shift instruction to obtain the k-th initial shifted vector, and the k-th initial shifted vector is filtered to obtain the k-th target shifted vector; when the sub-parameter of the k-th level shift operation indicates not to perform a shift operation, the (k - 1)-th target shifted vector is filtered to obtain the k-th target shifted vector. Where k is a positive integer greater than 1 and not greater than n, that is, the k-th level shift operation is any one of the second level shift operation to the n-th level shift operation.
[0180] Optionally, the vector execution module filters the k-th initial shifted vector to obtain the k-th target shifted vector, including: determining a third quantity, where the third quantity is equal to the number of bits required for the k-th level shift operation, and removing the first third quantity of values and the last third quantity of values in the k-th initial shifted vector to obtain the k-th target shifted vector.
[0181] For example, n is equal to 5, and the number of bits required for the i-th level shift operation is 8 * 2 5-i . Then the number of bits required for the first, second, third, fourth, and fifth level shift operations are 128, 64, 32, 16, and 8 respectively. Taking the second shift operation as an example, the number of bits of the second initial shifted vector is 368, denoted as lv2[367:0], the third quantity is 64, then the first 64 values and the last 64 values in lv1[367:0] are removed to obtain the second target shifted vector, and the number of bits of the second target shifted vector is 240, denoted as lv2_[239:0].
[0182] Figure 11 is a schematic diagram of another multi-level shift operation provided by an embodiment of the present application. As Figure 11 shown, n is equal to 5, the bit widths of the first vector, the preset vector, and the second vector are all 256, and the bit width of the first concatenated vector is 512, denoted as the first concatenated vector[511:0]. The number of bits required for the first, second, third, fourth, and fifth level shift operations are 128, 64, 32, 16, and 8 respectively, and the number of bits responsible for a vector execution module is 128. The shift sub-parameter is used to indicate whether to perform a shift operation, and the instruction type identifier is used to indicate the type of the vector shift instruction. The shift direction of the vector shift instruction can be determined according to the instruction type identifier.
[0183] As Figure 11As shown, first, a first-level shift operation is performed on the first splicing vector [511:0] to obtain the first initial shift vector [511:0]. The first initial shift vector [511:0] is filtered to obtain the first target shift vector [367:0] with 368 bits. A second-level shift operation is performed on the first target shift vector [367:0] to obtain the second initial shift vector [367:0]. The second initial shift vector [367:0] is filtered to obtain the second target shift vector [239:0] with 240 bits. A third-level shift operation is performed on the second target shift vector [239:0] to obtain the third initial shift vector [239:0]. The third initial shift vector [239:0] is filtered to obtain the third target shift vector [175:0] with 176 bits. A fourth-level shift operation is performed on the third target shift vector [175:0] to obtain the fourth initial shift vector [175:0]. The fourth initial shift vector [175:0] is filtered to obtain the fourth target shift vector [143:0] with 144 bits. A fifth-level shift operation is performed on the fourth target shift vector [143:0] to obtain the fifth initial shift vector [143:0]. The fifth initial shift vector [143:0] is filtered to obtain the fifth target shift vector [127:0] with 128 bits.
[0184] It should be noted that when the shift sub-parameter is equal to 1, it means that a shift operation needs to be performed. When the shift sub-parameter is equal to 0, it means that no shift operation needs to be performed.
[0185] 1006. The vector execution module determines the nth target shift vector as the execution result of the vector shift instruction.
[0186] The nth target shift vector is also the result of filtering the shift vector obtained from the last-level shift operation.
[0187] In the embodiment of the present application, since part of the data responsible for the current vector execution module has been selected after the first-level shift operation, the nth target shift vector is the data responsible for the current vector execution module. The shift unit in the vector execution module can directly write the nth target shift vector into the register of the vector execution module, and there is no need to perform the step of selecting data in step 705 above.
[0188] For the method provided in the embodiment of the present application, whether it is a single-vector shift instruction or a two-vector shift instruction, the logic of splicing, shifting, and determining the execution result needs to be executed. That is, the entire execution logic is the same. Therefore, different types of shift instructions can share the same shift logic circuit, reducing the number of shift logic circuits required to be deployed in the execution module, saving the area of the execution module, and further reducing the difficulty of layout and wiring in the execution module.
[0189] Moreover, after completing the first-level shift operation, data at the position where the current vector required for the module processing is selected in the first initial shift vector in advance, so as to reduce the bit width of the vectors to be shifted, thus simplifying the subsequent shift process and being conducive to improving the shift efficiency.
[0190] The embodiment of the present application further provides a processor, and the processor includes a vector execution module.
[0191] The vector execution module is configured to receive a vector shift instruction, and the vector shift instruction carries a data identifier, and the data identifier indicates the vector for which the shift is requested. The vector shift instruction is a single-vector shift instruction or a two-vector shift instruction.
[0192] The vector execution module is configured to, when the vector shift instruction belongs to a single-vector shift instruction, obtain the first vector indicated by the data identifier, obtain a preset vector, splice the first vector and the preset vector to obtain a first spliced vector, or when the vector shift instruction belongs to a two-vector shift instruction, obtain the first vector and the second vector indicated by the data identifier, and splice the first vector and the second vector to obtain a first spliced vector, and the value in the preset vector is 0.
[0193] The vector execution module is configured to shift the first spliced vector to obtain a shifted vector, and determine the execution result of the vector shift instruction based on the shifted vector.
[0194] In a possible implementation manner, the vector shift instruction further includes shift parameters, and the shift parameters include sub-parameters for n-level shift operations. The sub-parameters are used to indicate whether to perform a shift operation. n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted.
[0195] The vector execution module is configured to, when the sub-parameter of the i-th level of shift operation indicates to perform a shift operation, move the (i - 1)-th initial shift vector by the number of bits required for the i-th level of shift operation in the moving direction indicated by the vector shift instruction to obtain the i-th initial shift vector; when the sub-parameter of the i-th level of shift operation indicates not to perform a shift operation, use the (i - 1)-th initial shift vector as the i-th initial shift vector; where i is a positive integer not greater than n.
[0196] The vector execution module is configured to determine the execution result of the vector shift instruction based on the n-th initial shift vector.
[0197] In a possible implementation manner, the number of bits of the n-th initial shift vector is equal to the number of bits of the first spliced vector.
[0198] A vector execution module, which is used to determine, when the vector shift instruction belongs to a single vector shift instruction, the vector formed by the values at the first position in the nth initial shift vector as the execution result of the single vector shift instruction, where the first position is the position of the first vector in the first concatenated vector;
[0199] A vector execution module, which is used to determine, when the vector shift instruction belongs to a two-vector shift instruction, the vector formed by the values at the second position in the nth initial shift vector as the execution result of the two-vector shift instruction, where the second position is the position indicated by the two-vector shift instruction to be retained.
[0200] In a possible implementation, the vector shift instruction further includes a shift parameter, and the shift parameter includes sub-parameters for n-level shift operations. The sub-parameters are used to indicate whether to perform a shift operation. n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted;
[0201] A vector execution module, which is used to, when the sub-parameter of the ith-level shift operation indicates to perform a shift operation, move the (i - 1)th target shift vector by the number of bits required for the ith-level shift operation in the moving direction indicated by the vector shift instruction to obtain the ith initial shift vector, and filter the ith initial shift vector to obtain the ith target shift vector; when the sub-parameter of the ith-level shift operation indicates not to perform a shift operation, filter the (i - 1)th target shift vector to obtain the ith target shift vector; where i is a positive integer not greater than n;
[0202] A vector execution module, which is used to determine the nth target shift vector as the execution result of the vector shift instruction.
[0203] In a possible implementation, the vector execution module is used for:
[0204] Determine the target quantity, where the target quantity is the sum of the first quantity and the preset quantity. The first quantity is equal to the number of bits required for the (i + 1)th-level to the nth-level shift operations, and the preset quantity is equal to the number of bits to be retained for the execution result of the vector movement instruction;
[0205] Determine the vector formed by the first target quantity of values in the ith initial shift vector as the ith target shift vector.
[0206] In a possible implementation, the vector execution module is used for:
[0207] When the shift direction indicated by the vector shift instruction is to the right, obtain the first target quantity of values in the ith initial shift vector;
[0208] Replace the last first quantity of the obtained values with 0, and determine the vector formed by the current target quantity of values as the ith target shift vector.
[0209] In a possible implementation, the number of vector execution modules is multiple, and each vector execution module includes a shift unit and a register;
[0210] For any shift unit, when the vector shift instruction belongs to a single-vector shift instruction, the shift unit is configured to respectively read the first sub-vectors indicated by the data identifier in the registers of each vector execution module, splice the multiple read first sub-vectors to obtain a first vector, and splice the first vector and a preset vector to obtain a first spliced vector; or,
[0211] For any shift unit, when the vector shift instruction belongs to a two-vector shift instruction, the shift unit is configured to respectively read the first sub-vector and the second sub-vector indicated by the data identifier in the registers of each vector execution module, splice the multiple read first sub-vectors to obtain a first vector, splice the multiple read second sub-vectors to obtain a second vector, and splice the first vector and the second vector to obtain a first spliced vector.
[0212] In a possible implementation, each vector execution module corresponds to a different position identifier, and the position identifier indicates a position in the vector; the shift unit is configured to:
[0213] Determine the position identifier of each first sub-vector, where the position identifier of the first sub-vector is the position identifier of the vector execution module to which the register where the first sub-vector is located belongs;
[0214] Splice the multiple first sub-vectors according to the position identifiers of the multiple first sub-vectors to obtain a first vector, so that the position of each first sub-vector in the first vector is the position indicated by the position identifier of the first sub-vector.
[0215] In a possible implementation, for any vector execution module, the shift unit in the vector execution module is configured to determine the sub-vector at the position indicated by the position identifier in the execution result, and write the determined sub-vector into the register of the vector execution module.
[0216] In a possible implementation, the number of vector execution modules is multiple, and each vector execution module corresponds to a different position identifier, and the position identifier indicates a position in the vector; the vector shift instruction further includes a shift parameter, and the shift parameter includes sub-parameters for n-level shift operations, and the sub-parameters are used to indicate whether to perform a shift operation, n is a positive integer, and each level of shift operation further corresponds to the number of bits to be shifted;
[0217] A vector execution module, configured to, when a sub-parameter of a first-level shift operation indicates to perform a shift operation, shift a first concatenated vector by the number of bits required for the first-level shift operation in the moving direction indicated by a vector shift instruction to obtain a first initial shifted vector; and when the sub-parameter of the first-level shift operation indicates not to perform a shift operation, use the first concatenated vector as the first initial shifted vector.
[0218] A vector execution module, configured to filter the first initial shifted vector based on a position identifier of the vector execution module to obtain a first target shifted vector, where the first target shifted vector includes values at positions indicated by the position identifier in the first initial shifted vector.
[0219] A vector execution module, configured to perform a second-level shift operation to an n-level shift operation on the first target shifted vector to obtain an nth target shifted vector, and determine the nth target shifted vector as the execution result of the vector shift instruction.
[0220] In a possible implementation, the vector execution module is configured to:
[0221] Determine a second quantity, where the second quantity is equal to the sum of the number of bits required for the second-level shift operation to the n-level shift operation;
[0222] Determine target values at multiple positions indicated by the position identifier in the first initial shifted vector;
[0223] Form a first target shifted vector with the second quantity of pre-order values, multiple target values, and the second quantity of post-order values in the first initial shifted vector, and pad with zeros when the number of pre-order values or post-order values is less than the second quantity; where the pre-order values refer to the values in front of the multiple target values in the first initial shifted vector, and the post-order values refer to the values behind the multiple target values in the first initial shifted vector.
[0224] In a possible implementation, the vector execution module is configured to:
[0225] When a sub-parameter of a k-level shift operation indicates to perform a shift operation, shift the (k - 1)th target shifted vector by the number of bits required for the k-level shift operation in the moving direction indicated by the vector shift instruction to obtain a kth initial shifted vector, and filter the kth initial shifted vector to obtain a kth target shifted vector; when the sub-parameter of the k-level shift operation indicates not to perform a shift operation, filter the (k - 1)th target shifted vector to obtain a kth target shifted vector; where k is a positive integer greater than 1 and not greater than n.
[0226] In a possible implementation, the vector execution module is configured to:
[0227] Determine a third quantity, where the third quantity is equal to the number of bits to be shifted for the k-th level shift operation;
[0228] Remove the first third quantity of values and the last third quantity of values in the k-th initial shift vector to obtain the k-th target shift vector.
[0229] It should be noted that: the processor provided in the above embodiment and the embodiment of the vector shift method belong to the same concept. For the specific implementation process, please refer to the method embodiment, which will not be elaborated here.
[0230] The embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the vector shift method in the above embodiment.
[0231] Optionally, the computer device is provided as a terminal. Figure 12 The structural schematic diagram of a terminal 1200 provided by an exemplary embodiment of the present application is shown. The terminal 1200 includes: a processor 1201 and a memory 1202.
[0232] The processor 1201 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1201 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The processor 1201 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1201 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1201 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process the computational operations related to machine learning.
[0233] The memory 1202 may include one or more computer-readable storage media, which may be non-transitory. The memory 1202 may further include high-speed random access memory, as well as non-volatile memory, such as one or more magnetic disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one computer program, and the at least one computer program is used to be possessed by the processor 1201 to implement the vector shift method provided in the method embodiments of the present application.
[0234] In some embodiments, the terminal 1200 may further optionally include: a peripheral device interface 1203 and at least one peripheral device. The processor 1201, the memory 1202, and the peripheral device interface 1203 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1203 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 1204, a display screen 1205, a camera assembly 1206, an audio circuit 1207, and a power supply 1208.
[0235] The peripheral device interface 1203 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1201 and the memory 1202. In some embodiments, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 may be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0236] The radio frequency circuit 1204 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1204 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1204 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1204 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1204 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1204 may further include a circuit related to NFC (Near Field Communication), which is not limited in this application.
[0237] The display screen 1205 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1205 is a touch display screen, the display screen 1205 also has the ability to collect touch signals on or above the surface of the display screen 1205. The touch signals can be input as control signals to the processor 1201 for processing. At this time, the display screen 1205 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, the display screen 1205 can be one, provided on the front panel of the terminal 1200; in other embodiments, the display screen 1205 can be at least two, respectively provided on different surfaces of the terminal 1200 or in a folding design; in other embodiments, the display screen 1205 can be a flexible display screen, provided on the curved surface or folding surface of the terminal 1200. Even, the display screen 1205 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1205 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0238] The camera component 1206 is used to capture images or videos. Optionally, the camera component 1206 includes a front camera and a rear camera. The front camera is disposed on the front panel of the terminal 1200, and the rear camera is disposed on the back of the terminal 1200. In some embodiments, there are at least two rear cameras, which can be any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera component 1206 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0239] The audio circuit 1207 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1201 for processing, or input to the radio frequency circuit 1204 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively disposed at different parts of the terminal 1200. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1201 or the radio frequency circuit 1204 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1207 may further include a headphone jack.
[0240] The power supply 1208 is used to supply power to each component in the terminal 1200. The power supply 1208 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1208 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0241] Those skilled in the art can understand that Figure 12 the structure shown in
[0242] does not limit the terminal 1200, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements. Figure 13It is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1300 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1301 and one or more memories 1302. Among them, at least one computer program is stored in the memory 1302, and the at least one computer program is loaded and executed by the processor 1301 to implement the methods provided by the above various method embodiments. Of course, the server may also have components such as wired or wireless network interfaces, keyboards, and input / output interfaces for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0243] An embodiment of the present application also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the vector shift method in the above embodiment.
[0244] An embodiment of the present application also provides a computer program product, including a computer program, and the computer program is loaded and executed by a processor to implement the operations performed by the vector shift method in the above embodiment.
[0245] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.
[0246] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.
Claims
1. A vector shift method, characterized in that, Executed by a processor, the processor including a vector execution module, the method comprising: The vector execution module receives a vector shift instruction, the vector shift instruction carrying a data identifier, the data identifier indicating a vector for which a shift is requested, the vector shift instruction being a single-vector shift instruction or a two-vector shift instruction; When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains a first vector indicated by the data identifier, obtains a preset vector, and splices the first vector and the preset vector to obtain a first spliced vector, or when the vector shift instruction belongs to a two-vector shift instruction, the vector execution module obtains a first vector and a second vector indicated by the data identifier, and splices the first vector and the second vector to obtain the first spliced vector, the value in the preset vector being 0; The vector execution module shifts the first spliced vector to obtain a shifted vector, and determines an execution result of the vector shift instruction based on the shifted vector.
2. The method according to claim 1, characterized in that, The vector shift instruction further includes a shift parameter, the shift parameter including sub-parameters for n levels of shift operations, the sub-parameters being used to indicate whether to perform a shift operation, n being a positive integer, and each level of shift operation further corresponding to the number of bits to be shifted; The vector execution module shifts the first spliced vector to obtain a shifted vector, and determines an execution result of the vector shift instruction based on the shifted vector, including: When the sub-parameter of the i-th level of shift operation indicates to perform a shift operation, the vector execution module shifts the (i - 1)-th initial shifted vector by the number of bits required for the i-th level of shift operation in the direction indicated by the vector shift instruction to obtain the i-th initial shifted vector; when the sub-parameter of the i-th level of shift operation indicates not to perform a shift operation, the (i - 1)-th initial shifted vector is used as the i-th initial shifted vector; where i is a positive integer not greater than n; The vector execution module determines an execution result of the vector shift instruction based on the n-th initial shifted vector.
3. The method according to claim 2, wherein The number of bits of the n-th initial shifted vector is equal to the number of bits of the first spliced vector; the vector execution module determines an execution result of the vector shift instruction based on the n-th initial shifted vector, including: When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module determines a vector formed by the values at a first position in the n-th initial shifted vector as the execution result of the single-vector shift instruction, the first position being the position of the first vector in the first spliced vector; When the vector shift instruction belongs to a two-vector shift instruction, the vector execution module determines a vector formed by the values at a second position in the n-th initial shifted vector as the execution result of the two-vector shift instruction, the second position being the position indicated to be retained by the two-vector shift instruction.
4. The method according to claim 1, wherein The vector shift instruction further includes a shift parameter, the shift parameter including sub-parameters for n levels of shift operations, the sub-parameters being used to indicate whether to perform a shift operation, n being a positive integer, and each level of shift operation further corresponding to the number of bits to be shifted; The vector execution module shifts the first concatenated vector to obtain a shifted vector, and determines the execution result of the vector shift instruction based on the shifted vector, including: When the sub-parameter of the i-th level shift operation indicates to perform a shift operation, the vector execution module shifts the (i - 1)-th target shifted vector by the number of bits required for the i-th level shift operation in the moving direction indicated by the vector shift instruction to obtain the i-th initial shifted vector, and filters the i-th initial shifted vector to obtain the i-th target shifted vector; when the sub-parameter of the i-th level shift operation indicates not to perform a shift operation, the (i - 1)-th target shifted vector is filtered to obtain the i-th target shifted vector; where i is a positive integer not greater than n; The vector execution module determines the n-th target shifted vector as the execution result of the vector shift instruction.
5. The method according to claim 4, characterized in that The filtering the i-th initial shifted vector to obtain the i-th target shifted vector includes: Determining a target quantity, where the target quantity is the sum of a first quantity and a preset quantity, the first quantity is equal to the number of bits required for the shift operations from the (i + 1)-th level to the n-th level, and the preset quantity is equal to the number of bits to be retained for the execution result of the vector movement instruction; Determining the vector formed by the first target quantity of values in the i-th initial shifted vector as the i-th target shifted vector.
6. The method according to claim 5, characterized in that The determining the vector formed by the first target quantity of values in the i-th initial shifted vector as the i-th target shifted vector includes: When the shift direction indicated by the vector shift instruction is to the right, obtaining the first target quantity of values in the i-th initial shifted vector; Replacing the last first quantity of values among the obtained values with 0, and determining the vector formed by the current target quantity of values as the i-th target shifted vector.
7. The method according to any one of claims 1-6, characterized in that, The number of the vector execution modules is multiple, and each vector execution module includes a shift unit and a register; When the vector shift instruction belongs to a single-vector shift instruction, the vector execution module obtains the first vector indicated by the data identifier, obtains a preset vector, and concatenates the first vector and the preset vector to obtain a first concatenated vector; or when the vector shift instruction belongs to a two-vector shift instruction, the vector execution module obtains the first vector and the second vector indicated by the data identifier, and concatenates the first vector and the second vector to obtain the first concatenated vector, including: For any shift unit, when the vector shift instruction belongs to a single-vector shift instruction, the shift unit reads the first sub-vectors indicated by the data identifier from the registers of each vector execution module respectively, concatenates the read multiple first sub-vectors to obtain the first vector, and concatenates the first vector and the preset vector to obtain the first concatenated vector; or, For any shift unit, when the vector shift instruction belongs to a two-vector shift instruction, the shift unit reads the first sub-vector and the second sub-vector indicated by the data identifier from the registers of each vector execution module respectively, concatenates the read multiple first sub-vectors to obtain the first vector, concatenates the read multiple second sub-vectors to obtain the second vector, and concatenates the first vector and the second vector to obtain the first concatenated vector.
8. The method according to claim 7, wherein Each vector execution module corresponds to a different position identifier, and the position identifier indicates the position in the vector; The step of concatenating the read multiple first sub-vectors to obtain the first vector includes: Determining the position identifier of each first sub-vector, where the position identifier of the first sub-vector is the position identifier of the vector execution module to which the register where the first sub-vector is located belongs; According to the position identifiers of the multiple first sub-vectors, concatenating the multiple first sub-vectors to obtain the first vector, so that the position of each first sub-vector in the first vector is the position indicated by the position identifier of the first sub-vector.
9. The method according to claim 8, wherein After the vector execution module shifts the first concatenated vector to obtain a shifted vector and determines the execution result of the vector shift instruction based on the shifted vector, the method further includes: For any vector execution module, the shift unit in the vector execution module determines the sub-vector at the position indicated by the position identifier of the vector execution module in the execution result, and writes the determined sub-vector into the register of the vector execution module.
10. The method according to claim 1, characterized in that, The number of vector execution modules is multiple, and each vector execution module corresponds to a different position identifier, and the position identifier indicates the position in the vector; the vector shift instruction further includes a shift parameter, and the shift parameter includes sub-parameters for n-level shift operations, and the sub-parameters are used to indicate whether to perform a shift operation, n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted; The vector execution module shifts the first concatenated vector to obtain a shifted vector and determines the execution result of the vector shift instruction based on the shifted vector, including: When the sub-parameter of the first-level shift operation indicates to perform a shift operation, the vector execution module moves the first concatenated vector by the number of bits required for the first-level shift operation in the moving direction indicated by the vector shift instruction to obtain the first initial shifted vector; when the sub-parameter of the first-level shift operation indicates not to perform a shift operation, the first concatenated vector is used as the first initial shifted vector; The vector execution module filters the first initial shifted vector based on the position identifier of the vector execution module to obtain the first target shifted vector, and the first target shifted vector includes the value at the position indicated by the position identifier in the first initial shifted vector; The vector execution module performs the second-level shift operation to the n-level shift operation on the first target shifted vector to obtain the nth target shifted vector, and determines the nth target shifted vector as the execution result of the vector shift instruction.
11. The method according to claim 10, wherein The vector execution module filters the first initial shift vector based on the position identifier of the vector execution module to obtain a first target shift vector, including: The vector execution module determines a second quantity, where the second quantity is equal to the sum of the number of bits to be shifted for the second-level to n-level shift operations; The vector execution module determines target values at multiple positions indicated by the position identifier in the first initial shift vector; The vector execution module forms the first target shift vector with the second quantity of pre-order values, multiple target values, and the second quantity of post-order values in the first initial shift vector, and pads with zeros when the number of pre-order values or post-order values is less than the second quantity; where the pre-order values refer to the values in front of the multiple target values in the first initial shift vector, and the post-order values refer to the values behind the multiple target values in the first initial shift vector.
12. The method according to claim 10, wherein The vector execution module performs the second-level to n-level shift operations on the first target shift vector to obtain an nth target shift vector, including: When the sub-parameter of the kth-level shift operation indicates that a shift operation is to be performed, the vector execution module moves the (k - 1)th target shift vector by the number of bits to be shifted in the kth-level shift operation in the moving direction indicated by the vector shift instruction to obtain a kth initial shift vector, and filters the kth initial shift vector to obtain a kth target shift vector; when the sub-parameter of the kth-level shift operation indicates that no shift operation is to be performed, the vector execution module filters the (k - 1)th target shift vector to obtain a kth target shift vector; where k is a positive integer greater than 1 and not greater than n.
13. The method according to claim 12, characterized in that, The filtering of the kth initial shift vector to obtain a kth target shift vector includes: Determining a third quantity, where the third quantity is equal to the number of bits to be shifted in the kth-level shift operation; Removing the first third quantity of values and the last third quantity of values in the kth initial shift vector to obtain the kth target shift vector.
14. A processor, characterized in that, The processor includes a vector execution module; The vector execution module is configured to receive a vector shift instruction, where the vector shift instruction carries a data identifier indicating the vector for which a shift is requested, and the vector shift instruction is a single-vector shift instruction or a two-vector shift instruction; The vector execution module is configured to, when the vector shift instruction belongs to a single-vector shift instruction, obtain a first vector indicated by the data identifier, obtain a preset vector, and splice the first vector and the preset vector to obtain a first spliced vector, or when the vector shift instruction belongs to a two-vector shift instruction, obtain a first vector and a second vector indicated by the data identifier, and splice the first vector and the second vector to obtain the first spliced vector, where the values in the preset vector are 0; The vector execution module is configured to shift the first spliced vector to obtain a shifted vector, and determine the execution result of the vector shift instruction based on the shifted vector.
15. The processor according to claim 14, wherein The vector shift instruction further includes a shift parameter, the shift parameter includes sub-parameters for n-level shift operations, the sub-parameters are used to indicate whether to perform a shift operation, n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted; The vector execution module is configured to, when the sub-parameter of the i-th level shift operation indicates to perform a shift operation, shift the (i - 1)-th initial shifted vector by the number of bits required for the i-th level shift operation in the moving direction indicated by the vector shift instruction to obtain the i-th initial shifted vector; when the sub-parameter of the i-th level shift operation indicates not to perform a shift operation, use the (i - 1)-th initial shifted vector as the i-th initial shifted vector; where i is a positive integer not greater than n; The vector execution module is configured to determine the execution result of the vector shift instruction based on the n-th initial shifted vector.
16. The processor according to claim 15, wherein The number of bits of the n-th initial shifted vector is equal to the number of bits of the first spliced vector; The vector execution module is configured to, when the vector shift instruction belongs to a single vector shift instruction, determine the vector formed by the values at the first position in the n-th initial shifted vector as the execution result of the single vector shift instruction, where the first position is the position of the first vector in the first spliced vector; The vector execution module is configured to, when the vector shift instruction belongs to a two-vector shift instruction, determine the vector formed by the values at the second position in the n-th initial shifted vector as the execution result of the two-vector shift instruction, where the second position is the position indicated to be retained by the two-vector shift instruction.
17. The processor according to claim 14, wherein The vector shift instruction further includes a shift parameter, the shift parameter includes sub-parameters for n-level shift operations, the sub-parameters are used to indicate whether to perform a shift operation, n is a positive integer, and each level of shift operation also corresponds to the number of bits to be shifted; The vector execution module is configured to, when the sub-parameter of the i-th level shift operation indicates to perform a shift operation, shift the (i - 1)-th target shifted vector by the number of bits required for the i-th level shift operation in the moving direction indicated by the vector shift instruction to obtain the i-th initial shifted vector, and filter the i-th initial shifted vector to obtain the i-th target shifted vector; when the sub-parameter of the i-th level shift operation indicates not to perform a shift operation, filter the (i - 1)-th target shifted vector to obtain the i-th target shifted vector; where i is a positive integer not greater than n; The vector execution module is configured to determine the n-th target shifted vector as the execution result of the vector shift instruction.
18. The processor according to claim 17, wherein, The vector execution module is configured to: Determine a target quantity, where the target quantity is the sum of a first quantity and a preset quantity, the first quantity is equal to the number of bits required for the (i + 1)-th level shift operation to the n-th level shift operation, and the preset quantity is equal to the number of bits to be retained for the execution result of the vector movement instruction; Determine the vector formed by the first target number of values in the i-th initial shift vector as the i-th target shift vector.
19. The processor according to claim 18, wherein The vector execution module is configured to: When the shift direction indicated by the vector shift instruction is to the right, obtain the first target number of values in the i-th initial shift vector; Replace the last first number of values among the obtained values with 0, and determine the vector formed by the current target number of values as the i-th target shift vector.
20. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor to implement the operations performed by the vector shift method according to any one of claims 1 to 13.
Citation Information
Cited By
Data shift operation method and device, electronic equipment and storage medium
CN122261518A