Vector shift method, processor, and electronic device
The vector shift method simplifies the operation of SIMD processors by allowing a vector shift operation to be performed with a single instruction, thereby improving execution efficiency.
Patent Information
- Application Number
- JP2024534617
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-10
- Filing Date
- 2022-12-08
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2042-12-08
AI Technical Summary
Conventional SIMD processors require multiple instructions to implement a vector shift operation, leading to complex operation methods and reduced execution efficiency for specific functions.
A vector shift method that includes receiving an instruction with a register identifier and a shift parameter, executing the vector shift operation on source elements, and writing the target elements into a destination register, thereby simplifying the operation and improving efficiency.
The proposed method allows for a vector shift operation to be implemented with a single instruction, simplifying the operation method and enhancing the execution efficiency for specific functions.
Smart Images

Figure 0007699301000001 
Figure 0007699301000002 
Figure 0007699301000003
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a vector shift method, a processor, and an electronic device.
Background Art
[0002] With the development of multimedia applications, an increasing number of computing tasks of processors originate from the field of digital image processing, and image-based applications are becoming an inescapable workload in servers, desktop computers, personal mobile devices, i.e., embedded devices. In line with the actual situation of digital image processing software, updating the instruction set architecture and adding instruction support for operations frequently used in applications to the processor are the main directions of processor development, and it is also a simple and effective method to improve the performance of the processor for specific applications. Therefore, the single instruction multiple data (SIMD) structure has been increasingly added to more and more processors so that similar operations on a regular set of data are supported.
[0003] Currently, shift instructions are generally introduced in SIMD processors, and different shift instructions can meet different needs. However, in the conventional technical means, when implementing a vector shift operation for a specific function, multiple instructions are required to implement a series of operations, the operation method is complex, and the execution efficiency of a specific function is reduced.
Summary of the Invention
Problems to be Solved by the Invention
[0004] This application provides a vector shift method, a processor, and an electronic device to solve the problems in the prior art where multiple instructions are required to implement a vector shift operation, the operation method is complex, and the execution efficiency of a specific function is reduced.
Means for Solving the Problem
[0005] To solve the above problems, this application provides a vector shift method, and the method includes: Receiving an instruction including a register identifier and a shift parameter, where the register identifier includes a source register identifier and a destination register identifier, the source register identifier is for indicating a source register, the source register is a register for storing source elements to be operated during the execution of a vector shift operation, the destination register identifier is for indicating a destination register, the destination register is a register for storing target elements obtained after the execution of the vector shift operation, and the shift parameter is for indicating a rule based on which a vector shift operation is executed on the source elements; Executing the instruction to perform a vector shift operation on the source elements obtained from the source register according to the shift parameter, and obtaining target elements after the vector shift operation; Writing the target elements into the destination register.
[0006] To solve the above problems, this application provides a processor, and the processor includes a plurality of vector registers, an instruction decoding unit, and an execution unit. The plurality of vector registers include Source elements operated during vector shift operation a source register for storing, and a destination register. The command decoding unit is for decoding a vector shift command including a register identifier and a shift parameter. The register identifier includes a source register identifier for indicating a source register and a destination register identifier for indicating a destination register. In response to the vector shift command, the execution unit executes a vector shift operation on source elements obtained from the source register according to the shift parameter, obtains target elements after the vector shift operation, and writes the target elements into the destination register.
[0007] In order to solve the above problems, the present application provides an electronic device including a memory and one or more programs. The one or more programs are stored in the memory and are configured to be executed by one or more processors to perform the one or more of the above-described vector shift methods.
Advantages of the Invention
[0008] Compared with the prior art, the present application has the following advantages.
[0009] The vector shift method, processor, and electronic device provided by the embodiments of the present application add a register identifier and a shift parameter to the vector shift command, and use the register identifier to indicate the register storing the source elements to be operated during the execution of the vector shift operation and the register storing the target elements obtained after the execution of the vector shift operation, and use the shift parameter to indicate the rule based on which the vector shift operation is performed on the source elements. By doing so, the vector shift operation for a specific function can be implemented by one instruction instead of multiple instructions, the operation method is simple, and the execution efficiency for a specific function is improved.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Best Mode for Carrying Out the Invention
[0011] To further clarify and make it easier to understand the above objects, features, and advantages of the present application, the present application will be described in more detail below with reference to the drawings and specific embodiments.
[0012] The terms "first", "second", "third", etc. in the specification, claims, and drawings of the present application are for distinguishing similar objects and do not necessarily explain a specific order or sequence. It should be understood that such data can be exchanged when appropriate, such as being implemented in an order other than the order given in the illustration or description of the embodiments of the present application.
[0013] The following examples are described with reference to one processor, but other examples are applicable to other types of integrated circuits and logic devices. The foregoing technology and teachings of this application can be more readily applied to other types of circuits or semiconductor devices that benefit from higher pipeline throughput rates and performance improvements. The examples of this application are applicable to any processor or machine that performs data operations. However, this application is not limited to processors or machines that perform 256-bit, 128-bit, 64-bit, 32-bit, or 16-bit data operations, but is applicable to any processor and machine where the operation of combined data is desired.
[0014] In the following description, for the purpose of explanation to fully understand this application, a number of specific details are shown. However, those skilled in the art should recognize that these specific details are not necessary to practice this application. In other instances, some well-known electrical structures and circuits are not shown in detail to avoid unnecessarily obscuring this application. Further, the following description provides multiple examples, and the accompanying drawings show various examples for purposes of illustration. However, these examples are intended to provide some examples of this invention and are not intended to provide an exhaustive list of all possible embodiments of this invention, and thus should not be construed in a limiting sense.
[0015] In the following examples, instruction processing and distribution in the context of an execution unit will be described. However, other embodiments of the present invention may be implemented in the form of software. In one embodiment, the method of the present application is expressed as machine-executable instructions. These instructions can be used to cause a general-purpose processor or a dedicated processor programmed with these instructions to execute the steps of the present application. The present application may be provided as a computer program product or software, and the product or software may include a machine or computer-readable medium and instructions stored therein that can be used to program a computer (or other electronic device) to be executed by a processor according to the present application. Alternatively, the steps of the present invention can be executed by a dedicated hardware assembly including hardwired logic for executing the steps, or by any combination of a programmed computer assembly and a custom hardware assembly. These software can be stored in the memory within the system.
[0016] The vector shift method provided by the embodiment of the present application may have a CPU (Central Processing Unit) as the execution entity.
[0017] Example 1
[0018] Please refer to FIG. 1. FIG. 1 shows a flowchart of the steps of the vector shift method according to Embodiment 1 of the present application, and the vector shift process includes the following steps.
[0019] In step 101, an instruction including a register identifier and a shift parameter is received.
[0020] In an embodiment of the present application, an instruction refers to an instruction for executing a vector shift operation, and the instruction is an instruction executed by a processor. The instruction includes a register identifier and a shift parameter. The register identifier includes a source register identifier and a destination register identifier. The source register identifier is for indicating a source register, and the source register is a register that stores source elements to be operated during the execution of the vector shift operation. The destination register identifier is for indicating a destination register, and the destination register is a register that stores target elements obtained after the execution of the vector shift operation. The shift parameter is for indicating a rule based on which a vector shift operation is executed on the source elements.
[0021] Optionally, the number of source registers may be one or two, that is, the source elements are derived from one or two registers. Specifically, the number of source registers can be set according to business requirements, and the embodiments of the present application do not limit this.
[0022] Optionally, decode the received instruction, obtain the shift parameter included in the instruction. The shift parameter is for indicating a rule based on which a vector shift operation is executed on the source elements. In this example, the shift parameter can include parameters such as a shift amount and an opcode. Optionally, the opcode is a code represented in binary form, or the opcode is an identifier convertible to a binary code.
[0023] After decoding the instruction, execute step 102.
[0024] In step 102, execute the instruction, perform a vector shift operation on the source elements obtained from the source register according to the shift parameter, and obtain the target elements after the vector shift operation.
[0025] In an embodiment of the present application, after the CPU receives an instruction for executing a vector shift operation, the CPU executes the instruction to perform a vector shift operation on the source elements obtained from the source register according to the shift parameter, and can obtain the target elements after the vector shift operation.
[0026] After the target elements after the execution of the vector shift operation are obtained, step 103 is executed.
[0027] In step 103, the target elements are written into the destination register.
[0028] In an embodiment of the present application, after the target elements after the vector shift operation are obtained, the target elements can be written into the destination register.
[0029] Optionally, the shift amount and the shift operation rule can be determined according to the shift parameter, and the vector shift operation can be executed according to the shift amount and the shift operation rule. Specifically, it can be described in detail with reference to the following specific implementation forms.
[0030] Please refer to FIG. 2. FIG. 2 shows a flowchart of the steps of the target element acquisition method provided by the embodiment of the present application. The target element acquisition process includes the following steps.
[0031] In step 201, according to the shift parameter, the shift amount and the shift operation rule are determined, and there is at least one source element for which the vector shift operation is to be executed.
[0032] In the embodiments of the present application, the number of source registers may be one or more, the number of destination registers is one, the source register identifier and the destination register identifier may be the same or different, and the data type of the source element is any one of half-word, word, double-word, and quad-word. The shift amount can be used to indicate the number of shift bits of the source element that is operated when the vector shift Operation is executed. The shift amount can be derived from either an immediate number or a shift amount register. The immediate number is a parameter defined by the opcode among the shift parameters. The immediate number can determine its value with reference to the data type of the source element defined by the above opcode. The shift amount register is a register for storing the shift amount. When the shift amount is derived from the shift amount register, the shift amount is a group of data. For example, the shift amount can represent the shift situations of different source elements by including different bits. The shift operation rule refers to one or more operations executed on the source element.
[0033] After determining the shift amount and the shift operation rule according to the shift parameters, step 202 is executed.
[0034] In step 202, according to the shift amount and the shift operation rule, a corresponding shift operation is executed on the source element in the source register to generate a shift operation result.
[0035] In the embodiments of the present application, the shift operation rule refers to the shift operation method and / or constraint conditions for the elements in the source register.
[0036] After executing the corresponding shift operation on the source element in the source register according to the shift amount and the shift operation rule to generate a shift operation result, step 203 is executed.
[0037] In step 203, the shift operation result is determined as the target element.
[0038] In the embodiments of the present application, the shift parameter may include an opcode, and the opcode can be used to select a source element from the source register and to indicate the storage form of the target element in the destination register. Specifically, the vector shift Operation process can be described in detail with reference to the following implementation forms.
[0039] Please refer to FIG. 3. FIG. 3 shows a flowchart of the steps of the shift operation result generation method provided by the embodiments of the present application. The shift operation result generation process includes the following steps.
[0040] In step 301, according to the opcode, a source element for executing the vector shift Operation is selected from the source register, and the selected source element is determined as an operand.
[0041] In the embodiments of the present application, the shift parameter may include a shift amount and an opcode. The shift amount can be used to indicate the number of shift bits of the source element to be operated when the vector shift Operation is executed. The opcode can be used to indicate the shift operation rule executed on the source element in the source register and the target element in the destination register.
[0042] Alternatively, the format of the instruction is "opcode destination register, source register, shift amount". Exemplarily, the instruction can be represented as "[X]VSSR. {B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} vd / xd,vj / xj,ui]", where [X]VSSR represents the instruction name within the opcode, [X] is optional and is determined according to the register type, the part before the "." in {B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} represents the data type of the target element within the opcode, the part after the "." represents the data type of the source element within the opcode, B represents byte, H represents half-word, W represents word, D represents double-word, Q represents quad-word, U represents unsigned, vd / xd represents the destination register, vj / xj represents the source register, also, vd / xd can represent the source register, vj and vd are registers of the same bit width, xj and xd are registers of the same bit width, ui represents the immediate number, that is, the immediate number is the shift amount. Exemplarily, the instruction can also be represented as "[X]VSSR.{B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} vd / xd,vj / xj,vk / xk", [X]VSSR, {B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q}, vd / xd, and vj / xj represent the same meanings as in the previous example, and vk / xk represents the shift amount register.
[0043] Select a source element for performing a vector shift from the source registers according to the opcode, and after determining the selected source element as an operand, execute step 302. Operation After determining the selected source element as an operand, execute step 302.
[0044] In step 302, according to the opcode, perform a corresponding shift operation on the operand to generate a shift operation result.
[0045] In the embodiment of the present application, after the operand is determined, a corresponding shift operation can be performed on the operand according to the opcode to generate a shift operation result, and the elements included in the shift operation result become target elements.
[0046] After the shift operation result is generated, step 401 is executed.
[0047] Please refer to FIG. 4. FIG. 4 shows a flowchart of the steps of the target element storage method provided by the embodiment of the present application, and the target element storage process includes the following steps.
[0048] In step 401, according to the opcode, the storage form of the target element in the destination register is determined.
[0049] In the embodiment of the present application, the storage form refers to the rule for storing the target element in the destination register. Optionally, the storage form mainly represents the position rule for storing the target element in the destination register. Exemplarily, the storage form can include a form in which the upper half of the data in the target element is stored in the upper half of the position where the target element is arranged in the destination register, or a form in which the lower half of the data in the target element is stored in the lower half of the position where the target element is arranged in the destination register, or a form in which the data within the specified range of the target element is stored within the specified address range of the position where the target element is arranged in the destination register.
[0050] After determining the storage form of the target element in the destination register according to the opcode, step 402 is executed.
[0051] In step 402, according to the storage form, the target element is stored in the destination register.
[0052] In an embodiment of the present application, after the shift operation result is obtained, the shift operation result can be determined as a target element. After determining the storage form of the target element in the destination register according to the opcode, the target element can be stored in the destination register.
[0053] In the prior art, when implementing a vector shift operation, it is necessary to implement the vector shift with multiple instructions according to the vector shift needs, and the vector shift needs are determined according to the actual application. For example, if the vector shift needs are to right-shift the operands of two vector registers and round down to a half-width value, at least two right-shift instructions, two round-down instructions, and one saturation instruction to half-width are required to achieve the vector shift needs. In an embodiment of the present application, an instruction including a shift parameter is implemented, and different shift needs can be achieved with different shift parameters. Thereby, multiple vector shift needs can be achieved with one shift instruction, effectively reducing the system overhead and improving the execution efficiency of the vector shift for specific functions.
[0054] Hereinafter, the implementation process of the vector shift instruction in the case of different opcodes and different shift amounts will be described in detail with reference to Embodiments 2 to 4.
[0055] Embodiment 2
[0056] In an embodiment of the present application, the opcode can be a first type of vector opcode, and the source register includes a first source register and a second source register. With the first type of vector opcode, a source element can be obtained from the source register and a vector shift operation can be performed on the source element. As shown in FIG. 5, the processing method of the vector shift instruction can include the following steps.
[0057] In step 501, an instruction including a register identifier and a shift parameter is received.
[0058] In the embodiments of the present application, the meaning of the instruction and the parameters included in the instruction are as described in Embodiment 1, and will not be repeatedly described herein.
[0059] Optionally, there are two source registers, that is, the source elements are derived from two different registers. When there are multiple source registers, each source register identifier in all the source registers shall be different from the destination register identifier, or when there are multiple source registers, there is one source register identifier in all the source registers that is the same as the destination register identifier.
[0060] Optionally, decode the received instruction, obtain the shift parameter included in the instruction. The shift parameter is used to indicate the rule for performing a vector shift operation on the source element. In this example, the shift parameter can include parameters such as the shift amount and the opcode.
[0061] In step 502, according to the shift parameter, determine the shift amount and the shift operation rule.
[0062] In the embodiments of the present application, there is at least one source element on which the vector shift operation is performed. The shift amount is an immediate number, the shift operation rule is an opcode, the opcode is a first type of vector opcode, and the immediate number is a positive integer greater than or equal to 0.
[0063] Optionally, the first type of vector opcode is a code represented in binary form, or the opcode is an identifier that can be converted into a binary code. The format of the instruction is "opcode destination register, source register, shift amount". When the opcode is the first type of vector opcode, in actual implementation, the instruction is "VSSR 第1のタイプ.{B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} vd,vj,ui 第1のタイプ can be represented as VSSR 第1のタイプ is the instruction name within the first type of vector opcode, and {B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} is a parameter for indicating the data types of source and target elements within the first type of opcode. B represents byte, H represents halfword, W represents word, D represents doubleword, Q represents quadword, U represents unsigned, vd represents both the destination register and the source register, vj represents the source register, and ui 第1のタイプ represents the immediate value included in the instruction when the opcode is the first type of vector opcode. Exemplarily, VSSR 第1のタイプ1 .B.H is the first type of vector opcode convertible to binary form. For example, VSSR 第1のタイプ1 .B.H being the first type of vector opcode that is converted to the binary form 011100110101000001 can be cited.
[0064] Furthermore, ui 第1のタイプ is defined by the opcode, and ui 第1のタイプ can have its value determined by referring to the data types of the source and target elements. The ui 第1のタイプ is a parameter within a predetermined range (i.e., ui 第1のタイプ ∈ [minimum value, maximum value]), that is, ui 第1のタイプ has its minimum value determined according to the data types of the source and target elements, and ui 第1のタイプ has its maximum value being infinity. Exemplarily, when the first type of vector opcode is VSSR 第1のタイプ1 .B.H, ui 第1のタイプ has its minimum value as ui4, and when the first type of vector opcode is VSSR 第1のタイプ1 .H.W, ui 第1のタイプ has its minimum value as ui5, and when the first type of vector opcode is VSSR 第1のタイプ1 .W.D, ui第1のタイプ has a minimum value of ui6, and when the first type of vector operation code is VSSR 第1のタイプ1 .D.Q, then ui 第1のタイプ has a minimum value of ui7, and when the first type of vector operation code is VSSR 第1のタイプ1 .BU.H, then ui 第1のタイプ has a minimum value of ui4, and when the first type of vector operation code is VSSR 第1のタイプ1 .HU.W, then ui 第1のタイプ has a minimum value of ui5, and when the first type of vector operation code is VSSR 第1のタイプ1 .WU.D, then ui 第1のタイプ has a minimum value of ui6, and when the first type of vector operation code is VSSR 第1のタイプ1 .DU.Q, then ui 第1のタイプ has a minimum value of ui7. Thus, when the data type of the source element is a half-word and the data type of the target element is a byte, ui 第1のタイプ ∈ [ui4, infinity), and when the data type of the source element is a word and the data type of the target element is a half-word, ui 第1のタイプ ∈ [ui5, infinity), and when the data type of the source element is a double-word and the data type of the target element is a word, ui 第1のタイプ ∈ [ui6, infinity), and when the data type of the source element is a quad-word and the data type of the target element is a double-word, ui 第1のタイプ ∈ [ui7, infinity), and it can be seen that the target element may be an unsigned numerical value or a signed numerical value.
[0065] After determining the shift amount and shift operation rule according to the shift parameter, step 503 is executed.
[0066] In step 503, according to the first type of vector operation code, all source elements in the first source register are determined as one operand, and all source elements in the second source register are determined as another operand.
[0067] In the embodiments of the present application, in accordance with the first type of vector opcode, all elements in the first source register may be used as source elements, or some elements in the first source register may be used as source elements. Similarly, all elements in the second source register may be used as source elements, or some elements in the second source register may be used as source elements. Exemplarily, the first source register is register vd, the second source register is register vj. All elements in the first register vd are determined as source elements, and elements in the second register vj are determined as source elements. All source elements in the first source register are determined as one operand, and all source elements in the second source register are determined as another operand.
[0068] Optionally, the source elements respectively selected from the first source register and the second source register have the same data type, and the data type of the source elements is any one of half-word, word, double-word, and quad-word.
[0069] After determining the source elements obtained from the first source register as the operand of the second source register in accordance with the first type of vector opcode, step 504 is executed.
[0070] In step 504, in accordance with the first type of vector opcode, Operand in the first source register and operand in the second source register After splicing, a first spliced vector is generated.
[0071] In an embodiment of the present application, after splicing the operands in the first source register and the operands in the second source register left and right, a first splice vector is generated. The position setting for splicing the operands in the first source register and the operands in the second source register left and right is determined according to the position of the source register identifier in the instruction. That is, when the first source register identifier is the source register identifier immediately after the vector opcode of the first type in the instruction, and the second source register identifier is the source register identifier arranged after the first source register identifier in the instruction, the operand in the first source register is arranged on the left side, the operand in the second source register is arranged on the right side, and the first splice vector is generated. However, when the second source register identifier is the source register identifier immediately after the vector opcode of the first type in the instruction, and the first source register identifier is the source register identifier arranged after Second the source register identifier in the instruction, the operand in the second source register is arranged on the left side, the operand in the first source register is arranged on the right side, and the first splice vector is generated. Exemplarily, when the format of the instruction is "vector opcode of the first type vd, vj, immediate value", it is represented that the first source register is vd, the second source register is vj, and the destination register is vd. All the source elements in the first source register are regarded as one operand and denoted as operand vd and all in the second source register Source element are regarded as another operand and denoted as operand vj and the first splice vector is "operand vd operand vj ".
[0072] Optionally, the operand in the first source register and the operand in the second source register may also be cross-spliced in units of elements to generate a first splice vector. During cross-splicing, source elements with the same position information in the source register are cross-spliced as one group, and the positions of different groups in the first splice vector are arranged in descending order according to the addresses of the source elements. When splicing two elements in different registers as one group, the setting of the left and right positions is determined according to the position of the source register identifier in the instruction. Since this is the same as the above example, it will not be repeated here.
[0073] Exemplarily, the operand in the first source register includes "source element 1 (position information is a), source element 2 (position information is b), and source element 3 (position information is c)", and the operand in the second source register is "source element 4 (position information is a), source element 5 (position information is b), and source element 6 (position information is c)". Operand in the first source register The corresponding source register identifier is arranged at the left position of the instruction. Operand in the second source register Assuming that the corresponding source register identifier is arranged at the right position of the instruction, the left and right sides are the opposite positions of the two source register position identifiers. The result obtained by splicing the two operands left and right may be "source element 1 source element 2 source element 3 source element 4 source element 5 source element 6", or by considering source element 1 and source element 4 with the same position information a as one group for cross-splicing, source element 2 and source element 5 with the same position information b as one group for cross-splicing, and source element 3 and source element 6 with the same position information c as one group for cross-splicing, the finally obtained first splice vector will be "source element 1 source element 4 source element 2 source element 5 source element 3 source element 6".
[0074] Optionally, Operand in the first source registerWhen the operand in the first source register is N bits and the operand in the second source register is N bits, the first splice vector is 2N bits, where N is a positive integer greater than 0. The number of bits of the operand in the first source register may be determined according to the elements included in the first source register and the bits corresponding to the data type of the element, and the number of bits of the operand in the second source register may be determined according to the elements included in the second source register and the bits corresponding to the data type of the element.
[0075] After the first splice vector is generated, step 505 is executed.
[0076] In step 505, according to the immediate number, for each source element in the first splice vector, an operation of shifting, rounding, and saturating to half-width is performed to generate a first initial shift operation result.
[0077] In the embodiment of the present application, the first splice vector includes a plurality of elements (source elements). According to the immediate number, for each element in the first splice vector, an operation of shifting, rounding, and saturating to half-width is performed. The shift amount is the immediate number, and a first initial shift operation result is generated. The shift operation is a right shift operation, and the shift operation includes a logical shift and an arithmetic shift. The process of saturating to half-width means performing a numerical saturation process according to the value range that can be represented by the binary data after reducing the width of the data bits of the data to be processed to half. The processed data is correlated with the width before processing. Optionally, the processed data is still a multiple (e.g., 1 / 2) of the width before processing.
[0078] Optionally, each source element in the first splice vector is shifted, and the shift amount of each element is the same and is an immediate value. Exemplarily, the first splice vector includes element 1, element 2, and element 3. If the shift amount is ui4, shifting the first splice vector means shifting element 1 by ui4 bits, shifting element 2 by ui4 bits, and shifting element 3 by ui4 bits. The element includes multiple bits. Shifting the element to the right means shifting each bit of the element to the right by a preset number of bits, discarding the number of bits shifted to the right in the element, and setting a specified value in the empty position on the left side of the element. Both the preset number of bits and the specified value are values set according to specific situations.
[0079] Optionally, the shift rounding operations performed on each source element in the first splice vector include four rounding operations: rounding to an even number, rounding to zero, rounding up, and rounding down. It is preferable that the shift rounding operation performed on the first splice vector is a shift rounding up operation on the first splice vector.
[0080] Optionally, for any element x with a bit number of 2N and a shift amount of sa, the process of performing a right logical shift, rounding, and saturating to half-width on the element x includes the following first step to second step.
[0081] In the first step, an operation result A is obtained according to the shift amount. Specifically, when the shift amount is 0, the obtained operation result A is the element x. When the shift amount is a positive integer greater than 0, an intermediate operation result is set. The intermediate operation result is such that the lower bits are the data from the sa-bit to the 2N - 1-bit of the element x, the upper bits of the remaining sa bits are 0, the bit number of the intermediate operation result is 2N, and the intermediate operation result is added to the sa - 1-bit of the element x to obtain the operation result A. Here, N is a positive integer greater than 0, and sa is an immediate value.
[0082] In the second step, the value of the operation result is obtained, compared with the specified data, and the final operation result is obtained according to the comparison result. Specifically, the value of the operation result A is 2 N-1 is compared, and if the operation result A is 2 N-1 is greater than, the final operation result is data where all N bits are 1; otherwise, the final operation result is the data from the 0th bit to the (N - 1)th bit of the operation result A. Element x is a signed vector or an unsigned vector.
[0083] Optionally, for any element x with 2N bits and a shift amount of sa, the process of performing a right arithmetic shift, rounding, and saturating to the half-value width on the element x includes the following first step to second step.
[0084] In the first step, an operation result A is obtained according to the shift amount. Specifically, when the shift amount is 0, the obtained operation result A is the element x; when the shift amount is a positive integer greater than 0, an intermediate operation result is set. The lower bits of the intermediate operation result are the data from the sa-th bit to the (2N - 1)th bit of the element x, and the upper bits of the remaining sa bits are all the data of the (2N - 1)th bit of the element x. The number of bits of the intermediate operation result is 2N, and the intermediate operation result A is added to the (sa - 1)th bit of the element x to obtain the operation result A. Here, N is a positive integer greater than 0, and sa is an immediate value.
[0085] In the second step, the value of the operation result is obtained, compared with the specified data, and the final operation result is obtained according to the comparison result. Specifically, the comparison between the value of the operation result A and 2 N-1 and the comparison between the value of the operation result A and -2 N-1 are each performed. When the operation result A is greater than 2 N-1 , the final operation result has the most significant bit as 0, the remaining lower bits as 1, and the number of bits of the final operation result is N. When the operation result A is less than -2 N-1 , the final operation result has the most significant bit as 1, the remaining lower bits as 0, and the number of bits of the final operation result is N. When the operation result A is 2N-1 smaller than -2 N-1 If it is larger, the final operation result is the data from the 0th bit to the (N - 1)th bit of operation result A, and the number of bits of the final operation result is N.
[0086] The process of saturating to the half-value width for the rounded data includes the process of saturating to the half-value width for the rounded signed data and the process of saturating to the half-value width for the rounded unsigned data.
[0087] Referring to the embodiments of the present application, the final operation result according to the above example becomes an element in the first initial shift operation result according to the embodiments of the present application, and operation result A becomes an element in the first splice vector according to the embodiments of the present application.
[0088] According to the immediate value number, perform an operation of shifting, rounding, and saturating to the half-value width on the first splice vector to generate the first initial shift operation result, and then execute step 506.
[0089] In step 506, perform a bit selection operation on the first initial shift operation result to generate a shift operation result.
[0090] In the embodiments of the present application, the bit selection operation includes an operation of selecting consecutive lower half data for each element included in the first initial shift operation result, an operation of selecting consecutive upper half data for each element included in the first initial shift operation result, consecutive intermediate specified bits de - an operation of selecting data, and an operation of selecting non-consecutive specified bit data for each element included in the first initial shift operation result, including any one of them.
[0091] After performing a bit selection operation on the first initial shift operation result to generate a shift operation result, execute step 507.
[0092] In step 507, the elements in the shift operation result are used as target elements and are sequentially written into the destination register.
[0093] In the embodiment of the present application, the data type of the target element is determined according to the data type of the source element. Optionally, the number of bits corresponding to the data type of the target element is half of the number of bits corresponding to the data type of the source element. Exemplarily, when the data type of the source element is half-word, the data type of the target element is byte; when the data type of the source element is word, the data type of the target element is half-word; when the data type of the source element is double-word, the data type of the target element is word; when the data type of the source element is quad-word, the data type of the target element is double-word. The source element may be signed data or unsigned data.
[0094] Optionally, after the target element is determined, the method of sequentially writing the target element into the destination register includes: determining the position information of each target element in the first initial shift operation result; and sequentially writing the target element into the position in the destination register that matches the position information corresponding to the target element. Here, the position information indicates the order of the elements in the first initial shift operation result. Sequentially writing the target element into the position in the destination register that matches the position information corresponding to the target element means writing the target element from the most significant bit to the least significant bit into the positions from the (N / 2 - 1)-th bit to the 0-th bit of the destination register, or writing the target element from the least significant bit to the most significant bit into the positions from the 0-th bit to the (N / 2 - 1)-th bit of the destination register.
[0095] Referring to the process of obtaining target elements according to the first type of vector opcode in the embodiments of the present application, the first type of vector opcode includes four types of vector opcodes (i.e., the first vector opcode, the second vector opcode, the third vector opcode, and the fourth be vector opcode) for instructing different vector shift operations, and specifically, it can be described in detail with reference to the following four specific implementation forms.
[0096] In the first specific implementation form of the present application, the first type of vector opcode is the first vector opcode, and the specific processing method can include the following sub-steps A1 to A2.
[0097] In sub-step A1, according to the immediate value, for each source element in the first splice vector, perform a logical right shift, rounding, and signed saturation to half-width operation to generate a first initial shift operation result.
[0098] In the embodiments of the present application, the first type of vector opcode may be the first vector opcode, and the first vector opcode can be used to instruct the first splice vector to perform a logical right shift, rounding, and signed saturation to half-width operation. The first splice vector is 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits.
[0099] Optionally, a logical right shift refers to a method of shifting elements without considering the sign bit. That is, for each 1-bit right shift of an element, 0 can be padded as the most significant bit. Signed saturation means that for a 16-bit numerical value, saturation is performed according to the range of the signed value of an 8-bit numerical value (-128 to +127). Half-width means half of the bit-width segment. For any vector, perform a logical right shift and RoundedSince the process of performing an operation of saturating with a sign to the half-value width has been described above, it will not be repeatedly described here.
[0100] When the first type of vector opcode is the first vector opcode, in order to generate the first initial shift operation result, for each source element in the first splice vector according to the immediate number, a right logical shift, rounding, and saturating with a sign to the half-value width operation can be performed.
[0101] In sub-step A2, for each element included in the first initial shift operation result, successively select the lower half of the data, and determine the element after the selection operation as the shift operation result.
[0102] In an embodiment of the present application, after generating the first initial shift operation result, successively select the lower half of the data for each element included in the first initial shift operation result, and after the selection operation yao the element may be determined as the shift operation result.
[0103] Optionally, when both the first source element and the second source element included in the first source register are N / 2 bits, and the third source element and the fourth source element included in the second source register are N / 2 bits, the first initial shift operation result is N bits, and the shift operation result is the data indicated by the first initial shift operation result from bit 0 to bit N / 2 - 1.
[0104] Furthermore, for each target element in the shift operation result, successively select its lower half, and sequentially write it to the storage position corresponding to each target element of the destination register. Exemplarily, the first source register is the vector register vd, the second source register is the vector register vj, the vector register vd includes the first source element and the second source element, the vector register vj includes the third source element and the fourth source element, the bit numbers of the first source element, the second source element, the third source element, and the fourth source element are N / 2, and the first source element、 The second source element The third source element and the fourth source element can be spliced left and right using a vector shift instruction to form one 2N (2N = 256)-bit vector. For each source element in the vector, perform an operation of right logical shift, rounding, and signed saturation to half value width. The shift amount is derived from an immediate number. Select the lower half elements of the shift result of the first source element as one target element, select the lower half elements of the shift result of the second source element as one target element, select the lower half elements of the shift result of the third source element as one target element, select the lower half elements of the shift result of the fourth source element as one target element, and sequentially write them into the vector register vd.
[0105] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0106] In the second specific implementation form of the present application, the first type of vector opcode is the second vector opcode, and the specific processing method can include the following sub-steps B1 to B2.
[0107] In sub-step B1, according to the immediate number, perform an operation of right arithmetic shift, rounding, and signed saturation to half value width for each source element in the first splice vector to generate a first initial shift operation result.
[0108] In the embodiments of the present application, the first type of vector opcode may be the second vector opcode. The second vector opcode can be used to instruct to perform an operation of right arithmetic shift, rounding, and signed saturation to half value width on the first splice vector. The first splice vector is 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits.
[0109] Optionally, an arithmetic right shift refers to a shift method for elements that requires considering the sign bit. That is, each time an element is shifted one bit to the right, if the sign bit is 1, a 1 is filled in the leftmost most significant bit; otherwise, a 0 is filled in the leftmost most significant bit. Since the meanings of signed saturation and half-width are the same as those described above, they will not be repeatedly explained here. For any vector, perform an arithmetic right shift and Rounded Since the process of performing an operation of saturating to half-width with a signed value for a given vector has been described above, it will not be repeatedly explained here.
[0110] When the first type of vector opcode is Second the vector opcode of, in order to generate the first initial shift operation result, according to the immediate value, for each source element in the first splice vector, an operation of arithmetic right shift, rounding, and saturating to half-width with a signed value can be performed.
[0111] According to the immediate value, perform an operation of arithmetic right shift, rounding, and saturating to half-width with a signed value on the first splice vector to generate the first initial shift operation result, and then execute sub-step B2. In sub-step B2, respectively select the consecutive lower half data of each element included in the first initial shift operation result, and determine the element after the selection operation as the shift operation result.
[0112] In the embodiment of the present application, after generating the first initial shift operation result, respectively select the consecutive lower half data of each element included in the first initial shift operation result, and after the selection operation yao the element can be determined as the shift operation result.
[0113] Alternatively, when both the first source element and the second source element included in the first source register are N / 2 bits, and both the third source element and the fourth source element included in the second source register are N / 2 bits, the first initial shift operation result is N bits, and the shift operation result is the data indicated by the first initial shift operation result from bit 0 to bit N / 2 - 1.
[0114] Furthermore, for each target element in the shift operation result, its lower half is respectively selected and sequentially written into the storage positions corresponding to each target element of the destination register. Exemplarily, the first source register is the vector register vd, the second source register is the vector register vj, the vector register vd includes the first source element and the second source element, the vector register vj includes the third source element and the fourth source element, the number of bits of the first source element, the second source element, the third source element, and the fourth source element is N / 2, and the first source element 、 The second source element The third source element and the fourth source element can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) - bit vector. For each source element in the vector, an operation of right - arithmetic shift, rounding, and saturating with a signed half - value width is respectively performed. The shift amount is derived from an immediate number. For the shift result of the first source element, its lower - half element is selected as one target element, for the shift result of the second source element, its lower - half element is selected as one target element, for the shift result of the third source element, its lower - half element is selected as one target element, for the shift result of the fourth source element, its lower - half element is selected as one target element, and they are sequentially written into the vector register vd.
[0115] It should be understood that the above example is only an example illustrated to better understand the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0116] In the third specific implementation form of the present application, the first type of vector opcode is the third vector opcode, and the specific processing method can include the following sub-steps C1 to C2.
[0117] In sub-step C1, according to the immediate value, for each source element in the first splice vector, perform a logical right shift, rounding, and unsigned saturation to half-width operation to generate a first initial shift operation result.
[0118] In an embodiment of the present application, the first type of vector opcode may be the third vector opcode, and the third vector opcode can be used to instruct to perform a logical right shift, rounding, and unsigned saturation to half-width operation on the first splice vector. The first splice vector is 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits.
[0119] Optionally, rounding and unsigned saturation means that for a 16-bit numerical value, saturation is performed according to the range of signed values (0 to 255) of an 8-bit numerical value. The meaning of the logical right shift and saturation to half-width processing is the same as that described above, so it will not be repeatedly described here. For any vector, the process of performing a logical right shift and unsigned saturation to half-width has been described above, so it will not be repeatedly described here. Rounded Since the process of performing a logical right shift and unsigned saturation to half-width operation on any vector has been described above, it will not be repeatedly described here.
[0120] When the first type of vector opcode is the third vector opcode, in order to generate the first initial shift operation result, according to the immediate value, for each source element in the first splice vector, a logical right shift, rounding, and unsigned saturation to half-width operation can be performed.
[0121] After performing a logical right shift, rounding, and unsigned saturation to half-width operation on the first splice vector according to the immediate value to generate the first initial shift operation result, execute sub-step C2.
[0122] In sub-step C2, consecutive lower half data of each element included in the first initial shift operation result are respectively selected, and the element after the selection operation is determined as the shift operation result.
[0123] In the embodiment of the present application, after generating the first initial shift operation result, consecutive lower half data of each element included in the first initial shift operation result are respectively selected, and after the selection operation yao the element can be determined as the shift operation result.
[0124] Optionally, when both the first source element and the second source element included in the first source register are N / 2 bits, and both the third source element and the fourth source element included in the second source register are N / 2 bits, the first initial shift operation result is N bits, and the shift operation result is the data indicated by the first initial shift operation result from the 0th bit to the N / 2 - 1th bit.
[0125] Furthermore, for each target element in the shift operation result, its lower half is respectively selected, and written sequentially to the storage positions corresponding to each target element of the destination register. Exemplarily, the first source register is the vector register vd, the second source register is the vector register vj, the vector register vd includes the first source element and the second source element, the vector register vj includes the third source element and the fourth source element, the number of bits of the first source element, the second source element, the third source element, and the fourth source element is N / 2, and the first source element 、 the second source element The third source element and the fourth source elementIt can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256)-bit vector. For each source element in the vector, perform an operation of right logical shift, rounding, and unsigned saturation to half the width. The shift amount is derived from an immediate number. For the shift result of the first source element, select its lower half elements as one target element. For the shift result of the second source element, select its lower half elements as one target element. For the shift result of the third source element, select its lower half elements as one target element. For the shift result of the fourth source element, select its lower half elements as one target element, and sequentially write them into the vector register vd.
[0126] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0127] In the fourth specific implementation form of the present application, the first type of vector opcode is the fourth be vector opcode, and the specific processing method can include the following sub-steps D1 to D2.
[0128] In sub-step D1, according to the immediate number, perform an operation of right arithmetic shift, rounding, and unsigned saturation to half the width on each source element in the first splice vector to generate a first initial shift operation result.
[0129] In the embodiments of the present application, the first type of vector opcode may be the fourth be vector opcode, and the fourth be vector opcode can be used to instruct to perform an operation of right arithmetic shift, rounding, and unsigned saturation to half the width on the first splice vector. The first splice vector is 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits.
[0130] Optionally, since the meanings of right arithmetic shift, rounding, and half-width are the same as those described above, they will not be repeatedly described here. Right arithmetic shift And since the meaning of the process of saturating to the half-width is the same as that described above, it will not be repeatedly described here. Since the process of performing a right arithmetic shift on an arbitrary vector and saturating it without sign to the half-width has been described above, it will not be repeatedly described here.
[0131] When the first type of vector opcode is the fourth be vector opcode, in order to generate the first initial shift operation result, according to the immediate value, for each source element in the first splice vector, an operation of right arithmetic shift, rounding, and saturating without sign to the half-width can be performed.
[0132] According to the immediate value, perform an operation of right arithmetic shift, rounding, and saturating without sign to the half-width on the first splice vector to generate the first initial shift operation result, and then execute sub-step D2.
[0133] In sub-step D2, select the consecutive lower half data of each element included in the first initial shift operation result, and determine the element after the selection operation as the shift operation result.
[0134] In the embodiment of the present application, after generating the first initial shift operation result, select the consecutive lower half data of each element included in the first initial shift operation result, and after the selection operation yao the element can be determined as the shift operation result.
[0135] Optionally, when both the first source element and the second source element included in the first source register are N / 2 bits, and both the third source element and the fourth source element included in the second source register are N / 2 bits, the first initial shift operation result is N bits, and the shift operation result is the data indicated by the first initial shift operation result from bit 0 to bit N / 2 - 1.
[0136] Further, for each target element in the shift operation result, its lower half is respectively selected, and sequentially written to the storage positions corresponding to each target element of the destination register. Exemplarily, the first source register is the vector register vd, the second source register is the vector register vj, the vector register vd includes the first source element and the second source element, the vector register vj includes the third source element and the fourth source element, the number of bits of the first source element, the second source element, the third source element, and the fourth source element is N / 2, and the first source element 、 The second source element The third source element and the fourth source element can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) -bit vector. For each source element in the vector, an operation of right arithmetic shift, rounding, and unsigned saturation to half-width is respectively performed. The shift amount is derived from an immediate number. For the shift result of the first source element, its lower half element is selected as one target element, for the shift result of the second source element, its lower half element is selected as one target element, for the shift result of the third source element, its lower half element is selected as one target element, for the shift result of the fourth source element, its lower half element is selected as one target element, and they are sequentially written to the vector register vd.
[0137] It should be understood that the above example is only an example illustrated to better understand the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0138] By adopting the technical solution according to the present application and executing an instruction including a first vector opcode and an immediate value, a series of operations such as a logical shift on two source elements, rounding, and saturating with sign to half value width are implemented. By executing an instruction including a second vector opcode and an immediate value, a series of operations such as an arithmetic shift on two source elements, rounding, and saturating with sign to half value width are implemented. By executing an instruction including a third vector opcode and an immediate value, a series of operations such as a logical shift on two source elements, rounding, and saturating without sign to half value width are implemented. By executing an instruction including a fourth be vector opcode and an immediate value, a series of operations such as an arithmetic shift on two source elements, rounding, and saturating without sign to half value width are implemented. Therefore, by adopting the technical solution according to the present invention, different shift needs can be achieved using different shift parameters, whereby multiple vector shift needs can be achieved with one shift instruction, effectively reducing system overhead and improving the execution efficiency of vector shifts for specific functions.
[0139] Embodiment 3
[0140] In an embodiment of the present application, the opcode may be a second type of vector opcode, the source register includes a first source register and a second source register, and the second type of vector opcode can be used to execute a selection operation on the first source register and the second source register respectively to instruct to execute a corresponding vector shift operation. As shown in FIG. 6, the processing method of the vector shift instruction may include the following steps 601 to 609.
[0141] In step 601, an instruction including a register identifier and a shift parameter is received.
[0142] In an embodiment of the present application, the meaning of the instruction and the parameters included in the instruction have been described in Embodiment 1 and Embodiment 2, and thus will not be repeatedly described here.
[0143] Optionally, there are two source registers, that is, the source elements are derived from two different registers. When there are multiple source registers, each source register identifier in all the source registers is different from the destination register identifier, or when there are multiple source registers, there is one source register identifier in all the source registers that is the same as the destination register identifier. Compared with Example 2, the number of bits of the source register and the destination register in the embodiment of the present application is twice the number of bits of the source register and the destination register bits in Example 2. Exemplarily, when the number of bits of the source register in the embodiment of the present application is 256 bits, the number of bits of the source register in Example 2 is 128 bits.
[0144] Optionally, decode the received instruction, obtain the shift parameter included in the instruction. The shift parameter is used to indicate the rule based on which the vector shift operation is performed on the source element. In this example, the shift parameter can include parameters such as the shift amount and the opcode.
[0145] In step 602, determine the shift amount and the shift operation rule according to the shift parameter.
[0146] In the embodiment of the present application, there is at least one source element on which the vector shift operation is performed. The shift amount is an immediate number, the shift operation rule is an opcode, the opcode is a second type of vector opcode, and the immediate number is a positive integer greater than or equal to 0.
[0147] Optionally, the second type of vector opcode is a code represented in binary form, or the opcode is an identifier convertible to binary code. The instruction format is "opcode destination register, source register, shift amount". When the opcode is the second type of vector opcode, when actually implemented, the instruction can be represented as "XVSSR 第2のタイプ .{B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} xd,xj,ui 第2のタイプ ", where XVSSR 第2のタイプ is the instruction name within the second type of vector opcode, and {B.H / H.W / W.D / D.Q / BU.H / HU.W / WU.D / DU.Q} is a parameter for indicating the data types of the source and target elements among Second type of vector opcode . B represents byte, H represents half-word, W represents word, D represents double-word, Q represents quad-word, U represents unsigned, xd represents both the destination register and the source register, xj represents the source register, and ui 第2のタイプ represents the immediate value included in the instruction when the opcode is the second type of vector opcode. Exemplarily, XVSSR 第2のタイプ1 .B.H is a second type of vector opcode convertible to binary form. For example, the second type of vector opcode where XVSSR 第2のタイプ1 .B.H is converted to the binary form 011101110101000001 can be cited.
[0148] Furthermore, ui 第2のタイプ is a parameter defined according to the data types of the source and target elements. The value of the ui 第2のタイプ and the way of taking values are the same as those of ui 第1のタイプ in Embodiment 2, so it will not be repeatedly described here.
[0149] After determining the shift amount and shift operation rule according to the shift parameter, step 603 is executed.
[0150] In step 603, according to the second type of vector operation code, selection operations are respectively performed on the first source register and the second source register to obtain a first operand and a second operand.
[0151] In the embodiment of the present application, according to the second type of vector operation code, a selection operation is performed on the first source register to obtain a first operand, and a selection operation is performed on the second source register to obtain a second operand. Here, the first operand and the second operand have the same data type, and the source elements in the first operand and the elements in the second operand have a data type of either half-word, word, double-word, or quad-word.
[0152] Optionally, the selection operation includes any one of an operation of selecting consecutive lower half data for each element of the first source register and the second source register, an operation of selecting consecutive upper half data for each element of the first source register and the second source register, an operation of selecting consecutive specified bit data in the middle for each element of the first source register and the second source register, and an operation of selecting non-consecutive specified bit data for each element of the second source register. ta and
[0153] Optionally, the selection operation performed on the first source register and the selection operation performed on the second source register are the same. Exemplarily, for example, performing a selection operation on the first source register means respectively selecting consecutive lower half data from each element included in the first source register, and performing a selection operation on the second source register means respectively selecting consecutive lower half Data from each element included in the second source register. Also, for example, performing a selection operation on the first source register means selecting consecutive upper half Datawhich is to select each of them, and for the second source register Selected by Performing a selection operation means selecting the upper half of the consecutive elements among each element included in the second source register Data as Each selected. Further, for example, performing a selection operation on the first source register Selected by means selecting the middle consecutive specified bits among each element included in the first source register Data as Each selected, and for the second source register To ta performing a selection operation on Selected by it means selecting the middle consecutive specified bits in the second source register From each element included in as Data selected. Further, for example, performing a selection operation on the first source register Each means selecting the non - consecutive specified bits among each element included in the first source register Selected by as Data selected, and for the second source register Each performing a selection operation on Selected by it means selecting the non - consecutive specified bits among each element included in the second source register Data as Each selected.
[0154] According to the second type of vector opcode, perform selection operations on the first source register and the second source register respectively, obtain the first operand and the second operand, and then execute step 604.
[0155] In step 604, determine the data other than the first operand in the first source register as the third operand, and determine the data other than the second operand in the second source register as the fourth operand.
[0156] In the embodiment of the present application, the third operand and the fourth operand have the same data type as the first operand and the second operand, and the data types of the source element in the third operand and the element in the fourth operand are any of half-word, word, double-word, and quad-word.
[0157] Optionally, determine the data other than the first operand in the first source register as the third operand, and determine the data other than the second operand in the second source register as the fourth operand. Exemplarily, when the first operand is the continuous lower half data of each element included in the first source register, the third operand is the continuous upper half data of each element included in the first source register. Similarly, the second operand is the continuous lower half of each element of the first source register Data and the fourth operand is the continuous upper half of each element included in the first source register Data to be.
[0158] After the first operand, the second operand, the third operand, and the fourth operand are obtained, step 605 is executed.
[0159] In step 605, after splicing the first operand and the second operand, a second splice vector is generated, and after splicing the third operand and the fourth operand, a third splice vector is generated.
[0160] In the embodiment of the present application, after splicing the first operand and the second operand left and right, a second splice vector is generated. Here, the setting of the position for splicing the first operand and the second operand left and right is determined according to the position of the source register identifier in the instruction. That is, when the first source register identifier is the source register identifier immediately after the vector opcode of the second type in the instruction, and the second source register identifier is the source register identifier arranged after the first source register identifier in the instruction, since the first operand is derived from the first source register and the second operand is derived from the second source register, the first operand is arranged on the left side, the second operand is arranged on the right side, and the second splice vector is generated. However, when the second source register identifier is the source register identifier immediately after the vector opcode of the second type in the instruction, and the first source register identifier is Second the source register identifier arranged after the source register identifier in the instruction, since the first operand is derived from the first source register and the second operand is derived from the second source register, the second operand is arranged on the left side, the first operand is arranged on the right side, and the second splice vector is generated. Exemplarily, when the format of the instruction is "vector opcode of the second type vd, vj, immediate value", it is represented that the first source register is vd, the second source register is vj, and the destination register is vd. In this case, the first operand is derived from vd (the first operand vd is denoted as such), and the second Operand is derived from vj (the second operand vj is denoted as such), and the second splice vector is "the first operand vd the second operand vj ". Similarly, when the format of the instruction is "vector opcode of the second type vd, vj, immediate value", it is represented that the second source register is vd, the first source register is vj, and the destination register is vd. In this case, the second operand is vd (the second Operandvd is derived from (denoted as), and the first operand is vj (the first Operand vj is derived from (denoted as), and the second splice vector is “Di two operands vd the first operand vj becomes. The method of generating the third splice vector by the third operand and the fourth operand is the same as the above method of generating the second splice vector by the first operand and the second operand, so it will not be repeatedly described here.
[0161] Optionally, the first operand and the second operand may be further cross-spliced in units of elements to generate the second splice vector. During cross-splicing, source elements having the same address in the source register are cross-spliced as one group, and the positions of different groups in the second splice vector are arranged in descending order according to the addresses of the source elements. The setting of the left and right positions when splicing two elements respectively derived from two different registers as one group is determined according to the position of the source register identifier in the instruction. Since it is the same as the above example here, it will not be repeatedly described here. Exemplarily, the first operand includes "source element 1 (address a), source element 2 (address b), and source element 3 (address c)", and the second operand is "source element 4 (address a), source element 5 (address b), and source element 6 (address c)". Source element 1 and source element 4 with the same address a are cross-spliced as one group, source element 2 and source element 5 with the same address b are cross-spliced as one group, and source element 3 and source element 6 with the same address c are cross-spliced as one group. The Operand corresponding source register identifier is arranged at the left position of the instruction, and the second OperandAssuming that the source register identifier corresponding to [[ID=]] is arranged at the right position of the instruction, the finally obtained second splice vector will be "source element 1 source element 4 source element 2 source element 5 source element 3 source element 6". Since the method of generating the third splice vector by the third operand and the fourth operand is the same as the method of generating the second splice vector by the first operand and the second operand, it will not be repeatedly described here.
[0162] Optionally, when the total number of bits of the first operand is N bits and the total number of bits of the second operand is N bits, the second splice vector will be 2N bits. Similarly, when the total number of bits of the third operand is N bits and the total number of bits of the fourth operand is N bits, the third splice vector will be 2N bits, where N is a positive integer greater than 0. The total number of bits of the first operand can be determined according to each element included in the first operand and the bits corresponding to the data type of the element, and the total number of bits of the second operand can be determined according to each element included in the second operand and the bits corresponding to the data type of the element.
[0163] After the second splice vector and the third splice vector are obtained, step 606 is executed.
[0164] In step 606, according to the immediate value number, for each element in the second splice vector, an operation of shifting, rounding, and saturating to half value width is performed to generate a second initial shift operation result, and according to the immediate value number, for each element in the third splice vector, an operation of shifting, rounding, and saturating to half value width is performed to generate a third initial shift operation result.
[0165] In the embodiment of the present application, the second splice vector contains a plurality of elements. According to the immediate value number, performing an operation of shifting, rounding, and saturating to the half-value width on the second splice vector, that is, performing an operation of shifting, rounding, and saturating to the half-value width on each element in the second splice vector, using the shift amount as the immediate value number, and generating the second initial shift operation result. Similarly, the third splice vector contains a plurality of elements. According to the immediate value number, performing an operation of shifting, rounding, and saturating to the half-value width on the third splice vector, that is, performing an operation of shifting, rounding, and saturating to the half-value width on each element in the third splice vector, using the shift amount as the immediate value number, and generating the third initial shift operation result. Here, the shift operation is a right shift operation, and the shift operation includes a logical shift and an arithmetic shift.
[0166] Optionally, shifting the second splice vector includes performing a shift operation on each element in the second splice vector and using the shift amount as the immediate value number, that is, the shift amount of each element is the same and is the immediate value number. Exemplarily, the second splice vector includes element 1, element 2, and element 3. If the shift amount is ui4, Second shifting the splice vector means, that is, shifting element 1 by ui4 bits, shifting element 2 by ui4 bits, and shifting element 3 by ui4 bits respectively.
[0167] Optionally, shifting the third splice vector includes performing a shift operation on each element in the third splice vector and using the shift amount as the immediate value number, that is, the shift amount of each element is the same and is the immediate value number. Exemplarily, the third splice vector includes element 4, element 5, and element 6. If the shift amount is ui4, shifting the third splice vector means, that is, shifting element 4 by ui4 bits, shifting element 5 by ui4 bits, and shifting element 6 by ui4 bits respectively.
[0168] Optionally, the shift rounding operation performed on the second splice vector / third splice vector includes four rounding operations: rounding to an even number, rounding to zero, rounding up, and rounding down. It is preferable to perform a shift rounding up operation on the second splice vector / third splice vector.
[0169] The method of performing a right logical shift and rounding to saturate to the half-value width on the second splice vector / third splice vector is the same as the method described in Example 2, so it will not be repeatedly described here. Similarly, the method of performing a right arithmetic shift and rounding to saturate to the half-value width on the second splice vector / third splice vector is the same as the method described in Example 2, so it will not be repeatedly described here.
[0170] After the second initial shift operation result and the third initial shift operation result are generated, step 607 is executed.
[0171] In step 607, a bit selection operation is performed on the second initial shift operation result to generate a first shift operation result, and a bit selection operation is performed on the third initial shift operation result to generate a second shift operation result.
[0172] In an embodiment of the present application, performing a bit selection operation includes any one of the following operations: selecting consecutive lower half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result; selecting consecutive upper half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result; selecting data of consecutive specified bits in the middle for each element included in the second initial shift operation result and each element included in the third initial shift operation result; and selecting data of non-consecutive specified bits for each element included in the second initial shift operation result and each element included in the third initial shift operation result.
[0173] After the first shift operation result and the second shift operation result are generated, steps 608 and 609 are executed.
[0174] In step 608, according to the bit selection operation position of the first shift operation result, the first shift operation result is written into the corresponding storage position in the destination register.
[0175] In step 609, according to the bit selection operation position of the second shift operation result, the second shift operation result is written into the corresponding storage position in the destination register.
[0176] In an embodiment of the present application, using the elements in the first shift operation result and the second shift operation result as target elements, they are written into the corresponding storage positions in the destination register.
[0177] Optionally, the data type of the target element is determined according to the data type of the source element, and optionally, the number of bits corresponding to the data type of the target element is half of the number of bits corresponding to the data type of the source element. Exemplarily, when the data type of the source element is a half word, the data type of the target element is a byte; when the data type of the source element is a word, the data type of the target element is a half word; when the data type of the source element is a double word, the data type of the target element is a word; and when the data type of the source element is a quad word, the data type of the target element is a double word. The source element may be signed data or unsigned data.
[0178] Optionally, after determining the target element, the method of sequentially writing the target element into the destination register includes determining the position information of each target element in the second shift operation result and the third shift operation result, and sequentially writing the target element to the position that matches the position information corresponding to the target element in the destination register. Here, the position information indicates the order of the elements in the second shift operation result and the third shift operation result. Sequentially writing the target element to the position that matches the position information corresponding to the target element in the destination register means determining the storage position of each target element in the destination register. For each target element, the target element derived from the second shift operation result is stored in the upper half of the storage position where the target element is located, and the target element derived from the third shift operation result is stored in the lower half of the storage position where the target element is located, or for each target element, the target element derived from the second shift operation result is stored in the lower half of the storage position where the target element is located, and the target element derived from the third shift operation result is stored in the upper half of the storage position where the target element is located.
[0179] Referring to the process of obtaining target elements according to the second type of vector opcode in the embodiments of the present application, the second type of vector opcode includes a fifth vector opcode, a sixth vector opcode, a seventh vector opcode, and an eighth vector opcode to respectively indicate different vector shift operations, and specifically, it can be described in detail with reference to the following specific implementation forms.
[0180] In the first specific implementation form of the embodiments of the present application, the second type of vector opcode is the fifth vector opcode, the first operand is the data of consecutive lower halves of each element of the first source register, the second operand is the data of consecutive lower halves of each element of the second source register, the third operand is the data of consecutive upper halves of each element of the first source register, and the fourth operand is the data of consecutive upper halves of each element of the second source register. The specific processing method can include the following sub-steps E1 to E4.
[0181] In sub-step E1, according to the immediate value, for each element in the second splice vector, perform a logical right shift, rounding, and saturating with sign to half value width operation to generate a second initial shift operation result. According to the immediate value, for each element in the third splice vector, perform a logical right shift, rounding, and saturating with sign to half value width operation to generate a third initial shift operation result.
[0182] In an embodiment of the present application, the second type of vector opcode may be the fifth vector opcode. The fifth vector opcode is used to perform an operation of right logical shift, rounding, and saturating with a sign to a half-width value on the second splice vector and a selection operation on the data of each element, and to perform an operation of right logical shift, rounding, and saturating with a sign to a half-width value on the third splice vector and a selection operation on the data of each element. The second splice vector and the third splice vector are 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits. The meanings of right logical shift, rounding, half-width value, and saturating with a sign are the same as those in Embodiment 2. Since the process of performing the operation of right logical shift, rounding, and saturating with a sign to a half-width value is the same as that in Embodiment 2, it will not be repeatedly described here.
[0183] After the second initial shift operation result and the third initial shift operation result are generated, sub-step E2 is executed.
[0184] In sub-step E2, the consecutive lower half data of each element included in the second initial shift operation result are respectively selected, and the selected data is determined as the first shift operation result. At the same time, the consecutive upper half data of each element included in the third initial shift operation result are respectively selected, and the selected data is determined as the second shift operation result.
[0185] In an embodiment of the present application, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the elements included in the second shift operation result.
[0186] After the first shift operation result and the second shift operation result are obtained, sub-step E3 and sub-step E4 are executed.
[0187] In sub-step E3, each first target element included in the first shift operation result is written to the lower half of the position where each first target element of the destination register is arranged, respectively.
[0188] In sub-step E4, each second target element included in the second shift operation result is written to the upper half of the position where each second target element of the destination register is arranged, respectively.
[0189] Optionally, when the total number of bits of the first operand is N bits and the total number of bits of the second operand is N bits, the second initial shift operation result is N bits, First The shift operation result is the data indicated by the second initial shift operation result from bit 0 to bit N / 2 - 1. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj, the lower half data of each element in xd is used as the first operand, the lower half data of each element in xj is used as the second operand, and the first operand and the second operand are spliced left and right using a vector shift instruction to form a single 2N (2N = 256) - bit vector. For each element in the vector, an operation of logical right shift, rounding, and saturating with a signed half - value width is performed respectively. The shift amount is derived from an immediate number. SecondAfter selecting the lower half of each element included in the initial shift operation result, using each element for which the lower half selection operation has been performed as a first target element, write to the lower half of the position where each first target element of the vector register xd is arranged. Here, the data types of the source elements in the first operand and the source elements in the second operand are any one of half word, word, double word, and quad word. The data types of the source elements in the first operand and the source elements in the second operand are the same. The data types of the target elements written to the vector register xd are byte, half word, word, and double word so as to correspond to the data types of the above source elements. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0190] Optionally, when the total number of bits of the third operand is N bits and the total number of bits of the fourth operand is N bits, the third initial shift operation result is N bits. Second The shift operation result is the data indicated by the third initial shift operation result from the N / 2th bit to the N - 1th bit. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj. The upper half data of each element in xd is used as the third operand, the upper half data of each element in xj is used as the fourth operand, and the third operand and the fourth operand can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) - bit vector. For each element in the vector, perform an operation of logical right shift, rounding, and signed saturation to half - value width respectively. The shift amount is derived from an immediate number. ThirdAfter selecting the upper half of each element included in the initial shift operation result, using each element on which the upper half selection operation has been performed as a second target element, write to the upper half of the position where each second target element of the vector register xd is arranged. Here, the data types of the source element in the third operand and the source element in the fourth operand are any one of half word, word, double word, and quad word. The source element in the third operand and the source element in the fourth operand have the same data type. The data type of the target element written to the vector register xd is byte, half word, word, or double word so as to correspond to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0191] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0192] In the second specific implementation form of the embodiment of the present application, the second type of vector operation code is the sixth vector operation code, the first operand is the consecutive lower half data of each element of the first source register, and the second operand is the consecutive Lower lower half data of each element of the second source register, the third operand is the consecutive upper half data of each element of the first source register, and the fourth operand is the consecutive upper half data of each element of the second source register. The specific processing method can include the following sub-steps F1 to F4.
[0193] In sub-step F1, according to the immediate value, for each element in the second splice vector, perform an operation of right arithmetic shift, rounding, and saturating with a sign to the half-value width to generate a second initial shift operation result. According to the immediate value, for each element in the third splice vector, perform an operation of right arithmetic shift, rounding, and saturating with a sign to the half-value width to generate a third initial shift operation result.
[0194] In an embodiment of the present application, the second type of vector opcode may be the sixth vector opcode. The sixth vector opcode performs an operation of right arithmetic shift, rounding, and saturating with a sign to the half-value width on the second splice vector and a data selection operation for each element, and an operation of right arithmetic shift, rounding, and saturating with a sign to the half-value width on the third splice vector and Selection operation of data of each element is for instructing to perform the above operations. Both the second splice vector and the third splice vector are 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits. The meanings of right arithmetic shift, rounding, half-value width, and saturating with a sign are the same as those in Embodiment 2. Since the process of performing the operation of right arithmetic shift, rounding, and saturating with a sign to the half-value width is the same as that in Embodiment 2, it will not be repeatedly described here.
[0195] After the second initial shift operation result and the third initial shift operation result are generated, sub-step F2 is executed.
[0196] In sub-step F2, select the consecutive lower half data of each element included in the second initial shift operation result respectively, determine the selected data as the first shift operation result, select the consecutive upper half data of each element included in the third initial shift operation result, and determine the selected data as the second shift operation result.
[0197] In the embodiment of the present application, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the elements included in the second shift operation result.
[0198] After the first shift operation result and the second shift operation result are obtained, sub-step F3 and sub-step F4 are executed.
[0199] In sub-step F3, each first target element included in the first shift operation result is respectively written to the lower half of the position where each first target element of the destination register is arranged.
[0200] In sub-step F4, each second target element included in the second shift operation result is respectively written to the upper half of the position where each second target element of the destination register is arranged.
[0201] Optionally, when the total number of bits of the first operand is N bits and the total number of bits of the second operand is N bits, the second initial shift operation result is N bits, First The shift operation result is the data indicated by the second initial shift operation result from bit 0 to bit N / 2 - 1. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj, the lower half data of each element in xd is used as the first operand, the lower half data of each element in xj is used as the second operand, and the first operand and the second operand are spliced left and right using a vector shift instruction to form a single 2N (2N = 256) - bit vector. For each element in the vector, an operation of right arithmetic shift, rounding, and saturating with a signed half - value width is respectively performed. The shift amount is derived from an immediate number, SecondAfter selecting the lower half of each element included in the initial shift operation result, each element for which the lower half selection operation has been performed is used as a first target element, and is written to the lower half of the position where each first target element of the vector register xd is arranged. Here, the data types of the source elements in the first operand and the source elements in the second operand are any one of half word, word, double word, and quad word. The source elements in the first operand and the source elements in the second operand have the same data type. The data types of the target elements written to the vector register xd are byte, half word, word, and double word so as to correspond to the data types of the above source elements. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0202] Optionally, when the total number of bits of the third operand is N bits and the total number of bits of the fourth operand is N bits, the third initial shift operation result is N bits. Second The shift operation result is the data indicated by the third initial shift operation result from the N / 2-th bit to the N - 1-th bit. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj. The upper half data of each element in xd is used as the third operand, and the upper half data of each element in xj is used as the fourth operand. The third operand and the fourth operand can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) -bit vector. For each element in the vector, an operation of performing a right arithmetic shift, rounding, and saturating with a signed value to a half value width is performed respectively. The shift amount is derived from an immediate number. ThirdAfter selecting the upper half of each element included in the initial shift operation result, using each element on which the upper half selection operation has been performed as a second target element, write to the upper half of the position where each second target element of the vector register xd is arranged. Here, the data types of the source element in the third operand and the source element in the fourth operand are any one of half word, word, double word, and quad word. The source element in the third operand and the source element in the fourth operand have the same data type. The data types of the target elements written to the vector register xd are byte, half word, word, and double word so as to correspond to the data types of the above source elements. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0203] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiment of the present application, and does not uniquely limit the embodiment of the present application.
[0204] In the third specific implementation form of the embodiment of the present application, the second type of vector operation code is the seventh vector operation code, the first operand is the consecutive lower half data of each element of the first source register, the second operand is the consecutive lower half data of each element of the second source register, the third operand is the consecutive upper half data of each element of the first source register, the fourth operand is the consecutive upper half data of each element of the second source register, and the specific processing method can include the following sub-steps G1 to G4.
[0205] In sub-step G1, according to the immediate value number, for each element in the second splice vector, perform an operation of logical right shift, rounding, and unsigned saturation to half value width to generate a second initial shift operation result. According to the immediate value number, for each element in the third splice vector, perform an operation of logical right shift, rounding, and unsigned saturation to half value width to generate a third initial shift operation result.
[0206] In an embodiment of the present application, the second type of vector opcode may be the seventh vector opcode. The seventh vector opcode is used to execute an operation of logical right shift, rounding, and unsigned saturation to half value width on the second splice vector and a data selection operation for each element, and to execute an operation of logical right shift, rounding, and unsigned saturation to half value width on the third splice vector and a data selection operation for each element. Both the second splice vector and the third splice vector are 2N bits, where N is a positive integer greater than 0, and preferably N is 128 bits. The meanings of logical right shift, rounding, and half value width are the same as those in Embodiment 2. Since the process of performing the operation of logical right shift, rounding, and unsigned saturation to half value width is the same as that in Embodiment 2, it will not be repeatedly described here.
[0207] After the second initial shift operation result and the third initial shift operation result are generated, sub-step G2 is executed.
[0208] In sub-step G2, select the consecutive lower half data of each element included in the second initial shift operation result respectively, determine the selected data as the first shift operation result, and select the consecutive upper half data of each element included in the third initial shift operation result, and determine the selected data as the second shift operation result.
[0209] In an embodiment of the present application, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the elements included in the second shift operation result.
[0210] After the first shift operation result and the second shift operation result are determined, sub-step G3 and sub-step G4 are executed.
[0211] In sub-step G3, each first target element included in the first shift operation result is respectively written to the lower half of the position where each first target element of the destination register is arranged.
[0212] In sub-step G4, each second target element included in the second shift operation result is respectively written to the upper half of the position where each second target element of the destination register is arranged.
[0213] Optionally, when the total number of bits of the first operand is N bits and the total number of bits of the second operand is N bits, the second initial shift operation result is N bits, First The shift operation result is the data indicated by the second initial shift operation result from bit 0 to bit N / 2 - 1. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj, the lower half data of each element in xd is used as the first operand, the lower half data of each element in xj is used as the second operand, and the first operand and the second operand are spliced left and right using a vector shift instruction to form a single 2N (2N = 256) -bit vector. For each element in the vector, a logical right shift, rounding, and unsigned saturation to half value width operation are respectively performed. The shift amount is derived from an immediate number, SecondAfter selecting the lower half of each element included in the initial shift operation result, using each element for which the lower half selection operation has been performed as a first target element, sequentially write to the lower half of the position where each first target element of the vector register xd is arranged. Here, the data types of the source elements in the first operand and the source elements in the second operand are any one of half word, word, double word, and quad word. The source elements in the first operand and the source elements in the second operand have the same data type. The data types of the target elements written to the vector register xd are byte, half word, word, and double word so as to correspond to the data types of the above source elements. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0214] Optionally, when the total number of bits of the third operand is N bits and the total number of bits of the fourth operand is N bits, the initial shift operation result of the third is N bits. Second The shift operation result is the data indicated by the N / 2th bit to the N - 1th bit of the initial shift operation result of the third. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj. The upper half data of each element in xd is used as the third operand, the upper half data of each element in xj is used as the fourth operand. The third operand and the fourth operand can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) - bit vector. For each element in the vector, perform an operation of logical right shift, rounding, and unsigned saturation to half value width respectively. The shift amount is derived from an immediate number. ThirdAfter selecting the upper half of each element included in the initial shift operation result, using each element for which the upper half selection operation has been performed as a second target element, write to the upper half of the position where each second target element of the vector register xd is arranged. Here, the data types of the source element in the third operand and the source element in the fourth operand are any one of half word, word, double word, and quad word. The source element in the third operand and the source element in the fourth operand have the same data type. The data type of the target element written to the vector register xd is byte, half word, word, or double word so as to correspond to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0215] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0216] In the fourth specific implementation form of the embodiment of the present application, the second type of vector operation code is the eighth vector operation code, the first operand is the consecutive lower half data of each element of the first source register, the second operand is the consecutive lower half data of each element of the second source register, the third operand is the consecutive upper half data of each element of the first source register, the fourth operand is the consecutive upper half data of each element of the second source register, and the specific processing method can include the following sub-steps H1 to H4.
[0217] In sub-step H1, according to the immediate value, for each element in the second splice vector, perform an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width, generate a second initial shift operation result, and according to the immediate value, for each element in the third splice vector, perform an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width, and generate a third initial shift operation result.
[0218] In an embodiment of the present application, the second type of vector opcode may be the eighth vector opcode. The eighth vector opcode performs an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width on the second splice vector and data selection for each element Selection operation operation, and for the third splice vector, performs an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width and data selection Selection operation operation, and can be used to instruct to perform the operations. The second splice vector and the third splice vector are 2N bits, where N is a positive integer greater than 0, but preferably N is 128 bits. The meanings of right arithmetic shift, rounding, and half-value width are the same as those in Embodiment 2, and the process of performing the operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width is the same as that in Embodiment 2, so it will not be repeated here.
[0219] After the second initial shift operation result and the third initial shift operation result are generated, sub-step H2 is executed.
[0220] In sub-step H2, respectively select the consecutive lower half data of each element included in the second initial shift operation result, determine the selected data as the first shift operation result, and respectively select the consecutive upper half data of each element included in the third initial shift operation result, and determine the selected data as the second shift operation result.
[0221] In the embodiment of the present application, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the elements included in the second shift operation result.
[0222] After the first shift operation result and the second shift operation result are obtained, sub-step H3 and sub-step H4 are executed.
[0223] In sub-step H3, each first target element included in the first shift operation result is respectively written to the lower half of the position where each first target element of the destination register is arranged.
[0224] In sub-step H4, each second target element included in the second shift operation result is respectively written to the upper half of the position where each second target element of the destination register is arranged.
[0225] Optionally, when the total number of bits of the first operand is N bits and the total number of bits of the second operand is N bits, First The initial shift operation result is N bits, and the second shift operation result is the data shown from bit 0 to bit N / 2 - 1 of the second initial shift operation result. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj, the lower half data of each element in xd is used as the first operand, the lower half data of each element in xj is used as the second operand, and the first operand and the second operand are spliced left and right using a vector shift instruction to form a single 2N (2N = 256) -bit vector. For each element in the vector, an operation of right arithmetic shift, rounding, and saturating without sign to half value width is respectively performed, and the shift amount is derived from an immediate number. SecondAfter selecting the lower half of each element included in the initial shift operation result, each element on which the lower half selection operation has been performed is used as a first target element, and is written to the lower half of the position where each first target element of the vector register xd is arranged. Here, the data types of the source elements in the first operand and the source elements in the second operand are any one of half word, word, double word, and quad word. The source elements in the first operand and the source elements in the second operand have the same data type. The data types of the target elements written to the vector register xd are byte, half word, word, and double word so as to correspond to the data types of the above source elements. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeated here.
[0226] Optionally, when the total number of bits of the third operand is N bits and the total number of bits of the fourth operand is N bits, the third initial shift operation result is N bits, Second The shift operation result is the data indicated by the third initial shift operation result from the N / 2th bit to the N-1th bit. Exemplarily, the first source register is the vector register xd, the second source register is the vector register xj. The upper half data of each element in xd is used as the third operand, and the upper half data of each element in xj is used as the fourth operand. The third operand and the fourth operand can be spliced left and right using a vector shift instruction to form a single 2N (2N = 256) bit vector. For each element in the vector, an operation of performing a right arithmetic shift, rounding, and saturating without sign to half value width is performed respectively. The shift amount is derived from an immediate number, ThirdAfter selecting the upper half of each element included in the initial shift operation result, using each element for which the upper half selection operation has been performed as a second target element, write to the upper half of the position where each second target element of the vector register xd is arranged. Here, the data type of the source element in the third operand and the source element in the fourth operand is any one of half word, word, double word, and quad word, the source element in the third operand and the source element in the fourth operand have the same data type, and the data type of the target element written to the vector register xd is byte, half word, word, or double word so as to correspond to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0227] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiment of the present application, and does not uniquely limit the embodiment of the present application.
[0228] By adopting the technical solution according to the present application and executing an instruction including a fifth vector opcode and an immediate value, a series of operations such as a logical shift on two Of so - source elements and an operation of rounding and saturating with a signed saturation to a half value width are implemented. By executing an instruction including a sixth vector opcode and an immediate value, a series of operations such as an arithmetic shift on two Of so - source elements and an operation of rounding and saturating with a signed saturation to a half value width are implemented. By executing an instruction including a seventh vector opcode and an immediate value, a series of operations such as a logical shift on two Of so - source elements and an operation of rounding and saturating without a sign to a half value width are implemented. By executing an instruction including an eighth vector opcode and an immediate value, a series of operations such as a logical shift on two Of soImplement a series of operations such as arithmetic shift on the - s element, rounding, and unsigned saturation to half - value width. Therefore, when adopting the technical solution according to the present invention, different shift needs can be achieved using different shift parameters. As a result, multiple vector shift needs can be achieved with one shift instruction, effectively reducing the system overhead and improving the execution efficiency of vector shift for specific functions.
[0229] Example 4
[0230] In the embodiment of the present application, the opcode is a third - type vector opcode, and the source register is a first - source register. This Third type of vector opcode can be used to perform a selection operation in the first - source register Execute and to instruct to execute the corresponding vector shift operation. As shown in FIG. 7, the processing method of the vector shift instruction can include the following steps 701 - 707.
[0231] In step 701, an instruction including a register identifier and a shift parameter is received.
[0232] In the embodiment of the present application, the meaning of the instruction and the parameters included in the instruction have been described in Examples 1 - 3, so they will not be repeatedly described here.
[0233] Optionally, the number of source registers is one, that is, all source elements are from the same source register, the number of destination registers is one, and the source register and the destination register can be the same or different. Preferably, the number of bits of the first - source register is 128 bits or 256 bits.
[0234] Optionally, decode the received instruction, obtain the shift parameter included in the instruction, where the shift parameter is for indicating a rule to be followed when performing a vector shift operation on a source element. In this example, the shift parameter can include parameters such as a shift amount and an opcode.
[0235] In step 702, according to the shift parameter, determine a shift amount and a shift operation rule.
[0236] In the embodiment of the present application, there is at least one source element on which a vector shift operation is performed. The shift amount is derived from a shift amount register, the shift operation rule is an opcode, the opcode is a third type of vector opcode, and each element shift amount included in the shift amount register is a positive integer greater than or equal to 0.
[0237] Optionally, the third type of vector opcode is a code represented in binary form, or the opcode is an identifier convertible to a binary code. The format of the instruction is "opcode destination register, source register, shift amount". When the opcode is the third type of vector opcode, in actual implementation, the instruction is "[X]VSSR 第3のタイプ .{B.H / H.W / W.D / BU.H / HU.W / WU.D} vd / xd,vj / xj,vk 第3のタイプ / xk 第3のタイプ ", where [X]VSSR 第3のタイプ is the instruction name within the third type of vector opcode, and {B.H / H.W / W.D / BU.H / HU.W / WU.D} is a parameter for indicating the data types of the source element and the target element in Third type of vector opcode . B represents a byte, H represents a half word, W represents a word, D represents a double word, vd / xd represents both the destination register and the source register, vj / xj represents the source register, and vk 第3のタイプ / xk 第3のタイプrepresents the shift amount register identifier included in an instruction when the opcode is a third type of vector opcode. The binary contained in the shift amount register is regarded as an array. The number of parameters contained in the array is the same as the number of target elements, and the parameters contained in the data can be the same or different. Exemplarily, VSSR 第3のタイプ1 .B.H is a third type of vector opcode that can be converted into binary form. For example, VSSR 第3のタイプ1 .B.H being converted into the binary form of 01110001000000001 is an example of the third type of vector opcode.
[0238] After determining the shift amount and shift operation rule according to the shift parameter, step 703 is executed.
[0239] In step 703, according to the third type of vector opcode, a selection operation is performed on the first source register to obtain a fifth operand.
[0240] In the embodiment of the present application, the selection operation includes any one of the operations of selecting the consecutive lower half Data of each element of the first source register, selecting the consecutive upper half Data of each element of the first source register, selecting the consecutive specified bits in the middle Data of each element of the first source register, and selecting the non-consecutive specified bits Data of each element of the first register. Here, the data type of the source element in the fifth operand is any one of half-word, word, and double-word.
[0241] Optionally, the first source register contains 2N bits of data, and the 2N bits of data can correspond to a plurality of half-word elements, word elements, or double-word elements. The step of performing a selection operation on the first source register to obtain a fifth operand includes treating the data in the first source register in groups of M bits each as one group of data, where each group of data contains at least one source element, and determining all the source elements corresponding to all the groups of data as the fifth operand. Here, M and N are positive integers greater than 0, with M ≤ N, and the correspondence between the source element and the data is determined according to the conversion relationship between the data type of the source element and the number of data bits.
[0242] Optionally, there is no data at the same address between different groups of data, or there is a part of the data at the same address between different groups of data. Here, the address is the position information of the data in the first source register in the first source register, and in the first source register, the address of each data is uniquely identified.
[0243] Preferably, N is a multiple of M. Exemplarily, when N = 128, M = 128, and the first source register contains 256 bits of data, the data in the first source register is divided into a total of two groups (the first data group and the second data group) as one group every 128 bits. The first data group is the data from bit 0 to bit 127 of the first source register, and the second data group is the data from bit 128 to bit 255 of the first source register. There is no data at the same address between the first data group and the second data group. When the data type of the source element is a half-word, the first data group contains 8 half-word source elements; when the data type of the source element is a word, the first data group contains 4 word source elements; when the data type of the source element is a double-word, the first data group contains 2 double-word source elements.
[0244] After performing a selection operation on the first source register according to the third type of vector opcode to obtain a fifth operand, step 704 is executed.
[0245] In step 704, according to the third type of vector opcode and the shift amount, an operation of shifting, rounding, and saturating to a half value width is performed on the fifth operand to generate a fourth initial shift operation result.
[0246] In an embodiment of the present application, the shift amount is derived from a shift amount register, and the content stored in the shift amount register can be a group of data. The group of data includes a plurality of shift values, and each shift value is related to each source element in the fifth operand. The shift values of different source elements may be the same or different, and the number of shift values is the same as the number of source elements of the fifth operand, or the number of shift values is made the same as the fourth initial shift operation result.
[0247] Optionally, when the number of shift values is the same as the number of source elements of the fifth operand, each shift value corresponds to one source element in the fifth operand, and for each source element in the fifth operand, an operation of shifting, rounding, and saturating to a half value width is performed according to the shift amount to generate a fourth initial shift operation result. Exemplarily, the fifth operand includes source element 1, source element 2, source element 3, and source element 4, the shift amount register includes four shift values (shift value 1, shift value 2, shift value 3, and shift value 4), and if the shift amount corresponding to source element 1 is shift value 1, the shift amount corresponding to source element 2 is shift value 2, the shift amount corresponding to source element 3 is shift value 3, and the shift amount corresponding to source element 4 is shift value 4, then source element 1 is shifted according to shift value 1, source element 2 is shifted according to shift value 2, source element 3 is shifted according to shift value 3, and source element 4 is shifted according to shift value 4.
[0248] Alternatively, when the number of shift values is the same as the fourth initial shift operation result, the fifth operand is divided into a plurality of element groups according to the number of shift values such that the number of element groups is the same as the number of shift values. That is, each shift value corresponds to one element group, and for each source element in the fifth operand, an operation of shifting, rounding, and saturating to the half-value width is performed according to the shift amount corresponding to the arranged element group, and the fourth initial shift operation result is generated. Exemplarily, the fifth operand includes source element 1, source element 2, source element 3, and source element 4, and the shift amount register includes 2 two shift values (shift value 1 and shift value 2). If source element 1 and source element 2 form the first element group, the shift amount corresponding to the first element group is shift value 1, source element 3 and source element 4 form the second element group, and the shift amount corresponding to the second element group is shift value 2, then source element 1 and source element 2 are shifted according to shift value 1, and source element 3 and source element 4 are shifted according to shift value 2.
[0249] After the fourth initial shift operation result is generated, step 705 is executed.
[0250] In step 705, a bit selection operation is performed on the fourth initial shift operation result to generate a shift operation result.
[0251] In the embodiment of the present application, the bit selection operation includes any one of an operation of selecting consecutive lower halves of the fourth initial shift operation result, an operation of selecting consecutive upper halves of the fourth initial shift operation result, an operation of selecting consecutive specified bits in the middle of the fourth initial shift operation result, and an operation of selecting non-consecutive specified bits of the fourth initial shift operation result. Data an operation of selecting consecutive upper halves of the fourth initial shift operation result, an operation of selecting consecutive specified bits in the middle of the fourth initial shift operation result, and an operation of selecting non-consecutive specified bits of the fourth initial shift operation result. Data an operation of selecting consecutive specified bits in the middle of the fourth initial shift operation result, and an operation of selecting non-consecutive specified bits of the fourth initial shift operation result. Data an operation of selecting non-consecutive specified bits of the fourth initial shift operation result. Data an operation of selecting non-consecutive specified bits of the fourth initial shift operation result.
[0252] After the shift operation result is generated, step 706 is executed.
[0253] In step 706, the data in the shift operation result is sequentially written to the corresponding positions in the destination register.
[0254] In the embodiment of the present application, after the shift operation result is generated, a target element corresponding to each data in the shift operation result is determined, and each data is sequentially written to the storage position of each target element in the destination register.
[0255] Optionally, the data type of the target element is determined according to the data type of the source element. Optionally, the number of bits corresponding to the data type of the target element is half of the number of bits corresponding to the data type of the source element. Exemplarily, when the data type of the source element is half-word, the data type of the target element is byte; when the data type of the source element is word, the data type of the target element is half-word; when the data type of the source element is double-word, the data type of the target element is word. The source element may be signed data or unsigned data.
[0256] Alternatively, after the target element is determined, the method of sequentially writing the target element into the destination register includes determining the position information of each target element in the shift operation result, and sequentially writing the target element to a position that matches the position information corresponding to the target element in the destination register. Here, the position information indicates the order of elements in the shift operation result. Sequentially writing the target element to a position that matches the position information corresponding to the target element in the destination register means writing the target element from the most significant bit to the least significant bit to the positions from the (N / 2 - 1)-th bit to the 0-th bit of the destination register, or writing the target element from the least significant bit to the most significant bit to the positions from the 0-th bit to the (N / 2 - 1)-th bit of the destination register.
[0257] After sequentially writing the elements in the shift operation result as target elements into the destination register, step 707 is executed.
[0258] In step 707, according to the third type of vector opcode, the value of the position in the destination register where the data is not written is set.
[0259] After sequentially writing the data in the shift operation result to the corresponding positions in the destination register, according to the third type of vector opcode, the value of the position in the destination register where the data is not written can be set.
[0260] Referring to the process of obtaining the target element according to the third type of vector opcode in the embodiments of the present application, the third type of vector opcode can include the ninth vector opcode, the tenth vector opcode, the eleventh vector opcode, and the twelfth vector opcode to respectively indicate different vector shift operations. Specifically, it can be described in detail with reference to the following specific implementation forms.
[0261] In the first specific implementation form of the embodiment of the present application, the third type of vector opcode is the ninth vector opcode, and the fifth operand is any consecutive source element of the first source register. The specific processing method can include the following sub-steps K1 to K5.
[0262] In sub-step K1, according to the shift amount, an operation of logically shifting the fifth operand to the right, rounding, and saturating with a signed half-width is performed on the fifth operand to generate a fourth initial shift operation result.
[0263] In the embodiment of the present application, after the ninth vector opcode is obtained, according to the shift amount, an operation of logically shifting each element included in the fifth operand to the right, rounding, and saturating with a signed half-width is performed on each element included in the fifth operand to generate a fourth initial shift operation result. Here, the shift amount is derived from a shift amount register.
[0264] Optionally, the first source register includes 2N-bit data, and the 2N-bit data can correspond to a plurality of half-word elements, word elements, or double-word elements. The step of performing a selection operation on the first source register to obtain the fifth operand includes determining all source elements corresponding to each data group as the fifth operand with the data of every M bits of the first source register as one data group. The step of performing an operation of logically shifting each element included in the fifth operand to the right, rounding, and saturating with a signed half-width according to the shift amount to generate a fourth initial shift operation result includes determining a shift value in a shift amount register corresponding to each source element in the fifth operand, and performing an operation of logically shifting each source element to the right, rounding, and saturating with a signed half-width according to the shift value corresponding to each source element respectively to obtain a fourth initial shift operation result.
[0265] Furthermore, the meanings of right logical shift, rounding, half-value width, and signed saturation are the same as those in Embodiment 2. Since the process of performing the operation of right logical shifting, rounding, and signed saturation to the half-value width is the same as that in Embodiment 2, it will not be repeatedly described here.
[0266] After the fourth initial shift operation result is generated, sub-step K2 is executed.
[0267] In sub-step K2, the consecutive lower half data of each element included in the fourth initial shift operation result are respectively selected, and the elements after the selection operation are determined as the shift operation result.
[0268] In the embodiment of the present application, after the fourth initial shift operation result is generated, the consecutive lower half data of each element in the fourth initial shift operation result are respectively selected, and the consecutive lower half data of each element after the selection operation are determined as the shift operation result.
[0269] After the shift operation result is obtained, sub-step K3 is executed.
[0270] In sub-step K3, the storage positions in the destination register are divided according to a preset value to generate a plurality of storage areas.
[0271] In the embodiment of the present application, the preset value refers to a value for dividing the storage positions in the vector register. The preset value is, that is, the data bit width occupied by the target element. The specific numerical value of the preset value can be determined according to business requirements, and the embodiment of the present application does not limit it. It is preferable that the preset value is a value such that each storage area divided according to the preset value has the same size (the data bit widths stored in each storage area are the same).
[0272] After a plurality of storage areas are generated, sub-step K4 is executed.
[0273] In sub-step K4, the data in the shift operation result is sequentially written into the lower half of each storage area.
[0274] In the embodiment of the present application, when the fifth operand is M bits, the fourth initial shift operation result is M / 2 bits. In this case, the shift operation result is the data shown from bit 0 to bit M / 4 - 1 in the fourth initial shift operation result. Exemplarily, the first source register is the vector register xj, the first source register is 2M, and for each M-bit corresponding source element of xj, using the fifth operand, and using the vector shift instruction, for each source element every M bits included in the fifth operand, perform a logical right shift, round, and saturate with a signed half-width operation to obtain the fourth initial shift operation result. The shift amount is derived from the shift amount register, and the lower half elements selected from each element in the fourth initial shift operation result are sequentially written into the lower half of each target element every M bits of the vector register xd as the lower half of each target element, and the data in the upper half of each target element every M bits is set to 0. Here, the data type of the source element in the fifth operand is any one of half-word, word, and double-word, and the data type of the target element written into the vector register xd is byte, half-word, word corresponding to the data type of the above source element. Since the correspondence between the data type of the source element and the data type of the target element has been described before, it will not be repeated here.
[0275] After sequentially writing the data in the shift operation result into the lower half of each storage area, sub-step K5 is executed.
[0276] In sub-step K5, the values at the positions where the data in each storage area has not been written are each set to zero.
[0277] In the embodiment of the present application, taking the elements in the shift operation result as target elements, after sequentially writing them into the lower half of each memory area, the values at the positions where the target elements in each memory area are not written are set to zero.
[0278] Exemplarily, if the first source register is vj / xj and the third type of vector opcode is the 10 vector opcode, a vector shift instruction is executed. That is, for each source element every 128 bits of the first source register vj / xj, a logical right shift, rounding, and signed saturation to half the width operation is performed, and the lower half of each element in the shift result is selected and sequentially written into the lower half of each target element every 128 bits of the destination register vd / xd, and the upper half of each target element every 128 bits of the destination register is set to 0. The shift amount of each element is derived from the shift amount register vk / xk, and the data type of the source element is any one of half-word, word, and double-word.
[0279] It should be understood that the above example is only an example illustrated to better understand the technical solution according to the embodiment of the present application, and does not uniquely limit the embodiment of the present application.
[0280] In the second specific implementation form of the embodiment of the present application, the third type of vector opcode is the 10th vector opcode, and the fifth operand is any consecutive source element of the first source register. The specific processing method can include the following sub-steps M1 to M5.
[0281] In sub-step M1, according to the shift amount, a right arithmetic shift, rounding, and signed saturation to half the width operation is performed on the fifth operand to generate a fourth initial shift operation result.
[0282] In the embodiment of the present application, after the 10th vector opcode is obtained, according to the shift amount, for each element included in the 5th operand, an operation of performing an arithmetic right shift, rounding, and saturating with a signed half-width is performed to generate a 4th initial shift operation result. Here, the shift amount is derived from a shift amount register.
[0283] Optionally, the first source register contains 2N-bit data, and the 2N-bit data can correspond to a plurality of half-word elements, word elements, or double-word elements. The step of performing a selection operation in the first source register to obtain a 5th operand includes determining all source elements corresponding to each data group as a 5th operand by taking the data of every M bits in the first source register as one data group. The step of performing an arithmetic right shift, rounding, and saturating with a signed half-width operation on each element included in the 5th operand according to the shift amount to generate a 4th initial shift operation result includes determining a shift value in a shift value register corresponding to each source element in the 5th operand, and performing an arithmetic right shift, rounding, and saturating with a signed half-width operation on each source element according to the shift value corresponding to each source element respectively to obtain a 4th initial shift operation result.
[0284] Furthermore, the meanings of arithmetic right shift, rounding, half-width, and signed saturation are the same as those in Embodiment 2, and the process of performing an arithmetic right shift, rounding, and saturating with a signed half-width operation is the same as that in Embodiment 2, so it will not be repeated here.
[0285] After the 4th initial shift operation result is generated, sub-step M2 is executed.
[0286] In sub-step M2, the consecutive lower half data of each element included in the 4th initial shift operation result is selected respectively, and the element after the selection operation is determined as the shift operation result.
[0287] In the embodiment of the present application, after the fourth initial shift operation result is generated, the consecutive lower half data of each element is respectively selected from the fourth initial shift operation result, and the consecutive lower half data of each element after the selection operation can be determined as the shift operation result.
[0288] After the shift operation result is obtained, sub-step M3 is executed.
[0289] In sub-step M3, the storage positions in the destination register are divided according to a preset value to generate a plurality of storage areas.
[0290] In the embodiment of the present application, the preset value refers to a numerical value for dividing the storage positions in the destination register into areas. The preset value is, that is, the data bit width occupied by the target element. The specific numerical value of the preset value can be determined according to business requirements, and the embodiment of the present application does not limit it. Preferably, the preset value is a value such that each storage area divided according to the preset value has the same size (the data bit widths stored in each storage area are the same).
[0291] After dividing the storage positions in the destination register according to the preset value to generate a plurality of storage areas, sub-step M4 is executed.
[0292] In sub-step M4, the data in the shift operation result is sequentially written into the lower half of each storage area.
[0293] In the embodiment of the present application, after a plurality of storage areas are generated and the shift operation result is obtained, the elements in the shift operation result can be sequentially written into the lower half of each storage area with the target element as the target element.
[0294] Optionally, if the fifth operand is M bits, the fourth initial shift operation result is M / 2 bits. In this case, the shift operation result is the data shown from bit 0 to bit M / 4 - 1 in the fourth initial shift operation result. Exemplarily, the first source register is the vector register xj, the first source register is 2M, and for each M-bit portion of xj, the corresponding source element is used as the fifth operand. Using a vector shift instruction, for each source element of every M bits included in the fifth operand, a right arithmetic shift, rounding, and saturating operation with a half value width and sign is performed to obtain the fourth initial shift operation result. The shift amount is derived from the shift amount register. The lower half of the elements selected from each element in the fourth initial shift operation result are sequentially written to the lower half of each target element of every M bits of the vector register xd, and the upper half data of each target element of every M bits is set to 0. Here, the data type of the source element in the fifth operand is any one of half word, word, and double word. The data type of the target element written to the vector register xd is byte, half word, or word corresponding to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeatedly described here.
[0295] After sequentially writing the data in the shift operation result to the lower half of each storage area, sub-step M5 is executed.
[0296] In sub-step M5, the values of the positions in each storage area where the data has not been written are each set to zero.
[0297] In the embodiment of the present application, after sequentially writing the elements in the shift operation result to the lower half of each storage area as target elements, the values of the positions in each storage area where the target elements have not been written can each be set to 0.
[0298] Exemplarily, when the first source register is vj / xj and the third type of vector operation code is the ninth vector operation code, a vector shift instruction is executed. That is, for each source element every 128 bits of the first source register vj / xj, a right arithmetic shift, rounding, and signed saturation to half-width operation is performed. The lower half of each element in the shift result is selected and sequentially written to the lower half of each target element every 128 bits of the destination register vd / xd, and the upper half of each target element every 128 bits of the destination register is set to 0. The shift amount of each element is derived from the shift amount register vk / xk, and the data type of the source element is any one of half-word, word, and double-word.
[0299] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0300] In a third specific implementation form of the embodiment of the present application, the third type of vector operation code is the eleventh vector operation code, and the fifth operand is any consecutive source element of the first source register. The specific processing method can include the following sub-steps N1 to N5.
[0301] In sub-step N1, according to the shift amount, a right logical shift, rounding, and unsigned saturation to half-width operation is performed on the fifth operand to generate a fourth initial shift operation result.
[0302] In the embodiment of the present application, after the eleventh vector operation code is obtained, according to the shift amount, a right logical shift, rounding, and unsigned saturation to half-width operation can be performed on each element included in the fifth operand to generate a fourth initial shift operation result. Here, the shift amount is derived from the shift amount register.
[0303] Optionally, the first source register contains 2N-bit data, and the 2N-bit data can correspond to a plurality of half-word elements, word elements, or double-word elements. The step of performing a selection operation on the first source register to obtain a fifth operand includes determining all source elements corresponding to each data group as a fifth operand by taking the data every M bits of the first source register as one group of data groups. According to the shift amount, for each element included in the fifth operand, performing a logical right shift, rounding, and unsigned saturation to half the width operation to generate a fourth initial shift operation result includes determining a shift value in a shift value register corresponding to each source element in the fifth operand, and performing a logical right shift, rounding, and unsigned saturation to half the width operation on each source element according to the shift value corresponding to each source element respectively to obtain a fourth initial shift operation result.
[0304] Furthermore, the meanings of logical right shift, rounding, and half the width are the same as those in Embodiment 2, and the process of performing a logical right shift, rounding, and unsigned saturation to half the width operation is the same as that in Embodiment 2, so it will not be repeatedly described here.
[0305] According to the shift amount, after performing a logical right shift, rounding, and unsigned saturation to half the width operation on the fifth operand to generate a fourth initial shift operation result, sub-step N2 is executed.
[0306] In sub-step N2, the consecutive lower half data of each element included in the fourth initial shift operation result is selected respectively, and the element after the selection operation is determined as the shift operation result.
[0307] In the embodiment of the present application, after the fourth initial shift operation result is generated, the consecutive lower half data of each element is selected from the fourth initial shift operation result respectively, and the consecutive lower half data of each selected element is determined as the shift operation result.
[0308] After the shift operation result is obtained, sub-step N3 is executed.
[0309] In sub-step N3, the storage positions in the destination register are divided according to a preset value to generate a plurality of storage areas.
[0310] In the embodiment of the present application, the preset value refers to a value for dividing the storage positions in the destination register into regions. The preset value is, that is, the data bit width occupied by the target element. The specific numerical value of the preset value can be determined according to business requirements, and the embodiment of the present application does not limit it. The preset value is preferably a value such that each storage area divided according to the preset value has the same size (the data bit widths stored in each storage area are the same).
[0311] After a plurality of storage areas are generated, sub-step N4 is executed.
[0312] In sub-step N4, the data in the shift operation result is sequentially written into the lower half of each storage area.
[0313] In the embodiment of the present application, after a plurality of storage areas are generated and the shift operation result is obtained, the elements in the shift operation result can be sequentially written into the lower half of each storage area with the target element as the target.
[0314] Optionally, if the fifth operand is M bits, the fourth initial shift operation result is M / 2 bits. In this case, the shift operation result is the data shown from bit 0 to bit M / 4 - 1 in the fourth initial shift operation result. Exemplarily, the first source register is the vector register xj, the first source register is 2M, the source elements corresponding to each M bits of xj are used as the fifth operand, and using the vector shift instruction, for each source element of every M bits included in the fifth operand, a logical right shift, rounding, and unsigned saturation to half value width operation is performed to obtain the fourth initial shift operation result. The shift amount is derived from the shift amount register, and the lower half of the elements selected from each element in the fourth initial shift operation result are sequentially written to the lower half of each target element of every M bits of the vector register xd, and the upper half data of each target element of every M bits is set to 0. Here, the data type of the source element in the fifth operand is any one of half word, word, and double word, and the data type of the target element written to the vector register xd is byte, half word, word corresponding to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeated here.
[0315] After sequentially writing the data in the shift operation result to the lower half of each storage area, sub-step N5 is executed.
[0316] In sub-step N5, the values at the positions where the data in each storage area has not been written are each set to zero.
[0317] After sequentially writing the elements in the shift operation result to the lower half of each storage area as target elements, the values at the positions where the target elements in each storage area have not been written can each be set to zero.
[0318] Exemplarily, when the first source register is vj / xj and the third type of vector operation code is the vector operation code of 11 a vector shift instruction is executed, that is, for each source element every 128 bits of the first source register vj / xj, a logical right shift, rounding, and unsigned saturation to half-width operation is performed, the lower half of each element in the shift result is selected and sequentially written to the lower half of each target element every 128 bits of the destination register vd / xd, the upper half of each target element every 128 bits of the destination register is set to 0, the shift amount of each element is derived from the shift amount register vk / xk, and the data type of the source element is any one of half-word, word, and double-word.
[0319] It should be understood that the above example is only an example illustrated for better understanding of the technical solution according to the embodiments of the present application, and does not uniquely limit the embodiments of the present application.
[0320] In the fourth specific implementation form of the present application, the third type of vector operation code is the 12th vector operation code, and the fifth operand is any consecutive source element of the first source register. The specific processing method may include the following sub-steps S1 to S5.
[0321] In sub-step S1, according to the shift amount, a right arithmetic shift, rounding, and unsigned saturation to half-width operation is performed on the fifth operand to generate a fourth initial shift operation result.
[0322] In the embodiments of the present application, after the 12th vector operation code is obtained, according to the shift amount, a right arithmetic shift, rounding, and unsigned saturation to half-width operation is performed on each element included in the fifth operand to generate a fourth initial shift operation result, where the shift amount is derived from the shift amount register.
[0323] After the fourth initial shift operation result is generated, sub-step S2 is executed.
[0324] In sub-step S2, the consecutive lower half data of each element included in the fourth initial shift operation result are respectively selected, and the elements after the selection operation are determined as the shift operation result.
[0325] In the embodiment of the present application, after the fourth initial shift operation result is generated, the consecutive lower half data of each element can be respectively selected from the fourth initial shift operation result, and the consecutive lower half data of each selected element can be determined as the shift operation result.
[0326] Optionally, the first source register includes 2N-bit data, and the 2N-bit data can correspond to a plurality of half-word elements, word elements, or double-word elements. The step of performing a selection operation in the first source register to obtain a fifth operand includes determining all source elements corresponding to each data group as a fifth operand by taking the data every M bits of the first source register as a group of data groups. According to the shift amount, for the fifth operand, perform a right arithmetic shift, rounding, and unsigned saturation to half value width operation. The step of generating the fourth initial shift operation result includes determining the shift value in the shift value register corresponding to each source element in the fifth operand, and respectively performing a right arithmetic shift, rounding, and unsigned saturation to half value width operation on each source element according to the shift value corresponding to each source element to obtain the fourth initial shift operation result.
[0327] Furthermore, the meanings of right arithmetic shift, rounding, half value width, Unsigned and saturation are the same as those in Embodiment 2. Since the process of performing a right arithmetic shift, rounding, and unsigned saturation to half value width operation is the same as that in Embodiment 2, it will not be repeatedly described here.
[0328] After the shift operation result is obtained, sub-step S3 is executed.
[0329] In sub-step S3, the storage positions in the destination register are divided according to preset values to generate a plurality of storage areas.
[0330] In the embodiment of the present application, the preset value refers to a numerical value for dividing the storage positions in the destination register into areas. The preset value is, that is, the data bit width occupied by the target element. The specific numerical value of the preset value can be determined according to business requirements, and the embodiment of the present application does not limit it. It is preferable that the preset value is a value such that each storage area divided according to the preset value has the same size (the data bit widths stored in each storage area are the same).
[0331] After a plurality of storage areas are generated, sub-step S4 is executed.
[0332] In sub-step S4, the data in the shift operation result is sequentially written into the lower half of each storage area.
[0333] In the embodiment of the present application, after a plurality of storage areas are generated and the shift operation result is obtained, the elements in the shift operation result can be sequentially written into the lower half of each storage area with the target element as the target.
[0334] Optionally, if the fifth operand is M bits, the fourth initial shift operation result is M / 2 bits. In this case, the shift operation result is the data shown from bit 0 to bit M / 4 - 1 in the fourth initial shift operation result. Exemplarily, the first source register is the vector register xj, the first source register is 2M, the source elements corresponding to each M bits of xj are used as the fifth operand, and using the vector shift instruction, for each source element of each M bits included in the fifth operand, perform a right arithmetic shift, round, and unsigned saturation to half the width operation to obtain the fourth initial shift operation result. The shift amount is derived from the shift amount register, and the lower half of the elements selected from each element in the fourth initial shift operation result are sequentially written to the lower half of each target element of each M bits of the vector register xd, and the upper half data of each target element of each M bits is set to 0. Here, the data type of the source element in the fifth operand is any one of half-word, word, and double-word, and the data type of the target element written to the vector register xd is byte, half-word, word corresponding to the data type of the above source element. Since the correspondence relationship between the data type of the source element and the data type of the target element has been described before, it will not be repeated here.
[0335] After sequentially writing the data in the shift operation result to the lower half of each storage area, execute sub-step S5.
[0336] In sub-step S5, set the values of the positions in each storage area where the data has not been written to zero respectively.
[0337] After sequentially writing the elements in the shift operation result to the lower half of each storage area as target elements, the values of the positions in each storage area where the target elements have not been written can be set to 0 respectively.
[0338] It should be understood that the above examples are merely examples for better understanding the technical solution according to the embodiments of the present application, and do not uniquely limit the embodiments of the present application.
[0339] By adopting the technical solution according to the present application and executing an instruction including a ninth vector opcode and a shift amount, a series of operations such as a logical shift on two source elements, rounding, and saturating with sign to half-width are implemented. By executing an instruction including a tenth vector opcode and a shift amount, the shift amount is derived from a register, and a series of operations such as an arithmetic shift on two source elements, rounding, and saturating with sign to half-width are implemented. By executing an instruction including an eleventh vector opcode and a shift amount, a series of operations such as a logical shift on two source elements, rounding, and saturating without sign to half-width are implemented. By executing an instruction including a twelfth vector opcode and a shift amount, a series of operations such as an arithmetic shift on two source elements, rounding, and saturating without sign to half-width are implemented. Therefore, by adopting the technical solution according to the present invention, different shift needs can be achieved using different shift parameters, whereby multiple vector shift needs can be achieved with one shift instruction, effectively reducing the system overhead and improving the execution efficiency of vector shifts for specific functions.
[0340] Embodiment 5
[0341] Please refer to FIG. 8. FIG. 8 shows a structural block diagram of a processor according to Embodiment 2 of the present application.
[0342] As shown in FIG. 8, the processor may include a plurality of vector registers, an instruction decoding unit 83, and an execution unit 84. The plurality of vector registers include a source register 81 for storing source elements to be operated during the execution of a vector shift operation, and a destination register 82. The command decoding unit 83 is for decoding a vector shift command including a register identifier and a shift parameter. The register identifier includes a source register identifier and a destination register identifier. The source register identifier is for indicating the source register 81, and the destination register identifier is for indicating the destination register 82. In response to the vector shift command, the execution unit 84 performs a vector shift operation on the source elements obtained from the source register 81 according to the shift parameter, obtains the target elements after the vector shift operation, and writes the target elements into the destination register 82.
[0343] Preferably, the execution unit 84 determines a shift amount and a shift operation rule according to the shift parameter. Here, there is at least one source element on which the vector shift operation is performed. According to the shift amount and the shift operation rule, a corresponding shift operation is performed on the source elements in the source register to generate a shift operation result, and the shift operation result Elements in is determined as the target element.
[0344] Preferably, the shift parameter includes a shift amount and an opcode. The shift amount is for indicating the number of shift bits of the source elements to be operated when the vector shift Operation is executed, and the opcode is for indicating the shift operation rule to be performed on the source elements in the source register and the target elements in the destination register. The execution unit 84 selects from the source register according to the opcode the vector shift OperationSelect a source element to execute, determine the selected source element as an operand, execute a corresponding shift operation on the operand according to the opcode, generate a shift operation result, determine the storage form of the target element in the destination register according to the opcode, and store the target element in the destination register according to the storage form.
[0345] Preferably, the shift amount is an immediate number, and the source register includes a first source register and a second source register.
[0346] Preferably, the opcode is a first type of vector opcode, The execution unit 84 determines all source elements in the first source register as one operand and all source elements in the second source register as operands according to the first type of vector opcode, After splicing the operand in the first source register and the operand in the second source register according to the first type of vector opcode, generate a first spliced vector, According to the immediate number, perform an operation of shifting, rounding and saturating to half value width on each source element in the first spliced vector, and generate a first initial shift operation result, Execute a bit selection operation on the first initial shift operation result to generate a shift operation result. Here, the bit selection operation includes an operation of selecting consecutive lower half data for each element included in the first initial shift operation result, an operation of selecting consecutive upper half data for each element included in the first initial shift operation result, an operation of selecting intermediate consecutive specified bits de - An operation of selecting data, and any one of an operation of selecting non-consecutive specified bit data for each element included in the first initial shift operation result.
[0347] Preferably, the first type of vector opcode is the first vector opcode, The execution unit 84 performs an operation of right logical shift, rounding, and saturating with sign to half value width on each source element in the first splice vector according to the immediate value number, and generates a first initial shift operation result. Selects the consecutive lower half data of each element included in the first initial shift operation result, and determines the element after the selection operation as the shift operation result.
[0348] Preferably, the first type of vector opcode is the second vector opcode, The execution unit 84 performs an operation of right arithmetic shift, rounding, and saturating with sign to half value width on each source element in the first splice vector according to the immediate value number, and generates a first initial shift operation result. Selects the consecutive lower half data of each element included in the first initial shift operation result, and determines the element after the selection operation as the shift operation result.
[0349] Preferably, the first type of vector opcode is the third vector opcode, The execution unit 84 performs an operation of right logical shift, rounding, and saturating without sign to half value width on each source element in the first splice vector according to the immediate value number, and generates a first initial shift operation result. Selects the consecutive lower half data of each element included in the first initial shift operation result, and determines the element after the selection operation as the shift operation result.
[0350] Preferably, the first type of vector opcode is the fourth vector opcode, The execution unit 84 performs an operation of right arithmetic shift, rounding, and saturating without sign to half value width on each source element in the first splice vector according to the immediate value number, and generates a first initial shift operation result. Select each of the consecutive lower half data among the elements included in the first initial shift operation result, and determine the elements after the selection operation as the shift operation result.
[0351] Preferably, the opcode is a second type of vector opcode, The execution unit 84 executes a selection operation on the first source register and the second source register according to the second type of vector opcode to obtain a first operand and a second operand. Here, the selection operation includes any one of an operation of selecting consecutive lower half data for each element of the first source register and the second source register, an operation of selecting consecutive upper half data for each element of the first source register and the second source register, an operation of selecting consecutive specified bit data in the middle for each element of the first source register and the second source register, and an operation of selecting non-consecutive specified bit data for each element of the first source register and the second source register. The execution unit 84 also Determine the data other than the first operand in the first source register as a third operand, and determine the data other than the second operand in the second source register as a fourth operand. After splicing the first operand and the second operand, generate a second splice vector. After splicing the third operand and the fourth operand, generate a third splice vector. Here, the data type of the elements included in the second splice vector and the third splice vector is any one of half word, word, double word, and quad word. According to the immediate value number, perform an operation of shifting, rounding, and saturating to half value width on each element in the second splice vector to generate a second initial shift operation result. At the same time, according to the immediate value number, perform an operation of shifting, rounding, and saturating to half value width on each element in the third splice vector to generate a third initial shift operation result. Perform a bit selection operation on the second initial shift operation result to generate a first shift operation result, perform a bit selection operation on the third initial shift operation result to generate a second shift operation result, where To execute the bit selection operation, Include any one of the operations of selecting consecutive lower half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, selecting consecutive upper half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, selecting consecutive specified bit data in the middle for each element included in the second initial shift operation result and each element included in the third initial shift operation result, and selecting non-consecutive specified bit data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, Write the first shift operation result to the corresponding storage location in the destination register according to the bit selection operation position of the first shift operation result, Write the second shift operation result to the corresponding storage location in the destination register according to the bit selection operation position of the second shift operation result.
[0352] Preferably, the second type of vector opcode is the fifth vector opcode, the first operand is data consisting of consecutive lower halves of each element of the first source register, the second operand is data consisting of consecutive lower halves of each element of the second source register, the third operand is data consisting of consecutive upper halves of each element of the first source register, and the fourth operand is data consisting of consecutive upper halves of each element of the second source register, The execution unit 84 performs an operation of right logical shift, rounding, and signed saturation to half value width on each element in the second splice vector according to the immediate value, generates a second initial shift operation result, and performs an operation of right logical shift, rounding, and signed saturation to half value width on each element in the third splice vector according to the immediate value, and generates a third initial shift operation result. Selects the consecutive lower half data of each element included in the second initial shift operation result respectively, determines the selected data as the first shift operation result, and selects the consecutive upper half data of each element included in the third initial shift operation result respectively, and determines the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and Data At least one second target element is determined according to Writes each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively. Writes each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged respectively.
[0353] Preferably, the second type of vector opcode is the sixth vector opcode, the first operand is data consisting of the consecutive lower half of each element of the first source register, the second operand is data consisting of the consecutive lower half of each element of the second source register, the third operand is data consisting of the consecutive upper half of each element of the first source register, and the fourth operand is data consisting of the consecutive upper half of each element of the second source register Data And The execution unit performs an operation of right arithmetic shift, rounding, and saturating with sign to half value width on each element in the second splice vector according to the immediate value, generates a second initial shift operation result, and performs an operation of right arithmetic shift, rounding, and saturating with sign to half value width on each element in the third splice vector according to the immediate value, generates a third initial shift operation result, selects the consecutive lower half data of each element included in the second initial shift operation result respectively, determines the selected data as the first shift operation result, and selects the consecutive upper half data of each element included in the third initial shift operation result respectively, determines the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and Data at least one second target element is determined according to the each First of the target elements in the first shift operation result is written to the lower half of the position where each First of the target elements in the destination register is arranged respectively, each Second of the target elements in the second shift operation result is written to the upper half of the position where each Second of the target elements in the destination register is arranged respectively.
[0354] Preferably, the second type of vector opcode is the seventh vector opcode, the first operand is data consisting of the consecutive lower half of each element of the first source register, the second operand is data consisting of the consecutive lower half of each element of the second source register, the third operand is data consisting of the consecutive upper half of each element of the first source register, and the fourth operand is data consisting of the consecutive upper half of each element of the second source register. The execution unit 84 performs an operation of right logical shift, rounding, and unsigned saturation to half value width on each element in the second splice vector according to the immediate value number, generates a second initial shift operation result, and performs an operation of right logical shift, rounding, and unsigned saturation to half value width on each element in the third splice vector according to the immediate value number, generates a third initial shift operation result, selects the consecutive lower half data of each element included in the second initial shift operation result respectively, determines the selected data as the first shift operation result, selects the consecutive upper half data of each element included in the third initial shift operation result respectively, determines the selected data as the second shift operation result, determines at least one first target element according to the data included in the first shift operation result, and determines at least one second target element according to the data included in the second shift operation result, Data and is configured to determine at least one second target element according to the data included in the second shift operation result, writes each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively, and writes each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged respectively.
[0355] Preferably, the second type of vector opcode is the eighth vector opcode, the first operand is data consisting of the consecutive lower half of each element of the first source register, the second operand is data consisting of the consecutive lower half of each element of the second source register, the third operand is data consisting of the consecutive upper half of each element of the first source register, and the fourth operand is data consisting of the consecutive upper half of each element of the second source register. The execution unit 84 performs an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each element in the second splice vector according to the immediate value, and generates a second initial shift operation result. At the same time, Performs an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each element in the third splice vector according to the immediate value, and generates a third initial shift operation result. Selects the consecutive lower half data of each element included in the second initial shift operation result respectively, determines the selected data as the first shift operation result, and selects the consecutive upper half data of each element included in the third initial shift operation result respectively, and determines the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. Data It is arranged to determine at least one second target element according to the data included in the second shift operation result. Write each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively. Write each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged respectively.
[0356] Preferably, the instruction further includes a shift amount register identifier for indicating a shift amount register, which is a register for storing a shift amount.
[0357] Preferably, the opcode is a third type of vector opcode, and the source register includes a first source register. The execution unit executes a selection operation on the first source register according to the third type of vector operation code to obtain a fifth operand, where the selection operation includes any one of the following operations: an operation of selecting consecutive lower half data for each element of the first source register, an operation of selecting consecutive upper half data for each element of the first source register, an operation of selecting consecutive specified bits of data in the middle for each element of the first source register, and an operation of selecting non-consecutive specified bits of data for each element of the first source register. According to the third type of vector operation code and the shift amount, perform an operation of shifting, rounding, and saturating to a half value width on the fifth operand to generate a fourth initial shift operation result. Execute a bit selection operation on the fourth initial shift operation result to generate a shift operation result, where the bit selection operation includes any one of the following operations: an operation of respectively selecting consecutive lower half data for each element included in the fourth initial shift operation result, an operation of respectively selecting consecutive upper half data for each element included in the fourth initial shift operation result, an operation of respectively selecting consecutive specified bits of data in the middle for each element included in the fourth initial shift operation result, and an operation of respectively selecting non-consecutive specified bits of data for each element included in the fourth initial shift operation result. Sequentially write the data in the shift operation result to the corresponding positions in the destination register. According to the third type of vector operation code, set the value of the position in the destination register where the data has not been written.
[0358] Preferably, the third type of vector operation code is the ninth vector operation code, and the fifth operand is any consecutive source element of the first source register. The execution unit 84 performs a right logical shift, rounding, and signed saturation to half-width on each element included in the fifth operand according to the shift amount to generate a fourth initial shift operation result; The successive lower halves of each element included in the fourth initial shift operation result Data Each of the elements is selected, and the element after the selection is determined as the shift operation result. Dividing the storage locations in the destination register according to a preset value to determine a storage area for each target element; The data in the shift operation result is written sequentially into the lower half of each storage area; The values of the locations in each storage area where no data is written are set to zero.
[0359] Preferably, the third type vector opcode is a tenth vector opcode, and the fifth operand is any successive source element of the first source register; The execution unit 84 performs a right arithmetic shift, rounding, and half-width signed saturation operation on each element included in the fifth operand according to the shift amount to generate a fourth initial shift operation result; The successive lower halves of each element included in the fourth initial shift operation result Data Each of the elements is selected, and the element after the selection is determined as the shift operation result. Dividing the storage locations in the destination register according to a preset value to determine a storage area for each target element; The data in the shift operation result is written sequentially into the lower half of each storage area; The values of the locations in each storage area where no data is written are set to zero.
[0360] Preferably Before the third type vector opcode is an eleventh vector opcode, and the fifth operand is any consecutive source element of the first source register; The execution unit 84 performs an operation of right logical shift, rounding, and unsigned saturation to half value width on each element included in the fifth operand according to the shift amount, and generates a fourth initial shift operation result. Selects consecutive lower halves of each element included in the fourth initial shift operation result, Data Determines the element after the selection operation as the shift operation result, Divides the storage positions in the destination register according to a preset value, determines the storage area for each target element, Sequentially writes the data in the shift operation result to the lower halves of each storage area, Sets the values of the positions where the data has not been written in each storage area to zero.
[0361] Preferably, the third type of vector operation code is the twelfth vector operation code, and the fifth operand is any consecutive source element of the first source register. The execution unit 84 performs an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each element included in the fifth operand according to the shift amount, and generates a fourth initial shift operation result. Selects consecutive lower halves of each element included in the fourth initial shift operation result, Data Determines the element after the selection operation as the shift operation result, Divides the storage positions in the destination register according to a preset value, determines the storage area for each target element, Sequentially writes the data in the shift operation result to the lower halves of each storage area, Sets the values of the positions where the data has not been written in each storage area to zero.
[0362] Preferably, the number of source registers is one or more, the number of destination registers is one, and the source register identifier is the same as or different from the destination register identifier.
[0363] Preferably, there are a plurality of the source registers and one destination register. Each source register identifier in all of the source registers is different from the destination register identifier, or there is one source register identifier in all of the source registers that is the same as the destination register identifier.
[0364] Example 6
[0365] Please refer to FIG. 9. FIG. 9 shows a schematic structural diagram of an electronic device for performing a vector shift operation according to the present application. Example 6
[0366] As shown in FIG. 9, the electronic device can include one or more of a processing assembly 902, a memory 904, a power supply assembly 906, a multimedia assembly 908, an audio assembly 910, an input / output (I / O) interface 912, a sensor assembly 914, and a communication assembly 916.
[0367] The processing assembly 902 generally controls the overall operations of the electronic device, such as operations related to display, data communication, camera operations, and recording operations. Processing assembly 902 can include one or more processors 920 that execute instructions to complete all or part of the steps of the above-described method. Further, the processing assembly 902 can include one or more modules for facilitating the interaction between the processing assembly 902 and other assemblies. For example, Processing assembly 902 can include a multimedia module for facilitating the interaction between the multimedia assembly 908 and the processing assembly 902.
[0368] Memory 904 is configured to store various types of data for supporting operations in the electronic device. Examples of such data include instructions for any application or method operated on the electronic device, contact data, phone book data, messages, images, videos, etc. Memory 904 can be implemented by any type of volatile or non-volatile storage device such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk, or a combination thereof.
[0369] Power supply assembly 906 supplies power to various assemblies of the electronic device. The power supply assembly 906 can include a power management system, one or more power supplies, and Electronic device other assemblies related to the generation, management, and distribution of power to be provided to 900.
[0370] The multimedia assembly 908 includes a screen that provides an output interface between the electronic device and the user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). When the screen includes a touch panel, the screen can be implemented as a touch screen for receiving input signals from the user. The touch panel includes one or more touch sensors for sensing touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or swipe operation but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia assembly 908includes one front camera and / or one rear camera. The front camera and / or the rear camera can receive external multimedia data when the electronic device is in an operation mode such as a shooting mode or a video mode. Each front camera and rear camera may be a fixed optical lens system, or may have a focal length and an optical zoom function.
[0371] The audio assembly 910 is configured to output and / or input an audio signal. For example, the audio assembly 910 includes a microphone (MIC) configured to receive an external audio signal when the terminal is in an operation mode such as a call mode, a recording mode, a voice recognition mode, etc. The received audio signal can be further stored in the memory 904 or transmitted by the communication assembly 916. In some embodiments, the audio assembly 910 further includes a speaker for outputting an audio signal.
[0372] The I / O interface 912 provides an interface between the processing assembly 902 and the peripheral interface module, and the peripheral interface module may be a keypad, a click wheel, buttons, etc. These buttons can include, but are not limited to, a home button, a volume button, a start button, a lock button.
[0373] The sensor assembly 914 includes one or more sensors for providing an assessment of various aspects of the state of the electronic device 900. For example, the sensor assembly 914 can detect the open / closed state of the electronic device 900, such as the relative positioning of an assembly where the assembly is a display and keypad of a terminal, and the sensor assembly 914 can also detect changes in the position of the terminal or one assembly of the terminal, the presence or absence of user contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and changes in the temperature of the electronic device. The sensor assembly 914 can include a proximity sensor configured to detect the presence of nearby objects even without any physical contact. The sensor assembly 914 can also include an optical sensor such as a CMOS or Charge coupled device, CCD ) image sensor for imaging applications. In some embodiments, the sensor assembly 914 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0374] The communication assembly 916 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard such as WiFi, 2G or 3G, or a combination thereof. In one exemplary embodiment, the communication component 916 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component 916 further includes a Near Field Communication (NFC) module for facilitating short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0375] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic elements to execute the vector shift method described above.
[0376] The electronic device according to the embodiment of the present application is for implementing the vector shift method executed using corresponding instructions in the above-described embodiments of the plurality of methods, and since it has the beneficial effects of the implementation of the corresponding method, it will not be repeatedly described here.
[0377] Each embodiment in this specification is described progressively, and each embodiment focuses on the differences from other embodiments and proceeds with the description. For the same or similar parts of each embodiment, reference may be made to each other. For the embodiment of the apparatus, since it is basically similar to the embodiment of the method, the description is relatively simple, and the relevant parts may be described with reference to the embodiment of the method.
[0378] The above has described in detail the vector shift method, processor, electronic device, and readable storage medium provided by the present application. In this specification, although the principles and embodiments of the present invention are described by applying specific examples, the above embodiments are only described to help understand the method and its main concept of the present invention. At the same time, those skilled in the art can make changes to specific embodiments and their application scopes based on the concept of the present invention. As described above, the description in this specification should not be understood as limiting the present invention.
[0379] The algorithms and displays provided in this specification are not inherently related to a particular computer, electronic system, or other device. A variety of general-purpose systems can also be used in combination with the teachings based on this specification. The structures required to construct such systems will be apparent in light of the above description. Further, this application is not directed to a particular programming language. To implement the content of this application described in this specification, a variety of programming languages can be utilized, and it should be understood that the above description regarding a particular language is provided to disclose the best mode of this application.
[0380] In the specification provided by this application, a large number of specific details are described. However, it can be understood that the embodiments of this application can be implemented without these specific details. In some examples, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.
[0381] Similarly, to simplify this disclosure and to aid in the understanding of one or more aspects of various inventions, in the above description of the exemplary embodiments of this application, it should be understood that various features of this application may be grouped together in a single embodiment, drawing, or description thereof. However, the methods according to this disclosure should not be construed as reflecting an intention that the claimed application requires more features than are explicitly recited in each claim. More precisely, as reflected in the following claims, aspects of the present invention have a smaller number of features than all the features of the individual embodiments disclosed previously. Accordingly, the claims according to a particular embodiment thereby expressly incorporate that particular embodiment, and each claim itself functions as a separate embodiment of this application.
[0382] Those skilled in the art can understand that the modules within the devices in the embodiments can be adaptively modified and arranged in one or more devices different from the embodiments. The modules or units or assemblies in the embodiments can be combined into a single module or unit or assembly, and further, can be divided into a plurality of sub-modules or sub-units or sub-assemblies. Except for the case where at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the appended claims, abstract, and appended drawings) and all processes or units of the methods or devices thus disclosed can be combined in any combination. Each feature disclosed in this specification (including the appended claims, abstract, and appended drawings) may be replaced by alternative features that provide the same, equivalent, or similar purposes, unless otherwise specified.
[0383] In addition, those skilled in the art can understand that some of the embodiments described in this specification include some features rather than other features included in other embodiments, but the combinations of features of different embodiments are within the scope of this application and mean that different embodiments are formed. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0384] Embodiments of the various components of this application can be implemented in hardware, or software modules executed on one or more processors, or a combination thereof. Those skilled in the art know that a microprocessor or a digital signal processor (DSP) is according to the embodiments of this application Electronic deviceIt should be understood that some or all of the functions of some or all of the components in [the relevant context] can be actually used for implementation. Also, this application can be implemented as a device or apparatus program (for example, a computer program and a computer program product) for executing some or all of the methods described in this specification. Such a program for implementing this application may be stored in a computer-readable medium or may be in the form of one or more signals. Such signals may be downloadable from an Internet site, may be provided by a carrier signal, or may be provided in any other form.
[0385] The above embodiments are for illustrative purposes of this application and do not limit this application. It should be noted that those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, reference signs located between parentheses should not be construed as limiting the scope of the claims. The term "comprising" does not exclude the presence of elements or steps not recited in the claims. The term "a" or "one" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by hardware including several different elements and a properly programmed computer. In the individual claims listing several devices, some of these devices can be embodied by the same hardware. The use of terms such as first, second, and third does not indicate any order. These terms may be construed as names.
[0386] This application claims the priority of a Chinese patent application with an application number of 202111509173.2 and an invention title of "Vector Shift Method, Processor, and Electronic Device", which was filed with the China National Intellectual Property Administration on December 10, 2021, and the entire content thereof is incorporated herein by reference.
Claims
Claim 1 Receiving an instruction including a register identifier and a shift parameter, where the register identifier includes a source register identifier and a destination register identifier, the source register identifier is for indicating a source register, the source register is a register for storing source elements to be operated during execution of a vector shift operation, the destination register identifier is for indicating a destination register, the destination register is a register for storing target elements obtained after execution of the vector shift operation, the shift parameter is for indicating a rule based on which a vector shift operation is to be performed on the source elements, the shift parameter includes a shift amount and an opcode, the shift amount is for indicating the number of shift bits of the source elements to be operated when the vector shift operation is executed, the opcode is for indicating a shift operation rule to be executed on the source elements in the source register and the target elements in the destination register, and the opcode includes an instruction name and a parameter for indicating the data types of the source elements and the target elements; Executing the instruction to determine a shift amount and a shift operation rule according to the shift parameter, where there is at least one source element on which the vector shift operation is to be executed; Selecting, according to the opcode, source elements for performing the vector shift operation from the source register and determining the selected source elements as operands; Executing a corresponding shift operation on the operands according to the opcode to generate a shift operation result; Determining elements in the shift operation result as target elements; Determining, according to the opcode, a storage form of the target elements in the destination register; and Storing the target elements in the destination register according to the storage form. A vector shift method characterized by the above is provided. Claim 2 The shift amount is an immediate value, and the source register includes a first source register and a second source register. The vector shift method according to claim 1 is characterized by this.
3. The opcode is a first type of vector opcode. According to the opcode, the step of selecting a source element for executing the vector shift operation from among the source registers and determining the selected source element as an operand is as follows: According to the first type of vector opcode, it includes the step of determining all source elements in the first source register as one operand and determining all source elements in the second source register as an operand. According to the opcode, the step of performing a corresponding shift operation on the operand and generating a shift operation result is as follows: According to the first type of vector opcode, after splicing the operand in the first source register and the operand in the second source register, the step of generating a first spliced vector is as follows: According to the immediate value, for each source element in the first spliced vector, perform an operation of shifting, rounding, and saturating to a half value width to generate a first initial shift operation result. The step of performing a bit selection operation on the first initial shift operation result and generating a shift operation result includes, here, the bit selection operation being any one of the following operations: an operation of selecting consecutive lower half data for each element included in the first initial shift operation result, an operation of selecting consecutive upper half data for each element included in the first initial shift operation result, an operation of selecting consecutive intermediate specified bit data for each element included in the first initial shift operation result, and an operation of selecting non-consecutive specified bit data for each element included in the first initial shift operation result. The vector shift method according to claim 2 is characterized by this.
4. The first type of vector opcode is a first vector opcode. According to the immediate value, for each source element in the first spliced vector, perform an operation of shifting, rounding, and saturating to a half value width to generate a first initial shift operation result. The step is as follows: According to the immediate value number, for each source element in the first splice vector, perform an operation of right logical shift, rounding, and saturating with a signed half-width, and include a step of generating a first initial shift operation result. The step of performing a bit selection operation on the first initial shift operation result to generate a shift operation result is including the step of respectively selecting consecutive lower half data in each element included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result. The vector shift method according to claim 3 is characterized by this.
5. The first type of vector opcode is the second vector opcode, According to the immediate value number, for each source element in the first splice vector, perform an operation of shifting, rounding, and saturating with a half-width, and the step of generating a first initial shift operation result is According to the immediate value number, for each source element in the first splice vector, perform an operation of right arithmetic shift, rounding, and saturating with a signed half-width, and include a step of generating a first initial shift operation result. The step of performing a bit selection operation on the first initial shift operation result to generate a shift operation result is including the step of respectively selecting consecutive lower half data in each element included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result. The vector shift method according to claim 3 is characterized by this.
6. The first type of vector opcode is the third vector opcode, According to the immediate value number, for each source element in the first splice vector, perform an operation of shifting, rounding, and saturating with a half-width, and the step of generating a first initial shift operation result is According to the immediate value number, for each source element in the first splice vector, perform an operation of right logical shift, rounding, and saturating without a sign with a half-width, and include a step of generating a first initial shift operation result. The step of performing a bit selection operation on the first initial shift operation result to generate a shift operation result is including the step of respectively selecting consecutive lower half data in each element included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result. The vector shift method according to claim 3 is characterized by this.
7. The first type of vector opcode is a fourth vector opcode, and according to the immediate value, for each source element in the first splice vector, performing an operation of shifting, rounding, and saturating to a half-width value to generate a first initial shift operation result includes according to the immediate value, for each source element in the first splice vector, performing an operation of right arithmetic shift, rounding, and unsigned saturation to a half-width value to generate a first initial shift operation result, performing a bit selection operation on the first initial shift operation result to generate a shift operation result includes selecting each of the consecutive lower half data among the elements included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result, The vector shift method according to claim 3, characterized in that
8. The opcode is a second type of vector opcode, and according to the opcode, selecting a source element for performing the vector shift operation from the source register and determining the selected source element as an operand includes according to the second type of vector opcode, performing a selection operation on the first source register and the second source register respectively to obtain a first operand and a second operand, where the selection operation is an operation of selecting consecutive lower half data for each element of the first source register and the second source register, an operation of selecting consecutive upper half data for each element of the first source register and the second source register, an operation of selecting data of consecutive specified bits in the middle for each element of the first source register and the second source register, and an operation of selecting data of non-consecutive specified bits for each element of the first source register and the second source register, including any one of the steps, determining the data other than the first operand in the first source register as a third operand, and determining the data other than the second operand in the second source register as a fourth operand, including according to the opcode, performing a corresponding shift operation on the operand to generate a shift operation result includes After splicing the first operand and the second operand, generating a second splice vector, and after splicing the third operand and the fourth operand, generating a third splice vector, wherein the data type of the source elements included in the second splice vector and the third splice vector is any one of half-word, word, double-word, and quad-word, the step of According to the immediate value number, performing an operation of shifting, rounding, and saturating to a half-value width on each source element in the second splice vector to generate a second initial shift operation result, and according to the immediate value number, performing an operation of shifting, rounding, and saturating to a half-value width on each source element in the third splice vector to generate a third initial shift operation result, the step of Performing a bit selection operation on the second initial shift operation result to generate a first shift operation result, and performing a bit selection operation on the third initial shift operation result to generate a second shift operation result, including, here, performing the bit selection operation is an operation of selecting consecutive lower half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, an operation of selecting consecutive upper half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, an operation of selecting consecutive middle specified bit data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, and an operation of selecting non-consecutive specified bit data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, including any one of According to the storage form, the step of storing the target element in the destination register is According to the bit selection operation position of the first shift operation result, writing the first shift operation result to the corresponding storage position in the destination register, the step of writing the second shift operation result to a corresponding storage location in the destination register according to a bit selection operation position of the second shift operation result, characterized in that, The vector shift method according to claim 2.
9. The second type of vector opcode is a fifth vector opcode, the first operand is data consisting of consecutive lower halves of each element of the first source register, the second operand is data consisting of consecutive lower halves of each element of the second source register, the third operand is data consisting of consecutive upper halves of each element of the first source register, and the fourth operand is data consisting of consecutive upper halves of each element of the second source register. According to the immediate value number, for each source element in the second splice vector, perform an operation of shifting, rounding, and saturating to a half-value width to generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform an operation of shifting, rounding, and saturating to a half-value width to generate a third initial shift operation result, the steps are: According to the immediate value number, for each source element in the second splice vector, perform a right logical shift, rounding, and signed saturation to a half-value width to generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform a right logical shift, rounding, and signed saturation to a half-value width to generate a third initial shift operation result, including the steps of performing a bit selection operation on the second initial shift operation result to generate a first shift operation result, and performing a bit selection operation on the third initial shift operation result to generate a second shift operation result, the steps are: Selecting each consecutive lower half of the data of each element included in the second initial shift operation result, determining the selected data as a first shift operation result, and selecting each consecutive upper half of the data of each element included in the third initial shift operation result, and determining the selected data as a second shift operation result, wherein at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result; The step of writing the first shift operation result to a corresponding storage location in the destination register according to the bit selection operation position of the first shift operation result is: including the step of writing each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged; The step of writing the second shift operation result to a corresponding storage location in the destination register according to the bit selection operation position of the second shift operation result is: including the step of writing each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged, and the vector shift method according to claim 8 is characterized in that. [
10. ] The second type of vector opcode is the sixth vector opcode, the first operand is data consisting of each consecutive lower half of the elements of the first source register, the second operand is data consisting of each consecutive lower half of the elements of the second source register, the third operand is data consisting of each consecutive upper half of the elements of the first source register, and the fourth operand is data consisting of each consecutive upper half of the elements of the second source register; According to the immediate value number, for each source element in the second splice vector, perform an operation of shifting, rounding, and saturating to the half-value width to generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform an operation of shifting, rounding, and saturating to the half-value width to generate a third initial shift operation result. The steps are as follows: According to the immediate value number, for each source element in the second splice vector, perform an operation of right arithmetic shift, rounding, and signed saturation to the half-value width to generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform an operation of right arithmetic shift, rounding, and signed saturation to the half-value width to generate a third initial shift operation result. The steps include: Execute a bit selection operation on the second initial shift operation result to generate a first shift operation result, and execute a bit selection operation on the third initial shift operation result to generate a second shift operation result. The steps are as follows: Select the consecutive lower half data of each element included in the second initial shift operation result respectively, and determine the selected data as the first shift operation result. At the same time, select the consecutive upper half data of each element included in the third initial shift operation result respectively, and determine the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. According to the bit selection operation position of the first shift operation result, write the first shift operation result to the corresponding storage position in the destination register. The steps are as follows: Include the step of writing each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively. According to the bit selection operation position of the second shift operation result, write the second shift operation result to the corresponding storage position in the destination register. The steps are as follows: A step of writing each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged, The vector shift method according to claim 8, characterized in that it includes this.
11. The second type of vector opcode is the seventh vector opcode, the first operand is data consisting of consecutive lower halves of each element of the first source register, the second operand is data consisting of consecutive lower halves of each element of the second source register, the third operand is data consisting of consecutive upper halves of each element of the first source register, and the fourth operand is data consisting of consecutive upper halves of each element of the second source register. According to the immediate value number, for each source element in the second splice vector, perform an operation of shifting, rounding, and saturating to half value width to generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform an operation of shifting, rounding, and saturating to half value width to generate a third initial shift operation result. The steps are: According to the immediate value number, for each source element in the second splice vector, perform an operation of right logical shift, rounding, and unsigned saturation to half value width to generate a second initial shift operation result, According to the immediate value number, for each source element in the third splice vector, perform an operation of right logical shift, rounding, and unsigned saturation to half value width to generate a third initial shift operation result, Including these steps. Execute a bit selection operation on the second initial shift operation result to generate a first shift operation result, and execute a bit selection operation on the third initial shift operation result to generate a second shift operation result. The steps are: Selecting each consecutive lower half of the data of each element included in the second initial shift operation result, determining the selected data as a first shift operation result, and selecting each consecutive upper half of the data of each element included in the third initial shift operation result, and determining the selected data as a second shift operation result, where at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result; The step of writing the first shift operation result to a corresponding storage position in the destination register according to the bit selection operation position of the first shift operation result is including the step of writing each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged; The step of writing the second shift operation result to a corresponding storage position in the destination register according to the bit selection operation position of the second shift operation result is including the step of writing each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged, and the vector shift method according to claim 8, characterized in that.
12. The second type of vector opcode is the eighth vector opcode, the first operand is data consisting of each consecutive lower half of the elements of the first source register, the second operand is data consisting of each consecutive lower half of the elements of the second source register, the third operand is data consisting of each consecutive upper half of the elements of the first source register, and the fourth operand is data consisting of each consecutive upper half of the elements of the second source register. According to the immediate value number, for each source element in the second splice vector, perform an operation of shifting, rounding, and saturating to the half-value width, generate a second initial shift operation result, and according to the immediate value number, for each source element in the third splice vector, perform an operation of shifting, rounding, and saturating to the half-value width, and generate a third initial shift operation result. The steps are as follows: According to the immediate value number, for each source element in the second splice vector, perform an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width to generate a second initial shift operation result; According to the immediate value number, for each source element in the third splice vector, perform an operation of right arithmetic shift, rounding, and unsigned saturation to the half-value width to generate a third initial shift operation result, and the steps include: Execute a bit selection operation on the second initial shift operation result to generate a first shift operation result, and execute a bit selection operation on the third initial shift operation result to generate a second shift operation result. The steps are as follows: Select the consecutive lower half data of each element included in the second initial shift operation result respectively, determine the selected data as the first shift operation result, and select the consecutive upper half data of each element included in the third initial shift operation result respectively, and determine the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. According to the bit selection operation position of the first shift operation result, write the first shift operation result to the corresponding storage position in the destination register. The steps are as follows: Include the step of writing each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively. According to the bit selection operation position of the second shift operation result, write the second shift operation result to the corresponding storage position in the destination register. The steps are as follows: A step of writing each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged, the vector shift method according to claim 8, characterized in that it includes this.
13. The instruction further includes a shift amount register identifier for indicating a shift amount register which is a register for storing a shift amount, the vector shift method according to claim 1, characterized in that it includes this.
14. The opcode is a third type of vector opcode, and the source register includes a first source register. According to the opcode, the step of selecting a source element for executing the vector shift operation from among the source registers and determining the selected source element as an operand. According to the third type of vector opcode, the step of performing a selection operation in the first source register to obtain a fifth operand, the selection operation including any one of an operation of selecting consecutive lower half data for each element of the first source register, an operation of selecting consecutive upper half data for each element of the first source register, an operation of selecting consecutive specified bit data in the middle for each element of the first source register, and an operation of selecting non-consecutive specified bit data for each element of the first source register. According to the opcode, the step of performing a corresponding shift operation on the operand to generate a shift operation result. According to the third type of vector opcode and the shift amount, the step of performing an operation of shifting, rounding, and saturating to a half value width on the fifth operand to generate a fourth initial shift operation result. Performing a bit selection operation on the fourth initial shift operation result to generate a shift operation result, where the bit selection operation includes any one of the following operations: selecting the consecutive lower half data of each element included in the fourth initial shift operation result, selecting the consecutive upper half data of each element included in the fourth initial shift operation result, selecting the consecutive specified bits in the middle of each element included in the fourth initial shift operation result, and selecting the non-consecutive specified bits of each element included in the fourth initial shift operation result. According to the storage form, the step of storing the target element in the destination register is sequentially writing the data in the shift operation result to the corresponding positions in the destination register, and setting the values of the positions in the destination register where the data has not been written according to the third type of vector operation code. The vector shift method according to claim 13 is characterized by including the above steps.
15. The third type of vector operation code is the ninth vector operation code, and the fifth operand is any consecutive source element of the first source register. According to the third type of vector operation code and the shift amount, performing an operation of shifting, rounding, and saturating to half value width on the fifth operand. The step of generating the fourth initial shift operation result includes performing an operation of logically shifting to the right, rounding, and saturating with a signed half value width on each source element included in the fifth operand according to the shift amount, to generate the fourth initial shift operation result. Performing a bit selection operation on the fourth initial shift operation result to generate a shift operation result. The step includes selecting the consecutive lower half data of each element included in the fourth initial shift operation result, and determining the element after the selection operation as the shift operation result. sequentially writing the data in the shift operation result to corresponding positions in the destination register, and setting values of positions in the destination register where the data is not written according to the third type of vector operation code; dividing storage positions in the destination register according to a preset value, and determining a storage area for each target element; sequentially writing the data in the shift operation result to the lower halves of the respective storage areas; setting the values of positions in each storage area where the data is not written to zero, the vector shift method according to claim 14, characterized by including:
16. the third type of vector operation code is the tenth vector operation code, and the fifth operand is any consecutive source element of the first source register; performing an operation of shifting, rounding, and saturating to a half value width on the fifth operand according to the third type of vector operation code and the shift amount to generate a fourth initial shift operation result; including performing an operation of right arithmetic shifting, rounding, and saturating to a signed half value width on each source element included in the fifth operand according to the shift amount to generate a fourth initial shift operation result; executing a bit selection operation on the fourth initial shift operation result to generate a shift operation result; including selecting the consecutive lower half data of each element included in the fourth initial shift operation result, and determining the element after the selection operation as the shift operation result; sequentially writing the data in the shift operation result to corresponding positions in the destination register, and setting values of positions in the destination register where the data is not written according to the third type of vector operation code; dividing storage positions in the destination register according to a preset value, and determining a storage area for each target element; sequentially writing the data in the shift operation result to the lower halves of the respective storage areas; setting the values of positions in each storage area where the data is not written to zero, the vector shift method according to claim 14, characterized by including:
17. The third type of vector opcode is the 11th vector opcode, and the fifth operand is any consecutive source element of the first source register. The step of performing an operation of shifting, rounding, and saturating to half-width on the fifth operand according to the third type of vector opcode and the shift amount to generate a fourth initial shift operation result is: The step includes performing an operation of right logical shift, rounding, and unsigned saturation to half-width on each source element included in the fifth operand according to the shift amount to generate a fourth initial shift operation result. The step of performing a bit selection operation on the fourth initial shift operation result to generate a shift operation result is: The step includes selecting the consecutive lower half data of each element included in the fourth initial shift operation result respectively, and determining the element after the selection operation as the shift operation result. The step of sequentially writing the data in the shift operation result to the corresponding positions in the destination register, and setting the values of the positions in the destination register where the data is not written according to the third type of vector opcode is: The step of dividing the storage positions in the destination register according to a preset value and determining the storage area of each target element, The step of sequentially writing the data in the shift operation result to the lower half of each storage area, And the step of setting the values of the positions in each storage area where the data is not written to zero respectively. The vector shift method according to claim 14 is characterized by including the above steps.
18. The third type of vector opcode is the 12th vector opcode, and the fifth operand is any consecutive source element of the first source register. The step of performing an operation of shifting, rounding, and saturating to half-width on the fifth operand according to the third type of vector opcode and the shift amount to generate a fourth initial shift operation result is: The step includes performing an operation of right arithmetic shift, rounding, and unsigned saturation to half-width on each source element included in the fifth operand according to the shift amount to generate a fourth initial shift operation result. The step of performing a bit selection operation on the fourth initial shift operation result to generate a shift operation result is: including a step of selecting each of the consecutive lower half data of the elements included in the fourth initial shift operation result and determining the element after the selection operation as the shift operation result, sequentially writing the data in the shift operation result to the corresponding positions in the destination register, and according to the third type of vector opcode, setting the values of the positions in the destination register where the data is not written, the step is, dividing the storage positions in the destination register according to a preset value and determining the storage area of each target element, sequentially writing the data in the shift operation result to the lower half of each storage area, including a step of setting the values of the positions in each storage area where the data is not written to zero, the vector shift method according to claim 14, characterized in that.
19. The number of source registers is one or more, the number of destination registers is one, and the source register identifier is the same as or different from the destination register identifier. The vector shift method according to any one of claims 1, 14 to 18, characterized in that.
20. There are a plurality of source registers and one destination register, each source register identifier in all the source registers is different from the destination register identifier, or there is one source register identifier in all the source registers that is the same as the destination register identifier. The vector shift method according to any one of claims 1 to 13, characterized in that.
21. A processor including a plurality of vector registers, an instruction decoding unit, and an execution unit, the plurality of vector registers include a source register for storing source elements to be operated during the execution of a vector shift operation and a destination register, The command decoding unit is for decoding a vector shift command including a register identifier and a shift parameter. The register identifier includes a source register identifier for indicating a source register and a destination register identifier for indicating a destination register. The shift parameter includes a shift amount and an opcode. The shift amount is for indicating the number of shift bits of source elements to be operated when the vector shift operation is executed. The opcode is for indicating a shift operation rule to be executed on source elements in the source register and target elements in the destination register. The opcode includes a command name and a parameter for indicating the data types of the source elements and the target elements. The execution unit determines a shift amount and a shift operation rule according to the shift parameter. There is at least one source element for which the vector shift operation is executed. According to the opcode, a source element for executing the vector shift operation is selected from the source register, and the selected source element is determined as an operand. According to the opcode, a corresponding shift operation is executed on the operand to generate a shift operation result. An element in the shift operation result is determined as a target element. According to the opcode, a storage form of the target element in the destination register is determined, and the target element is stored in the destination register according to the storage form. The processor is characterized by the above. Claim 22 The processor according to claim 21, wherein the shift amount is an immediate value, and the source register includes a first source register and a second source register. Claim 23 The opcode is a first type of vector opcode. The execution unit determines all source elements in the first source register as one operand and all source elements in the second source register as operands according to the first type of vector opcode. According to the first type of vector opcode, after splicing the operand in the first source register and the operand in the second source register, a first spliced vector is generated. According to the immediate value number, for each source element in the first splice vector, perform an operation of shifting, rounding, and saturating to half value width, and generate a first initial shift operation result. Execute a bit selection operation on the first initial shift operation result to generate a shift operation result. Here, the bit selection operation includes any one of the following operations: an operation of selecting consecutive lower half data for each element included in the first initial shift operation result, an operation of selecting consecutive upper half data for each element included in the first initial shift operation result, an operation of selecting consecutive specified bits of data in the middle for each element included in the first initial shift operation result, and an operation of selecting non-consecutive specified bits of data for each element included in the first initial shift operation result. The processor according to claim 22 is characterized by this.
24. The first type of vector opcode is the first vector opcode. According to the immediate value number, the execution unit performs an operation of right logical shift, rounding, and saturating with sign to half value width on each source element in the first splice vector, and generates a first initial shift operation result. Select consecutive lower half data for each element included in the first initial shift operation result, and determine the element after the selection operation as the shift operation result. The processor according to claim 23 is characterized by this.
25. The first type of vector opcode is the second vector opcode. According to the immediate value number, the execution unit performs an operation of right arithmetic shift, rounding, and saturating with sign to half value width on each source element in the first splice vector, and generates a first initial shift operation result. Select consecutive lower half data for each element included in the first initial shift operation result, and determine the element after the selection operation as the shift operation result. The processor according to claim 23 is characterized by this.
26. The first type of vector opcode is the third vector opcode. According to the immediate value number, the execution unit performs an operation of right logical shift, rounding, and saturating without sign to half value width on each source element in the first splice vector, and generates a first initial shift operation result. Selecting each of the consecutive lower half data among the elements included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result, The processor according to claim 23, characterized in that.
27. The first type of vector opcode is the fourth vector opcode, The execution unit performs an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each source element in the first splice vector according to the immediate value number, and generates a first initial shift operation result, Selecting each of the consecutive lower half data among the elements included in the first initial shift operation result, and determining the element after the selection operation as the shift operation result, The processor according to claim 23, characterized in that.
28. The opcode is a second type of vector opcode, The execution unit performs a selection operation on the first source register and the second source register according to the second type of vector opcode to obtain a first operand and a second operand, where the selection operation is an operation of selecting consecutive lower half data for each element of the first source register and the second source register, an operation of selecting consecutive upper half data for each element of the first source register and the second source register, an operation of selecting consecutive specified bits of intermediate data for each element of the first source register and the second source register, and an operation of selecting non-consecutive specified bits of data for each element of the first source register and the second source register, including any one of them, The execution unit also Determining the data other than the first operand in the first source register as a third operand, and determining the data other than the second operand in the second source register as a fourth operand, After splicing the first operand and the second operand, generating a second splice vector, and after splicing the third operand and the fourth operand, generating a third splice vector, where the data type of the source elements included in the second splice vector and the third splice vector is any one of half word, word, double word, and quad word, According to the immediate value number, for each source element in the second splice vector, perform an operation of shifting, rounding, and saturating to the half-value width to generate a second initial shift operation result. Also, according to the immediate value number, for each source element in the third splice vector, perform an operation of shifting, rounding, and saturating to the half-value width to generate a third initial shift operation result. Execute a bit selection operation on the second initial shift operation result to generate a first shift operation result. Execute a bit selection operation on the third initial shift operation result to generate a second shift operation result. Here, the bit selection operation is executed by selecting the continuous lower half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, selecting the continuous upper half data for each element included in the second initial shift operation result and each element included in the third initial shift operation result, selecting the data of the intermediate continuous specified bits for each element included in the second initial shift operation result and each element included in the third initial shift operation result, and selecting the data of the non-continuous specified bits for each element included in the second initial shift operation result and each element included in the third initial shift operation result, and includes any one of them. Write the first shift operation result to the corresponding storage position in the destination register according to the bit selection operation position of the first shift operation result. Write the second shift operation result to the corresponding storage position in the destination register according to the bit selection operation position of the second shift operation result. The processor according to claim 22, characterized in that. [
29. ] The second type of vector operation code is the fifth vector operation code. The first operand is data consisting of the continuous lower half of each element of the first source register. The second operand is data consisting of the continuous lower half of each element of the second source register. The third operand is data consisting of the continuous upper half of each element of the first source register. The fourth operand is data consisting of the continuous upper half of each element of the second source register. The execution unit performs an operation of right logical shift, rounding, and signed saturation to a half value width on each source element in the second splice vector according to the immediate value, and generates a second initial shift operation result. The execution unit performs an operation of right logical shift, rounding, and signed saturation to a half value width on each source element in the third splice vector according to the immediate value, and generates a third initial shift operation result. Each of the consecutive lower half data of each element included in the second initial shift operation result is selected, and the selected data is determined as a first shift operation result. Each of the consecutive upper half data of each element included in the third initial shift operation result is selected, and the selected data is determined as a second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. Each first target element included in the first shift operation result is written to the lower half of the position where each first target element of the destination register is arranged. Each second target element included in the second shift operation result is written to the upper half of the position where each second target element of the destination register is arranged. The processor according to claim 28, characterized in that.
30. The second type of vector opcode is a sixth vector opcode. The first operand is data consisting of the consecutive lower half of each element of the first source register. The second operand is data consisting of the consecutive lower half of each element of the second source register. The third operand is data consisting of the consecutive upper half of each element of the first source register. The fourth operand is data consisting of the consecutive upper half of each element of the second source register. The execution unit performs an operation of right arithmetic shift, rounding, and signed saturation to half value width on each source element in the second splice vector according to the immediate value, and generates a second initial shift operation result. The execution unit performs an operation of right arithmetic shift, rounding, and signed saturation to half value width on each source element in the third splice vector according to the immediate value, and generates a third initial shift operation result. Each of the consecutive lower half data of each element included in the second initial shift operation result is selected, and the selected data is determined as the first shift operation result. Each of the consecutive upper half data of each element included in the third initial shift operation result is selected, and the selected data is determined as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. Each first target element included in the first shift operation result is written to the lower half of the position where each first target element of the destination register is arranged. Each second target element included in the second shift operation result is written to the upper half of the position where each second target element of the destination register is arranged. The processor according to claim 28, characterized in that.
31. The second type of vector opcode is the seventh vector opcode. The first operand is data consisting of the consecutive lower half of each element of the first source register. The second operand is data consisting of the consecutive lower half of each element of the second source register. The third operand is data consisting of the consecutive upper half of each element of the first source register. The fourth operand is data consisting of the consecutive upper half of each element of the second source register. The execution unit performs an operation of right logical shift, rounding, and unsigned saturation to half value width on each source element in the second splice vector according to the immediate value, and generates a second initial shift operation result. According to the immediate value number, perform an operation of right logical shift, rounding, and unsigned saturation to half value width on each source element in the third splice vector to generate a third initial shift operation result. Select the consecutive lower half data of each element included in the second initial shift operation result respectively, determine the selected data as the first shift operation result, and select the consecutive upper half data of each element included in the third initial shift operation result respectively, determine the selected data as the second shift operation result. Here, at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. Write each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged respectively. Write each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged respectively. The processor according to claim 28, characterized in that.
32. The second type of vector opcode is the eighth vector opcode, the first operand is data consisting of the consecutive lower half of each element of the first source register, the second operand is data consisting of the consecutive lower half of each element of the second source register, the third operand is data consisting of the consecutive upper half of each element of the first source register, and the fourth operand is data consisting of the consecutive upper half of each element of the second source register. According to the immediate value number, perform an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each source element in the second splice vector to generate a second initial shift operation result. According to the immediate value number, perform an operation of right arithmetic shift, rounding, and unsigned saturation to half value width on each source element in the third splice vector to generate a third initial shift operation result. Selecting each consecutive lower half of the data of each element included in the second initial shift operation result, determining the selected data as a first shift operation result, and selecting each consecutive upper half of the data of each element included in the third initial shift operation result, determining the selected data as a second shift operation result, where at least one first target element is determined according to the data included in the first shift operation result, and at least one second target element is determined according to the data included in the second shift operation result. Writing each first target element included in the first shift operation result to the lower half of the position where each first target element of the destination register is arranged. Writing each second target element included in the second shift operation result to the upper half of the position where each second target element of the destination register is arranged. The processor according to claim 28, characterized in that.
33. The instruction further includes a shift amount register identifier for indicating a shift amount register that stores a shift amount. The processor according to claim 21, characterized in that.
34. The opcode is a third type of vector opcode, and the source register includes a first source register. The execution unit executes a selection operation in the first source register according to the third type of vector opcode to obtain a fifth operand, where the selection operation includes an operation of selecting consecutive lower half of the data for each element of the first source register, an operation of selecting consecutive upper half of the data for each element of the first source register, an operation of selecting consecutive intermediate specified bits of data for each element of the first source register, and an operation of selecting non-consecutive specified bits of data for each element of the first source register, including any one of them. Performing an operation of shifting, rounding, and saturating to a half value width on the fifth operand according to the third type of vector opcode and the shift amount to generate a fourth initial shift operation result. Perform a bit selection operation on the fourth initial shift operation result to generate a shift operation result. Here, the bit selection operation includes any one of the following operations: selecting the consecutive lower half data for each element included in the fourth initial shift operation result, selecting the consecutive upper half data for each element included in the fourth initial shift operation result, selecting the data of consecutive specified bits in the middle for each element included in the fourth initial shift operation result, and selecting the data of non-consecutive specified bits for each element included in the fourth initial shift operation result. Sequentially write the data in the shift operation result to the corresponding positions in the destination register. Set the values of the positions in the destination register where the data is not written according to the third type of vector operation code. The processor according to claim 33, characterized in that.
35. The third type of vector operation code is the ninth vector operation code, and the fifth operand is any consecutive source element of the first source register. According to the shift amount, the execution unit performs an operation of right logical shift, rounding, and saturating with a sign to a half value width for each source element included in the fifth operand, and generates a fourth initial shift operation result. Select the consecutive lower half elements of each source element included in the fourth initial shift operation result, and determine the element after the selection operation as the shift operation result. Divide the storage positions in the destination register according to a preset value, and determine the storage area for each target element. Sequentially write the data in the shift operation result to the lower half of each storage area. Set the values of the positions in each storage area where the data is not written to zero respectively. The processor according to claim 34, characterized in that.
36. The third type of vector operation code is the tenth vector operation code, and the fifth operand is any consecutive source element of the first source register. The execution unit performs an operation of right arithmetic shift, rounding, and saturating with sign to half value width on each source element included in the fifth operand according to the shift amount, and generates a fourth initial shift operation result. Selects the consecutive lower half data of each element included in the fourth initial shift operation result respectively, and determines the element after the selection operation as the shift operation result. Divides the storage positions in the destination register according to a preset value, and determines the storage area of each target element. Writes the data in the shift operation result sequentially to the lower half of each storage area. Sets the values of the positions in each storage area where the data has not been written to zero respectively. The processor according to claim 34, characterized by the above.
37. The third type of vector opcode is the eleventh vector opcode, and the fifth operand is any consecutive source element of the first source register. The execution unit performs an operation of right logical shift, rounding, and saturating without sign to half value width on each source element included in the fifth operand according to the shift amount, and generates a fourth initial shift operation result. Selects the consecutive lower half data of each element included in the fourth initial shift operation result respectively, and determines the element after the selection operation as the shift operation result. Divides the storage positions in the destination register according to a preset value, and determines the storage area of each target element. Writes the data in the shift operation result sequentially to the lower half of each storage area. Sets the values of the positions in each storage area where the data has not been written to zero respectively. The processor according to claim 34, characterized by the above.
38. The third type of vector opcode is the twelfth vector opcode, and the fifth operand is any consecutive source element of the first source register. The execution unit performs an operation of right arithmetic shift, rounding, and saturating without sign to half value width on each source element included in the fifth operand according to the shift amount, and generates a fourth initial shift operation result. Selects the consecutive lower half data of each element included in the fourth initial shift operation result respectively, and determines the element after the selection operation as the shift operation result. Divide the storage positions in the destination register according to the preset value to determine the storage area for each target element. Sequentially write the data in the shift operation result to the lower half of each storage area. The processor according to claim 34, wherein the values of the positions in each storage area where the data is not written are each set to zero.
39. The number of source registers is one or more, the number of destination registers is one, and the source register identifier is the same as or different from the destination register identifier. The processor according to any one of claims 21, 34 to 38, characterized in that.
40. There are a plurality of source registers and one destination register. The processor according to any one of claims 21 to 33, characterized in that each source register identifier in all the source registers is different from the destination register identifier, or there is one source register identifier in all the source registers that is the same as the destination register identifier.
41. An electronic device including a memory and one or more processors, wherein one or more programs are stored in the memory and are configured to be executed by the one or more processors according to the vector shift method according to any one of claims 1 to 18. An electronic device, characterized in that.
Citation Information
Patent Citations
Vector Frequency Compress Instruction
CN104011673A
Apparatus and method for right shifting packed quadwords and extracting packed doublewords
CN109947697A
Data shifting method, device and equipment and computer readable storage medium
CN110221807A
Multiplexing operation for SIMD processing
JP2005174297A
Vector arithmetic and logical instructions performing operations on different first and second data element widths from corresponding first and second vector registers
US11003447B2